Systems and methods for masking treatment-affected regions in the genome to improve classifier performance
By identifying and masking treatment-affected features in nucleic acid methylation data, the system enhances the accuracy of cancer detection classifiers, addressing the challenge of distinguishing between cancer-induced and treatment-induced methylation alterations.
Patent Information
- Application Number
- PCT/US2024/060122
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2024-12-13
- Publication Date
- 2025-06-19
AI Technical Summary
Current cancer detection classifiers face challenges in distinguishing between methylation alterations caused by cancer and those induced by cancer treatment, leading to confounding factors that reduce the accuracy and reliability of disease detection.
The system and method involve comparing nucleic acid methylation data from pre-treatment and post-treatment samples to identify treatment-affected features, which are then excluded or masked during the training process of machine learning classifiers to improve their performance.
By masking treatment-affected regions, the classifiers become more robust and accurate in differentiating between cancer-related signals and treatment-induced changes, leading to improved disease detection and tumor fraction estimation.
Smart Images

Figure US2024060122_19062025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR MASKING TREATMENT-AFFECTED REGIONS IN THE GENOME TO IMPROVE CLASSIFIER PERFORMANCECROSS REFERNCE TO RELATED APPLICATION
[0001] This application claims benefit of priority from U.S. Provisional Patent Application No. 63 / 610,143, filed December 14, 2023, which is hereby incorporated by reference in its entirety.TECHNICAL FIELD
[0002] The present disclosure relates generally to the field of bioinformatics and genomics and, more specifically, to systems and methods for improving the performance of a disease state classifier.BACKGROUND
[0003] Cancer is a complex disease characterized by uncontrolled cell growth, and its effective treatment often involves the implementation of one or more different types of therapies, such as chemoradiotherapy (CRT). More particularly, CRT is a widely used approach that combines chemotherapy and radiation therapy to target and eradicate cancerous cells. Although effective in its ability to destroy cancer cells, CRT can induce changes in cell-free DNA (cfDNA) and DNA methylation patterns in cancer subjects. The altered methylation patterns in cfDNA resulting from cancer treatment may introduce confounding factors that can adversely affect the performance of cancer detecting classifiers. One or more aspects of this disclosure may address one or more of the issues described above.
[0004] The background description provided herein is for the purpose of generally presenting the context of the disclosure. Unless otherwise indicated herein,the materials described in this section are not prior art to the claims in this application and are not admitted to be prior art, or suggestions of the prior art, by inclusion in this section.SUMMARY OF THE DISCLOSURE
[0005] According to certain aspects of the disclosure, systems and methods are described for masking treatment-affected genomic regions in training data to improve the performance of a machine learning classifier trained to predict a disease state in a sample.
[0006] In one aspect, a computer-implemented method is provided. The computer-implemented method may include: receiving, at a computing device, a first set of nucleic acid methylation data and a second set of nucleic acid methylation data, wherein the first set of nucleic acid methylation data is associated with a pretreatment sample and wherein the second set of nucleic acid methylation data is associated with a post-treatment sample; comparing, using a processor of the computing device, a first feature set of the first set of nucleic acid methylation data against a second feature set of the second set of nucleic acid methylation data; determining, based on the comparing, at least one treatment affected feature in the second feature set of the second set of nucleic acid methylation data; and implementing, based on the determining, an exclusion process on the at least one treatment affected feature in the second feature set of the second set of nucleic acid methylation data.
[0007] In another aspect, a system is provided. The system may include: one or more processors; one or more computer readable media storing instructions that are executable by the one or more processors to perform operations to: receive, at a computing device associated with the system, a first set of nucleic acid methylationdata and a second set of nucleic acid methylation data, wherein the first set of nucleic acid methylation data is associated with a pre-treatment sample and wherein the second set of nucleic acid methylation data is associated with a post-treatment sample; compare a first feature set of the first set of nucleic acid methylation data against a second feature set of the second set of nucleic acid methylation data; determine, based on the comparing, at least one treatment affected feature in the second feature set of the second set of nucleic acid methylation data; and implement, based on the determining, an exclusion process on the at least one treatment affected feature in the second feature set of the second set of nucleic acid methylation data.
[0008] In yet another aspect, a non-transitory computer-readable medium storing computer-executable instructions is provided. The non-transitory computer- readable medium stores computer-executable instructions which, when executed by a system, may cause the system to perform operations comprising: receiving, at a computing device, a first set of nucleic acid methylation data and a second set of nucleic acid methylation data, wherein the first set of nucleic acid methylation data is associated with a pre-treatment sample and wherein the second set of nucleic acid methylation data is associated with a post-treatment sample; comparing, using a processor of the computing device, a first feature set of the first set of nucleic acid methylation data against a second feature set of the second set of nucleic acid methylation data; determining, based on the comparing, at least one treatment affected feature in the second feature set of the second set of nucleic acid methylation data; and implementing, based on the determining, an exclusion process on the at least one treatment affected feature in the second feature set of the second set of nucleic acid methylation data.
[0009] In yet another aspect, a computer-implemented method is provided.The computer-implemented method may include: receiving, at a computing device, a first set of nucleic acid methylation data and a second set of nucleic acid methylation data, wherein the first set of nucleic acid methylation data is associated with a first sample associated with a first treatment condition and wherein the second set of nucleic acid methylation data is associated with a second sample associated with a second treatment condition; comparing, using a processor of the computing device, a first feature set of the first set of nucleic acid methylation data against a second feature set of the second set of nucleic acid methylation data; determining, based on the comparing, at least one treatment affected feature in the second feature set of the second set of nucleic acid methylation data; and implementing, based on the determining, an exclusion process on the at least one treatment affected feature in the second feature set of the second set of nucleic acid methylation data.
[0010] Additional objects and advantages of the disclosed embodiments will be set forth in part in the description that follows, and in part will be apparent from the description, or may be learned by practice of the disclosed embodiments. The objects and advantages of the disclosed embodiments will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims.
[0011] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosed embodiments, as claimed.BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments and together with the description, serve to explain the principles of the disclosure.
[0013] FIG. 1A depicts an exemplary computer system for executing the methods described herein.
[0014] FIG. 1 B depicts an exemplary software platform for executing the methods described herein.
[0015] FIG. 2 depicts an exemplary workflow for identifying and removing treatment-affected features, according to one or more embodiments of the present disclosure.
[0016] FIG. 3 depicts an exemplary diagram illustrating a process for removing and down-weighting treatment-affected features, according to one or more embodiments of the present disclosure.
[0017] FIG. 4 depicts a flowchart of an exemplary method of masking treatment-affected regions, according to one or more embodiments of the present disclosure.
[0018] FIG. 5 depicts an example computing system, according to one or more embodiments of the present disclosure.DETAILED DESCRIPTION OF EMBODIMENTS
[0019] The terminology used below may be interpreted in its broadest reasonable manner, even though it is being used in conjunction with a detailed description of certain specific examples of the present disclosure. Indeed, certain terms may even be emphasized below; however, any terminology intended to be interpreted in any restricted manner will be overtly and specifically defined as such inthis Detailed Description section. Both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the features, as claimed.
[0020] Cancer is a significant health concern worldwide, and various treatment modalities are employed to combat this disease. One of the most common approaches to fight cancer is chemoradiotherapy (CRT). CRT combines the use of chemotherapy, which involves medications that target cancer cells throughout the body, and radiation therapy, which employs high-energy radiation beams, to destroy cancerous cells. While CRT is effective in many cases, it may affect the subject’s body, including modifying cell-free DNA (cfDNA) and the release of cfDNA into the bloodstream.
[0021] CfDNA is genetic material that originates from cells and can be found circulating in the plasma of blood, among other biofluids. Samples of cfDNA may carry information about the disease state of a subject from which the sample was extracted. In the context of cancer, cfDNA can carry information about the presence of genetic mutations associated with cancer. Detecting disease-related, e.g., cancer- related, signals in cfDNA has shown promise as a non-invasive method for disease, e.g., cancer, diagnosis and monitoring. For example, one observation in cancer subjects undergoing CRT is increased levels of cfDNA in their plasma. This phenomenon is attributed to the destruction of cancer cells and the release of their DNA into the bloodstream.
[0022] In some instances, the treatment for a disease — or another treatment that a subject may undergo — may cause changes in cfDNA that confounds the ability to use cfDNA to detect the disease. One challenge in cancer detection lies in distinguishing the presence of cancer-specific signals in cfDNA from the effectsinduced by cancer treatment, such as CRT. For example, the quantity of cfDNA may change during cancer treatment and / or CRT may induce changes to the DNA methylation patterns within the human genome.
[0023] Changes in methylation patterns may affect cancer detection and diagnosis, as they can alter the genetic markers used by diagnostic classifiers trained to identify cancer. More particularly, current state-of-the-art classifiers may analyze cfDNA to identify the presence of cancer and estimate tumor fraction. The altered methylation patterns in cfDNA resulting from cancer treatment may introduce confounding factors that can adversely affect the performance of these classifiers and may hinder the accurate estimation of tumor fractions in subjects undergoing treatment. These effects can limit the ability of cancer classifiers to predict and detect the progress of the subjects’ treatment. This in turn can limit the ability of practitioners and researchers to accurately tailor treatment to their patients’ specific disease progress and response to the treatment.
[0024] In view of the foregoing, it can be appreciated that distinguishing between methylation alterations caused by a disease itself, e.g., cancer, and those caused by treatment, may be challenging. When these treatment-induced methylation changes are not adequately accounted for, confounding factors may be introduced into training data that a disease-detecting classifier is trained on, thereby reducing its accuracy and reliability in disease detection. Moreover, accurate estimation of tumor fractions in subjects undergoing treatment may facilitate monitoring disease progression and / or making better-informed treatment decisions. Inaccuracies in this estimation may impact disease detection.
[0025] Accordingly, the present disclosure is designed to address the impact of treatment-induced changes in methylation patterns of nucleic acids, includingcfDNA, in biopsy samples on the performance of disease, e.g., cancer, detecting classifiers. More particularly, the concepts described herein may be utilized to identify methylation patterns specifically affected by disease treatment, such as cancer treatment, and then “mask” these regions during classifier training. By systematically identifying and masking these treatment-affected methylation regions, the described concepts enable the classifier to differentiate the genuine disease- related, e.g., cancer-related, signals in cfDNA from those influenced by treatment. This process may be facilitated by first comparing cfDNA samples from subjects undergoing treatment (“on-treatment” subjects) with those who have not yet undergone treatment (“treatment-naive” subjects). One or more statistical tests may be performed to evaluate the methylation differences between these two groups of subjects. Through this comparison, differentially methylated regions may be identified (e.g., genomic regions with test statistics that pass a predefined significance threshold and fold-change cutoff may be classified as differentially methylated regions). The identified treatment-affected regions may thereafter be masked from the training input that the classifier is trained on.
[0026] The concepts described herein integrate various technological improvements and correspondingly improve the functionality of a computing device as used in the detection and treatment of cancer in several ways. For instance, in an aspect, the described concepts make the computer-based disease detection system more robust to the confounding effects of treatment on cfDNA methylation patterns. By masking treatment-affected regions, the computer executing and implementing a classifier model trained and used according to the techniques described herein, is better equipped to provide reliable disease, e.g., diagnoses, and tumor fraction estimates, even when treatment has altered the genomic methylation landscape.More particularly, the described concepts improve the model’s pattern recognition capability, thereby enabling it to extract meaningful insights from biological data that would be difficult, if not impossible, for humans to do. Additionally, the improved model resulting from the implementation of the described concepts may be more suitable for clinical applications, such as minimal residual disease (MRD) detection and monitoring changes in ctDNA levels over time. This improved functionality benefits healthcare professionals by providing them with more accurate and clinically relevant information for subject care. Furthermore, the processes executed by the computer to improve the model involve complex calculations and data manipulations on a large amount of biological data that a human individual could not reasonably complete on their own or in their mind. Specifically, computationally intensive statistical tests are leveraged by the computer to evaluate methylation differences between sample sets, processes which cannot be completed by a human.
[0027] The subject matter of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, which form a part hereof, and which show, by way of illustration, specific exemplary embodiments. An embodiment or implementation described herein as “exemplary” is not to be construed as preferred or advantageous, for example, over other embodiments or implementations; rather, it is intended to reflect or indicate that the embodiment(s) is / are “example” embodiment(s). Subject matter may be embodied in a variety of different forms and, therefore, covered or claimed subject matter is intended to be construed as not being limited to any exemplary embodiments set forth herein; exemplary embodiments are provided merely to be illustrative. Likewise, a reasonably broad scope for claimed or covered subject matter is intended. Among other things, for example, subject matter may be embodied as methods, devices,components, or systems. Accordingly, embodiments may, for example, take the form of hardware, software, firmware, or any combination thereof. The following detailed description is, therefore, not intended to be taken in a limiting sense.
[0028] Throughout the specification and claims, terms may have nuanced meanings suggested or implied in context beyond an explicitly stated meaning. Likewise, the phrase “in one embodiment” or “in some embodiments,” or “in one aspect” or “in some aspects” as used herein does not necessarily refer to the same embodiment or aspect, and the phrase “in another embodiment” or “in another aspect” as used herein does not necessarily refer to a different embodiment or aspect. It is intended, for example, that claimed subject matter include combinations of exemplary embodiments in whole or in part.
[0029] Non-limiting cancer types that the concepts described herein may be applied to include, for example, breast cancer, lung cancer (e.g., non-small cell lung cancer (NSCLC)), prostate cancer, colorectal cancer, renal cancer, uterine cancer, pancreatic cancer, cancer of the esophagus, a lymphoma, head and neck cancer, ovarian cancer, a hepatobiliary cancer, a melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, and gastric cancer. Additionally, it is also important to note that although the concepts described throughout this disclosure are made in reference to cancer, these designations are for exemplary purposes only and are not intended to be limiting. Specifically, the concepts described herein may be applicable to other disease types and other diseasedetecting classifiers. More generally, the approaches described herein may be applied to scenarios in which subjects are subject to conditions other than cancer treatment, e.g., a treatment relevant for the disease state being detected, that mayalso introduce confounding methylation patterns in nucleic acids, including plasma cfDNA, and affect the performance of a relevant disease-detecting classifier.
[0030] FIG. 1 A depicts an exemplary system for masking treatment-affected regions to enhance the performance of a classifier. Exemplary system 100 includes a data collection component 10, a database 20, and device data intelligence component 30, operably connected to each other via network 40. Alternatively, or additionally, one or more of the components may be connected with another component locally without reliance on network connection; e.g., through a wired connection. In many aspects described herein, sequencing data of cell-free nucleic acids are used to illustrate the concepts. However, one of skill in the art would understand that the current method may be applied to sequencing data of DNA, RNA, or other materials, as well from a variety of sample types, e.g., a blood sample (e.g., a serum sample, a plasma sample, a whole blood sample), a urine sample, a saliva sample, a tissue sample, a bone marrow sample, etc.
[0031] As disclosed herein, data collection component 10 may include a device or machine with which sequencing data may be generated. In some embodiments, data collection component 10 may include one or more sequencing devices or a facility that uses one or more sequencing devices to generate nucleic acid (e.g., DNA or RNA) sequence data of biological samples. In some aspects, data collection 10 may be a database that receives sequencing information generated from one or more sequencing devices. Any suitable liquid or solid biological samples may be used for sequencing. In some embodiments, a biological sample may be cell-based, for example, one or more types of tissue. In some embodiments, a biological sample may be a sample that includes cell-free nucleic acid fragments. Examples of biological samples include, but are not limited to, a blood sample (e.g.,a cell-free DNA (cfDNA) sample, a cell-free RNA (cfRNA) sample, a serum sample, a plasma sample, a whole blood sample), a urine sample, a saliva sample, a tissue sample, a bone marrow sample, etc. Further, although sequencing of DNA from these samples is discussed herein, RNA from these samples may alternatively or additionally be sequenced.
[0032] Examples of sequencing data may include, but are not limited to, sequence read data of targeted genomic locations, partial or whole genome sequencing data of the genome represented by nucleic acid fragments in cell-free or cell-based samples, partial or whole genome sequencing data including one or more types of epigenetic modifications (e.g., methylation), or combinations thereof.
[0033] Data acquired by the data collection component 10 may be transferred to database 20 via network 40 or a local or network connection. In some embodiments, data collection component 10 may alternatively receive data from one or more sequencing devices. In some embodiments, the collected data may be analyzed by data intelligence component 30, via network 40 or a local or network connection. FIG. 1 B depicts exemplary functional modules that may be implemented to perform tasks of data intelligence component 30.
[0034] FIG. 1 B depicts an exemplary computer system 110 for masking treatment-affected regions to enhance the performance of a classifier. Exemplary system 110 achieves such functionalities by implementing, on one or more computer devices, user input and output (I / O) module 120, memory or database 130, data processing module 140, data analysis module 150, classification module 160, network communication module 170, and any other functional modules that may be needed for carrying out a particular task (e.g., an error correction or compensation module, a data compression module, etc.). As disclosed herein, user I / O module 120may further include an input sub-module, such as a keyboard, and an output submodule, such as a display (e.g., a printer, a monitor, or a touchpad). In some embodiments, all functionalities may be performed by one computer system. In some embodiments, the functionalities are performed by more than one computer system.
[0035] Also disclosed herein, a particular task may be performed by implementing one or more functional modules. In particular, each of the enumerated modules itself may, in turn, include multiple sub-modules. For example, data processing module 140 may include a sub-module for data quality evaluation (e.g., for discarding very short sequence reads or sequence reads including obvious errors), a sub-module for normalizing numbers of sequence reads that align to different regions of a reference genome, a sub-module to compensate / correct guanine-cytosine (GC) biases, a sub-module for matching data associated with a cancer sample with other data associated with one or more non-cancer samples, etc.
[0036] In some embodiments, a user may use I / O module 120 to manipulate data that is available either on a local device or can be obtained via a network connection from a remote service device or another user device. For example, I / O module 120 may allow a user, e.g., via a keyboard, a mouse, or a touchpad, to initiate or perform data analysis via a graphical user interface (GUI). In some embodiments, a user may manipulate data via voice control. In some embodiments, user authentication may be required before a user is granted access to the data being requested. In some embodiments, user I / O module 120 may be used to manage various functional modules. For example, a user may request via user I / O module 120 input data while an existing data processing session is in process. A user may do so by selecting a menu option or type in a command discretely without interrupting the existing process. In another example, a user may utilize user I / Omodule 120 to set various thresholds, configure sample matching settings, and / or provide other instructions to computer system 110 that dictate how treatment- affected regions are identified and / or masked. As disclosed herein, a user may use any type of input to direct and control data processing and analysis via I / O module 120.
[0037] In some embodiments, system 110 further comprises a memory or database 130. In some embodiments, database 130 comprises a local database that may be accessed via user I / O module 120. In some embodiments, database 130 comprises a remote database that may be accessed by user I / O module 120 via network connection. In some embodiments, database 130 is a local database that stores data retrieved from another device (e.g., a user device or a server). In some embodiments, memory or database 130 may store data retrieved in real-time from internet searches. In some embodiments, database 130 may send data to and receive data from one or more of the other functional modules, including, but not limited to, a data collection module (not shown), data processing module 140, data analysis module 150, classification module 160, network communication module 170, and etc. In some embodiments, some or all of the pre-treatment and posttreatment sample data may be stored on database 130.
[0038] In some embodiments, database 130 may be a database local to the other functional modules. In some embodiments, database 130 may be a remote database that may be accessed by the other functional modules via wired or wireless network connection (e.g., via network communication module 170). In some embodiments, database 130 may include a local portion and a remote portion.
[0039] In some embodiments, system 110 comprises a data processing module 140. Data processing module 140 may receive data from I / O module 120 ordatabase 130. In some embodiments, data processing module 140 may perform standard data processing algorithms, such as one or more of noise reduction, signal enhancement, normalization of counts of sequence reads, correction of GC bias, etc. In some embodiments, data processing module 140 may be configured to identify features in pre-treatment and post-treatment DNA methylation data. For example, computer system 110 may be able to identify one or more differentially methylated regions (DM Rs), which are regions where DNA methylation varies significantly between different biological samples, e.g., a pre-treatment and a post-treatment sample. From these identified features, data processing module 140 may be configured to identify one or more treatment-affected features and subsequently remove them from consideration in one or more downstream processes (e.g., feature selection, classifier training, etc.).
[0040] In some embodiments, system 110 comprises a data analysis module 150. In some embodiments, data analysis module 150 includes identifying and treating systematic errors in sequencing data, as described in connection with data processing module 140.
[0041] In some embodiments, system 110 comprises a classification module 160, which may embody a “machine-learning model” or “trained classifier.” As used herein, a “machine-learning model” or “trained classifier” generally encompasses instructions, data, and / or a model configured to receive input, and apply one or more of a weight, bias, classification, or analysis on the input to generate an output. The output may include, for example, a classification of the input, an analysis based on the input, a design, process, prediction, or recommendation associated with the input, or any other suitable type of output. A machine-learning model is generally trained using training data, e.g., experiential data and / or samples of input data,which are fed into the model in order to establish, tune, or modify one or more aspects of the model, e.g., the weights, biases, criteria for forming classifications or clusters, or the like. Aspects of a machine-learning model may operate on an input linearly, in parallel, via a network (e.g., a neural network), or via any suitable configuration. In the context of this disclosure, the machine-learning model may be trained on a combination of real and synthetic sample data.
[0042] The execution of the machine-learning model may include deployment of one or more machine-learning techniques, such as k-nearest neighbors, linear regression, logistic regression, random forest, gradient boosted machine (GBM), deep learning, a deep neural network, and / or any other suitable machine-learning technique that solves problems in the field of Natural Language Processing (NLP). Supervised, semi-supervised, and / or unsupervised training may be employed. For example, supervised learning may include providing training data and labels corresponding to the training data, e.g., as ground truth. Unsupervised approaches may include clustering, classification or the like. K-means clustering or K-Nearest Neighbors may also be used, which may be supervised or unsupervised. Combinations of K-Nearest Neighbors and an unsupervised cluster technique may also be used. Any suitable type of training may be used, e.g., stochastic, gradient boosted, random seeded, recursive, epoch or batch-based, etc.
[0043] In an exemplary use case, a machine-learning model may be trained to analyze data from a test sample from a test subject whose status with respect to a medical condition is unknown and subsequently classifies the unknown test sample from the test subject based on the likelihood of the subject fitting into a particular category. In some embodiments, the one or more parameters may include a binomial probability score that is calculated based on logistic regression analysis. Asdisclosed herein, the binomial probability score may correspond to the likelihood of a subject having a certain medical condition, such as cancer (e.g., NSCLC). For example, a score of over a predefined threshold may indicate that the subject associated with a test sample is more likely to have cancer than not have cancer. In some embodiments, the one or more parameters may include a sequencing or methylation data distribution pattern correlating with the presence of cancer. A subject associated with a test sample having sequencing or methylation data with a pattern resembling the cancer pattern to a sufficient degree may be predicted as having cancer. In some embodiments, a sequencing or methylation data distribution pattern may be identified in connection with a specific type of cancer, determining a tissue of origin or cancer signal origin, thus allowing a test sample to be classified as indicative of a certain cancer type.
[0044] As disclosed herein, network communication module 170 may be used to facilitate communications between a user device, one or more databases, and any other suitable system or device through a wired or wireless network connection. Any communication protocol / device may be used, including, without limitation, a modem, an Ethernet connection, a network card (wireless or wired), an infrared communication device, a wireless communication device, and / or a chipset (such as a Bluetooth™ device, an 802.11 device, a WiFi device, a WiMax device, cellular communication facilities, etc.), a near-field communication (NFC), a Zigbee communication, a radio frequency (RF) or radio-frequency identification (RFID) communication, a PLC protocol, a 3G / 4G / 5G / LTE based communication, and / or the like. For example, a user device having a user interface platform for processing / analyzing tumor fraction data may communicate with another user device with the same platform, a regular user device without the same platform (e.g., aregular smartphone), a remote server, a physical device of a remote loT local network, a wearable device, a user device communicably connected to a remote server, and etc.
[0045] The functional modules described herein are provided by way of example. It will be understood that different functional modules may be combined to create different utilities. It will also be understood that additional functional modules or sub-modules may be created to implement a certain utility.
[0046] Referring now to FIG. 2, an exemplary workflow 200 is provided for masking treatment-affect regions in training data that may be provided to a machine learning classifier. Aspects of the exemplary workflow 200 may be performed in accordance with some or all components described in FIG. 1 A and 1 B.
[0047] At step 205, biological samples may be collected from two distinct groups of subjects: on-treatment subjects and treatment-naive subjects. The former may include those individuals who may currently be undergoing cancer treatment, which may include chemotherapy, radiotherapy, or other therapeutic interventions or who may have previously undergone cancer treatment. The latter may include those individuals who have not yet received any cancer treatment or who have not yet received the relevant cancer treatment and serve as a reference group to compare against the on-treatment group. In an aspect, either the pre-treatment samples associated with the treatment-naive subjects or the post-treatment samples associated with the on-treatment subjects may be cfDNA samples or cfRNA samples. For simplicity purposes, the remainder of this disclosure is described with the biological samples being cfDNA samples. In an aspect, the collection of cfDNA samples may be carried out using minimally invasive methods, such as blood draws or other plasma collection procedures commonly used to obtain cfDNA. That said,any suitable method of sample collection and any suitable sample type may be collected at step 205. In an aspect, a single sample may be collected for a subject from each group. Alternatively, in another aspect, multiple samples for a subject may be collected (e.g., multiple samples may be collected for the subject at a single time point, multiple samples may be collected for the subject across two or more different time points, etc.). For example, an individual may provide a first sample at a first time, before cancer treatment, in which case the first sample may be included in the reference group and then later provide a second sample at a second time, after cancer treatment, in which case the second sample may be included in the on- treatment group. Further, step 205 may not include active sample collection and may instead refer to receipt of samples and / or data associated with samples that were previously collected. In one aspect, metadata associated with each sample may be collected, e.g., subject demographics, medical history, treatment regimens, and any relevant clinical information. This metadata may in some aspects provide context for interpreting methylation patterns and understanding the impact of cancer treatment on these patterns.
[0048] At step 210, a methylation analysis may be conducted on the collected samples. Methylation analysis may involve the assessment of DNA methylation patterns at specific genomic regions, for example, at cytosine- phosphate-guanine (CpG) sites. In an aspect, various high-throughput technologies may be employed for methylation profiling, including one or more of bisulfite sequencing, methylated DNA immunoprecipitation sequencing (MeDIP-seq), DNA methylation microarrays, and the like. For simplicity purposes, bisulfite sequencing is the methylation profiling technique described herein, however, this designation is not intended to be limiting.
[0049] In an aspect, bisulfite sequencing may involve the treatment of DNA with sodium bisulfite, which converts unmethylated cytosines (C) into uracils (U) while leaving methylated cytosines unchanged. After bisulfite treatment, the DNA may be subjected to high-throughput sequencing, such as next-generation sequencing (NGS), to determine the methylation status of individual CpG sites across the genome. Whole-genome bisulfite sequencing (WGBS) provides comprehensive coverage of CpG sites and allows for a detailed assessment of methylation patterns. The generated methylation data may undergo bioinformatics analysis. In this regard, the methylation data may first undergo one or more preprocessing steps to ensure the quality and integrity of the methylation data. These steps may include one or more of: data cleaning, quality control, and the removal of artifacts or outliers that may affect the accuracy of the analysis. Preprocessing may also involve the alignment of sequence reads to a reference genome. The ratio of C to T at each CpG site may be used to calculate the methylation level.
[0050] In an aspect, the methylation level at each CpG site may be represented by a beta value, which are typically reported as decimal values ranging from 0 to 1 . A beta value of 0 indicates that the CpG site is completely unmethylated. A beta value of 1 .0 indicates that the CpG site is completely methylated. A beta value of 0.50 indicates that the CpG site is 50% methylated. Beta values offer a straightforward interpretation of DNA methylation levels. For example, a beta value of 0.2 at a specific CpG site suggests that 20% of the DNA molecules at the site are methylated, while the remaining 80% are unmethylated.
[0051] At step 215, the results of the methylation analysis may be utilized to identify a plurality of features in the pre-treatment methylation data and the post-treatment methylation data that are informative for a particular task, such as disease classification or prediction. In the context of this application, each “feature” may be representative of a CpG site, a plurality of CpG sites, and / or a particular genomic region and may be a candidate for use in classifier training. For instance, each feature may correspond to a differentially methylated region (DMR), which are genomic segments that are characterized by statistically significant changes in DNA methylation levels (e.g., between pre-treatment subjects and post-treatment subjects).
[0052] To facilitate DMR identification, and correspondingly relevant feature selection for each sample set, one or more statistical tests may be employed to quantitatively measure the degree of methylation difference at each CpG site by comparing the methylation levels (e.g., beta values or methylation percentages) of on-treatment subjects with those of treatment-naive subjects. In an aspect, the choice of statistical test may depend, at least in part, on the specific research goals and characterizations of the data. Commonly used tests that may be utilized include t-tests, Analysis of Variance (ANOVA), Wilcoxon tests, heat mapping, or specialized tests for high-throughput DNA methylation data, such as differential methylation analysis packages like DMRcate or Differentially Methylated CpG Island (DMCI) analysis. These statistics quantify the magnitude and significance of methylation differences, providing a numerical measure of how different the methylation patterns are between the two groups.
[0053] To identify meaningful differences, predefined significance thresholds may be established. These thresholds may include p-values and adjusted p-values(e.g., FDR correction) to control for multiple testing. Methylation differences that pass these significance thresholds may be considered statistically significant for featureselection. Additionally or alternative to the foregoing, in some aspects, fold-change cutoffs may be utilized to determine whether a change in a particular measurement (e.g., methylation level) between pre-treatment and post-treatment methylation data is considered biologically significant. Fold-change cutoffs may ensure that only substantial methylation differences are considered for further analysis. In an aspect, the fold change may be calculated by comparing two measurements, such as methylation levels at a CpG site, between two groups. If the fold change is less than a predetermined threshold (e.g., 2-fold, 3-fold, etc.), then the relevant CpG site may be categorized as statistically insignificant. Conversely, if the fold change is greater than the predetermined threshold, then the relevant CpG site may be characterized as a DMR and / or a relevant feature.
[0054] In an aspect, DMRs may be annotated to provide information about their genomic location and context. Annotations may include details about whether the DMRs overlap with genes, promoters, enhancers, CpG islands, or other functional elements. In an aspect, not all identified DMRs may be biologically relevant or necessary for subsequent analysis. Therefore, priority may be assigned to those DMRs that are associated with specific genes or pathways of interest, e.g., those that are relevant for cancer detection. In some aspects, methylation difference evaluation may involve integrating methylation data with other relevant clinical or biological data, such as gene expression profiles, treatment history, or treatment response information, to gain a more comprehensive understanding of the observed methylation changes.
[0055] At step 220, those features that are representative of treatment- affected methylation patterns may be identified. More particularly, because these treatment-affected features are indicated as having a biological origin derived from atreatment effect (e.g., CRT in connection with cancer treatment), they are expected to be observed in the post-treatment methylation data and not in the pre-treatment methylation data. Treatment-affected features may be observed if there are statistically significant associations such that a feature is not necessarily present or active in pre-treatment methylation data but is absent or inactive in post-treatment methylation data, or vice versa. In an aspect, the identification of these treatment- affected features may be facilitated by identifying those features in the posttreatment methylation data that satisfy a certain set of criteria. For instance, one criteria is that these treatment-affected features may exhibit a specific methylation pattern. For example, a single feature may be defined as a set of a predefined number of consecutive CpG sites (e.g., 5 consecutive CpG sites). If this predefined number of consecutive CpG sites exhibit methylation states associated with previously identified treatment-affected regions (e.g., all of the CpG sites in the set are methylated) then this may provide an indication that the relevant feature is treatment affected. In general, methyl variants active in on-treatment subjects but not pre-treatment subjects are sought.
[0056] In an aspect, tumor fraction monitoring may be utilized to help distinguish whether a genomic feature is the result of cancer treatment or is indicative of a new cancer. For instance, before treatment, a baseline tumor fraction (i.e. , a reference point) may be established for a subject, which represents the amount of tumor-derived DNA in the subject’s bloodstream before any treatment is initiated. After treatment, the tumor fraction may be monitored over time. If the tumor fraction remains relatively stable or decreases during or after treatment, but specific features emerge or change significantly in post-treatment samples, it may suggest that these features are related to the treatment process. Stated differently, treatment-affected features may be expected to appear at a frequency or level that is inconsistent with the tumor fraction. Conversely, in the case of new cancer development or cancer recurrence, there is typically an increase in the tumor fraction, as the tumor grows or re-emerges. If the emergence of new features in cfDNA coincides with a significant increase in the tumor fraction, it may be more likely that these features are associated with the development of a new cancer or cancer recurrence.
[0057] Additionally or alternatively to the foregoing, other considerations may also be leveraged to determine whether a feature is treatment derived or cancer derived. For instance, as an example, the clinical history of a subject may be examined to determine whether the subject has undergone a recent cancer treatment. If they have, then the appearance of certain features in the post-treatment sample may be more likely to be linked to the treatment effect. However, if there is no history of treatment, the same features may raise suspicion of new cancer. In another example, the time frame of feature appearance may also be a consideration. More particularly, features appearing shortly after a treatment may be more likely to be attributed to treatment effects, while those developing later may be more indicative of new cancer. Other considerations, not explicitly disclosed here, may also be leveraged.
[0058] In some aspects, utilizing some or all of the considerations described above, computer system 110 may assign a score (e.g., represented by a numeral, a percentage, etc.) to each identified feature. The score may be representative of the determinations made by computer system 110 that the feature is either resultant from a treatment or is cancer-derived. In an aspect, computer system 110 may compare the score against a predetermined threshold (e.g., a treatment-affectedfeature determination threshold) that may be established to delineate between treatment-affected features and cancer-related features. More particularly, those features having a score above the threshold may be considered to be associated with cancer whereas those features having a score below the threshold may be considered to be treatment-derived, or vice versa.
[0059] At step 225, computer system 110 may be configured to selectively exclude, or “mask,” the genomic regions that have been identified as treatment- affected from the input genomic data. This exclusion process may be implemented to prevent the cancer detection algorithms associated with a machine-learning classifier from being influenced by methylation changes resulting from cancer treatment, which may introduce confounding factors.
[0060] In an aspect, a classifier may first be established that is configured to operate within predefined methyl variant regions. These regions may be composed of specific sets of 5 CpG sites (CpGs) (e.g., 5 CpGs, etc.) with characteristic methylation patterns indicative of cancer. The synchronized alignment of the classifier and the feature set may be performed to ensure that the classifier relies on the predetermined methyl variant regions as the basis for its predictions. Additionally or alternatively to the foregoing, in an aspect, the masking process may involve marking the genomic coordinates corresponding to specific CpG sites, DMRs, etc., which are associated with a treatment-affected feature. These coordinates may be flagged to the computer system as regions not to be considered during subsequent algorithmic analysis. Specifically, the computer algorithm or software utilized for analysis may be programmed to recognize and exclude data points falling within these masked regions.
[0061] In some aspects, each treatment-affected feature may be identified and masked before the formal feature selection process (i.e., in which the most informative features that contribute to accurate cancer detection are identified) to prevent their influence on a cancer-detecting classifier. Stated differently, the target methyl variant regions may be excluded from the dataset used to train the classifier. These regions are effectively removed from the training data to allow the classifier to focus on the CpGs outside of the treatment-affected regions. However, if residual bias is detected due to the presence of treatment-affected features after the masking step, additional steps may be taken to correct this bias. For example, an algorithm’s weights associated with these features may be minimized or set to zero. This may cause a classifier to focus more on correctly classifying non-treatment affected regions while giving less importance to the treatment-affected regions. In some aspects, the weighting process may be sample based. For instance, if individual samples are associated with different degrees of treatment effect (i.e., some samples are more strongly affected by treatment than others), then different weights may be assigned to each sample. For example, samples from subjects with minimal treatment effects may receive higher weights, while those from subjects with significant treatment effects may receive lower weights. In other aspects, any samples from subjects with any treatment effects may receive zero weight. In still other aspects, all data from a sample showing treatment effects may be weighted less or given zero weight, or only the portion of the sample data showing treatment effects may be weighted less or given zero weight.
[0062] In an optional aspect, the masking process may be initially implemented within a two-fold cross-validation approach. In cross-validation, the dataset is divided into two subsets: a training subset and a validation subset. Theclassifier may be trained on one subset and validated on the other. The subsets may be swapped to promote comprehensive training and validation.
[0063] Referring now to FIG. 3, diagram 300 provides a schematic diagram of treatment-affected feature identification and removal. Section 305 presents a first list of pre-treatment features and a second list of post-treatment features. The pretreatment features may be present in the samples of treatment-naive subjects whereas the post-treatment features may be present in the samples of on-treatment subjects. In an aspect, each feature set may contain features that are associated with different diseases (e.g., adenocarcinoma and squamous cell carcinoma (SCC)). More particularly, different diseases may have certain features that are shared and some features that are distinct. For instance, focusing on section 305, adenocarcinoma and SCC may each have a subset of features associated with their distinctive disease but may also have a plurality of shared features that are found in both diseases. In an aspect, some features between the pre-treatment features list and the post-treatment features list may be different, which is resultant from the implementation of the therapy.
[0064] As can be observed in section 305, the list of post-treatment features may also contain a plurality of features associated with confounding treatment effects. Specifically, these are treatment-affected features that originate as a result of being a byproduct of a certain cancer treatment (e.g., CRT). In an aspect, components of system 110 may be configured to identify and computationally mask these treatment-affected features, as previously described above, prior to the feature selection process at 310. In an aspect, after feature selection 310 is complete, a subset of the total features presented in section 305 may remain, as presented in section 315. In some situations, each treatment-affected feature may be correctlyidentified and masked by system 110 prior to feature selection 310. However, in other instances, as presented in section 315, a subset of treatment-affected features may remain. These instances may arise, e.g., when computer system 100 retains a feature because it cannot confidently conclude that the feature is exclusively derived from a treatment effect, rather than being associated with cancer status (e.g., as a result of sample size limitations). For example, a score assigned to a particular feature may fall into a designated “buffer” range situated around a threshold (e.g., where features having scores above the threshold are considered to be cancer- related and features having scores below the threshold are considered to be originate from a treatment). Computer system 100 may be configured to conclude that scores within this “buffer” range are too close to classify either way and may be configured to maintain these features in the selected feature pool. To address these situations, components of system 110 (e.g., data analysis module 150, etc.) may be configured to weigh some or all of the remaining treatment-affected features lower (e.g., assigned weights close to or at zero) than other features so that when a classifier model is trained on the selected features at 320, the impact of any treatment-affected feature is reduced. For example, section 325 illustrates that out of the two treatment-affected features that survived masking and feature selection (i.e. , features 32 and 34), feature 32 was lower weighted than feature 34. In an aspect, although not all treatment-affected features were identified, removed, and / or “zero- weighted,” the majority of them were, which should help to improve classifier performance. Additionally or alternatively to the foregoing, in an aspect, computer system 110 may further filter the selection of features that the classifier is trained on based on the identification that there is low noise associated with the feature. Moreparticularly, the lack of noise may provide an indication that the methylation pattern exhibited by the CpG sites associated with the feature are rare in non-cancer cfDNA.
[0065] Referring now to FIG. 4, an exemplary workflow for identifying and masking treatment-affected regions in the genome is disclosed. The exemplary workflow may be performed, e.g., by components of computer systems 100, 110 (shown in FIG. 1 A and 1 B).
[0066] At step 405, computer system 110 may receive a first and second set of sequencing data that is associated with a pre-treatment and a post-treatment sample, respectively. In an aspect, each pre-treatment sample may be associated with a treatment-naTve subject that has not undergone any type of therapy or treatment for a disease condition, or has not undergone any type of relevant therapy or treatment that may give rise to treatment-related features. Conversely, each posttreatment sample may be associated with an on-treatment subject that has received some type of therapy for their disease (e.g., CRT). In an aspect, each of the first and second set of sequencing data may be DNA methylation data derived, e.g., from a WGBS or cfDNA approach.
[0067] At step 410, computer system 110 may be configured to compare a first feature set in the first set of sequencing data to a second feature set in the second set of sequencing data. More particularly, in an aspect, methylation analysis may be conducted on each set of sequencing data. This analysis may be conducted to identify the methylation level at each CpG site (e.g., which may be represented by a “beta value” that ranges from 0 to 1 ). The results of the methylation analysis may be utilized to identify a plurality of features in the pre- and post-treatment methylation data sets that are representative of the genomic information conveyed therein and / or that most distinguish each data set from the other.
[0068] At step 415, at least one treatment-affected feature in the second feature set of the second set of sequencing data may be determined based on the comparison conducted at step 410. More particularly, in an aspect, the at least one treatment-affected feature may be identified as a genomic region that contains a specific methylation pattern (e.g., a methylation pattern that is emblematic of previously observed methylation patterns that originated in response to a treatment) and that was not also present in the first feature set. The coordinates of the genomic regions associated with the treatment-affected region(s) may be flagged and communicated to computer system 110. In another aspect, the treatment-affected feature may correspond to a genomic region containing a methylation pattern that is present in the first feature set but is absent in the second feature set. More particularly, as a result of a particular cancer treatment, a methylation state of one or more CpGs sites associated with a feature in the first feature set may be affected so that the feature no longer appears as such in the second feature set. For instance, this may result when an administered treatment eliminated a normally shedding population of non-cancer cells, thereby causing a feature associated with those noncancer cells in the untreated population to be absent in the treated population.
[0069] At step 420, upon identification of the at least one treatment-affected feature at step 415, computer system 110 may be configured to implement an exclusion process on the second feature set. More particularly, computer system 110 may leverage knowledge of the genomic locations associated with the treatment- affected regions to remove or computationally mask the treatment-affected region(s) from the second feature so that they are not considered or utilized in any type of downstream processes (e.g., formal feature selection, classifier training, etc.).
[0070] In an aspect, after a feature selection process has occurred (e.g., to select those features in the pre- and post-treatment data that are the most informative indicators of the presence or absence of a disease condition), computer system 110 may determine whether any treatment-affected regions still remain. Responsive to determining that at least one treatment-affected feature is still present in the formal feature set, computer system 110 may apply a weighting operation to the features wherein the treatment-affected features are down-weighted (e.g., to zero or close to zero) so that their subsequent impact on classifier training is minimized.
[0071] In an aspect, various applications exist in the post diagnostic space in which a classifier of the described embodiments may be leveraged to generate a quantitative estimate of the tumor fraction that is robust to other various biological processes occurring within a subject, e.g., inflammation, cells dying, etc. A classifier of the described embodiments may be capable of improving the limit of detection (LOD) for MRD analysis. More particularly, by masking treatment-affected regions in cfDNA methylation patterns, the MRD analysis may be more specific to cancer- related methylation changes. This improved specificity may inhibit the occurrence of false-positive results, meaning that subjects are not incorrectly identified as having residual disease when they do not, thereby improving LOD in terms of precision.
[0072] Although the foregoing aspects have been generally described with reference to analysis of nucleic acid methylation data from pre-treatment samples and post-treatment samples, such designations are not limiting. For instance, computer system 110 may receive a first and second set of sequencing data that is associated with a first sample associated with a first treatment condition and a second sample associated with a second treatment condition, respectively. In anaspect, the first sample associated with the first treatment condition may be associated with an on-treatment subject that has received a first type of therapy for their disease (e.g., CRT), and the second sample associated with the second treatment condition may also be associated with an on-treatment subject that has also received the first type of therapy for their disease, as well as a second type of therapy (e.g., immunotherapy, hormone therapy, etc.). In another aspect, the first sample associated with the first treatment condition may be associated with an on- treatment subject that has received a therapy for their disease at a first time (e.g., X amount of days prior to sample collection) and the second sample associated with the second treatment condition may be associated with an on-treatment subject that has received the same therapy for their disease at a second time (e.g., X+Y amount of days prior to sample collection).
[0073] In general, any process discussed in this disclosure that is understood to be computer-implementable may be performed by one or more processors of a computer system, such as system environment 110, as described above. A process or process step performed by one or more processors may also be referred to as an operation. The one or more processors may be configured to perform such processes by having access to instructions (e.g., software or computer-readable code) that, when executed by the one or more processors, cause the one or more processors to perform the processes. The instructions may be stored in a memory of the computer server. A processor may be a central processing unit (CPU), a graphics processing unit (GPU), or any suitable types of processing unit.
[0074] A computer system, such as system environment 110, may include one or more computing devices. If the one or more processors of the computersystem are implemented as a plurality of processors, the plurality of processors may be included in a single computing device or distributed among a plurality of computing devices. If a system environment comprises a plurality of computing devices, the memory of the computer system may include the respective memory of each computing device of the plurality of computing devices.
[0075] FIG. 5 is a simplified functional block diagram of a computer system 500 that may be configured as a computing device for executing the processes described herein, according to exemplary embodiments of the present disclosure. FIG. 5 is a simplified functional block diagram of a computer that may be configured according to exemplary embodiments of the present disclosure. In various embodiments, any of the systems herein may be an assembly of hardware including, for example, a data communication interface 520 for packet data communication. The platform also may include a central processing unit (“CPU”) 502, in the form of one or more processors, for executing program instructions. The platform may include an internal communication bus 508, and a storage unit 506 (such as ROM, HDD, SDD, etc.) that may store data on a computer readable medium 522, although the system 500 may receive programming and data via network communications via electronic network 525 (e.g., voice, video, audio, images, or any other data over the electronic network 525). The system 500 may also have a memory 504 (such as RAM) storing instructions 524 for executing techniques presented herein, although the instructions 524 may be stored temporarily or permanently within other modules of system 500 (e.g., processor 502 and / or computer readable medium 522). The system 500 also may include input and output ports 512 and / or a display 510 to connect with input and output devices such as keyboards, mice, touchscreens, monitors, displays, etc. The various system functions may be implemented in adistributed fashion on a number of similar platforms, to distribute the processing load. Alternatively, the systems may be implemented by appropriate programming of one computer hardware platform.
[0076] In this disclosure, the term “based on” means “based at least in part on.” The singular forms “a,” “an,” and “the” include plural referents unless the context dictates otherwise. The term “exemplary” is used in the sense of “example” rather than “ideal.” The terms “comprises,” “comprising,” “includes,” “including,” or other variations thereof, are intended to cover a non-exclusive inclusion such that a process, method, or product that comprises a list of elements does not necessarily include only those elements, but may include other elements not expressly listed or inherent to such a process, method, article, or apparatus. Relative terms, such as “about,” “approximately,” “substantially,” and “generally,” are used to indicate a possible variation of ±10% of a stated or understood value. In addition, the term “between” used in describing ranges of values is intended to include the minimum and maximum values described herein. The use of the term “or” in the claims and specification is used to mean “and / or” unless explicitly indicated to refer to alternatives only if the alternatives are mutually exclusive, although the disclosure supports a definition that refers to only alternatives and “and / or.” As used herein “another” may mean at least a second or more.
[0077] As used herein, the term “user” generally encompasses any person or entity, such as a researcher and / or a care provider (e.g., a doctor, etc.), that may desire information, resolution of an issue, or engage in any other type of interaction with a provider of the systems and methods described herein (e.g., via an application interface resident on their electronic device, etc.). The term “electronic application” or “application” may be used interchangeably with other terms like “program,” or thelike, and generally encompasses software that is configured to interact with, modify, override, supplement, or operate in conjunction with other software.
[0078] Program aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of executable code and / or associated data that is carried on or embodied in a type of machine-readable medium. “Storage” type media include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer of the mobile communication network into the computer platform of a server and / or from a server to the mobile device. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links, or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.
[0079] Furthermore, while some embodiments described herein include some but not other features included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the invention, and formdifferent embodiments, as would be understood by those skilled in the art. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0080] Thus, while certain embodiments have been described, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the invention, and it is intended to claim all such changes and modifications as falling within the scope of the invention. For example, functionality may be added or deleted from the block diagrams and operations may be interchanged among functional blocks. Steps may be added or deleted to methods described within the scope of the present invention.
[0081] The above disclosed subject matter is to be considered illustrative, and not restrictive, and the appended claims are intended to cover all such modifications, enhancements, and other implementations, which fall within the true spirit and scope of the present disclosure. Thus, to the maximum extent allowed by law, the scope of the present disclosure is to be determined by the broadest permissible interpretation of the following claims and their equivalents, and shall not be restricted or limited by the foregoing detailed description. While various implementations of the disclosure have been described, it will be apparent to those of ordinary skill in the art that many more implementations are possible within the scope of the disclosure. Accordingly, the disclosure is not to be restricted except in light of the attached claims and their equivalents.
Claims
WHAT IS CLAIMED IS:
1. A computer-implemented method, the computer-implemented method comprising: receiving, at a computing device, a first set of nucleic acid methylation data and a second set of nucleic acid methylation data, wherein the first set of nucleic acid methylation data is associated with a pre-treatment sample and wherein the second set of nucleic acid methylation data is associated with a post- treatment sample; comparing, using a processor of the computing device, a first feature set of the first set of nucleic acid methylation data against a second feature set of the second set of nucleic acid methylation data; determining, based on the comparing, at least one treatment affected feature in the second feature set of the second set of nucleic acid methylation data; and implementing, based on the determining, an exclusion process on the at least one treatment affected feature in the second feature set of the second set of nucleic acid methylation data.
2. The computer-implemented method of claim 1 , wherein the pre-treatment sample is derived from a subject diagnosed with a disease state, and wherein the post-treatment sample is derived from the subject after the subject has undergone a treatment for the disease state.
3. The method of claim 2, wherein the disease state is cancer.
4. The method of claim 3, wherein the treatment for the disease state is chemoradiotherapy (CRT).
5. The computer-implemented method of claim 1 , wherein the pre-treatment sample and the post-treatment sample are of a sample type and wherein the sample type is one of: a tissue sample, a urine sample, and a blood sample.
6. The computer-implemented method of claim 1 , wherein the pre-treatment sample and the post-treatment sample comprise cell-free DNA (cfDNA) or cell-free RNA (cfRNA).
7. The method of claim 1 , wherein the determining comprises: identifying that the at least one treatment affected feature comprises a methylation pattern; and determining that the identified methylation pattern is not present, or is present to a significantly lesser extent, in the first set of nucleic acid methylation data.
8. The method of claim 1 , wherein the at least one treatment affected feature corresponds to a methylation pattern present in the first set of nucleic acid methylation data that is not present in the second set of nucleic acid methylation data.
9. The method of claim 1 , wherein the implementing the exclusion process comprises:identifying a genomic location associated with each of the at least one treatment affected features; and excluding the genomic location from consideration in a feature selection process associated with machine-learning classifier training.
10. The method of claim 9, further comprising training, subsequent to the excluding, the machine-learning classifier based at least on the second feature set.11 . The method of claim 9, wherein the excluding comprises removing the at least one treatment affected feature from the second feature set of the second set of nucleic acid methylation data.
12. The method of claim 9, wherein the excluding comprises computationally masking the at least one treatment affected feature from a system associated with the computing device.
13. The method of claim 9, further comprising: identifying, after the excluding, one or more remaining treatment affected features in a feature selection set associated with the post-treatment sample; and weighting the one or more remaining treatment affected features in the feature selection set lower than other features in the feature selection set.
14. The method of claim 13, further comprising:applying, as training data input, the feature selection set containing the weighted one or more remaining treatment affected features to a machine-learning classifier.
15. The method of claim 1 , further comprising: selecting, via a feature selection process and subsequent to the exclusion process implemented on the at least one treatment affected feature, one or more training features from the second feature set; and utilizing the selected one or more training features to train a machine-learning classifier to at least identify a presence of a disease state.
16. The method of claim 15, further comprising: receiving a test sample of nucleic acid methylation data associated with a test subject; applying the test sample to the trained machine-learning classifier; and receiving, subsequent to the applying, an indication from the trained machinelearning classifier whether the test sample is associated with the disease state.
17. A system, the system comprising: one or more processors; one or more computer readable media storing instructions that are executable by the one or more processors to perform operations to: receive, at a computing device associated with the system, a first set of nucleic acid methylation data and a second set of nucleic acid methylation data, wherein the first set of nucleic acid methylation data is associated with apre-treatment sample and wherein the second set of nucleic acid methylation data is associated with a post-treatment sample; compare a first feature set of the first set of nucleic acid methylation data against a second feature set of the second set of nucleic acid methylation data; determine, based on the comparing, at least one treatment affected feature in the second feature set of the second set of nucleic acid methylation data; and implement, based on the determining, an exclusion process on the at least one treatment affected feature in the second feature set of the second set of nucleic acid methylation data.
18. The system of claim 17, wherein the pre-treatment sample is derived from a subject diagnosed with a disease state, and wherein the post-treatment sample is derived from the subject after the subject has undergone a treatment for the disease state.
19. The system of claim 18, wherein the treatment for the disease state is chemoradiotherapy (CRT).
20. The system of claim 17, wherein the pre-treatment sample and the posttreatment sample are of a sample type and wherein the sample type is one of: a tissue sample, a urine sample, and a blood sample.21 . The system of claim 17, wherein the pre-treatment sample and the posttreatment sample comprise cell-free DNA (cfDNA) or cell-free RNA (cfRNA).
22. The system of claim 17, wherein the operations to determine comprise operations to: identify that the at least one treatment affected feature comprises a methylation pattern; and determine that the identified methylation pattern is not present in the first set of nucleic acid methylation data.
23. The system of claim 17, wherein the at least one treatment affected feature corresponds to a methylation pattern present in the first set of nucleic acid methylation data that is not present in the second set of nucleic acid methylation data.
24. The system of claim 23, wherein the operations are further configured to: train, subsequent to the excluding, the machine-learning classifier based at least on the second feature set.
25. The system of claim 17, wherein the operations to implement the exclusion process comprise operations to: identify a genomic location associated with each of the at least one treatment affected features; and exclude the genomic location from consideration in a feature selection process associated with machine-learning classifier training.
26. The system of claim 25, wherein the operations to exclude comprise operations to: remove the at least one treatment affected feature from the second feature set of the second set of nucleic acid methylation data.
27. The system of claim 25, wherein the operations to exclude comprise operations to: computationally mask the at least one treatment affected feature from a system associated with the computing device.
28. The system of claim 25, wherein the operations are further configured to: identify, after the excluding, one or more remaining treatment affected features in a feature selection set associated with the post-treatment sample; and weight the one or more remaining treatment affected features in the feature selection set lower than other features in the feature selection set.
29. The system of claim 25, wherein the operations are further configured to: apply, as training data input, the feature selection set containing the weighted one or more remaining treatment affected features to a machine-learning classifier.
30. The system of claim 17, wherein the operations are further configured to: select, via a feature selection process and subsequent to the exclusion process implemented on the at least one treatment affected feature, one or more training features from the second feature set; andutilize the selected one or more training features to train a machine-learning classifier to at least identify a presence of a disease state.31 . The system of claim 30, wherein the operations are further configured to: receive a test sample of nucleic acid methylation data associated with a test subject; apply the test sample to the trained machine-learning classifier; and receive, subsequent to the applying, an indication from the trained machinelearning classifier whether the test sample is associated with the disease state.
32. A non-transitory computer-readable medium storing computer-executable instructions which, when executed by a system, cause the system to perform operations comprising: receiving, at a computing device, a first set of nucleic acid methylation data and a second set of nucleic acid methylation data, wherein the first set of nucleic acid methylation data is associated with a pre-treatment sample and wherein the second set of nucleic acid methylation data is associated with a post- treatment sample; comparing, using a processor of the computing device, a first feature set of the first set of nucleic acid methylation data against a second feature set of the second set of nucleic acid methylation data; determining, based on the comparing, at least one treatment affected feature in the second feature set of the second set of nucleic acid methylation data; andimplementing, based on the determining, an exclusion process on the at least one treatment affected feature in the second feature set of the second set of nucleic acid methylation data.
33. A computer-implemented method, the computer-implemented method comprising: receiving, at a computing device, a first set of nucleic acid methylation data and a second set of nucleic acid methylation data, wherein the first set of nucleic acid methylation data is associated with a first sample associated with a first treatment condition and wherein the second set of nucleic acid methylation data is associated with a second sample associated with a second treatment condition; comparing, using a processor of the computing device, a first feature set of the first set of nucleic acid methylation data against a second feature set of the second set of nucleic acid methylation data; determining, based on the comparing, at least one treatment affected feature in the second feature set of the second set of nucleic acid methylation data; and implementing, based on the determining, an exclusion process on the at least one treatment affected feature in the second feature set of the second set of nucleic acid methylation data.
34. The computer-implemented method of claim 33, wherein the first treatment condition corresponds to a pre-treatment condition and wherein the second treatment condition corresponds to a post-treatment condition.
35. The computer-implemented method of claim 33, wherein the first treatment condition corresponds to a first post-treatment condition and wherein the second treatment condition corresponds to a second post-treatment condition; wherein the first post-treatment condition is associated with administration of a first treatment; and wherein the second post-treatment condition is associated with administration of the first treatment and at least one other treatment.
36. The computer-implemented method of claim 33, wherein the first treatment condition corresponds to a first post-treatment condition and wherein the second treatment condition corresponds to a second post-treatment condition; wherein the first post-treatment condition is associated with administration of a treatment at a first time; and wherein the second post-treatment condition is associated with administration of the treatment at a second time, later than the first time.
Citation Information
Patent Citations
Tumor fraction estimation using methylation variants
US20230272486A1
AU2021245992A1