Systems and methods to identify clonal hematopoiesis related methylation signatures
A computer-implemented method using nucleic acid methylation analysis and machine learning models addresses the confounding effects of CHIP, enhancing the detection and quantification of clonal hematopoiesis and improving disease classifier accuracy.
Patent Information
- Application Number
- PCT/US2025/011969
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-27
- Filing Date
- 2025-01-17
- Publication Date
- 2025-07-24
AI Technical Summary
Clonal hematopoiesis of indeterminate potential (CHIP) introduces noise and complicates the accurate interpretation of true cancer signals in diagnostic technologies due to its confounding effects on methylation patterns, leading to increased false positives and challenges in generalizing diagnostic tools across diverse populations.
A computer-implemented method that analyzes nucleic acid methylation signatures using linear regression and beta-binomial models to identify significant loci associated with CHIP, incorporating factors like age, gender, and cell type proportion, and employs machine learning models to improve detection accuracy.
Enhances the detection and quantification of CHIP, reducing false positives and improving the performance of disease classifiers by providing a more accurate and direct measurement of clonal hematopoiesis through precise identification of abnormal methylation signals.
Smart Images

Figure US2025011969_24072025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS TO IDENTIFY CLONAL HEMATOPOIESIS RELATED METHYLATION SIGNATURESCROSS REFERENCE TO RELATED APPLICATIONS
[0001] The application claims priority to U.S. Provisional Application No. 63 / 622,343, filed on January 18, 2024, and U.S. Provisional Application No. 63 / 665,005, filed on June 27, 2024, which are both incorporated by reference herein in their entireties.TECHNICAL FIELD
[0002] The present disclosure relates generally to the field of bioinformatics and genomics and, more specifically, to systems and methods for detecting and quantifying clonal hematopoiesis based on the analysis of methylation signatures.BACKGROUND
[0003] Clonal Hematopoiesis of Indeterminate Potential (CHIP) represents a biological phenomenon that is characterized by somatic mutational processes during the clonal expansion of hematopoietic stem cells. Its occurrence is tied to both the normal aging process and the initiation of premalignant events, and its emergence in a subject has been recognized as a precursor to blood cancer and / or a potential risk factor for various other diseases. CHIP has also been found to be a confounding factor in diagnostic technologies. More particularly, the presence of CHIP may introduce noise and complicate the accurate interpretation of true cancer signal within assays and algorithms.
[0004] The background description provided herein is for the purpose of generally presenting the context of the disclosure. Unless otherwise indicated herein,the materials described in this section are not prior art to the claims in this application and are not admitted to be prior art, or suggestions of the prior art, by inclusion in this section.SUMMARY OF THE DISCLOSURE
[0005] According to certain aspects of the disclosure, systems and methods are described for identifying clonal hematopoiesis-related methylation signatures.
[0006] In one aspect, a computer-implemented method is provided. The computer-implemented method may contain steps including: receiving, at a computing device, a set of nucleic acid methylation data; receiving, at the computing device, a designation of one or more genomic regions; identifying, using a process of the computing device and within the set of nucleic acid methylation data, one or more abnormal methylation features; and mapping, using the processor, the one or more abnormal methylation features to the one or more genomic regions.
[0007] In another aspect, a computer-implemented method is provided. The computer-implemented method may contain steps including: receiving, at a computing device, nucleic acid methylation data associated with a sample; defining, using a processor of the computing device, a binary call for the sample; employing, using the processor, a linear regression model to identify a relationship between methylation levels for each CpG site in the nucleic acid methylation data and a designation associated with clonal hematopoiesis of indeterminate potential (CHIP); employing, using the processor, a beta-binomial model to assess the relationship; identifying, using the processor, significant loci associated with those CpG sites in the nucleic acid methylation data identified as having a CHIP designation; and associating, using the processor, the significant loci with one or more genomic features.
[0008] In yet another aspect, a system is provided. The system may include: one or more processors; one or more computer readable media storing instructions that are executable by the one or more processors to perform operations to: receive nucleic acid methylation data associated with a sample; define a binary call for the sample; employ a linear regression model to identify a relationship between methylation levels for each CpG site in the nucleic acid methylation data and a designation associated with clonal hematopoiesis of indeterminate potential (CHIP); employ a beta-binomial model to assess the relationship; identify significant loci associated with those CpG sites in the nucleic acid methylation data identified as having a CHIP designation; and associate the significant loci with one or more genomic features
[0009] Additional objects and advantages of the disclosed embodiments will be set forth in part in the description that follows, and in part will be apparent from the description, or may be learned by practice of the disclosed embodiments. The objects and advantages of the disclosed embodiments will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims.
[0010] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosed embodiments, as claimed.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments and together with the description, serve to explain the principles of the disclosure.
[0012] FIG. 1A depicts an exemplary computer system for executing the methods described herein.
[0013] FIG. 1 B depicts an exemplary software platform for executing the methods described herein.
[0014] FIG. 2 depicts an exemplary workflow for a method for identifying abnormal methylation features associated with clonal hematopoiesis, according to one or more embodiments of the present disclosure.
[0015] FIG. 3 depicts a graph that represents a validation of the performance of the CHIP analysis quantification processes outlined in the workflow depicted in FIG. 2, according to one or more embodiments of the present disclosure.
[0016] FIG. 4 depicts a graph that indicates a relationship between the numbers of detected abnormal methylation features in white blood cells (WBC) as individuals age, according to one or more embodiments of the present disclosure.
[0017] FIG. 5 depicts a graph indicating that the majority of WBC abnormal methylation features are sample-specific, according to one or more embodiments of the present disclosure.
[0018] FIG. 6 depicts an exemplary diagram that illustrates the overlap between CH IP-related methylation features with other cancer features that were previously identified in other analysis / studies, according to one or more embodiments of the present disclosure.
[0019] FIG. 7 depicts a graph cataloging the distribution of the number of sample-specific abnormal methylation features per sample, according to one or more embodiments of the present disclosure.
[0020] FIG. 8 depicts an exemplary workflow for performing an epigenomewide association study (EWAS) in relation to CHIP, according to one or more embodiments of the present disclosure.
[0021] FIG. 9 depicts a graph that provides information associated with cell type proportion, according to one or more embodiments of the present disclosure.
[0022] FIGS. 10A and 10B depict graphs that provide additional details related to the prevalence of CHIP, according to one or more embodiments of the present disclosure.
[0023] FIGS. 11 A and 11 B depict graphs that illustrate the distribution of the CHIP mutation, according to one or more embodiments of the present disclosure.
[0024] FIG. 12 depicts an exemplary graph that illustrates the distribution of p-values and q-values across the whole genome for CHIP overall, according to one or more embodiments of the present disclosure.
[0025] FIG. 13 depicts an exemplary plot associated with the analysis of CHIP overall, according to one or more embodiments of the present disclosure.
[0026] FIG. 14 depicts an exemplary plot associated with the analysis of CHIP overall, according to one or more embodiments of the present disclosure.
[0027] FIG. 15 depicts an exemplary graph that illustrates the WGBS feature distribution of significantly differentially methylated loci that are plotted against the background in CHIP overall, according to one or more embodiments of the present disclosure.
[0028] FIG. 16 depicts a graph that illustrates the WGBS feature distribution of significantly differentially methylated loci that are plotted against the background inCHIP overall, according to one or more embodiments of the present disclosure.
[0029] FIG. 17 depicts a graph that illustrates the results of a Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis conducted for CHIP overall, according to one or more embodiments of the present disclosure.
[0030] FIGS. 18A and 18B depict exemplary boxplots associated with CHIP overall, according to one or more embodiments of the present disclosure.
[0031] FIG. 19 depicts a table providing validation results associated with the analysis of CHIP overall, according to one or more embodiments of the present disclosure.
[0032] FIG. 20 depicts a graph providing validation results associated with the analysis of CHIP overall, according to one or more embodiments of the present disclosure.
[0033] FIG. 21 depicts a table that provides data derived from an overlap comparison between significant loci and a first exemplary set of targeted panel regions in CHIP overall, according to one or more embodiments of the present disclosure.
[0034] FIG. 22 depicts a table that provides data derived from an overlap comparison between significant loci and a second exemplary set of targeted panel regions in CHIP overall, according to one or more embodiments of the present disclosure.
[0035] FIG. 23 depicts a graph that illustrates the distribution of p-values and q-values across the whole genome for DNMT3A mutated CHIP, according to one or more embodiments of the present disclosure.
[0036] FIG. 24 depicts an exemplary plot associated with the analysis of DNMT3A mutated CHIP, according to one or more embodiments of the present disclosure.
[0037] FIG. 25 depicts an exemplary plot associated with the analysis of DNMT3A mutated CHIP, according to one or more embodiments of the present disclosure.
[0038] FIG. 26 depicts a graph that illustrates the WGBS feature distribution of significantly differentially methylated loci that are plotted against the background in DNMT3A mutated CHIP, according to one or more embodiments of the present disclosure.
[0039] FIG. 27 depicts a graph that illustrates the WGBS feature distribution of significantly differentially methylated loci that are plotted against the background in DNMT3A mutated CHIP, according to one or more embodiments of the present disclosure.
[0040] FIG. 28 depicts a graph that illustrates the results of a KEGG enrichment analysis conducted for DNMT3A mutated CHIP, according to one or more embodiments of the present disclosure.
[0041] FIGS. 29A and 29B depict exemplary boxplots associated with DNMT3A mutated CHIP, according to one or more embodiments of the present disclosure.
[0042] FIG. 30 depicts a table providing validation results associated with the analysis of DNMT3A mutated CHIP, according to one or more embodiments of the present disclosure.
[0043] FIG. 31 depicts a graph providing validation results associated with the analysis of DNMT3A mutated CHIP, according to one or more embodiments of the present disclosure.
[0044] FIG. 32 depicts a graph that illustrates the distribution of p-values and q-values across the whole genome for TET2 mutated CHIP, according to one or more embodiments of the present disclosure.
[0045] FIGS. 33 depicts an exemplary plot associated with the analysis of TET2 mutated CHIP, according to one or more embodiments of the present disclosure.
[0046] FIG. 34 depicts an exemplary plot associated with the analysis of TET2 mutated CHIP, according to one or more embodiments of the present disclosure.
[0047] FIG. 35 depicts an exemplary graph that illustrates the WGBS feature distribution of significantly differentially methylated loci that are plotted against the background in TET2 mutated CHIP, according to one or more embodiments of the present disclosure.
[0048] FIG. 36 depicts a graph that illustrates the WGBS feature distribution of significantly differentially methylated loci that are plotted against the background in TET2 mutated CHIP, according to one or more embodiments of the present disclosure.
[0049] FIG. 37 depicts a table that illustrates the results of a KEGG enrichment analysis conducted for TET2 mutated CHIP, according to one or more embodiments of the present disclosure.
[0050] FIGS. 38A and 38B depict exemplary boxplots associated with TET2 mutated CHIP, according to one or more embodiments of the present disclosure.
[0051] FIG. 39 depicts a table providing validation results associated with the analysis of TET2 mutated CHIP, according to one or more embodiments of the present disclosure.
[0052] FIG. 40 depicts a graph providing validation results associated with the analysis of TET2 mutated CHIP, according to one or more embodiments of the present disclosure.
[0053] FIG. 41 depicts an example computing system, according to one or more embodiments of the present disclosure.DETAILED DESCRIPTION OF EMBODIMENTS
[0054] The terminology used below may be interpreted in its broadest reasonable manner, even though it is being used in conjunction with a detailed description of certain specific examples of the present disclosure. Indeed, certain terms may even be emphasized below; however, any terminology intended to be interpreted in any restricted manner will be overtly and specifically defined as such in this Detailed Description section. Both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the features, as claimed.
[0055] Clonal hematopoiesis is a broad term that refers to the phenomenon when a population of blood cells, derived from a single progenitor cell, becomes dominant. This clonal expansion may occur due to various factors, including genetic mutations, aging, and / or exposure to certain environmental influences. One type of clonal hematopoiesis is Clonal Hematopoiesis of Indeterminate Potential (CHIP), which corresponds to the presence of genetic mutations in a subset of blood cells that leads to their clonal expansion. The mutations associated with CHIP may have implications for health, including an increased risk of developing blood cancers or other diseases.
[0056] CHIP is considered a confounding issue in various biomedical and / or diagnostic contexts due to its potential to introduce noise and complicate the interpretation of results. More particularly, in diagnostic tests that rely on detecting genetic mutations or aberrations, the presence of CHIP-related mutations may lead to an increased number of false positives. This is because the mutations associated with CHIP may be mistakenly interpreted as indicators of a disease or pathology. For instance, CHIP-related clonal expansions may exhibit specific nucleic acid methylation patterns. In studies focusing on methylation signatures, the presence of CHIP may introduce confounding signals that may impact the accuracy of analyses and classifications.
[0057] Additionally to the foregoing, CHIP involves the presence of diverse genetic mutations in blood cells, which typically affect genes associated with hematopoiesis. This genetic heterogeneity may lead to variations in clonal expansion patterns among individuals. Furthermore, CHIP is not uniform across populations - its prevalence and genetic characteristics may vary. More particularly, the population-specific effects of CHIP may result in challenges when generalizing findings or developing diagnostic tools applicable to diverse demographic groups.
[0058] Accordingly, the present disclosure may address one or more of the foregoing challenges associated with detecting and quantifying CHIP by introducing novel methods based on the analysis of nucleic acid methylation signatures. By focusing on methylation signatures, the methods described herein may provide a more accurate and direct measurement of clonal hematopoiesis, thereby improving the performance of existing disease classifiers, such as cancer classifiers.
[0059] One method described herein corresponds to the direct measurement of clonal hematopoiesis through the identification of abnormal methylation features.Utilizing specific target genomic regions and a bio-feature-extractor tool, the concepts described herein are configured to identify abnormal methylation signals above the non-disease, e.g., non-cancer, baseline noise background. This method ensures a one-to-one mapping of features in the target genomic regions, allowing for precise measurement of abnormal methylation signals across genomic regions. Additionally, this method may improve accuracy and sensitivity compared to traditional somatic mutation-based approaches. Another method described herein leverages a generalized linear regression model that correlates methylation variations across the entire genome with CHIP status. By incorporating factors such as age, gender, and cell type proportion, the analysis aims to identify significant loci and understand the epigenetic variations linked to CHIP. This is a population-level approach that may provide insights into the broader epigenetic landscape associated with clonal hematopoiesis. Other methods, not explicitly listed and described here, may also be utilized.
[0060] The concepts described herein integrate various technological improvements and correspondingly improve the functionality of a computing device as used in the detection of CHIP in several ways. For instance, both of the foregoing methods utilize computational tools and trained machine learning models to improve the ability of a computing device that supports these tools to more efficiently detect and analyze CHIP. Additionally, the concepts described herein apply computational techniques to specific biological problems related to clonal hematopoiesis. This practical application of computer technology may enhance the understanding of biological phenomena and improve diagnostic and / or research methodologies. For instance, in an aspect, the described concepts may make the computer-based CHIP detection system more robust to its confounding effects with respect to methylationpatterns that may be associated with a disease, e.g., cancer. The computer executing and implementing a classifier model trained and used according to the techniques described herein may be better equipped to provide reliable mutation- related information. Furthermore, the processes executed by the computer to improve the model involve complex calculations and data manipulations on a large amount of biological data that a human individual could not reasonably complete on their own or in their mind. Specifically, computationally intensive statistical tests are leveraged by the computer to evaluate methylation patterns between sample sets, processes which cannot be completed by a human.
[0061] The subject matter of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, which form a part hereof, and which show, by way of illustration, specific exemplary embodiments. An embodiment or implementation described herein as “exemplary” is not to be construed as preferred or advantageous, for example, over other embodiments or implementations; rather, it is intended to reflect or indicate that the embodiment(s) is / are “example” embodiment(s). Subject matter may be embodied in a variety of different forms and, therefore, covered or claimed subject matter is intended to be construed as not being limited to any exemplary embodiments set forth herein; exemplary embodiments are provided merely to be illustrative. Likewise, a reasonably broad scope for claimed or covered subject matter is intended. Among other things, for example, subject matter may be embodied as methods, devices, components, or systems. Accordingly, embodiments may, for example, take the form of hardware, software, firmware, or any combination thereof. The following detailed description is, therefore, not intended to be taken in a limiting sense.
[0062] Throughout the specification and claims, terms may have nuanced meanings suggested or implied in context beyond an explicitly stated meaning. Likewise, the phrase “in one embodiment” or “in some embodiments,” or “in one aspect” or “in some aspects” as used herein does not necessarily refer to the same embodiment or aspect, and the phrase “in another embodiment” or “in another aspect” as used herein does not necessarily refer to a different embodiment or aspect. It is intended, for example, that claimed subject matter include combinations of exemplary embodiments in whole or in part.
[0063] Diseases referred to herein may include cancer. Non-limiting cancer types that the concepts described herein may be applied to include, for example, breast cancer, lung cancer (e.g., non-small cell lung cancer (NSCLC)), prostate cancer, colorectal cancer, renal cancer, uterine cancer, pancreatic cancer, cancer of the esophagus, a lymphoma, head and neck cancer, ovarian cancer, a hepatobiliary cancer, a melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, and gastric cancer. Additionally, it is also important to note that although the concepts described throughout this disclosure are made in reference to cancer, these designations are for exemplary purposes only and are not intended to be limiting. Specifically, the concepts described herein may be applicable to other disease types and other disease-detecting machine-learning classifiers.
[0064] FIG. 1 A depicts an exemplary system for detecting and quantifying CHIP via the analysis of CH IP-related methylation signatures. Exemplary system 100 includes a data collection component 10, a database 20, and device data intelligence component 30, operably connected to each other via network 40.Alternatively, or additionally, one or more of the components may be connected with another component locally without reliance on network connection; e.g., through awired connection. In many aspects described herein, sequencing data of cell-free nucleic acids are used to illustrate the concepts. However, one of skill in the art would understand that the current method may be applied to sequencing data of DNA, RNA, or other materials, as well from a variety of sample types, e.g., a blood sample (e.g., a serum sample, a plasma sample, a whole blood sample), a urine sample, a saliva sample, a tissue sample, a bone marrow sample, etc.
[0065] As disclosed herein, data collection component 10 may include a device or machine with which sequencing data may be generated. In some embodiments, data collection component 10 may include one or more sequencing devices or a facility that uses one or more sequencing devices to generate nucleic acid (e.g., DNA or RNA) sequence data of biological samples. In some aspects, data collection component 10 may be a database that receives sequencing information generated from one or more sequencing devices. Any suitable liquid or solid biological samples may be used for sequencing. In some embodiments, a biological sample may be cell-based, for example, one or more types of tissue. In some embodiments, a biological sample may be a sample that includes cell-free nucleic acid fragments. Examples of biological samples include, but are not limited to, a blood sample (e.g., a cell-free DNA (cfDNA) sample, a cell-free RNA (cfRNA) sample, a serum sample, a plasma sample, a whole blood sample), a urine sample, a saliva sample, a tissue sample, a bone marrow sample, etc. Further, although sequencing of DNA from these samples is discussed herein, RNA from these samples may alternatively or additionally be sequenced.
[0066] Examples of sequencing data may include, but are not limited to, sequence read data of targeted genomic locations, partial or whole genome sequencing data of the genome represented by nucleic acid fragments in cell-free orcell-based samples, partial or whole genome sequencing data including one or more types of epigenetic modifications (e.g., methylation), or combinations thereof.
[0067] Data acquired by the data collection component 10 may be transferred to database 20 via network 40 or a local or network connection. In some embodiments, data collection component 10 may alternatively receive data from one or more sequencing devices. In some embodiments, the collected data may be analyzed by data intelligence component 30, via network 40 or a local or network connection. FIG. 1 B depicts exemplary functional modules that may be implemented to perform tasks of data intelligence component 30.
[0068] FIG. 1 B depicts an exemplary computer system 110 for measuring effects associated with CHIP through the analysis of CH IP-related methylation patterns. Exemplary system 110 achieves such functionalities by implementing, on one or more computer devices, user input and output (I / O) module 120, memory or database 130, data processing module 140, data analysis module 150, classification module 160, network communication module 170, and any other functional modules that may be needed for carrying out a particular task (e.g., an error correction or compensation module, a data compression module, etc.). As disclosed herein, user I / O module 120 may further include an input sub-module, such as a keyboard, and an output sub-module, such as a display (e.g., a printer, a monitor, or a touchpad). In some embodiments, all functionalities may be performed by one computer system. In some embodiments, the functionalities are performed by more than one computer system.
[0069] Also disclosed herein, a particular task may be performed by implementing one or more functional modules. In particular, each of the enumerated modules itself may, in turn, include multiple sub-modules. For example, dataprocessing module 140 may include a sub-module for data quality evaluation (e.g., for discarding very short sequence reads or sequence reads including obvious errors), a sub-module for normalizing numbers of sequence reads that align to different regions of a reference genome, a sub-module to compensate / correct guanine-cytosine (GC) biases, a sub-module for matching data associated with a cancer sample with other data associated with one or more non-cancer samples, etc.
[0070] In some embodiments, a user may use I / O module 120 to manipulate data that is available either on a local device or can be obtained via a network connection from a remote service device or another user device. For example, I / O module 120 may allow a user, e.g., via a keyboard, a mouse, or a touchpad, to initiate or perform data analysis via a graphical user interface (GUI). In some embodiments, a user may manipulate data via voice control. In some embodiments, user authentication may be required before a user is granted access to the data being requested. In some embodiments, user I / O module 120 may be used to manage various functional modules. For example, a user may request via user I / O module 120 input data while an existing data processing session is in process. A user may do so by selecting a menu option or type in a command discretely without interrupting the existing process. In another example, a user may utilize user I / O module 120 to set various thresholds, configure sample matching settings, and / or provide other instructions to computer system 110 that dictate how treatment- affected regions are identified and / or masked. As disclosed herein, a user may use any type of input to direct and control data processing and analysis via I / O module 120.
[0071] In some embodiments, system 110 further comprises a memory or database 130. In some embodiments, database 130 comprises a local database thatmay be accessed via user I / O module 120. In some embodiments, database 130 comprises a remote database that may be accessed by user I / O module 120 via network connection. In some embodiments, database 130 is a local database that stores data retrieved from another device (e.g., a user device or a server). In some embodiments, memory or database 130 may store data retrieved in real-time from internet searches. In some embodiments, database 130 may send data to and receive data from one or more of the other functional modules, including, but not limited to, a data collection module (not shown), data processing module 140, data analysis module 150, classification module 160, network communication module 170, and etc.
[0072] In some embodiments, database 130 may be a database local to the other functional modules. In some embodiments, database 130 may be a remote database that may be accessed by the other functional modules via wired or wireless network connection (e.g., via network communication module 170). In some embodiments, database 130 may include a local portion and a remote portion.
[0073] In some embodiments, system 110 comprises a data processing module 140. Data processing module 140 may receive data from I / O module 120 or database 130. In some embodiments, data processing module 140 may perform standard data processing algorithms, such as one or more of noise reduction, signal enhancement, normalization of counts of sequence reads, correction of GC bias, etc. In some embodiments, data processing module 140 may be configured to detect and measure methylation signatures, and specifically abnormal methylation features, associated with clonal hematopoiesis.
[0074] In some embodiments, system 110 comprises a data analysis module 150. In some embodiments, data analysis module 150 includes identifying andtreating systematic errors in sequencing data, as described in connection with data processing module 140.
[0075] In some embodiments, system 110 comprises a classification module 160, which may embody a “machine-learning model” or “trained classifier.” As used herein, a “machine-learning model” or “trained classifier” generally encompasses instructions, data, and / or a model configured to receive input, and apply one or more of a weight, bias, classification, or analysis on the input to generate an output. The output may include, for example, a classification of the input, an analysis based on the input, a design, process, prediction, or recommendation associated with the input, or any other suitable type of output. A machine-learning model is generally trained using training data, e.g., experiential data and / or samples of input data, which are fed into the model in order to establish, tune, or modify one or more aspects of the model, e.g., the weights, biases, criteria for forming classifications or clusters, or the like. Aspects of a machine-learning model may operate on an input linearly, in parallel, via a network (e.g., a neural network), or via any suitable configuration. In some aspects, the machine-learning model may be trained on a combination of real and synthetic sample data.
[0076] The execution of the machine-learning model may include deployment of one or more machine-learning techniques, such as k-nearest neighbors, linear regression, logistic regression, random forest, gradient boosted machine (GBM), deep learning, a deep neural network, and / or any other suitable machine-learning technique that solves problems in the field of Natural Language Processing (NLP). Supervised, semi-supervised, and / or unsupervised training may be employed. For example, supervised learning may include providing training data and labels corresponding to the training data, e.g., as ground truth. Unsupervisedapproaches may include clustering, classification or the like. K-means clustering or K-Nearest Neighbors may also be used, which may be supervised or unsupervised. Combinations of K-Nearest Neighbors and an unsupervised cluster technique may also be used. Any suitable type of training may be used, e.g., stochastic, gradient boosted, random seeded, recursive, epoch or batch-based, etc.
[0077] In an exemplary use case, a machine-learning model may be trained to analyze data from a test sample from a test subject whose status with respect to a medical condition is unknown and subsequently classifies the unknown test sample from the test subject based on the likelihood of the subject fitting into a particular category. In some embodiments, the one or more parameters may include a binomial probability score that is calculated based on logistic regression analysis. As disclosed herein, the binomial probability score may correspond to the likelihood of a subject having a certain medical condition, such as cancer. For example, a score of over a predefined threshold may indicate that the subject associated with a test sample is more likely to have cancer than not have cancer. In some embodiments, the one or more parameters may include a sequencing or methylation data distribution pattern correlating with the presence of cancer. A subject associated with a test sample having sequencing or methylation data with a pattern resembling the cancer pattern to a sufficient degree may be predicted as having cancer. In some embodiments, a sequencing or methylation data distribution pattern may be identified in connection with a specific type of cancer, determining a tissue of origin or cancer signal origin, thus allowing a test sample to be classified as indicative of a certain cancer type.
[0078] As disclosed herein, network communication module 170 may be used to facilitate communications between a user device, one or more databases,and any other suitable system or device through a wired or wireless network connection. Any communication protocol / device may be used, including, without limitation, a modem, an Ethernet connection, a network card (wireless or wired), an infrared communication device, a wireless communication device, and / or a chipset (such as a Bluetooth™ device, an 802.11 device, a WiFi device, a WiMax device, cellular communication facilities, etc.), a near-field communication (NFC), a Zigbee communication, a radio frequency (RF) or radio-frequency identification (RFID) communication, a PLC protocol, a 3G / 4G / 5G / LTE based communication, and / or the like. For example, a user device having a user interface platform for processing / analyzing CHIP-related methylation signature data may communicate with another user device with the same platform, a regular user device without the same platform (e.g., a regular smartphone), a remote server, a physical device of a remote loT local network, a wearable device, a user device communicably connected to a remote server, and etc.
[0079] The functional modules described herein are provided by way of example. It will be understood that different functional modules may be combined to create different utilities. It will also be understood that additional functional modules or sub-modules may be created to implement a certain utility.
[0080] Referring now to FIG. 2, an exemplary workflow 200 is provided for performing abnormal methylation feature modeling in relation to CHIP. Aspects of the exemplary workflow 200 may be performed in accordance with some or all components described in FIG. 1A and 1 B. In an aspect, abnormal methylation may signify deviations from the typical methylation patterns, e.g., in DNA or RNA, observed in non-cancerous or healthy cells. These features may include specific patterns or levels of methylation that are indicative of clonal hematopoiesis.
[0081] At step 205, one or more specific genomic regions may be selected to serve as the focus of the analysis. In an aspect, one or more specific genomic regions may serve as the focus of the analysis. The selection of these regions may allow for capturing the specific epigenetic changes linked to clonal expansion of hematopoietic stem cells. These regions may be chosen based on their relevance to clonal hematopoiesis or their known association with abnormal methylation patterns.
[0082] At step 210, abnormal methylation features associated with clonal hematopoiesis may be identified. To facilitate this process, methylation patterns within the selected genomic regions may be analyzed using specialized techniques, e.g., via high-throughput sequencing methods such as WGBS. This may allow for the comprehensive examination of DNA methylation at a per-base resolution. Thereafter, abnormal methylation signals may be identified by comparing the methylation patterns observed in the samples to a baseline or reference. This baseline may be derived from non-cancer samples orthose without clonal hematopoiesis, thereby serving as a representation of normal methylation levels. The quantification of abnormal methylation signals may involve the application of a probabilistic model, such as a one-beta-binomial model per region. These models may help to assess the likelihood of observing abnormal methylation fragments within each genomic region. In an aspect, to distinguish abnormal methylation signals from normal variation, a cutoff may be determined. This cutoff may be based on mutual information between cancer and non-cancer samples and may represent a threshold that separates abnormal signals from the background noise. The results from this step are a quantitative representation of abnormal methylation signals within each of the selected and targeted genomic regions of interest.
[0083] At step 215, once the abnormal methylation features are identified, they may be mapped to the corresponding genomic regions chosen as the focus for the study. This mapping may ensure a one-to-one correspondence between the identified features and the specific genomic locations targeted for analysis. Stated differently, it may establish a clear relationship between abnormal methylation patterns and the genomic regions of interest.
[0084] Referring now to FIG. 3, a graph 300 is provided that represents a validation of the performance of the CHIP analysis quantification processes outlined in workflow 200. The validation may determine whether the method represented by workflow 200 effectively captures and quantifies age-related abnormal methylation signatures associated with clonal hematopoiesis. In an aspect, to facilitate the validation, the methylation features generated in workflow 200 may be utilized to estimate tumor fraction in the targeted genomic regions that exhibit a strong cancer signal in both solid and heme cancers. Corresponding tumor fraction estimates may be obtained using a small variant mutation based method. The consistency between tumor fraction estimates based on methylation features and those based on small variants was then evaluated to show, in FIG. 3, that the results are highly consistent and that they align with the expectations of methylation classifier scores.
[0085] Referring now to FIG. 4, a graph 400 is provided that presents a relationship between the numbers of detected abnormal methylation features in white blood cells (WBC) as individuals age. Each plot point represents an individual sample, and the color is representative of the sample type (e.g., healthy / non-cancer or diseased / cancer). Examination of graph 400 suggests little difference between non-cancer derived WBC and cancer derived WBC as individuals age. However, a statistically strong relationship between an individual’s age and the number ofabnormally methylated features is identified in graph 400 (e.g., the number of abnormally methylated features increase as people age). Accordingly, it can be derived from the data associated with graph 400 that methylation signatures are associated with aging.
[0086] Referring now to FIG. 5, a graph 500 is provided that represents that the majority of WBC abnormal methylation features are sample-specific. In an aspect, recurrent methylation patterns refer to those DNA methylation alterations that are observed consistently across different individuals or samples. These patterns are important because they suggest commonalities or shared features in the methylation landscape across a population. If methylation alterations are recurrent, they can potentially be corrected or adjusted for by utilizing a reference set / profile. Understanding and correcting recurrent methylation patterns may promote the accuracy of the classifier by mitigating the impact of confounding signals. Conversely, sample-specific methylation patterns are alternations that are unique to individual samples and are not consistently observed across different individuals. These patterns pose challenges for correction, because each sample may exhibit distinct methylation changes, making it difficult to generalize correction strategies. Graph 500 indicates that a majority of the CHIP-related methylation features recur less frequently.
[0087] Referring now to FIG. 6, a diagram 600 is provided that illustrates the overlap between CHIP-related methylation features with other cancer features that were previously identified in other analyses / studies. Diagram 600 reveals that approximately 40% of the CHIP-related methylation features are shared with cancer features, and the majority of these are sample-specific. These results further emphasize the importance of performing a sample-specific correction to remove anyconfounding signal derived from the CHIP-related methylation features. Specifically, because a meaningful percentage of CHIP-related methylation features overlap with cancer features, a classifier may generate a false positive result unless those features are corrected for. FIG. 7 presents a graph cataloging the distribution of the number of sample-specific abnormal methylation features per sample. As the graph shows, the median number of sample-specific abnormal methylation features is 29 in this dataset.
[0088] In an aspect, cfDNA-WBC paired sequencing may be utilized to obtain genomic information from both the circulating cfDNA and the corresponding white blood cells. More particularly, when DNA or RNA in the blood is studied, there are signals present from both WBCs and other tissues, like tumors. In clonal hematopoiesis, certain WBCs may start growing and multiplying more than usual, which may lead to genetic changes in these cells. In medical tests that analyze DNA in the blood, it is important to distinguish between normal genetic variations and those changes that are linked to potential health issues, like cancer. WBC paired sequencing helps to improve the accuracy of these tests by providing a clearer picture of genetic changes that are specific to WBCs.
[0089] In view of the foregoing, WBC paired sequencing may be performed to obtain cfDNA and WBC genomic information. The methylation patterns specifically in WBCs may be analyzed to identify abnormal methylation features associated with CHIP. The information from both cfDNA and WBC sequencing may be utilized to develop individualized correction profiles. The methylation signals in cfDNA may be adjusted based on the identified sample-specific methylation alterations observed in WBCs. By distinguishing CH IP-associated methylation features in WBCs, the paired sequencing helps to reduce the noise introduced by these features in the cfDNAanalysis. The goal is to remove or account for the CHIP-induced false positives and reduce the limit of detection (LOD) of the assay. Accordingly, to enhance classifier performance, the corrected methylation signals derived from WBC paired sequencing may be incorporated into the training and validation of classifiers. This ensures that the classifier accounts for both recurrent and sample-specific methylation patterns associated with clonal hematopoiesis.
[0090] In view of all of the foregoing, it was found that there is a moderate accumulation of abnormal methylation features as people age (e.g., 168 sites per decade for individuals aged 50 years or older). The majority of WBC abnormal methylation features are sample-specific: 50% of the features have very low population recurrence (e.g., corresponding to <1 % of the study population). A large portion of these sample-specific WBC abnormal methylation features are also highly recurrent in cancers. Because of this, paired cfDNA-WBC sequencing should be performed to remove WBC-induced false positives and correspondingly reduce the limit-of-detection (LOD) of a relevant assay.
[0091] Turning now to FIG. 8, an exemplary workflow 800 is provided for performing an epigenome-wide association study (EWAS) in relation to CHIP. Aspects of the exemplary workflow 800 may be performed in accordance with some or all components described in FIG. 1 A and 1 B.
[0092] At step 805, DNA methylation data may be gathered from a comprehensive set of CpG sites across the genome or, alternatively, from specifically targeted genomic regions. In an aspect, DNA methylation data may be obtained from cfDNA biological samples. In an aspect, the collection of cfDNA samples may be carried out using minimally invasive methods, such as blood draws or other plasma collection procedures commonly used to obtain cfDNA. That said,any suitable method of sample collection and any suitable sample type may be collected at step 205. Further, step 805 may not include active sample collection and may instead refer to receipt of samples and / or data associated with samples that were previously collected (e.g., from a prior study). In one aspect, metadata associated with each sample may be collected, e.g., subject demographics, medical history, treatment regimens, and any relevant clinical information. This metadata may in some aspects provide context for interpreting methylation patterns and understanding the impact of cancer treatment on these patterns.
[0093] A methylation analysis may be conducted on the collected samples. Methylation analysis may involve the assessment of DNA methylation patterns at specific genomic regions, for example, at cytosine-phosphate-guanine (CpG) sites. In an aspect, various high-throughput technologies may be employed for methylation profiling, including one or more of bisulfite sequencing, methylated DNA immunoprecipitation sequencing (MeDIP-seq), DNA methylation microarrays, and the like. For simplicity purposes, bisulfite sequencing is the methylation profiling technique described herein, however, this designation is not intended to be limiting.
[0094] In an aspect, bisulfite sequencing may involve the treatment of DNA with sodium bisulfite, which converts unmethylated cytosines (C) into uracils (U) while leaving methylated cytosines unchanged. After bisulfite treatment, the DNA may be subjected to high-throughput sequencing, such as next-generation sequencing (NGS), to determine the methylation status of individual CpG sites across the genome. Whole-genome bisulfite sequencing (WGBS) may provide comprehensive coverage of CpG sites and allow for a detailed assessment of methylation patterns. The generated methylation data may undergo bioinformatics analysis. In this regard, the methylation data may first undergo one or morepreprocessing steps to ensure the quality and integrity of the methylation data. These steps may include one or more of: data cleaning, quality control, and the removal of artifacts or outliers that may affect the accuracy of the analysis. Preprocessing may also involve the alignment of sequence reads to a reference genome. The ratio of C to T at each CpG site may be used to calculate the methylation level.
[0095] In an aspect, the methylation level at each CpG site may be represented by a beta value, which are typically reported as decimal values ranging from 0 to 1 . A beta value of 0 indicates that the CpG site is completely unmethylated. A beta value of 1 .0 indicates that the CpG site is completely methylated. A beta value of 0.50 indicates that the CpG site is 50% methylated. Beta values offer a straightforward interpretation of DNA methylation levels. For example, a beta value of 0.2 at a specific CpG site suggests that 20% of the DNA molecules at the site are methylated, while the remaining 80% are unmethylated.
[0096] In an aspect, informative CpG sites may be selected for the EWAS analysis while those with low coverage, low variability, or other characteristics that may compromise the quality of the analysis may be discarded. More particularly, this filtering step may promote the reliability, informative nature, and relevance in a subsequent association analysis conducted on a set of CpG sites. To this extent, in one aspect, CpG sites associated with poor data quality, which may arise from technical artifacts or assay-specific issues, may be excluded. Additionally or alternatively, in another aspect, a coverage threshold may be employed to filter out CpG sites with low sequencing depth (e.g., low coverage may result in less reliable methylation measurements, and excluding such sites helps to ensure more accurateanalyses). Additionally or alternatively, in some aspects, CpGs with less than 10 samples in either group may be removed.
[0097] At step 810, a binary CHIP call may be defined for each sample based on the variant allele frequency (VAF) of somatic mutations associated with CHIP. In an aspect, a threshold for VAF that defines the presence or absence of a CHIP-associated mutation may be established. In an aspect, exemplary thresholds for CHIP calls may be VAF > 0.02 or > 0.01 , indicating that the mutation is present in at least 1 % or 2% of the cell, respectively. For each sample, an assessment may be made whether it has a somatic mutation with a VAF that exceeds the chosen threshold. If the VAF surpasses the threshold, the sample may be labeled as having CHIP (e.g., in which case a binary value of 1 may be assigned to it). Conversely, if the VAF is lower than the threshold, the sample may be labeled as not having CHIP (e.g., in which case a binary value of 0 may be assigned to it). As a non-limiting example of the foregoing, if the chosen VAF threshold is 0.02, a CpG site with a somatic mutation in which the variant allele is present in more than 2% of the cells would be classified as having CHIP and assigned a binary value of 1 . Otherwise, it would be classified as not having CHIP and assigned a binary value of 0.
[0098] At step 815, a generalized linear regression model may be utilized to assess the association between DNA methylation levels and CHIP status. More particularly, in an aspect, after a binary CHIP call is made for each sample, the analysis may determine whether there is a statistically significant association between clonal hematopoiesis and methylation levels at each CpG site. In this regard, a regression analysis may be performed for each CpG site to examine the relationship between the methylation level (considered a dependent variable) and certain predictors (independent variables), such as the binary CHIP status, age,gender, and cell type proportion. For instance, the age of the individuals may influence methylation patterns, and controlling for age may help assess the association between CHIP status and methylation while accounting for age-related effects. Gender differences may also impact methylation profiles, and including gender as a covariate may control for these potential variations. Additionally, the proportion of different cell types in the blood samples may affect methylation patterns. More particularly, in a biological sample, the proportion of different cell types can vary between individuals. DNA methylation patterns are cell-type specific, and variations in cell type proportions may confound the association between DNA methylation levels and a particular trait or condition. Without accounting for cell type proportion, associations observed in EWAS may be driven by differences in cell composition rather than true epigenetic changes related to the trait of interest.
[0099] At step 820, a beta-binomial model may be utilized to assess the relationship between DNA methylation levels at CpG sites and CHIP status.
[0100] At step 825, significant loci associated with CHIP may be identified. In the context of DNA methylation, a “locus” may refer to a specific genomic location where a CpG site is present. Each CpG locus represents a potential regulatory region in the genome. In an aspect, a significance threshold may be set to determine which loci reach statistical significance. For instance, the significance threshold may be a p-value of < 0.05.
[0101] At step 830, functional annotation may be performed to associate the identified significant loci with known genomic features to understand their biological significance. More particularly, after identifying significant loci from statistical analyses, these loci need to be associated with specific genomic regions. This may involve mapping genomic coordinates to reference genomes. In an aspect, theidentified loci may be overlaid with one or more known genomic features, such as promotors, exons, introns, enhancers, transcription factor binding sites, etc. In some aspects, biological pathways that are significantly enriched with genes associated with the significant loci may be identified.
[0102] Referring now to FIG. 9, a graph 900 is presented that provides information associated with cell type proportion. Examination of graph 900 indicates that a small number of cell types account for the majority of the cell type proportions. For instance, the cell types neutrophil and megakaryocyte alone account for over half (55%) of the cell type proportions in the samples.
[0103] Referring now collectively to FIGS. 10A and 10B, two graphs 1000 and 1005, respectively, are presented that provide additional details related to the prevalence of CHIP. FIG. 10A provides data points related to the relevance of CHIP based on gender, and FIG. 10B provides data points related to the relevance of CHIP based on age. As can be observed in FIG. 10B, as age increases, there is a greater prevalence of CHIP in the population. For example, at age 70, half of the population in this study group had CHIP.
[0104] Referring now collectively to FIGS. 11A and 11 B, two graphs 1100 and 1105, respectively, are presented that illustrate the distribution of the CHIP mutation. FIG. 11 A illustrates the proportions of specific CHIP mutations in individuals known to have clonal hematopoiesis. For instance, the majority of samples in the study that graph 1100 in FIG. 11A was based on have the DNMT3A mutation, with TET2 being the second leading mutation present. The distribution of the rest of the mutations are considered to be sample-specific. For example, referring now to FIG. 11 B, approximately 70% of individuals only have 1 mutated gene, 19% of the population has 2 mutated genes, and so on.
[0105] In an aspect, three targeted analyses may be performed in relation to the EWAS associated with workflow 800: one for CHIP overall (e.g., as defined as any mutation that has a VAF more than 2%), the second related to DNMT3A- mutated CHIP, and the third directed to TET2 mutated CHIP. FIGS. 12 - 22 correspond to the analysis of CHIP overall, FIGS. 23 - 30 correspond to the analysis of DNMT3A-mutated CHIP, and FIGS. 31 - 37(A-B) correspond to the analysis of TET2-mutated CHIP.
[0106] Referring now to FIG. 12, graph 1200 is presented that illustrates the distribution of p-values and q-values across the whole genome for CHIP overall. 1054 significant loci were identified having a false discovery rate (FDR) of less than 0.05 (5%) and an absolute effect size greater than 0.1 (10%). 24 loci were identified as being hypermethylated (e.g., having increased DNA methylation in CHIP), and 1030 loci were identified as being hypomethylated (e.g., having decreasing DNA methylation in CHIP).
[0107] Referring now to FIG. 13, graph 1300 illustrates a quantile-quantile (QQ) plot in the analysis of CHIP overall. In this plot, the observed p-values were plotted against expected p-values, and a genomic inflation factor was determined from the resultant data points. A well-controlled study generally has a genomic inflation factor close to 1. The data associated with graph 1300 indicates that the genomic inflation factor was found to be 0.92.
[0108] Referring now to FIG. 14, graph 1400 illustrates a “Manhattan” plot associated with CHIP overall. Manhattan plots are graphical representations commonly utilized in genomics to visualize the results of statistical tests conducted across the genome. Specifically, these plots may be utilized to identify specific regions of the genome where genetic variants are significantly associated with a traitor disease. For instance, graph 1400 indicates that certain genomic regions associated with chromosome 9 may be significant.
[0109] Referring now to FIG. 15, graph 1500 illustrates the WGBS feature distribution of significantly differentially methylated loci that are plotted against the background in CHIP overall. Analysis of graph 1500 reveals that the significant loci are arranged in the promotor and the coding sequencing area (CDS).
[0110] Referring now to FIG. 16, graph 1600 illustrates the WGBS feature distribution of significantly differentially methylated loci that are plotted against the background in CHIP overall. Analysis of graph 1600 reveals that the significant loci are largely arranged in the island.
[0111] Referring now to FIG. 17, graph 1700 illustrates the results of a Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis conducted for CHIP overall. In aspect, KEGG enrichment analysis may be performed to identify significant genetic pathways associated with CH IP-associated methylation patterns. Results from this analysis may be utilized to assess whether the identified pathways have implications for cancer or other relevant biological processes. With respect to graph 1700, significant genetic pathways are illustrated, e.g., hsa04928 parathyroid hormone synthesis, secretion and action, hsa040105 Rap1 signaling pathway, hsa04010 MAPK signaling pathway, hsa04330 Notich signaling pathway, hsa04927 Cortisol synthesis and secretion, and hsa04724 Glutamatergic synapse, that were found to have false discovery rates less than 0.05 (i.e. , 5%). These are common pathways related to cancer.
[0112] Referring now collectively to FIGS. 18A and 18B, boxplots 1800 and 1805 illustrate the methylation trends of two significant loci (e.g., ch 19: 10407287 inFIG. 18A and chr19:55950003 in FIG. 18B) plotted against the age category forCHIP overall.
[0113] Referring now collectively to FIGS. 19 and 20, table 1900 in FIG. 19 and graph 2000 in FIG. 20 present validation results from the analysis of CHIP overall. More particularly, the results associated with CHIP overall were compared against previous findings in the literature that were derived from a separate panel (e.g., a 450k Beadchip panel). This comparison revealed that the results derived by the concepts disclosed herein were consistent with the findings in the literature. More particularly, with respect to table 1900, 39 loci were identified as having the same direction in both analysis. Furthermore, with respect to graph 2000, the significant loci identified in the literature has similarly small p-values.
[0114] Referring now to FIG. 21 , table 2100 illustrates data derived from an overlap comparison between significant loci and a first exemplary set of targeted panel regions in CHIP overall. More particularly, an analysis was performed to identify how many of the first exemplary set of targeted panel regions overlapped with the significant loci. Results indicated that there are 30 regions overlapping with loci, and among those regions, 30 of them are used in classifier training. With respect to table 2100, additional analysis was performed to determine if those loci were related to heme features. Examination of the results in table 2100 indicate that the p-value is not significant, thereby implying that the heme feature is not enriched in the overlapping regions (i.e. , the regions that have the significant loci).
[0115] Referring now to FIG. 22, table 2200 illustrates data derived from an overlap comparison between significant loci and a second exemplary set of targeted panel regions in CHIP overall. More particularly, an analysis was performed to identify how many of the second exemplary set of targeted panel regions overlappedwith the significant loci. Results indicated that there are 20 regions overlapping with loci, and among those regions, 14 of them are used in classifier training. With respect to table 2200, additional analysis was performed to determine if those loci were related to heme features. Examination of the results in table 2100 indicate that the p-value is not significant, thereby implying that the heme feature is not enriched in the overlapping regions (i.e., the regions that have the significant loci).
[0116] Referring now to FIG. 23, graph 2300 is presented that illustrates the distribution of p-values and q-values across the whole genome for DNMT3A mutated CHIP. 62 significant loci were identified having an FDR of less than 0.05 (5%) and an absolute effect size greater than 0.1 (10%). 18 loci were identified as being hypermethylated (e.g., having increased DNA methylation in CHIP), and 44 loci were identified as being hypomethylated (e.g., having decreasing DNA methylation in CHIP).
[0117] Referring now to FIG. 24, graph 2400 illustrates a QQ plot in the analysis of DNMT3A mutated CHIP. In this plot, the observed p-values were plotted against expected p-values and a genomic inflation factor was determined from the resultant data points. As discussed above, a well-controlled study generally has a genomic inflation factor close 1 . The data associated with graph 2400 indicates that the genomic inflation factor was found to be 0.95.
[0118] Referring now to FIG. 25, graph 2500 illustrates a “Manhattan” plot associated with DNMT3A mutated CHIP. Manhattan plots are graphical representations commonly utilized in genomics to visualize the results of statistical tests conducted across the genome. Specifically, these plots may be utilized to identify specific regions of the genome where genetic variants are significantlyassociated with a trait or disease. For instance, graph 2500 indicates that certain genomic regions associated with chromosome 3 may be significant.
[0119] Referring now to FIG. 26, graph 2600 illustrates the WGBS feature distribution of significantly differentially methylated loci that are plotted against the background in DNMT3A mutated CHIP. Analysis of graph 2600 reveals that the significant loci may be arranged in the 3pUTR, exon_non_UTR_CDS, and intron regions.
[0120] Referring now to FIG. 27, graph 2700 illustrates the WGBS feature distribution of significantly differentially methylated loci that are plotted against the background in DNMT3A mutated CHIP. Analysis of graph 2700 reveals that the significant loci are arranged on the shelf and shore.
[0121] Referring now to FIG. 28, graph 2800 illustrates the results of a KEGG enrichment analysis conducted for DNMT3A mutated CHIP. With respect to graph 1700, the hsa04012ErbB signaling pathway and the hsa04660 T cell receptor signaling pathway were identified as significant genetic pathways potentially related to cancer.
[0122] Referring now collectively to FIGS. 29A and 29B, boxplots 2900 and 2905 illustrate the methylation trends of two significant loci (e.g., chr19:4259101 in FIG. 29A and chr19:30159761 in FIG. 29B) plotted against the age category for DNMT3A mutated CHIP.
[0123] Referring now collectively to FIGS. 30 and 31 , table 3000 in FIG. 30 and graph 3100 in FIG. 31 present validation results from the analysis of DNMT3A mutated CHIP. More particularly, the results associated with CHIP overall were compared against previous findings in the literature that were derived from a separate panel (e.g., a 450k Beadchip panel). This comparison revealed that theresults derived by the concepts disclosed herein were consistent with the findings in the literature. More particularly, with respect to table 3000, among the 39 loci having a p-value < 0.001 in both studies, 36 loci have the same directions and 3 have opposite directions. With respect to graph 3100, for those significant loci in the literature having FDR < 0.05 (5%), they also tend to have smaller p-value in the analysis conducted herein.
[0124] Referring now to FIG. 32, graph 3200 is presented that illustrates the distribution of p-values and q-values across the whole genome forTET2 mutated CHIP. 1655 significant loci were identified having an FDR of less than 0.05 (5%) and an absolute effect size greater than 0.1 (10%). 280 loci were identified as being hypermethylated (e.g., having increased DNA methylation in CHIP), and 1375 loci were identified as being hypomethylated (e.g., having decreasing DNA methylation in CHIP).
[0125] Referring now to FIG. 33, graph 3300 illustrates a QQ plot in the analysis of TET2 mutated CHIP. In this plot, the observed p-values were plotted against expected p-values, and a genomic inflation factor was determined from the resultant data points. A well-controlled study generally has a genomic inflation factor close 1 . The data associated with graph 3300 indicates that the genomic inflation factor was found to be 1 .00.
[0126] Referring now to FIG. 34, graph 3400 illustrates a “Manhattan” plot associated with TET2 mutated CHIP. Manhattan plots are graphical representations commonly utilized in genomics to visualize the results of statistical tests conducted across the genome. Specifically, these plots may be utilized to identify specific regions of the genome where genetic variants are significantly associated with a traitor disease. For instance, graph 3400 indicates that certain genomic regions associated with chromosome 16 may be significant.
[0127] Referring now to FIG. 35, graph 3500 illustrates the WGBS feature distribution of significantly differentially methylated loci that are plotted against the background in TET2 mutated CHIP. Analysis of graph 3500 reveals that the significant loci may be arranged in the 3pUTR, CDS, exon_non_UTR_CDS, and promoter regions.
[0128] Referring now to FIG. 36, graph 3600 illustrates the WGBS feature distribution of significantly differentially methylated loci that are plotted against the background in TET2 mutated CHIP. Analysis of graph 3600 reveals that the significant loci are generally arranged on the island and shore.
[0129] Referring now to FIG. 37, table 3700 illustrates the results of a KEGG enrichment analysis conducted for TET2 mutated CHIP. With respect to table 3700, the hsa04725 Cholinergic synapse, hsa04360 Axon guidance, hsa04080 Neuroactive ligand-receptor interaction, hsa04923 Regulation of lipolysis in adipocytes, hsa04550 Signaling pathways regulating pluripotency of stem cells, hsa0423 Longevity regulating pathway - multiple species, hsa04015 Rap1 signaling pathway, hsa04929 GnRH secretion, hsa04750 Inflammatory mediator regulation of TRP channel, hsa04935 Growth hormone synthesis, secretion and action, hsa04211 Longevity regulating pathway, and hsa04810 Regulation of actin cytoskeltonhsa04012ErbB signaling pathway were identified as significant genetic pathways potentially related to cancer.
[0130] Referring now collectively to FIGS. 38A and 38B, boxplots 3800 and 3805 illustrate the methylation trends of two significant loci (e.g., chr19:19221708 inFIG. 38A and chr19:5417118 in FIG. 38B) plotted against the age category for TET2 mutated CHIP.
[0131] Referring now collectively to FIGS. 39 and 40, table 3900 in FIG. 39 and graph 4000 in FIG. 40 present validation results from the analysis of TET2 mutated CHIP. More particularly, the results associated with CHIP overall were compared against previous findings in the literature that were derived from a separate panel (e.g., a 450k Beadchip panel). This comparison revealed that the results derived by the concepts disclosed herein were consistent with the findings in the literature. More particularly, with respect to table 3900, among the 48 loci having a p-value < 0.001 in both studies, 34 loci have the same directions, and 14 have opposite directions. With respect to graph 4000, for those significant loci in the literature having FDR < 0.05 (5%), they also tend to have smaller p-value in the analysis conducted herein.
[0132] In general, any process discussed in this disclosure that is understood to be computer-implementable may be performed by one or more processors of a computer system, such as system environment 110, as described above. A process or process step performed by one or more processors may also be referred to as an operation. The one or more processors may be configured to perform such processes by having access to instructions (e.g., software or computer-readable code) that, when executed by the one or more processors, cause the one or more processors to perform the processes. The instructions may be stored in a memory of the computer server. A processor may be a central processing unit (CPU), a graphics processing unit (GPU), or any suitable types of processing unit.
[0133] A computer system, such as system environment 110, may include one or more computing devices. If the one or more processors of the computer system are implemented as a plurality of processors, the plurality of processors may be included in a single computing device or distributed among a plurality of computing devices. If a system environment comprises a plurality of computing devices, the memory of the computer system may include the respective memory of each computing device of the plurality of computing devices.
[0134] FIG. 41 is a simplified functional block diagram of a computer system 4100 that may be configured as a computing device for executing the processes described herein, according to exemplary embodiments of the present disclosure. FIG. 41 is a simplified functional block diagram of a computer that may be configured according to exemplary embodiments of the present disclosure. In various embodiments, any of the systems herein may be an assembly of hardware including, for example, a data communication interface 4120 for packet data communication. The platform also may include a central processing unit (“CPU”) 4102, in the form of one or more processors, for executing program instructions. The platform may include an internal communication bus 4108, and a storage unit 4106 (such as ROM, HDD, SDD, etc.) that may store data on a computer readable medium 4122, although the system 4100 may receive programming and data via network communications via electronic network 4125 (e.g., voice, video, audio, images, or any other data over the electronic network 4125). The system 4100 may also have a memory 4104 (such as RAM) storing instructions 4124 for executing techniques presented herein, although the instructions 4124 may be stored temporarily or permanently within other modules of system 4100 (e.g., processor 4102 and / or computer readable medium 4122). The system 4100 also may include input andoutput ports 4112 and / or a display 4110 to connect with input and output devices such as keyboards, mice, touchscreens, monitors, displays, etc. The various system functions may be implemented in a distributed fashion on a number of similar platforms, to distribute the processing load. Alternatively, the systems may be implemented by appropriate programming of one computer hardware platform.
[0135] In this disclosure, the term “based on” means “based at least in part on.” The singular forms “a,” “an,” and “the” include plural referents unless the context dictates otherwise. The term “exemplary” is used in the sense of “example” rather than “ideal.” The terms “comprises,” “comprising,” “includes,” “including,” or other variations thereof, are intended to cover a non-exclusive inclusion such that a process, method, or product that comprises a list of elements does not necessarily include only those elements, but may include other elements not expressly listed or inherent to such a process, method, article, or apparatus. Relative terms, such as “about,” “approximately,” “substantially,” and “generally,” are used to indicate a possible variation of ±10% of a stated or understood value. In addition, the term “between” used in describing ranges of values is intended to include the minimum and maximum values described herein. The use of the term “or” in the claims and specification is used to mean “and / or” unless explicitly indicated to refer to alternatives only if the alternatives are mutually exclusive, although the disclosure supports a definition that refers to only alternatives and “and / or.” As used herein “another” may mean at least a second or more.
[0136] As used herein, the term “user” generally encompasses any person or entity, such as a researcher and / or a care provider (e.g., a doctor, etc.), that may desire information, resolution of an issue, or engage in any other type of interaction with a provider of the systems and methods described herein (e.g., via an applicationinterface resident on their electronic device, etc.). The term “electronic application” or “application” may be used interchangeably with other terms like “program,” or the like, and generally encompasses software that is configured to interact with, modify, override, supplement, or operate in conjunction with other software.
[0137] Program aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of executable code and / or associated data that is carried on or embodied in a type of machine-readable medium. “Storage” type media include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer of the mobile communication network into the computer platform of a server and / or from a server to the mobile device. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links, or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.
[0138] Furthermore, while some embodiments described herein include some but not other features included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the invention, and form different embodiments, as would be understood by those skilled in the art. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0139] Thus, while certain embodiments have been described, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the invention, and it is intended to claim all such changes and modifications as falling within the scope of the invention. For example, functionality may be added or deleted from the block diagrams and operations may be interchanged among functional blocks. Steps may be added or deleted to methods described within the scope of the present invention.
[0140] The above disclosed subject matter is to be considered illustrative, and not restrictive, and the appended claims are intended to cover all such modifications, enhancements, and other implementations, which fall within the true spirit and scope of the present disclosure. Thus, to the maximum extent allowed by law, the scope of the present disclosure is to be determined by the broadest permissible interpretation of the following claims and their equivalents, and shall not be restricted or limited by the foregoing detailed description. While various implementations of the disclosure have been described, it will be apparent to those of ordinary skill in the art that many more implementations are possible within the scope of the disclosure. Accordingly, the disclosure is not to be restricted except in light of the attached claims and their equivalents.
Claims
WHAT IS CLAIMED IS:
1. A computer-implemented method, the computer-implemented method comprising: receiving, at a computing device, a set of nucleic acid methylation data; receiving, at the computing device, a designation of one or more genomic regions; identifying, using a process of the computing device and within the set of nucleic acid methylation data, one or more abnormal methylation features; and mapping, using the processor, the one or more abnormal methylation features to the one or more genomic regions.
2. The computer-implemented method of claim 1 , wherein the nucleic acid methylation data is obtained from a cell-free DNA (cfDNA) sample derived from a blood sample.
3. The computer-implemented method of claim 1 , wherein the identifying the one or more abnormal methylation features comprises utilizing a probabilistic model to evaluate the likelihood of abnormal methylation within each of the one or more genomic regions.
4. The computer-implemented method of claim 1 , further comprising generating a report output of the mapped one or more abnormal methylation features.
5. The computer-implemented method of claim 1 , further comprising utilizing a trained machine learning model to identify clonal hematopoiesis-related methylation patterns based on the identified one or more abnormal methylation features.
6. The computer-implemented method of claim 1 , further comprising filtering, prior to identifying the one or more abnormal methylation features, the nucleic acid methylation data to exclude regions with low sequencing coverage or low variability.
7. The computer-implemented method of claim 1 , further comprising generating a binary classification for each of the one or more abnormal methylation features, wherein the binary classification corresponds to a positive or negative association with a clonal hematopoiesis status.
8. A computer-implemented method, the computer-implemented method comprising: receiving, at a computing device, nucleic acid methylation data associated with a sample; defining, using a processor of the computing device, a binary call for the sample; employing, using the processor, a linear regression model to identify a relationship between methylation levels for each CpG site in the nucleic acid methylation data and a designation associated with clonal hematopoiesis of indeterminate potential (CHIP); employing, using the processor, a beta-binomial model to assess the relationship;identifying, using the processor, significant loci associated with those CpG sites in the nucleic acid methylation data identified as having a CHIP designation; and associating, using the processor, the significant loci with one or more genomic features.
9. The computer-implemented method of claim 8, wherein the binary call for the sample is based on a variant allele frequency (VAF) threshold.
10. The computer-implemented method of claim 9, wherein the VAF threshold is 0.02.11 . The computer-implemented method of claim 8, wherein the employing the linear regression model comprises employing one or more covariates to account for population-specific effects.
12. The computer-implemented method of claim 11 , wherein the one or more covariates may include one or more of: age, gender, and cell type proportion.
13. The computer-implemented method of claim 8, wherein the identified significant loci are annotated with one or more functional elements.
14. The computer-implemented method of claim 13, wherein the one or more functional elements include one or more: promotors, enhancers, or transcription factor binding sites.
15. The computer-implemented method of claim 8, further comprising generating a graphical visualization of the identified significant loci and associated genomic features.
16. The computer-implemented method of claim 8, further comprising utilizing the identified significant loci as an input feature for training a machine learning model to classify samples based on CHIP status.
17. The computer-implemented method of claim 8, further comprising training a machine learning model to distinguish between different subtypes of CHIP based on the identified significant loci.
18. The computer-implemented method of claim 8, further comprising removing, via a preprocessing technique, low-coverage CpG sites from the nucleic acid methylation data prior to employing the linear regression model.
19. A system, comprising: one or more processors; one or more computer readable media storing instructions that are executable by the one or more processors to perform operations to: receive nucleic acid methylation data associated with a sample; define a binary call for the sample; employ a linear regression model to identify a relationship between methylation levels for each CpG site in the nucleic acid methylation data anda designation associated with clonal hematopoiesis of indeterminate potential (CHIP); employ a beta-binomial model to assess the relationship; identify significant loci associated with those CpG sites in the nucleic acid methylation data identified as having a CHIP designation; and associate the significant loci with one or more genomic features.
20. The system of claim 19, further comprising utilizing the identified significant loci as an input feature for training a machine learning model to classify samples based on CHIP status.
Citation Information
Patent Citations
Methods and systems for predicting an origin of a variant
WO2022046947A1
Cellular heterogeneity–adjusted clonal methylation (CHALM): a methylation quantification method
WO2022226229A1
Detecting the presence of a tumor based on methylation status of cell-free nucleic acid molecules
WO2023197004A1