Systems and methods for constructing and utilizing a plasma cell disorder classifier to perform informed feature analysis

A non-invasive liquid biopsy using cfDNA methylation sequencing and a two-stage classifier addresses the limitations of current PCD diagnostics, enabling early and accurate detection and characterization of PCD subtypes.

WO2025217381A1PCT designated stage Publication Date: 2025-10-16GRAIL INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/024035
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-12-19
Filing Date
2025-04-10
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Current diagnostic methods for plasma cell disorders (PCD) are invasive, risky, and lack sensitivity and specificity, making early detection and characterization challenging.

Method used

A non-invasive liquid biopsy approach using targeted methylation sequencing of cell-free DNA (cfDNA) to detect specific DNA methylation signatures indicative of PCD stages, employing a two-stage classifier to identify the presence of PCD and classify its subtype.

Benefits of technology

Provides early detection and accurate characterization of PCD subtypes with high sensitivity and specificity, improving diagnostic performance over traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025024035_16102025_PF_FP_ABST
    Figure US2025024035_16102025_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method for characterizing disease progression is provided. The computer-implemented method may include: receiving, at a computing device, a set of nucleic acid methylation data; receiving, at the computing device, a designation of one or more genomic regions; generating, using a processor of the computing device, a trajectory of disease progression; identifying, using the processor and within the set of nucleic acid methylation data, one or more temporal methylation features associated with progression along the trajectory; and mapping, using the processor, the one or more temporal methylation features to the one or more genomic regions.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR CONSTRUCTING AND UTILIZING A PLASMA CELL DISORDER CLASSIFIER TO PERFORM INFORMED FEATURE ANALYSISCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application No. 63 / 736,028 filed on December 19, 2024, U.S. Provisional Application No.63 / 676,112, filed on July 26, 2024, and U.S. Provisional Application No. 63 / 632,723, filed on April 11 , 2024, which are incorporated by reference herein in their entireties.TECHNICAL FIELD

[0002] The present disclosure relates generally to the field of bioinformatics and genomics and, more specifically, to systems and methods for detecting and characterizing disease progression in plasma cell disorders (PCD) based on the analysis of temporal methylation signatures.BACKGROUND

[0003] Plasma cell disorders (PCD) are diseases in which a population of clonal plasma cells begins to proliferate and accumulate in the body, and in particular, in the bone marrow. Plasma cells are typically present in bone marrow in small amounts, but when abnormally proliferating plasma cells fill up bone marrow, they can cause damage to bones, form multiple tumors, and lead to multiple myeloma cancer. PCD can be characterized by subtypes through which the disease progresses from an asymptomatic state to a symptomatic advanced state. Early detection and subtype characterization of PCD can be critical to improving prognoses, but even in the advanced stage of multiple myeloma, plasma cellscirculating in the blood are at such a low frequency that they cannot be readily detected morphologically.

[0004] The background description provided herein is for the purpose of generally presenting the context of the disclosure. Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted to be prior art, or suggestions of the prior art, by inclusion in this section.SUMMARY OF THE DISCLOSURE

[0005] According to certain aspects of the disclosure, systems and methods are described for identifying methylation signatures related to plasma cell disorders (PCD).

[0006] In one aspect, a computer-implemented method is provided. The computer-implemented method may include: receiving, at a computing device, a set of nucleic acid methylation data; receiving, at the computing device, a designation of one or more genomic regions; generating, using a processor of the computing device, a trajectory of disease progression; identifying, using the processor and within the set of nucleic acid methylation data, one or more temporal methylation features associated with progression along the trajectory; and mapping, using the processor, the one or more temporal methylation features to the one or more genomic regions.

[0007] In another aspect, a system is provided. The system may include: one or more processors; one or more computer readable media storing instructions that are executable by the one or more processors to perform operations to: receive a set of nucleic acid methylation data; receive a designation of one or more genomic regions; generate, using the one or more processors, a trajectory of diseaseprogression; identify, using the one or more processors and within the set of nucleic acid methylation data, one or more temporal methylation features associated with progression along the trajectory; and map, using the one or more processors, the one or more temporal methylation features to the one or more genomic regions.

[0008] In yet another aspect, a non-transitory computer-readable medium storing computer-executable instructions is provided. The non-transitory computer- readable medium stores computer-executable instructions which, when executed by a system, may cause the system to perform operations comprising: receiving, at a computing device, a set of nucleic acid methylation data; receiving, at the computing device, a designation of one or more genomic regions; generating, using a processor of the computing device, a trajectory of disease progression; identifying, using the processor and within the set of nucleic acid methylation data, one or more temporal methylation features associated with progression along the trajectory; and mapping, using the processor, the one or more temporal methylation features to the one or more genomic regions.

[0009] Additional objects and advantages of the disclosed embodiments will be set forth in part in the description that follows, and in part will be apparent from the description, or may be learned by practice of the disclosed embodiments. The objects and advantages of the disclosed embodiments will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims.

[0010] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosed embodiments, as claimed.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments and together with the description, serve to explain the principles of the disclosure.

[0012] FIG. 1A is a diagram depicting an exemplary computer system for executing the methods described herein.

[0013] FIG. 1 B a diagram depicting an exemplary software platform for executing the methods described herein.

[0014] FIG. 2 is a diagram depicting an exemplary workflow for a method for identifying abnormal methylation features associated with PCD, according to one or more aspects of the present disclosure.

[0015] FIG. 3A depicts a graph that presents the distribution of samples in a study of PCDs categorized by sex, according to one or more aspects of the present disclosure.

[0016] FIG. 3B depicts a graph that presents the distribution of samples in a study of PCDs categorized by sex, according to one or more aspects of the present disclosure.

[0017] FIG. 4 depicts a plot that presents the distribution of probability of cancer score across different disease states, including non-cancer, MGUS, SMM, and MM, according to one or more aspects of the present disclosure.

[0018] FIG. 5 depicts a graph that presents the cross-validated detection sensitivities of the PCD classifier for different subtypes, including MGUS, SMM, and MM, according to one or more aspects of the present disclosure.

[0019] FIG. 6 depicts a graph that displays the detection sensitivity of thePCD classifier on an independent holdout test set derived from the training data, according to one or more aspects of the present disclosure.

[0020] FIG. 7A depicts a confusion matrix that presents predictions of the classifier for the three subtypes, MGUS, SMM, and MM, based on a first sample set, according to one or more aspects of the present disclosure.

[0021] FIG. 7B depicts a confusion matrix that presents the classifier’s predictions using an independent holdout set of test samples, according to one or more aspects of the present disclosure.

[0022] FIG. 8 depicts a heatmap that displays the log odds ratios from pairwise feature activation comparisons across three groups, according to one or more aspects of the present disclosure.

[0023] FIG. 9 depicts a graph presenting results of a uniform distribution on a manifold (MAP) analysis that was utilized to perform dimensional reduction on the data produced by the differential feature analysis illustrated in FIG. 8, according to one or more aspects of the present disclosure.

[0024] FIG. 10A depicts a UMAP plot of tumor methylated fraction (TMeF) values that are displayed through heat-mapping of the data points on the graph depicted in FIG. 9, according to one or more aspects of the present disclosure.

[0025] FIG. 10B depicts a UMAP plot of the predicted probability of cancer scores that are displayed through heat-mapping of the data points on the graph depicted in FIG. 9, according to one or more aspects of the present disclosure.

[0026] FIG. 11 A depicts a plot that visualizes the progression trajectory from non-cancerous conditions through various stages of PCD, according to one or more aspects of the present disclosure.

[0027] FIG. 11 B depicts the plot depicted in FIG. 11 A but color-coded by psuedotime, according to one or more aspects of the present disclosure.

[0028] FIG. 12A depicts a graph that presents the relationship between pseudotime and log10_TMeF, according to one or more aspects of the present disclosure.

[0029] FIG. 12B depicts a graph that presents the relationship between psuedotime and probability of cancer score, according to one or more aspects of the present disclosure.

[0030] FIG. 13 represents the UMAP projection of first and second sample population types, according to one or more aspects of the present disclosure.

[0031] FIG. 14 depicts a histogram that shows the distribution of pseudotime values for the first and second population of samples, according to one or more aspects of the present disclosure.

[0032] FIG. 15 depicts a heatmap that presents a visualization and summary of the clustering analysis performed on the features that change as a function of disease progression in PCDs, according to one or more aspects of the present disclosure.

[0033] FIG. 16 presents the heatmap depicted in FIG. 15 with comparisons highlighted between the feature activation patterns between the primary project and the holdout dataset, according to one or more aspects of the present disclosure.

[0034] FIG. 17 is a diagram depicting an exemplary computing system, according to one or more embodiments of the present disclosure.DETAILED DESCRIPTION OF EMBODIMENTS

[0035] The terminology used below may be interpreted in its broadest reasonable manner, even though it is being used in conjunction with a detailed description of certain specific examples of the present disclosure. Indeed, certain terms may even be emphasized below; however, any terminology intended to be interpreted in any restricted manner will be overtly and specifically defined as such in this Detailed Description section. Both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the features, as claimed.

[0036] In this disclosure, the term “based on” means “based at least in part on.” The singular forms “a,” “an,” and “the” include plural referents unless the context dictates otherwise. The term “exemplary” is used in the sense of “example” rather than “ideal.” The terms “comprises,” “comprising,” “includes,” “including,” or other variations thereof, are intended to cover a non-exclusive inclusion such that a process, method, or product that comprises a list of elements does not necessarily include only those elements, but may include other elements not expressly listed or inherent to such a process, method, article, or apparatus. Relative terms, such as “about,” “approximately,” “substantially,” and “generally,” are used to indicate a possible variation of ±10% of a stated or understood value. In addition, the term “between” used in describing ranges of values is intended to include the minimum and maximum values described herein. The use of the term “or” in the claims and specification is used to mean “and / or” unless explicitly indicated to refer to alternatives only or the alternatives are mutually exclusive, although the disclosure supports a definition that refers to only alternatives and “and / or.” As used herein “another” may mean at least a second or more.

[0037] As used herein, the term “user” generally encompasses any person or entity, such as a researcher and / or a care provider (e.g., a doctor, etc ), that may desire information, resolution of an issue, or engage in any other type of interaction with a provider of the systems and methods described herein (e.g., via an application interface resident on their electronic device, etc.). The term “electronic application” or “application” may be used interchangeably with other terms like “program,” or the like, and generally encompasses software that is configured to interact with, modify, override, supplement, or operate in conjunction with other software.

[0038] Plasma cell disorders (PCD)s are a type of disease in which a population of plasma cells derived from a single progenitor cell (i.e., clonal plasma cells) begins to proliferate and accumulate in the body, and in particular, in the bone marrow. Plasma cells are typically present in bone marrow in small amounts, but when they proliferate and accumulate, they can fill up bone marrow, cause damage to bones, form multiple tumors, and lead to cancer. PCD can be characterized by subtypes through which the disease progresses. One subtype of PCD is monoclonal gammopathy of unknown significance (MGLIS), an asymptomatic condition which affects about 3.2% of people over age 50, and about 5.3% of people over age 70. A small portion of people with MGUS will see their disease progress into the next subtype, smoldering multiple myeloma (SMM), another asymptomatic condition which affects 0.53% of people over age 40. A portion of people with SMM will see their disease progress to the most advanced subtype of multiple myeloma (MM), a symptomatic and cancerous form of PCD.

[0039] MGUS, SMM, and MM can be considered various stages existing on a continuum in which the accumulation of plasma cells increases to the point of causing symptoms such as organ damage and bone lesions. MGUS is characterizedby fewer than 10% of plasma cells in bone marrow and no symptoms; SMM is characterized by greater than 10% of plasma cells in bone marrow and no symptoms; and MM is characterized by greater than 10% of plasma cells in bone marrow and the presence of symptoms. Each year, about 1 % of people with MGUS will see their disease progress to MM. This clonal proliferation may occur due to various factors, including genetic mutations, aging, and / or exposure to certain environmental influences. Early detection and subtype characterization of PCD can be critical to improving prognoses, but even in the advanced stage of multiple myeloma, plasma cells circulating in the blood are at such a low frequency that they cannot be readily detected morphologically. In addition to the challenge of detecting and characterizing a small population of cells, PCD presents a challenge due to the fact that it encompasses multiple heterogeneous disease subtypes and a broad spectrum of disease burden in a continuous manner.

[0040] Conventionally, the standard for diagnosing and monitoring PCDs relies heavily on invasive procedures such as bone marrow biopsies, which involve extracting marrow tissue from the bone to analyze the presence of abnormal plasma cells. Additionally, imaging modalities like PET / CT scans are used to detect bone lesions indicative of MM. However, these methods have significant drawbacks. Bone marrow biopsies are painful, carry risks of complications such as infection and bleeding, and may not accurately reflect the disease burden due to sampling errors. Imaging techniques, while non-invasive, expose patients to radiation and may fail to detect early-stage disease or small lesions. Furthermore, existing biomarkers used in these diagnostic methods, such as monoclonal protein levels in the blood, lack the specificity and the sensitivity required to predict disease progression accurately.

[0041] In light of these challenges, there is a need for a non-invasive, reliable, and highly sensitive method for detecting and characterizing PCDs and their progression. Accordingly, the concepts described herein address this need by introducing a novel liquid biopsy approach utilizing targeted methylation sequencing of cell-free DNA (cfDNA) derived from subject blood samples. This method enables the detection of specific DNA methylation signatures that are indicative of various stages of PCDs. By employing a classifier trained to differentiate between non- cancerous samples and different PCD stages (e.g., MGUS, SMM, MM), this approach offers a non-invasive alternative to current diagnostic methods and may provide early detection and characterization of PCD, thereby improving the performance of existing disease classifiers, such as cancer classifiers. More particularly, the classifier may be a “2-stage” classifier that is designed to detect and characterize PCDs by analyzing methylation patterns in cfDNA extracted from blood samples. The classifier may be configured to operate in two stages: a first stage that is focused on detecting the presence of any PCD in the cfDNA sample and a second stage that is focused on classifying the specific subtype of the disorder detected (e.g., whether it is MGUS, SMM, or MM).

[0042] One method described herein corresponds to direct measurement of cell-free DNA and detection of MGUS, SMM, and / or MM through the identification of abnormal methylation features. Utilizing specific target genomic regions and a bio- feature-extractor tool, the concepts described herein are configured to identify abnormal methylation signals above the non-disease, e.g., non-cancer, baseline noise background. This method ensures a one-to-one mapping of features in the target genomic regions, allowing for precise measurement of abnormal methylation signals across genomic regions. Additionally, this method may improve accuracy andsensitivity compared to traditional somatic mutation-based approaches. Another method described herein leverages a generalized linear regression model that correlates methylation variations across the entire genome with PCD status. By incorporating factors such as age, gender, and cell type proportion, the analysis aims to identify significant loci and understand the epigenetic variations linked to PCD. This is a population-level approach that may provide insights into the broader epigenetic landscape associated with clonal hematopoiesis. Other methods, not explicitly listed and described here, may also be utilized.

[0043] The concepts described herein integrate various technological improvements and correspondingly improve the functionality of a computing device as used in the detection and characterization of PCD in several ways. For instance, all of the foregoing methods utilize computational tools and trained machine learning models to improve the ability of a computing device that supports these tools to more efficiently detect and analyze PCD. Additionally, the concepts described herein apply computational techniques to specific biological problems related to PCD. This practical application of computer technology may enhance the understanding of biological phenomena and improve diagnostic and / or research methodologies. For instance, in an aspect, the described concepts may make the computer-based PCD detection system more robust in its capacity to detect and characterize methylation patterns that may be associated with different subtypes of PCD as well as transitional states between the PCD subtypes. The computer executing and implementing a classifier model trained and used according to the techniques described herein may be better equipped to provide reliable mutation-related information. Furthermore, the processes executed by the computer to improve the model involve complex calculations and data manipulations on a large amount ofbiological data that a human individual could not reasonably complete on their own or in their mind. Specifically, computationally intensive statistical tests are leveraged by the computer to evaluate methylation patterns between sample sets, processes which cannot be completed by a human.

[0044] The subject matter of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, which form a part hereof, and which show, by way of illustration, specific exemplary embodiments. An embodiment or implementation described herein as “exemplary” is not to be construed as preferred or advantageous, for example, over other embodiments or implementations; rather, it is intended to reflect or indicate that the embodiment(s) is / are “example” embodiment(s). Subject matter may be embodied in a variety of different forms and, therefore, covered or claimed subject matter is intended to be construed as not being limited to any exemplary embodiments set forth herein; exemplary embodiments are provided merely to be illustrative. Likewise, a reasonably broad scope for claimed or covered subject matter is intended. Among other things, for example, subject matter may be embodied as methods, devices, components, or systems. Accordingly, embodiments may, for example, take the form of hardware, software, firmware, or any combination thereof. The following detailed description is, therefore, not intended to be taken in a limiting sense.

[0045] Throughout the specification and claims, terms may have nuanced meanings suggested or implied in context beyond an explicitly stated meaning. Likewise, the phrase “in one embodiment” or “in some embodiments,” or “in one aspect” or “in some aspects” as used herein does not necessarily refer to the same embodiment or aspect, and the phrase “in another embodiment” or “in another aspect” as used herein does not necessarily refer to a different embodiment oraspect. It is intended, for example, that claimed subject matter include combinations of exemplary embodiments in whole or in part.

[0046] Diseases referred to herein may include plasma cell disorder (PCD), also known as plasma cell dyscrasia. Non-limiting subtypes of PCD include, for example, monoclonal gammopathy of undetermined significance (MGUS), smoldering multiple myeloma (SMM), and multiple myeloma (MM). Additionally, it is also important to note that although the concepts described throughout this disclosure are made in reference to cancer, these designations are for exemplary purposes only and are not intended to be limiting. Specifically, the concepts described herein may be applicable to other disease types and other diseasedetecting machine-learning classifiers.

[0047] FIG. 1 A depicts an exemplary system for detecting and quantifying clonal PCD via the analysis of PCD-related methylation signatures. Exemplary system 100 includes a data collection component 10, a database 20, and device data intelligence component 30, operably connected to each other via network 40. Alternatively, or additionally, one or more of the components may be connected with another component locally without reliance on network connection; e.g., through a wired connection. In many aspects described herein, sequencing data of cell-free nucleic acids are used to illustrate the concepts. However, one of skill in the art would understand that the current method may be applied to sequencing data of DNA, RNA, or other materials, as well from a variety of sample types, e.g., a blood sample (e.g., a serum sample, a plasma sample, a whole blood sample), a urine sample, a saliva sample, a tissue sample, a bone marrow sample, etc.

[0048] As disclosed herein, data collection component 10 may include a device or machine with which sequencing data may be generated. In someembodiments, data collection component 10 may include one or more sequencing devices or a facility that uses one or more sequencing devices to generate nucleic acid (e.g., DNA or RNA) sequence data of biological samples. In some aspects, data collection component 10 may be a database that receives sequencing information generated from one or more sequencing devices. Any suitable liquid or solid biological samples may be used for sequencing. In some embodiments, a biological sample may be cell-based, for example, one or more types of tissue. In some embodiments, a biological sample may be a sample that includes cell-free nucleic acid fragments. Examples of biological samples include, but are not limited to, a blood sample (e.g., a cell-free DNA (cfDNA) sample, a cell-free RNA (cfRNA) sample, a serum sample, a plasma sample, a whole blood sample), a urine sample, a saliva sample, a tissue sample, a bone marrow sample, etc. Further, although sequencing of DNA from these samples is discussed herein, RNA from these samples may alternatively or additionally be sequenced.

[0049] Examples of sequencing data may include, but are not limited to, sequence read data of targeted genomic locations, partial or whole genome sequencing data of the genome represented by nucleic acid fragments in cell-free or cell-based samples, partial or whole genome sequencing data including one or more types of epigenetic modifications (e.g., methylation), or combinations thereof.

[0050] Data acquired by the data collection component 10 may be transferred to database 20 via network 40 or a local or network connection. In some embodiments, data collection component 10 may alternatively receive data from one or more sequencing devices. In some embodiments, the collected data may be analyzed by data intelligence component 30, via network 40 or a local or networkconnection. FIG. 1 B depicts exemplary functional modules that may be implemented to perform tasks of data intelligence component 30.

[0051] FIG. 1 B depicts an exemplary computer system 110 for measuring effects associated with PCD through the analysis of PCD-related methylation patterns. Exemplary system 110 achieves such functionalities by implementing, on one or more computer devices, user input and output (I / O) module 120, memory or database 130, data processing module 140, data analysis module 150, classification module 160, network communication module 170, and any other functional modules that may be needed for carrying out a particular task (e.g., an error correction or compensation module, a data compression module, etc.). As disclosed herein, user I / O module 120 may further include an input sub-module, such as a keyboard, and an output sub-module, such as a display (e.g., a printer, a monitor, or a touchpad). In some embodiments, all functionalities may be performed by one computer system. In some embodiments, the functionalities are performed by more than one computer system.

[0052] Also disclosed herein, a particular task may be performed by implementing one or more functional modules. In particular, each of the enumerated modules itself may, in turn, include multiple sub-modules. For example, data processing module 140 may include a sub-module for data quality evaluation (e.g., for discarding very short sequence reads or sequence reads including obvious errors), a sub-module for normalizing numbers of sequence reads that align to different regions of a reference genome, a sub-module to compensate / correct guanine-cytosine (GC) biases, a sub-module for matching data associated with a cancer sample with other data associated with one or more non-cancer samples, etc.

[0053] In some embodiments, a user may use I / O module 120 to manipulate data that is available either on a local device or can be obtained via a network connection from a remote service device or another user device. For example, I / O module 120 may allow a user, e.g., via a keyboard, a mouse, or a touchpad, to initiate or perform data analysis via a graphical user interface (GUI). In some embodiments, a user may manipulate data via voice control. In some embodiments, user authentication may be required before a user is granted access to the data being requested. In some embodiments, user I / O module 120 may be used to manage various functional modules. For example, a user may request via user I / O module 120 input data while an existing data processing session is in process. A user may do so by selecting a menu option or type in a command discretely without interrupting the existing process. In another example, a user may utilize user I / O module 120 to set various thresholds, configure sample matching settings, and / or provide other instructions to computer system 110, e.g., that dictate how data may be analyzed. As disclosed herein, a user may use any type of input to direct and control data processing and analysis via I / O module 120.

[0054] In some embodiments, system 110 further comprises a memory or database 130. In some embodiments, database 130 comprises a local database that may be accessed via user I / O module 120. In some embodiments, database 130 comprises a remote database that may be accessed by user I / O module 120 via a network connection. In some embodiments, database 130 is a local database that stores data retrieved from another device (e.g., a user device or a server). In some embodiments, memory or database 130 may store data retrieved in real-time from internet searches. In some embodiments, database 130 may send data to and receive data from one or more of the other functional modules, including, but notlimited to, a data collection module (not shown), data processing module 140, data analysis module 150, classification module 160, network communication module 170, and etc.

[0055] In some embodiments, database 130 may be a database local to the other functional modules. In some embodiments, database 130 may be a remote database that may be accessed by the other functional modules via wired or wireless network connection (e.g., via network communication module 170). In some embodiments, database 130 may include a local portion and a remote portion.

[0056] In some embodiments, system 110 comprises a data processing module 140. Data processing module 140 may receive data from I / O module 120 or database 130. In some embodiments, data processing module 140 may perform standard data processing algorithms, such as one or more of noise reduction, signal enhancement, normalization of counts of sequence reads, correction of GC bias, etc. In some embodiments, data processing module 140 may be configured to detect and measure methylation signatures, and specifically abnormal methylation features, associated with clonal hematopoiesis.

[0057] In some embodiments, system 110 comprises a data analysis module 150. In some embodiments, data analysis module 150 includes identifying and treating systematic errors in sequencing data, as described in connection with data processing module 140.

[0058] In some embodiments, system 110 comprises a classification module 160, which may embody a “machine-learning model” or “trained classifier.” As used herein, a “machine-learning model” or “trained classifier” generally encompasses instructions, data, and / or a model configured to receive input, and apply one or more of a weight, bias, classification, or analysis on the input to generate an output. Theoutput may include, for example, a classification of the input, an analysis based on the input, a design, process, prediction, or recommendation associated with the input, or any other suitable type of output. A machine-learning model is generally trained using training data, e.g., experiential data and / or samples of input data, which are fed into the model in order to establish, tune, or modify one or more aspects of the model, e.g., the weights, biases, criteria for forming classifications or clusters, or the like. Aspects of a machine-learning model may operate on an input linearly, in parallel, via a network (e.g., a neural network), or via any suitable configuration. In some aspects, the machine-learning model may be trained on a combination of real and synthetic sample data.

[0059] The execution of the machine-learning model may include deployment of one or more machine-learning techniques, such as k-nearest neighbors, linear regression, logistic regression, random forest, gradient boosted machine (GBM), deep learning, a deep neural network, and / or any other suitable machine-learning technique that solves problems in the field of Natural Language Processing (NLP). Supervised, semi-supervised, and / or unsupervised training may be employed. For example, supervised learning may include providing training data and labels corresponding to the training data, e.g., as ground truth. Unsupervised approaches may include clustering, classification or the like. K-means clustering or K-Nearest Neighbors may also be used, which may be supervised or unsupervised. Combinations of K-Nearest Neighbors and an unsupervised cluster technique may also be used. Any suitable type of training may be used, e.g., stochastic, gradient boosted, random seeded, recursive, epoch or batch-based, etc.

[0060] In an exemplary use case, a machine-learning model may be trained to analyze data from a test sample from a test subject whose status with respect to amedical condition is unknown and subsequently classifies the unknown test sample from the test subject based on the likelihood of the subject fitting into a particular category. In some embodiments, the one or more parameters may include a binomial probability score that is calculated based on logistic regression analysis, which may be referred to as a “probability of cancer score” or a “p_cancer score.” As disclosed herein, the p_cancer score may correspond to the likelihood of a subject having a certain medical condition, such as PCD, or a certain category of PCD. For example, a score of over a predefined threshold may indicate that the subject associated with a test sample is more likely to have cancer than not have cancer, more likely to have PCD than not have PCD, or more likely to have a specific type of PCD than any other type of PCD.

[0061] In some embodiments, the one or more parameters may include a sequencing or methylation data distribution pattern correlating with the presence of PCD. A subject associated with a test sample having sequencing or methylation data with a pattern resembling a PCD pattern to a sufficient degree may be predicted as having PCD. In some embodiments, a sequencing or methylation data distribution pattern may be identified in connection with a specific subtype of PCD, determining a tissue of origin or PCD signal origin, thus allowing a test sample to be classified as indicative of a certain PCD subtype. In some embodiments, the one or more parameters may include an estimate of cell-free nucleic acid molecules in a sample derived from tumor cells, which may be referred to as a “tumor methylated fraction” or “TMeF.” As disclosed herein, TMeF values may be calculated by identifying differentially-methylated regions in a reference database and measuring the quantity of nucleic acid molecules or molecule fragments matching the reference methylation signatures.

[0062] As disclosed herein, network communication module 170 may be used to facilitate communications between a user device, one or more databases, and any other suitable system or device through a wired or wireless network connection. Any communication protocol / device may be used, including, without limitation, a modem, an Ethernet connection, a network card (wireless or wired), an infrared communication device, a wireless communication device, and / or a chipset (such as a Bluetooth™ device, an 802.11 device, a WiFi device, a WiMax device, cellular communication facilities, etc.), a near-field communication (NFC), a Zigbee communication, a radio frequency (RF) or radio-frequency identification (RFID) communication, a PLC protocol, a 3G / 4G / 5G / LTE based communication, and / or the like. For example, a user device having a user interface platform for processing / analyzing CHIP-related methylation signature data may communicate with another user device with the same platform, a regular user device without the same platform (e.g., a regular smartphone), a remote server, a physical device of a remote loT local network, a wearable device, a user device communicably connected to a remote server, and etc.

[0063] The functional modules described herein are provided by way of example. It will be understood that different functional modules may be combined to create different utilities. It will also be understood that additional functional modules or sub-modules may be created to implement a certain utility.

[0064] Referring now to FIG. 2, an exemplary workflow 200 is provided for constructing a two-stage classifier with a trajectory-based differential methylation region (DMR) analysis. Aspects of the exemplary workflow 200 may be performed in accordance with some or all components described in FIG. 1 A and 1 B.

[0065] At step 205, data may be collected and preprocessed. This step may first involve collecting or receiving a diverse set of cfDNA biological samples from a group of subjects. These subjects may include those individuals diagnosed with MGUS, SMM, MM, and non-cancer controls. This comprehensive dataset may help capture the heterogeneity inherent in PCDs and enables the identification of distinct methylation patterns that signify different stages of disease. In an aspect, blood samples may be collected or received from these participants, and cfDNA may be extracted, ensuring a wide representation of disease stages and minimizing invasiveness. In some aspects, the collection and sequencing may have occurred previously, and samples and / or genetic data from those samples may be received at step 205.

[0066] Following data collection, targeted methylation sequencing techniques may be employed to identify and quantify methylation patterns across the genome. Methylation, which involves the addition of methyl groups to DNA, plays a significant role in regulating gene expression and may provide valuable insights into the presence and progression of diseases, including PCDs. By targeting specific regions of the DNA known to be involved in the development and progression of PCD subtypes (e.g., MGUS, SMM, and MM), the sequencing process may capture the unique methylation signatures associated with each disease state. These signatures may reveal changes in DNA methylation that occur as the disease progresses, providing a molecular snapshot of the disease at various stages. In some aspects, quality control measures may be implemented to verify the integrity and reliability of the sequencing results. This may involve removing any samples that do not meet predefined quality standards, such as those with low DNA yield or poor sequencing quality, such that only quality data is used in further analysis.

[0067] In some aspects, the sequencing data obtained from targeted methylation may be normalized to correct for technical variations and extract relevant features that are most informative for distinguishing between different disease states. The features may be selected based on their ability to represent unique methylation signatures that correlate with disease progression, thereby forming the basis for the subsequent machine learning analysis.

[0068] At step 210, in an aspect, techniques such as principal component analysis (PCA) may be applied to reduce the high-dimensional data into a lowerdimensional space. This reduction preserves the most significant variance in the data while discarding redundant and less informative components. By focusing on the key features that capture the primary differences between samples, dimensionality reduction may help to reveal underlying patterns and relationships among the different disease states, from MGLIS through MM. Following dimensionality reduction, the clustering process may group the samples into distinct clusters based on their methylation profiles. Techniques like k-means clustering may be used to identify these clusters, which represent different stages of disease progression. Clustering allows for the identification of naturally occurring groupings within the data, helping to distinguish between various disease subtypes and understand the continuum of disease progression. Together, dimensionality reduction and clustering may enable the classifier to leverage high-dimensional methylation data effectively, enhancing its ability to detect subtle differences between disease states.

[0069] At step 215, after reducing the dimensionality of the methylation data and clustering samples into groups that represent different disease stages, a trajectory that reflects the likely biological progression from one stage to another maybe mapped out. To perform this mapping, a minimum spanning tree (MST) algorithm may be leveraged, which is a mathematical graph that connects all the data points (or clusters) with the minimum possible total edge weight. In this context, each node represents a cluster or a single sample, and the edges between them reflect the similarity or distance between these nodes based on their methylation profiles. Using MST to map the data provides a visual framework to infer the developmental trajectory of PCDs.

[0070] At step 220, following the construction of an MST to map disease progression, pseudotimes may be inferred and smooth lineages may be constructed. In an aspect, pseudotime is an abstract metric that represents the progression of a sample along a biological trajectory, effectively ordering the samples according to their stage of disease development rather than chronological time. This concept may be useful in studying PCDs because it captures the continuum of disease progression from non-cancerous state through MGUS, SMM, and finally MM. In an aspect, to infer pseudotime, the MST may be used as a backbone structure, upon which a series of algorithms, such as “Slingshot,” may be applied. These algorithms fit smooth, non-linear curves through the high-dimensional space of the data points, essentially modeling the gradual changes in methylation patterns as the disease progresses. The smooth lineages generated by these curves represent the most likely paths of disease evolution. This process allows for the identification of key transition states and highlights markers that change continuously along the length of the pseudotime. By constructing these smooth lineages, researchers may gain a nuanced understanding of the dynamic nature of PCDs, enabling the identification of temporal features that may serve as potential diagnostic or prognostic biomarkers.

[0071] At step 225, changes in DNA methylation patterns across the continuum of disease progression may be examined to identify key methylation features, as inferred from the pseudotime and smooth lineages. Differential methylation region (DMR) analysis traditionally identifies genomic regions where methylation levels significantly differ between distinct groups, such as cancerous versus non-cancerous states. However, when applied along a trajectory, DMR analysis becomes a dynamic tool that models methylation changes as a continuous function of disease progression rather than discrete comparisons. This approach may be particularly valuable for PCDs, where the transition from MGUS to MM is not abrupt, but occurs along a continuum.

[0072] By aligning methylation data along the pseudotime inferred from the constructed smooth lineages, trajectory-based DMR analysis can detect regions of the genome where methylation levels gradually increase or decrease as the disease progresses. In an aspect, the significance of these methylation changes over pseudotime may be assessed, which may help to identify not only stable methylation features that define distinct disease stages, but also transitional features that may signify early progression or disease onset. More particularly, genomic regions where methylation levels significantly shift in a manner that correlates with disease development may be identified. These shifts may reveal key methylation features that are indicative of various phases of the disease, e.g., such as early onset features in MGUS, progression features in SMM, and advanced disease features inMM. In an aspect, once these methylation features are identified, they may be clustered into groups that represent key events in disease progression. By categorizing these features into distinct clusters, a roadmap of disease progressionmay be created, identifying not only the key markers at each stage but also potential transitional states may be targeted for early intervention.

[0073] At step 230, the methylation signatures obtained at step 225 may be utilized to build a machine learning model that may accurately distinguish between individuals with a PCD and those without. This “first phase” of the “two-stage” classifier may be designed to non-invasively detect the presence of PCDs such as MGLIS, SMM, and MM. During training, the classifier may learn to recognize the unique methylation landscape associated with the PCDs by analyzing how the methylation features behave across different samples. As discussed, these features may be selected because they have shown to change in a consistent, significant manner along the disease trajectory, making them highly informative for detecting the presence of PCD. By inputting these specific methylation features into the model, the classifier may be able to learn patterns that differentiate PCD subjects from non- cancerous individuals.

[0074] At step 235, responsive to confirming that a subject has PCD (e.g., by identifying methylation patterns indicative of any PCD), the classifier may be trained to accurately categorize the PCD identified in stage one into a specific subtype (e.g., MGUS, SMM, and MM). In an aspect, training the classifier to accurately categorize PCD subtypes may involve utilizing a more specialized feature set that captures the unique methylation patterns associated with each subtype of PCD. Features associated with the foregoing subtypes may be carefully selected based on their ability to differentiate between the subtypes in the trajectory analysis. The model in stage two may leverage these subtype-specific methylation patterns to learn the subtle differences between MGUS, SMM, and MM, thereby enabling more accurate classification. Specifically, by focusing on these refined features, stage twoenhances the specificity and accuracy of the diagnostic process, providing a detailed and nuanced understanding of the subject’s condition.

[0075] Provided below are a variety of figures that present various types of data associated with the development of the novel classifier described herein. At least a subset of these figures present data that is derived from samples obtained from a study population. This study population consisted of 148 subjects, from which 47 subjects had MGLIS, 59 subjects had SMM, and 39 subjects had MM. There were 71 female subjects and 77 male subjects, which indicates a relatively balanced sex distribution among the study participants. Twenty-seven subjects were under the age of 50, 64 subjects were between the ages of 50 to 70, and 57 subjects were older than 70, a distribution that shows that a majority of the participants are middle-aged to older adults, reflecting the typical demographic affected by PCDs.

[0076] Referring now to FIG. 3A and FIG. 3B, graphs 300 and 305 are presented, respectively, that illustrate the distribution of samples in a study of PCDs (specifically the subtypes MGLIS, SMM, and MM) categorized by sex and age. Graph 300 in FIG. 3A displays the distribution of samples by biological sex across the different disease subtypes. For each subtype, the samples were almost equally distributed between females and males. Graph 305 in FIG. 3B shows the distribution of samples by age category (<50, 50-70, and > 70 years) across the disease subtypes. Age distribution across the various groups was comparable between MGUS and SMM, with the bulk of subjects in the two groups being 50 years or older. By contrast, in the MM group, the population of subjects was distributed equally across the various age brackets. This distribution highlights that MGUS and SMM predominately affect older adults, whereas MM has a more even spread across different age groups.

[0077] Referring now to FIG. 4, graph 400 presents a plot that illustrates the distribution of p-cancer across different disease states, including non-cancer, MGUS, SMM, and MM. The p-cancer score is a classifier output representing the probability of a sample being cancerous, with higher scores indicating a greater likelihood of cancer presence. Mean p_cancer scores for each disease state are indicated as open triangles. The mean p_cancer score for non-cancer samples was close to zero, while the mean p_cancer score for PCD samples increased from MGUS to SMM to MM groups, indicating increased disease burden as PCD progresses. More particularly, the distribution of p-cancer scores for MGUS is broader, with a significant range from low to moderate scores. The p-cancer scores for SMM show further elevation, with many samples having higher scores that approach the upper limit of the scale. The mean p-cancer scores for SMM is notably higher than MGUS, indicating an increased disease burden and a higher probability of cancer development as the condition progresses. For MM, the p-cancer scores are consistently high, with most samples nearing a score of 1 , indicating a strong probability of cancer.

[0078] Referring now to FIG. 5, graph 500 presents the cross-validated detection sensitivities of the PCD classifier for different subtypes, including MGUS, SMM, and MM. Additionally, Table 1 below lists the detection sensitivity percentages with corresponding 95% confidence intervals (Cl) and the number of correctly classified samples out of the total samples tested.Table 1

[0079] Collective examination of graph 500 and Table 1 reveals that the detection sensitivity for MM is reported as 100%, with a 95% Cl ranging from 91 .0 to 100.0. This indicates that the classifier was able to correctly identify all MM samples in the cross-validation set, demonstrating perfect sensitivity with the sampled population. Such a high sensitivity is very important for MM because it is a more advanced and symptomatic stage of PCD. Further examination reveals that the detection sensitivity for SMM is 91 .5% with a 95% Cl from 81.3 to 97.2. This high sensitivity indicates that the classifier is also effective in detecting SMM. As SMM is a precursor stage to MM, the ability to identify SMM with good sensitivity is important for early intervention and monitoring, potentially inhibiting progression to MM. Further examination reveals that the detection sensitivity for MGUS is 93.6% with a 95% Cl from 82.5 to 98.7. MGUS is the earliest and often asymptomatic stage in the spectrum of PCDs, and achieving a high sensitivity is beneficial for early detection. This allows for monitoring and possibly intervening before the disease progresses to more severe stages.

[0080] Referring now to FIG. 6, graph 600 displays the detection sensitivity of the PCD classifier on an independent holdout test set derived from the training data. Graph 600 specifically highlights the performance of the classifier when applied to new, unseen samples, specifically focusing on the detection of MGUS and MM. Additionally, Table 2 below provides a summary of the detection sensitivity of the PCD classifier for two subtypes: MGUS and MM.Table 2

[0081] Collective examination of graph 600 and Table 2 reveals that the detection sensitivity for MGUS in this holdout set is notably low at 16.4% with a 95% Cl ranging from 8.2 to 28.1 . For instance, out of 61 MGUS samples, only 10 were correctly identified by the classifier. This drop in sensitivity may be suggestive of one or more potential issues. For instance, if the MGUS samples in this holdout set were self-reported and lacked pathological confirmation, then they may not represent true MGUS cases or have inconsistent molecular profiles, leading to poor performance by the classifier. Additionally or alternatively, the lower sensitivity may also be due to different molecular characteristics or phenotype distributions of MGUS in the holdout set compared to the training data. Further examination reveals that the detection sensitivity for MM in the holdout set remains high at 96.9% with a 95% Cl from 83.8 to 99.9. Although there is a slight drop from the 100% sensitivity reported during cross-validation (e.g., as represented by graph 500 in FIG. 5 and Table 1), it still reflects a strong ability of the classifier to detect MM.

[0082] Referring now to FIG. 7A and FIG. 7B, confusion matrices for the PCD classifier are presented, depicting its performance in predicting discrete subtypes of PCDs. Collectively, these matrices are utilized to assess how accurately the classifier predicts each subtype based on the actual diagnosis. Each cell in the matrix indicates the number of samples classified as particular subtype versus the actual subtype. More particularly, rows represent the predicted labels and the columns represent the actual labels. As illustrated in FIGs. 7A and 7B below, theconfusion matrices 700 and 705 demonstrate that predicting discrete subtypes is challenging.

[0083] Confusion matrix 700 in FIG. 7A shows predictions of the classifier for three subtypes: MGUS, SMM, and MM, based on a first sample set. In confusion matrix 700, the overall precision of the PCD subtyping classifier was observed to be 51 .1 % in its predictions. Of 44 MGUS samples from the first sample population, 22 samples were correctly predicted, while the other 22 samples were incorrectly predicted to be SMM. Of 54 SMM samples, 29 samples were correctly predicted as SMM, while 12 samples were incorrectly predicted as MM and 13 samples were incorrectly predicted as MGUS. Of 39 MM samples, 19 samples were correctly predicted, while 15 samples were incorrectly predicted to be SMM and 5 samples were incorrectly predicted to be MGUS.

[0084] Confusion matrix 705 in FIG. 7B shows the classifier’s predictions using an independent holdout set of test samples. In confusion matrix 705, the PCD subtyping classifier was observed to be 82.9% precise in its predictions. Of 10 MGUS samples, 4 samples were correctly predicted, while 3 samples were incorrectly predicted as SMM and 3 were incorrectly predicted as MM. Of 31 MM samples, 30 samples were correctly predicted while 1 sample was incorrectly predicted as SMM. Overall, the precision with the holdout test samples is significantly higher at 82.9%, suggesting better performance compared to the first sample set. However, the absence of SMM cases in this dataset means the classifier performance on this subtype cannot be evaluated. Additionally, the small sample size limits the generalizability of these results.

[0085] The difficulty in accurately predicting discrete subtypes reflects the complexity of plasma cell disorders, where overlapping molecular and clinicalfeatures may lead to intermediate states that are hard to classify. This suggests that further refinement of the classifier is needed. Overall, the confusion matrices presented in FIGS. 7A and 7B emphasize the need for more robust classification tools that can accurately differentiate between closely related disease states.

[0086] The results shown in FIGS. 3A, 3B, 4, 5, 6, and 7 demonstrated that without the utilization of the novel classifier constructed via the steps outlined in the exemplary workflow shown in FIG. 2, the PCD subtyping classifier reliably predicted the most severe MM subtype of PCD, but struggled to distinguish between neighboring, discrete subtypes of PCD. This observation may be due to the fact that traditional differential methylation region analysis treats all samples within a group as homogenous, and may be inadequate to detect methylation signatures of transitional states in a heterogeneous group. Accordingly, the present disclosure contemplates the use of additional and / or alternative analysis techniques to identify methylation signatures associated with specific subtypes of PCD and associated with transitions from one PCD subtype to another.

[0087] Referring now to FIG. 8, heatmap 800 displays the log odds ratios from pairwise feature activation comparisons across three groups. More particularly, the top row indicates methylation feature activation from non-cancer to MGUS; the middle row indicates methylation feature activation from MGUS to SMM; and the bottom row indicates methylation feature activation from SMM to MM. For each row, red colors indicate that the features are more significantly activated in group 1 and blue colors indicate that the features are more significantly activated in group 2. For instance, in the second row, red colors are indicative of feature activation in SMM, and blue colors are indicative of feature activation in MGUS.

[0088] Heatmap 800 is divided into several blocks (1 , 2, 3, and 4), each representing distinct patterns of methylation feature activation as the disease progresses from non-cancer through MGUS, SMM, to MM. Differential feature activation analysis was used to identify feature activation patterns between PCD subtypes. A feature value matrix generated by the PCD subtyping classifier was transformed to binary format (1 = activated, 0 = non-activated). For each feature, a log odds ratio was used to assess feature activation between the paired groups. As shown in FIG. 8, heat map 800 depicts feature activation comparisons between paired groups demonstrated both linear feature activation patterns (e.g., patterns 2 and 4) and non-linear activation patterns (e.g., patterns 1 and 3) along the PCD subtype continuum.

[0089] Methylation feature activation pattern 1 , outlined and labeled as “1” on FIG. 8, represents initial activation of PCD-related methylation features from noncancer to the MGUS stage, with those features continuing to be stably expressed at comparable levels through all stages of PCD. Such a pattern suggests that certain methylation changes are initially more prominent in MGUS but become less pronounced in SMM before escalating again in MM.

[0090] Methylation feature activation pattern 4, outlined and labeled as “4” on FIG. 8, represents onset progressive activation of PCD-related methylation features from the MGUS stage to the SMM stage, with those features increasingly expressed from the SMM stage to the MM stage at a lower rate of increase as compared to the rate of increase from the MGUS stage to the SMM stage. This pattern indicates that certain features become more relevant only from the SMM stage onwards, increasing as the disease progresses to MM. It suggests that these methylation changes are more associated with disease progression rather than initiation.

[0091] Methylation feature activation pattern 2, outlined and labeled as “2” onFIG. 8, represents onset progressive activation of PCD-related methylation features from SMM to MM. This pattern supports the concept of continuous disease progression, where methylation changes steadily accumulate as the disease advances from a benign stage (MGUS) through intermediate (SMM) to a malignant state (MM).

[0092] Methylation feature activation pattern 3, outlined and labeled as “3” on FIG. 8, represents peak activation of PCD-related methylation features at the MGUS stage and deactivation at the SMM stage. This pattern suggests that certain methylation changes occur early in disease development (e.g., at the MGUS stage) and remain relatively stable without further significant changes as the diseases progresses to SMM and MM.

[0093] Referring now to FIG. 9, uniform distribution on a manifold (UMAP) analysis was used to perform dimensional reduction on the data produced by the differential feature analysis illustrated in FIG. 8. Parameters used to generate the UMAP graph include: an n_neighbors value of 10, a minimum distance of 0.3, and retaining 60% of variance. As shown in graph 900 in FIG. 9, UMAP analysis of the methylation features identified by the PCD subtyping classifier allowed for successful subtype separation and characterization of disease progression pathways (e.g., from non-cancerous states to advanced disease stages). For example, the majority of non-camper samples are represented by data points in the upper-left corner of the graph, and the majority of MM samples are represented by data points in the lower- right corner of the graph. The majority of MGUS samples are represented by data points in the center-left portion of the graph, upper to the majority of SMM samples, which are represented by data points in the center-right portion of the graph. Thus,disease progression is represented on the UMAP graph as a pathway from the upper-left corner to the lower-right corner. This visualization supports the notion that PCDs do not progress in discrete jumps but rather along a continuum, in which intermediate states blend into each other. Additionally, the proximity of MGUS, SMM, and MM samples in the UMAP space indicates that the boundaries between these subtypes are not sharp. This visual representation aligns with the understanding that these conditions share overlapping biological features and that PCD progression involves gradual changes in molecular characteristics. Additionally, the visualization highlights the potential for using dimensional reduction and clustering techniques like UMAP to identify specific progression markers and better understand the transitions between disease stages.

[0094] FIG. 10A and FIG. 10B present two UMAP plots, plot 1000 and plot 1005 respectively, that illustrate how two variables, tumor methylated fraction (TMeF) and p_cancer (probability of cancer), change as the disease progresses through different stages of plasma cell disorders. As shown in plot 1000 in FIG. 10A, predicted TMeF values are displayed through heat-mapping of the data points on the UMAP graph of FIG. 9. Dark blue represents low log10_TMeF values, indicating a lower fraction of tumor-specific methylation. Dark red represents high log10_TMeF values, indicating a high fraction of tumor-specific methylation. Light blue to light red represents intermediate log10_TMeF values, indicating moderate tumor methylation levels. As the plot progresses from top-left to bottom-right (moving along the UMAP axes), there is a notable increase from the upper-left corner to the lower-right corner of the graph, aligning with PCD disease progression as validated against actual PCD subtype. The clustering of darker red points towards the bottom-right indicates that more advanced stages of PCD, like MM, exhibit higher tumor methylation fractions,consistent with increased disease burden and more extensive tumor involvement. As shown in plot 1005 in FIG. 10B, predicted probability of cancer (p_cancer) scores are displayed through heat-mapping of the data points on the UMAP graph of FIG. 9. Blue indicates a low probability of cancer, red indicates a high probability of cancer, and light blue to light red indicate intermediate probabilities of cancer. Examination of plot 1005 reveals that probability of cancer scores, another measure of disease burden in cancers and related disease, increase from the upper-left corner to the lower-right corner of the graph, also aligning with PCD disease progression as validated against actual PCD subtype. The clustering of red points towards the bottom-right indicates that samples classified as more advanced stages of PCD, like MM, have a higher probability of being identified as cancer by the classifier.

[0095] Referring now to FIGS. 11 A and 11 B, the use of trajectory analysis to understand the progression of PCD from benign conditions through MGLIS and SMM to overt cancer (MM) is illustrated. Traditional DMR analysis often treats each disease stage as a separate, homogenous group and compares them against each other. This method may fail to capture the subtleties of disease progression, especially when dealing with heterogeneous groups that do not fit neatly into discrete categories. PCDs, like many cancers, do not progress in a simple, linear fashion but rather through a continuum with overlapping stages. Subjects can present at various points along this continuum, reflecting differences in disease burden and molecular changes. By using a trajectory-based approach, the analysis can map out the progression from benign to overt cancer conditions, thereby capturing the full spectrum of disease states. This method allows for a more nuanced understanding of how methylation changes accumulate over time and how different disease stages relate to each other. Additionally, trajectory analysis may identify molecularsignatures associated with transitional states that traditional DMR methods may miss. This is particularly valuable for understanding early markers of progression or features that distinguish slow-progressing MGUS from rapidly progressing MM.

[0096] In an aspect, the trajectory is constructed using algorithms like Slingshot, which first identify a global lineage structure by constructing a minimum spanning tree (MST) among clusters of samples. Then, smooth lineages are inferred to map out the progression pathway, creating a pseudotime variable that orders samples along this trajectory. This helps create a methylation-driven timeline that reflects the actual clinical progression from non-cancerous conditions to MGUS, SMM, and finally MM. This approach accommodates the heterogeneity of PCD and can track the evolution of disease in a continuous manner.

[0097] Referring now to FIG. 11A, plot 1100 visualizes the progression trajectory from non-cancerous conditions through various stages of PCD (MGUS, SMM, MM). Each dot represents a sample, colored according to its actual disease state, e.g., green corresponding to non-cancer, yellow corresponding to MGUS, orange corresponding to SMM, and red corresponding to MM. The black line running through the center represents the inferred trajectory of disease progression, showing a continuum from non-cancer to MM. The path reflects how samples transition through disease states in a way that mimics actual disease progression in subjects. Referring now to FIG. 11 B, plot 1105 shows the same UMAP projection but color- coded by psuedotime (which represents a measure of disease progression derived from the trajectory analysis). Blue to green represents the early stages of progression (non-cancer, early MGUS) and yellow to red represents later stages of progression (late MGUS, SMM, MM). The gradient from blue to red visuallyrepresents how samples evolve from early to late stages of disease, with pseudotime providing a continuous measure of this progression rather than discrete categories.

[0098] Referring now to FIGS. 12A and 12B, scatter plots 1200 and 1205 are presented that correlate pseudotime with two different metrics: the log-transformed TMeF (log10_TMeF) and the probability of cancer (p_cancer), respectively. Turning first to FIG. 12A, plot 1200 presents the relationship between pseudotime and log10_TMeF. Each dot represents a sample, color-coded according to its actual disease state: green corresponding to non-cancer, yellow corresponding to MGLIS, orange corresponding to SMM, and red corresponding to MM. A blue regression line is plotted to illustrate the trend of increasing log10_TMeF with increasing pseudotime. The correlation coefficient (R) is 0.62, with a p-value less than 2.2e-16, indicating a moderate positive correlation, suggesting that as the pseudotime increases, the tumor methylation fraction also increases. This increase in log10_TMeF reflects the accumulation of tumor-specific methylation changes, which is expected as the disease progress from benign to malignant stages. Turning now to FIG. 12B, plot 1205 presents the relationship between psuedotime and p_cancer. The color-coding of the data points is the same as in FIG. 12A. A blue regression line is plotted to illustrate a positive trend where the probability of cancer increases with increasing psuedotime. The R-value is 0.88 and the p-value is less than 2.2e-16, indicating a strong positive correlation, suggesting that as the pseudotime increases, the probability of cancer also increases. This may further suggest that the pseudotime effectively captures the progression from non-cancer to more malignant states.

[0099] FIG. 13 presents plot 1300 that represents the UMAP projection of first and second sample population types. Specifically, plot 1300 highlights howthese samples align with the established trajectory of disease progression from noncancer to MGLIS, SMM, and MM. The black line represents the inferred trajectory of disease progression from non-cancer to MM. The first population MM samples (red circles) align well with the second population (red triangles) along the trajectory towards the end of the progression line. This indicates that the first population samples are similar to the second population samples in terms of their molecular characteristics and disease progression stage, thereby validating the trajectory model for MM. The second population MGLIS samples (yellow triangles) are projected primarily in the region corresponding to non-cancer and early MGUS states, overlapping significantly with non-cancer samples. This clustering suggests that these second population MGUS samples may be more similar to non-cancer states or early-stage disease than to the MGUS samples from the first population set. This disparity may be attributed to the nature of the second population MGUS samples, which, as mentioned in the disclosure, are self-reported, without pathological confirmation.

[0100] FIG. 14 presents histogram 1400 that shows the distribution of pseudotime values for the first and second population of samples. Each bar represents the count of samples within specific pseudotime intervals, color-coded by their actual disease state. The distribution of the second sample population across pseudotime intervals shows a gradual shift from non-cancer (green) to MGUS (yellow), SMM (orange), and MM (red) as pseudotime increases. This panel reflects the continuous progression of disease from benign to malignant states. The distribution of the second population of samples (top panel) shows that MGUS samples are mostly positioned in the pseudotime intervals overlapping with non- cancer samples, with fewer samples extending into the higher pseudotime intervalstypical of more advanced stages. This further supports the observation that the second sample population is more similar to non-cancer samples and / or represents an early disease state with lower disease burden.

[0101] FIG. 15 illustrates a heatmap 1500 that presents a visualization and summary of the clustering analysis performed on the features that change as a function of disease progression in PCDs. A total of 1058 features were found to be significantly associated with disease progression. These features were clustered into three distinct groups using K-means clustering, based on their temporal activation patterns. The first is the “Early Initiator” cluster (cluster 3), that contains features that start to appear early in the MGLIS stage. The consistent upregulation of these features from MGUS through SMM to MM suggests that they play a role in the initial development of the disease and continue to be relevant as the disease progresses. Identifying early initiator features is important for understanding the early molecular changes that occur in MGUS and could serve as early biomarkers for disease detection and monitoring. The second is the “Indicator of Entering Advanced Disease” cluster (cluster 1), that contains features that are primarily upregulated during the late SMM and MM phases. These features may be indicative of important molecular changes that occur as the disease moves into more aggressive and symptomatic phases, making them potential targets for therapeutic intervention or markers for predicting disease progression. The third is the “Mid-SMM cluster” (cluster 2), that contains features that start to emerge during the mid-SMM phase and show increased upregulation in the MM phase. They represent molecular changes that occur as the disease transitions from an intermediate to an advanced state. Mid-SMM indicator features help in understanding the molecular dynamics thatoccur during this transitional phase and could provide insights into the mechanisms that drive disease progression.

[0102] FIG. 16 presents the heatmap 1500 illustrated in FIG. 15 comparing the feature activation patterns between the primary project and the holdout dataset. Heatmap 1600 of FIG. 16 shows that the feature activation patterns identified in the primary project dataset are largely consistent with those found in the holdout dataset. This consistency across datasets supports the robustness and reliability of the identified feature clusters as markers of disease progression in PCDs.

[0103] In general, any process discussed in this disclosure that is understood to be computer-implementable may be performed by one or more processors of a computer system, such as system environment 110, as described above. A process or process step performed by one or more processors may also be referred to as an operation. The one or more processors may be configured to perform such processes by having access to instructions (e.g., software or computer-readable code) that, when executed by the one or more processors, cause the one or more processors to perform the processes. The instructions may be stored in a memory of the computer server. A processor may be a central processing unit (CPU), a graphics processing unit (GPU), or any suitable types of processing unit.

[0104] A computer system, such as system environment 110, may include one or more computing devices. If the one or more processors of the computer system are implemented as a plurality of processors, the plurality of processors may be included in a single computing device or distributed among a plurality of computing devices. If a system environment comprises a plurality of computingdevices, the memory of the computer system may include the respective memory of each computing device of the plurality of computing devices.

[0105] FIG. 17 is a simplified functional block diagram of a computer system 1700 that may be configured as a computing device for executing the processes described herein, according to exemplary embodiments of the present disclosure. FIG. 17 is a simplified functional block diagram of a computer that may be configured according to exemplary embodiments of the present disclosure. In various embodiments, any of the systems herein may be an assembly of hardware including, for example, a data communication interface 1720 for packet data communication. The platform also may include a central processing unit (“CPU”) 1702, in the form of one or more processors, for executing program instructions. The platform may include an internal communication bus 1708, and a storage unit 1706 (such as ROM, HDD, SDD, etc.) that may store data on a computer readable medium 1722, although the system 1700 may receive programming and data via network communications via electronic network 1725 (e.g., voice, video, audio, images, or any other data over the electronic network 1725). The system 1700 may also have a memory 1704 (such as RAM) storing instructions 1724 for executing techniques presented herein, although the instructions 1724 may be stored temporarily or permanently within other modules of system 1700 (e.g., processor 1702 and / or computer readable medium 1722). The system 1700 also may include input and output ports 1712 and / or a display 1710 to connect with input and output devices such as keyboards, mice, touchscreens, monitors, displays, etc. The various system functions may be implemented in a distributed fashion on a number of similar platforms, to distribute the processing load. Alternatively, the systems may be implemented by appropriate programming of one computer hardware platform.

[0106] In this disclosure, the term “based on” means “based at least in part on.” The singular forms “a,” “an,” and “the” include plural referents unless the context dictates otherwise. The term “exemplary” is used in the sense of “example” rather than “ideal.” The terms “comprises,” “comprising,” “includes,” “including,” or other variations thereof, are intended to cover a non-exclusive inclusion such that a process, method, or product that comprises a list of elements does not necessarily include only those elements, but may include other elements not expressly listed or inherent to such a process, method, article, or apparatus. Relative terms, such as “about,” “approximately,” “substantially,” and “generally,” are used to indicate a possible variation of ±10% of a stated or understood value. In addition, the term “between” used in describing ranges of values is intended to include the minimum and maximum values described herein. The use of the term “or” in the claims and specification is used to mean “and / or” unless explicitly indicated to refer to alternatives only if the alternatives are mutually exclusive, although the disclosure supports a definition that refers to only alternatives and “and / or.” As used herein “another” may mean at least a second or more.

[0107] As used herein, the term “user” generally encompasses any person or entity, such as a researcher and / or a care provider (e.g., a doctor, etc.), that may desire information, resolution of an issue, or engage in any other type of interaction with a provider of the systems and methods described herein (e.g., via an application interface resident on their electronic device, etc.). The term “electronic application” or“application” may be used interchangeably with other terms like “program,” or the like, and generally encompasses software that is configured to interact with, modify, override, supplement, or operate in conjunction with other software.

[0108] Program aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of executable code and / or associated data that is carried on or embodied in a type of machine-readable medium. “Storage” type media include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer of the mobile communication network into the computer platform of a server and / or from a server to the mobile device. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links, or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.

[0109] Furthermore, while some embodiments described herein include some but not other features included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the invention, and form different embodiments, as would be understood by those skilled in the art. Forexample, in the following claims, any of the claimed embodiments can be used in any combination.

[0110] Thus, while certain embodiments have been described, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the invention, and it is intended to claim all such changes and modifications as falling within the scope of the invention. For example, functionality may be added or deleted from the block diagrams and operations may be interchanged among functional blocks. Steps may be added or deleted to methods described within the scope of the present invention.

[0111] The above disclosed subject matter is to be considered illustrative, and not restrictive, and the appended claims are intended to cover all such modifications, enhancements, and other implementations, which fall within the true spirit and scope of the present disclosure. Thus, to the maximum extent allowed by law, the scope of the present disclosure is to be determined by the broadest permissible interpretation of the following claims and their equivalents, and shall not be restricted or limited by the foregoing detailed description. While various implementations of the disclosure have been described, it will be apparent to those of ordinary skill in the art that many more implementations are possible within the scope of the disclosure. Accordingly, the disclosure is not to be restricted except in light of the attached claims and their equivalents.

Claims

CLAIMS1 . A computer-implemented method for characterizing disease progression, the computer-implemented method comprising: receiving, at a computing device, a set of nucleic acid methylation data; receiving, at the computing device, a designation of one or more genomic regions; generating, using a processor of the computing device, a trajectory of disease progression; identifying, using the processor and within the set of nucleic acid methylation data, one or more temporal methylation features associated with progression along the trajectory; and mapping, using the processor, the one or more temporal methylation features to the one or more genomic regions.

2. The computer-implemented method of claim 1 , wherein generating a trajectory of disease progression comprises: performing pseudotime trajectory analysis to generate the trajectory of disease progression.

3. The computer-implemented method of claim 1 , wherein identifying one or more methylation features associated with progression along the trajectory comprises: performing trajectory-based differential gene expression analysis.

4. The computer-implemented method of claim 1 , wherein the one or more temporal methylation features associated with progression along the trajectory comprises one or more clusters of methylation features.

5. The computer-implemented method of claim 4, wherein the one or more clusters of methylation features are defined by changes in gene expression as a function of disease progression.

6. The computer-implemented method of claim 1 , wherein the disease comprises at least two subtypes, and wherein a temporal methylation feature of the one or more temporal methylation features is associated with at least one subtype of the disease.

7. The computer-implemented method of claim 1 , wherein the disease comprises at least two subtypes, and wherein a temporal methylation feature of the one or more temporal methylation features is associated with transition from a first subtype of the disease to a second subtype of the disease.

8. The computer-implemented method of claim 1 , wherein the disease is plasma cell disease.

9. The computer-implemented method of claim 8, wherein the disease progression of plasma cell disease comprises progression from monoclonalgammopathy of unknown significance (MGUS), to smoldering multiple myeloma (SMM), to multiple myeloma (MM).

10. The computer-implemented method of claim 9, wherein the one or more temporal methylation features is associated with MGUS, SMM, MM, onset of MGUS, transition from MGUS to SMM, and / or transition from SMM to MM.11 . A system, comprising: one or more processors; and one or more computer readable media storing instructions that are executable by the one or more processors to: receive a set of nucleic acid methylation data; receive a designation of one or more genomic regions; generate, using the one or more processors, a trajectory of disease progression; identify, using the one or more processors and within the set of nucleic acid methylation data, one or more temporal methylation features associated with progression along the trajectory; and map, using the one or more processors, the one or more temporal methylation features to the one or more genomic regions.

12. The system of claim 11 , wherein the instructions that are executable by the one or more processors to generate comprise instructions that are executable by the one or more processors to:perform pseudotime trajectory analysis to generate the trajectory of disease progression.

13. The system of claim 11 , wherein the instructions that are executable by the one or more processors to identify comprise instructions that are executable by the one or more processors to: perform trajectory-based differential gene expression analysis.

14. The system of claim 11 , wherein the one or more temporal methylation features associated with progression along the trajectory comprises one or more clusters of methylation features.

15. The system of claim 14, wherein the one or more clusters of methylation features are defined by changes in gene expression as a function of disease progression.

16. The system of claim 11 , wherein the disease comprises at least two subtypes, and wherein a temporal methylation feature of the one or more temporal methylation features is associated with at least one subtype of the disease.

17. The system of claim 11 , wherein the disease comprises at least two subtypes, and wherein a temporal methylation feature of the one or more temporal methylation features is associated with transition from a first subtype of the disease to a second subtype of the disease.

18. The system of claim 11 , wherein the disease is a plasma cell disorder.

19. The system of claim 18, wherein the disease progression of plasma cell disease comprises progression from monoclonal gammopathy of unknown significance (MGUS), to smoldering multiple myeloma (SMM), to multiple myeloma (MM); and wherein the one or more temporal methylation features is associated with MGUS, SMM, MM, onset of MGUS, transition from MGUS to SMM, and / or transition from SMM to MM.

20. A non-transitory computer-readable medium storing computer-executable instructions which, when executed by a system, cause the system to perform operations comprising: receiving, at a computing device, a set of nucleic acid methylation data; receiving, at the computing device, a designation of one or more genomic regions; generating, using a processor of the computing device, a trajectory of disease progression; identifying, using the processor and within the set of nucleic acid methylation data, one or more temporal methylation features associated with progression along the trajectory; and mapping, using the processor, the one or more temporal methylation features to the one or more genomic regions.

Citation Information

Patent Citations

  • Cancer Classification with Synthetic Training Samples

    US20210310075A1

  • Systems and methods for diagnosing a disease or a condition

    WO2024050541A1