Systems and methods for synthetic data titration for data augmentation

By generating synthetic data sets through the combination of real cancer and non-cancer samples, the limitations of real data availability for training machine learning classifiers are addressed, resulting in improved performance and detection capabilities, especially in intermediate disease stages.

WO2025117652A1PCT designated stage expired Publication Date: 2025-06-05GRAIL INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/057630
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-11-27
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

The availability of diverse and comprehensive data sets for training machine learning disease classifiers, such as those for non-small cell lung cancer (NSCLC), is limited due to data scarcity, cost, and real-world constraints, leading to challenges in accurately detecting and classifying diseases, especially in intermediate stages.

Method used

The use of data augmentation techniques to generate synthetic data sets by combining real cancer and non-cancer samples in specific proportions, allowing for the creation of synthetic samples that mimic desired tumor methylation fractions, thereby expanding the training data and improving classifier performance.

Benefits of technology

The approach enhances the diversity and volume of training data, leading to improved detection and classification performance of machine learning classifiers, particularly in intermediate disease stages, and addresses the limitations of real data availability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024057630_05062025_PF_FP_ABST
    Figure US2024057630_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for generating synthetic data may include receiving an indication to generate a synthetic data set having characteristics associated with a designated tumor methylation fraction; identifying, based on the received indication, a first data set associated with a real cancer sample and a second data set associated with a real non-cancer sample; determining, based on the designated tumor methylation fraction, a first proportion of the first data set to combine with a second proportion of the second data set; selecting, based on the determining, a first subset of the first data set that corresponds to the first proportion and a second subset of the second data set that corresponds to the second proportion; and generating, using the processor, at least one synthetic data sample in the synthetic data set by combining the first subset of the first data set with the second subset of the second data set.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR SYNTHETIC DATA TITRATION FOR DATA AUGMENTATIONCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of priority from U.S. Provisional Patent Application No. 63 / 604,064, filed November 29, 2023, which is hereby incorporated by reference in its entirety.TECHNICAL FIELD

[0002] The present disclosure relates generally to the field of medical data analysis and, more specifically, to systems and methods for improving the training and performance assessment of a machine learning disease classifier.BACKGROUND

[0003] Non-small cell lung cancer (NSCLC) is a prevalent and lifethreatening disease with limited treatment options, particularly in advanced stages. The early detection and accurate assessment of minimal residual disease (MRD) in NSCLC patients, as in many cancers, is critical for improving treatment outcomes and patient survival rates. The development of machine learning classifiers has helped aid in the identification and monitoring of diseases such as NSCLC in subjects. These classifiers are most effective when they are trained on sufficiently large and diverse data sets (e.g., pluralities of samples exhibiting a diverse array of cancer signal intensity). However, various factors such as data scarcity, cost, and / ortemporal or other real-world constraints may hinder the ability to train classifiers using a diverse data set.

[0004] The background description provided herein is for the purpose of generally presenting context of the disclosure. Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted to be prior art, or suggestions of the prior art, by inclusion in this section.SUMMARY OF THE DISCLOSURE

[0005] According to certain aspects of the disclosure, systems and methods are described for utilizing data augmentation techniques to artificially expand the training data that may be used to train a machine learning classifier.

[0006] In one aspect, a computer-implemented method is provided. The computer-implemented includes: receiving, at a computing device, an indication to generate a synthetic data set having characteristics associated with a designated tumor methylation fraction; identifying, based on the received indication and using a processor of the computing device, a first data set associated with a real cancer sample and a second data set associated with a real non-cancer sample; determining, based on the designated tumor methylation fraction and using the processor, a first proportion of the first data set to combine with a second proportion of the second data set; selecting, based on the determining and using the processor, a first subset of the first data set that corresponds to the first proportion and a second subset of the second data set that corresponds to the second proportion; and generating, using the processor, at least one synthetic data sample satisfying the designated tumor methylation fraction in the synthetic data set by combining the first subset of the first data set with the second subset of the second data set.

[0007] In another aspect, a system is provided. The system may include: one or more processors; one or more computer readable media storing instructions that are executable by the one or more processors to perform operations to: receive, at a computing device associated with the system, an indication to generate a synthetic data set having characteristics associated with a designated tumor methylation fraction; identify, based on the received indication, a first data set associated with a real cancer sample and a second data set associated with a real non-cancer sample; determine, based on the designated tumor methylation fraction and using the processor, a first proportion of the first data set to combine with a second proportion of the second data set; select, based on the determining, a first subset of the first data set that corresponds to the first proportion and a second subset of the second data set that corresponds to the second proportion; and generate at least one synthetic data sample satisfying the designated tumor methylation fraction in the synthetic data set by combining the first subset of the first data set with the second subset of the second data set.

[0008] In yet another aspect, a non-transitory computer-readable medium storing computer-executable instructions is provided. The non-transitory computer- readable medium stores computer-executable instructions which, when executed by a system, cause the system to perform operations comprising: receiving, at a computing device, an indication to generate a synthetic data set having characteristics associated with a designated tumor methylation fraction; identifying, based on the received indication and using a processor of the computing device, a first data set associated with a real cancer sample and a second data set associated with a real non-cancer sample; determining, based on the designated tumor methylation fraction and using the processor, a first proportion of the first data set tocombine with a second proportion of the second data set; selecting, based on the determining and using the processor, a first subset of the first data set that corresponds to the first proportion and a second subset of the second data set that corresponds to the second proportion; and generating, using the processor, at least one synthetic data sample satisfying the designated tumor methylation fraction in the synthetic data set by combining the first subset of the first data set with the second subset of the second data set.

[0009] Additional objects and advantages of the disclosed embodiments will be set forth in part in the description that follows, and in part will be apparent from the description, or may be learned by practice of the disclosed embodiments. The objects and advantages of the disclosed embodiments will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims.

[0010] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosed embodiments, as claimed.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments and together with the description, serve to explain the principles of the disclosure.

[0012] FIG. 1 depicts a graph illustrating challenges associated with conventional real data collection.

[0013] FIG. 2A depicts an exemplary computer system for executing the methods described herein.

[0014] FIG. 2B depicts an exemplary software platform for executing the methods described herein.

[0015] FIG. 3 depicts an exemplary workflow for creating synthetic data to enhance the training of disease classifiers, according to one or more embodiments of the present disclosure.

[0016] FIG. 4 depicts an exemplary diagram illustrating a process for combining real cancer data and real non-cancer data, according to one or more embodiments of the present disclosure.

[0017] FIG. 5 depicts a table that illustrates an example of how metrics associated with real data and synthetic data may be organized and stored, according to one or more embodiments of the present disclosure.

[0018] FIG. 6 depicts an exemplary diagram that illustrates a process of combining real cancer data and a plurality of real non-cancer data sets based on a designated targeted tumor methylation fraction, according to one or more embodiments of the present disclosure.

[0019] FIG. 7 depicts an exemplary diagram that illustrates a training and validation process for a machine learning classifier, according to one or more embodiments of the present disclosure.

[0020] FIG. 8 depicts a graph that illustrates the performance improvement of a classifier that is trained on a mix of real and synthetic sample data, according to one or more embodiments of the present disclosure.

[0021] FIGS. 9A and 9B depict graphs illustrating the improvement in model performance in two different NSCLC subtypes realized by using synthetically- expanded training data, according to one or more embodiments of the present disclosure.

[0022] FIGS. 10A - 10D depict graphs illustrating the improvement in model performance across different cancer stages realized by using synthetically-expanded training data, according to one or more embodiments of the present disclosure.

[0023] FIG. 11 A - 11 D depict graphs illustrating the improvement to the detection rate that can be observed in NSCLC subtypes across different cancer stages, according to one or more embodiments of the present disclosure.

[0024] FIG. 12 depicts a graph illustrating the improvement in the Limit of Detection (LOD) in two different NSCLC subtypes realized by using synthetically- expanded training data set, according to one or more embodiments of the present disclosure.

[0025] FIG. 13 depicts a flowchart of an exemplary method of generating synthetic data, according to one or more embodiments of the present disclosure.

[0026] FIG. 14 depicts a graph that illustrates specificity levels across different exemplary cross-validation models for evaluating classifier performance, according to one or more embodiments of the present disclosure.

[0027] FIG. 15 depicts a diagram that represents an exemplary scenario illustrating how the inclusion of synthetic mixtures affects the weighting of real samples during classifier training, according to one or more embodiments of the present disclosure.

[0028] FIG. 16 depicts a diagram that further highlights how the inclusion of synthetic mixtures may affect the weighting of non-cancer samples during the specificity determination phase of classifier training, according to one or more embodiments of the present disclosure.

[0029] FIG. 17 depicts a graph that illustrates the empirical specificity of a classifier trained using a first sample weighting approach during classifier training, according to one or more embodiments of the present disclosure.

[0030] FIG. 18 depicts a graph that illustrates the improvement in the empirical specificity of a classifier after implementing a second sample weighting approach during classifier training, according to one or more embodiments of the present disclosure.

[0031] FIG. 19 depicts a graph that represents how the implementation of a new sample weighting approach resolves issues of underrepresentation for noncancer samples used as diluents during the specificity determination process, according to one or more embodiments of the present disclosure.

[0032] FIG. 20 depicts an example computing system, according to one or more embodiments of the present disclosure.DETAILED DESCRIPTION OF EMBODIMENTS

[0033] The terminology used below may be interpreted in its broadest reasonable manner, even though it is being used in conjunction with a detailed description of certain specific examples of the present disclosure. Indeed, certain terms may even be emphasized below; however, any terminology intended to be interpreted in any restricted manner will be overtly and specifically defined as such in this Detailed Description section. Both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the features, as claimed.

[0034] A fundamental challenge in developing effective disease-detecting classifiers (e.g., cancer detecting classifiers, such as an NSCLC classifier) is theavailability of a sufficiently large, comprehensive, and diverse data set for training and evaluation. For instance, NSCLC is a complex disease, and accurate classifiers require a substantial volume of data to train on. However, the availability of NSCLC samples for research purposes is often limited, making it difficult and timeconsuming to build robust classifiers. For instance, acquiring real NSCLC patient data may be an expensive and resource-intensive process, which may be further exacerbated if negotiations need to be conducted with third-party entities.

[0035] Further, for certain disease states, the type of sample data that is needed may not be readily available. For example, the majority of existing NSCLC sample data generally falls into two categories: samples with very low cancer signal (e.g., procured from subjects during the early stages of the disease) and samples with very high cancer signal (e.g., procured from subjects during the advanced stages of the disease). These sample extremes provide valuable insights into the early and late stages of NSCLC but fail to adequately represent the intermediate ranges of cancer signal. Data from these middle ranges may be useful to have in order to further train the classifier and improve its disease recognition capabilities.

[0036] The lack of intermediate samples for some diseases may be due, in part, to the fleeting nature of the signal in disease subjects. As used herein, a “disease signal” refers to changes in nucleic acid methylation that are associated with the development or progression of a disease. For example, a “cancer signal” refers to changes in nucleic acid methylation that are associated with the development or progression of cancer. More particularly, the cancer “dwell time,” i.e., the phase when a detectable cancer signal in a patient sample transitions from low to high, or the subject transitions from disease free to a more advanced disease state, may not be a lengthy temporal window. During this brief intermediate phase,the cancer signal may exhibit characteristics that are valuable for early detection and MRD assessment. If data collected is not synchronized with the dwell time (e.g., due to the lack of relevant subject availability and / or a lack of knowledge that a subject even has the disease, etc.) then opportunities may be missed to collect this intermediate data for classifier training. Moreover, it can be further difficult to target data collection efforts to rectify this data collection gap because it can be very difficult to identify prospective disease subjects during the intermediate phase.

[0037] An illustration of the foregoing deficiencies with existing sample data collection is depicted via graph 100 in FIG. 1. More particularly, graph 100 highlights the distinctions between a desired fit line 2 (i.e., a line generated when an abundant / ideal amount of data is available across all stages of disease progression) and a best-fit line 4 for available data (i.e., a line generated based on the amount of real data available). Specifically, it can be observed that the desired fit line 2 and the best-fit line 4 for available data are substantially similar at the early and late stages of the disease, because these are the stages where real data (as represented by dots) has conventionally been the most available. However, it can further be observed that there is a large departure between the two lines during the middle stages of disease progression, as encompassed by box 6. This departure is resultant from a lack of available real data during these intermediate disease stages, for reasons such as those described above.

[0038] Accordingly, the present disclosure is designed to address one or more of the foregoing challenges by leveraging various data augmentation techniques to artificially increase the size and diversity of the training data set(s) that can be used to train a disease classifier. More particularly, the concepts described herein may be utilized to computationally titrate two or more samples together,thereby generating synthetic data that covers the intermediate disease signals, such as those highlighted in box 6 of FIG. 1 , where real data may be lacking. Specifically, the data set for a “spike” sample from a subject known to have a disease, e.g., cancer, may be combined, in designated fractions, with a corresponding data set for a non-disease, e.g., a non-cancer, background sample from a subject known to not have a disease, referred to herein as a “diluent.” In other words, a data set from a non-disease sample falling on the left side of the curve in FIG. 1 may be mixed in various concentrations with a data set from a disease sample falling on the right side of the curve in FIG. 1 to create a synthetic titration intended to mimic a sample that would fall along the central region of the curve in FIG. 1 , or along a further left side of the curve, for which no real sample, or insufficient real samples, may be available. In various aspects, this mixing may occur dynamically, e.g., in response to detection of predetermined criteria by the computer system (e.g., upon determining that the amount of available training data is below a predetermined volume, upon determining that the amount of real data at a specific disease stage is below a predetermined volume, etc.).

[0039] Aside from randomly sampling to create the synthetic mixture, this disclosure describes a synthetic titration process that targets specific metrics for a given synthetic sample. As an example, in the case of cancer, the process creates a mixture sample that mimics a desired tumor methylation fraction (“TMeF”) or “tumor fraction”, a measure of the proportion of abnormally methylated, tumor-derived cfDNA in a cancer patient’s sample. As described herein, it is advantageous to target a TMeF sample exhibiting a middling cancer signal. The generated synthetic samples may be used, in conjunction with real sample data, to train a diseaseclassifier, thereby improving its detection and classification performance capabilities by improving the availability of data that is otherwise missing from the training set.

[0040] In an aspect, the concepts described herein integrate various technological elements, including: data processing, bioinformatics, computational biology, and machine learning. These elements are manipulated to create synthetic data, which is then used to optimize classifier training. Specifically, the generation of the synthetic data bridges the data gap that conventionally hinders the optimal training of traditional disease classifiers, e.g., cancer classifiers such as NSCLC classifiers. Correspondingly, this synthetic data generation may enhance a computer’s data processing capabilities. For instance, by creating synthetic data to augment real-world data sets, the computer’s functionality may be enhanced in terms of data diversity (e.g., computers may be enabled to work with larger, more varied data sets, which, in turn, enhances the performance of machine learning algorithms). Specifically, the resultant data diversity leads to more robust and adaptable classifiers that can correspondingly generate more accurate and reliable results. Additionally, because the sample “mixing” occurs entirely in the digital realm, synthetic samples may be generated “on-the-fly,” leading to improved efficiency and scalability. More particularly, the ability to generate synthetic data digitally in realtime is a practical solution to the challenges associated with obtaining and processing real data. Based on the foregoing, it should be apparent that the concepts described herein cannot reasonably be performed mentally by a user due to several significant limitations. For instance, biomedical data, especially in the context of cancer detection and MRD assessment, can be highly complex and multidimensional, often involving a large number of variables, such as genetic markers, clinical data, and patient histories. It is not plausible for an individual tomentally, or physically, manipulate and generate synthetic data with the necessary precision and complexity needed to optimally train a classifier. Moreover, because the synthetic samples are generated on-demand by the classifier based on a variety of target variables, it is further not plausible for an individual to generate the large volumes of synthetic samples according to the methods described herein without significantly slowing the training process for the disease classifier. Therefore, the techniques described herein are directed to providing a technological solution to a problem native to the technological field of training high-performance disease classifiers.

[0041] The subject matter of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, which form a part hereof, and which show, by way of illustration, specific exemplary embodiments. An embodiment or implementation described herein as “exemplary” is not to be construed as preferred or advantageous, for example, over other embodiments or implementations; rather, it is intended to reflect or indicate that the embodiment(s) is / are “example” embodiment(s). Subject matter may be embodied in a variety of different forms and, therefore, covered or claimed subject matter is intended to be construed as not being limited to any exemplary embodiments set forth herein; exemplary embodiments are provided merely to be illustrative. Likewise, a reasonably broad scope for claimed or covered subject matter is intended. Among other things, for example, subject matter may be embodied as methods, devices, components, or systems. Accordingly, embodiments may, for example, take the form of hardware, software, firmware, or any combination thereof. The following detailed description is, therefore, not intended to be taken in a limiting sense.

[0042] Throughout the specification and claims, terms may have nuanced meanings suggested or implied in context beyond an explicitly stated meaning. Likewise, the phrase “in one embodiment” or “in some embodiments,” or “in one aspect” or “in some aspects” as used herein does not necessarily refer to the same embodiment or aspect, and the phrase “in another embodiment” or “in another aspect” as used herein does not necessarily refer to a different embodiment or aspect. It is intended, for example, that claimed subject matter include combinations of exemplary embodiments in whole or in part.

[0043] Non-limiting cancer types that the concepts described herein may be applied to include, for example, breast cancer, lung cancer (e.g., non-small cell lung cancer (NSCLC)), prostate cancer, colorectal cancer, renal cancer, uterine cancer, pancreatic cancer, cancer of the esophagus, a lymphoma, head and neck cancer, ovarian cancer, a hepatobiliary cancer, a melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, and gastric cancer. Additionally, it is also important to note that although the concepts described throughout this disclosure are made in reference to cancer, e.g., lung cancer and an NSCLC classifier, these designations are for exemplary purposes only and are not intended to be limiting. Specifically, the concepts described herein may be applicable to other disease types and other disease-detecting classifiers. Additionally, although “synthetic data” is described throughout the specification as a mixing of data associated with two samples (e.g., a real cancer sample and a real non-cancer sample), such a designation is not limiting and is made for simplicity purposes only. More particularly, data from multiple samples may be mixed together to form synthetic data. For example, data from one real cancer sample and two real noncancer samples may be mixed together to generate a synthetic data set. As anotherexample, data from one real cancer sample, one real non-cancer sample, and one previously generated synthetic sample may be mixed together to create another synthetic sample.

[0044] FIG. 2A depicts an exemplary system for creating synthetic data to enhance the training of disease classifiers. Exemplary system 200 includes a data collection component 10, a database 20, and device data intelligence component 30, operably connected to each other via network 40. Alternatively, or additionally, one or more of the components may be connected with another component locally without reliance on network connection; e.g., through a wired connection. In many aspects described herein, sequencing data of cell-free nucleic acids are used to illustrate the concepts. However, one of skill in the art would understand that the current method may be applied to sequencing data of DNA, RNA, or other materials as well from a variety of sample types, e.g., a blood sample (e.g., a serum sample, a plasma sample, a whole blood sample), a urine sample, a saliva sample, a tissue sample, a bone marrow sample, etc.

[0045] As disclosed herein, data collection component 10 may include a device or machine with which sequencing data may be generated. In some embodiments, data collection component 10 may include one or more sequencing devices or a facility that uses one or more sequencing devices to generate nucleic acid (e.g., DNA or RNA) sequence data of biological samples. In some aspects, data collection 10 may be a database that receives sequencing information generated from one or more sequencing devices. Any suitable liquid or solid biological samples may be used for sequencing. In some embodiments, a biological sample may be cell-based, for example, one or more types of tissue. In some embodiments, a biological sample may be a sample that includes cell-free nucleic acid fragments.Examples of biological samples include, but are not limited to, a blood sample (e.g., a cell-free DNA (cfDNA) sample, a serum sample, a plasma sample, a whole blood sample), a urine sample, a saliva sample, a tissue sample, a bone marrow sample, etc. Further, although sequencing of DNA from these samples is discussed herein, RNA from these samples may alternatively or additionally be sequenced.

[0046] Examples of sequencing data may include, but are not limited to, sequence read data of targeted genomic locations, partial or whole genome sequencing data of the genome represented by nucleic acid fragments in cell-free or cell-based samples, partial or whole genome sequencing data including one or more types of epigenetic modifications (e.g., methylation), or combinations thereof.

[0047] Data acquired by the data collection component 10 may be transferred to database 20 via network 40 or a local or network connection. In some embodiments, the collected data may be analyzed by data intelligence component 30, via network 40 or a local or network connection. FIG. 2B depicts exemplary functional modules that may be implemented to perform tasks of data intelligence component 30.

[0048] FIG. 2B depicts an exemplary computer system 210 for generating synthetic sample data that may be used to train a machine-learning classifier. Exemplary system 210 achieves such functionalities by implementing, on one or more computer devices, user input and output (I / O) module 120, memory or database 130, data processing module 140, data analysis module 150, classification module 160, network communication module 170, and any other functional modules that may be needed for carrying out a particular task (e.g., an error correction or compensation module, a data compression module, etc.). As disclosed herein, userI / O module 120 may further include an input sub-module, such as a keyboard, andan output sub-module, such as a display (e.g., a printer, a monitor, or a touchpad). In some embodiments, all functionalities may be performed by one computer system. In some embodiments, the functionalities are performed by more than one computer system. The various modules (e.g., for data processing, analysis, classification, communication, etc.) may be one or more processes executing in a distributed computing environment. For instance, in some embodiments, one or more components of the computer system 210 may be network accessible via cloud infrastructure. For example, the database 130 used to store data may be stored in one or more remote cloud servers. In this regard, the database may be one or more large storage buckets (e.g., cloud-based storage buckets such as simple storage service “S3” buckets, etc.) from which data may be retrieved on demand. As another example, data processing, analysis, and classification may be performed in cloudbased environments using services like cloud-based data processing platforms, serverless computing, cloud-based machine learning platforms, and the like.

[0049] Also disclosed herein, a particular task may be performed by implementing one or more functional modules. In particular, each of the enumerated modules itself may, in turn, include multiple sub-modules. For example, data processing module 140 may include a sub-module for data quality evaluation (e.g., for discarding very short sequence reads or sequence reads including obvious errors), a sub-module for normalizing numbers of sequence reads that align to different regions of a reference genome, a sub-module to compensate / correct GC biases, a sub-module for matching data associated with a cancer sample with other data associated with one or more non-cancer samples, etc.

[0050] In some embodiments, a user may use I / O module 120 to manipulate data that is available either on a local device or can be obtained via a networkconnection from a remote service device or another user device. For example, I / O module 120 may allow a user, e.g., via a keyboard, a mouse, or a touchpad, to perform data analysis via a graphical user interface (GUI). In some embodiments, a user may manipulate data via voice control. In some embodiments, user authentication may be required before a user is granted access to the data being requested. In some embodiments, user I / O module 120 may be used to manage various functional modules. For example, a user may request via user I / O module 120 input data while an existing data processing session is in process. A user may do so by selecting a menu option or type in a command discretely without interrupting the existing process. In another example, a user may utilize user I / O module 120 to set various thresholds, configure sample matching settings, and / or provide other instructions to computer system 210 that dictate how synthetic mixtures are generated and subsequently utilized. As disclosed herein, a user may use any type of input to direct and control data processing and analysis via I / O module 120.

[0051] In some embodiments, system 210 further comprises a memory or database 130. In some embodiments, database 130 comprises a local database that may be accessed via user I / O module 120. In some embodiments, database 130 comprises a remote database that may be accessed by user I / O module 120 via network connection. In some embodiments, database 130 is a local database that stores data retrieved from another device (e.g., a user device or a server). In some embodiments, memory or database 130 may store data retrieved in real-time from internet searches. In some embodiments, database 130 may send data to and receive data from one or more of the other functional modules, including, but not limited to, a data collection module (not shown), data processing module 140, dataanalysis module 150, classification module 160, network communication module 170, and etc. In some embodiments, some or all real-sample data and / or synthetic sample data may be stored on database 130.

[0052] In some embodiments, database 130 may be a database local to the other functional modules. In some embodiments, database 130 may be a remote database that may be accessed by the other functional modules via wired or wireless network connection (e.g., via network communication module 170). In some embodiments, database 130 may include a local portion and a remote portion.

[0053] In some embodiments, system 210 comprises a data processing module 140. Data processing module 140 may receive the real-time data, from I / O module 120 or database 130. In some embodiments, data processing module 140 may perform standard data processing algorithms, such as one or more of noise reduction, signal enhancement, normalization of counts of sequence reads, correction of GC bias, etc. In some embodiments, data processing module 140 may be configured to digitally generate synthetic sample data according to embodiments of the disclosure. For example, real sample data, e.g., containing biological sample sequencing data, may be received from one or more assay panels. The totality of real sample data may be composed of both cancer data (e.g., those samples exhibiting a cancer signal above a certain threshold) and non-cancer data (e.g., those samples exhibiting a cancer signal below a certain threshold). Data processing module 140 may be configured to pair data from a real cancer sample with data from one or more non-cancer samples to generate a synthetic sample mixture that may exhibit characteristics of a real sample having a certain coverage and tumor fraction(e.g., tumor methylation fraction). In various embodiments, data processing module140 may additionally create a training data set, on which a machine-learning classifier may be trained, that constitutes both real data and synthetic data.

[0054] In some embodiments, system 210 comprises a data analysis module 150. In some embodiments, data analysis module 150 includes identifying and treating systematic errors in sequencing data, as described in connection with data processing module 140.

[0055] In some embodiments, system 210 comprises a classification module 160, which may embody a “machine-learning model” or “trained classifier.” As used herein, a “machine-learning model” or “trained classifier” generally encompasses instructions, data, and / or a model configured to receive input, and apply one or more of a weight, bias, classification, or analysis on the input to generate an output. The output may include, for example, a classification of the input, an analysis based on the input, a design, process, prediction, or recommendation associated with the input, or any other suitable type of output. A machine-learning model is generally trained using training data, e.g., experiential data and / or samples of input data, which are fed into the model in order to establish, tune, or modify one or more aspects of the model, e.g., the weights, biases, criteria for forming classifications or clusters, or the like. Aspects of a machine-learning model may operate on an input linearly, in parallel, via a network (e.g., a neural network), or via any suitable configuration. In the context of this disclosure, the machine-learning model may be trained on a combination of real and synthetic sample data.

[0056] The execution of the machine-learning model may include deployment of one or more machine-learning techniques, such as k-nearest neighbors, linear regression, logistic regression, random forest, gradient boosted machine (GBM), deep learning, a deep neural network, and / or any other suitablemachine-learning technique that solves problems in the field of Natural Language Processing (NLP). Supervised, semi-supervised, and / or unsupervised training may be employed. For example, supervised learning may include providing training data and labels corresponding to the training data, e.g., as ground truth. Unsupervised approaches may include clustering, classification or the like. K-means clustering or K-Nearest Neighbors may also be used, which may be supervised or unsupervised. Combinations of K-Nearest Neighbors and an unsupervised cluster technique may also be used. Any suitable type of training may be used, e.g., stochastic, gradient boosted, random seeded, recursive, epoch or batch-based, etc.

[0057] In an exemplary use case, a machine-learning model may be trained to analyze data from a test sample from a test subject whose status with respect to a medical condition is unknown and subsequently classifies the unknown test sample from the test subject based on the likelihood of the subject fitting into a particular category. In some embodiments, the one or more parameters may include a score (e.g., a binomial probability score that is calculated based on logistic regression analysis). As disclosed herein, the binomial probability score may correspond to the likelihood of a subject having a certain medical condition, such as cancer (e.g., NSCLC). For example, a score of over a predefined threshold may indicate that the subject associated with a test sample is more likely to have cancer than not have cancer. In some embodiments, the one or more parameters may include a sequencing data distribution pattern correlating with the presence of cancer. A subject associated with a test sample having sequencing data with a pattern resembling the cancer pattern may be diagnosed as having cancer. In some embodiments, a sequencing data distribution pattern may be identified in connectionwith a specific type of cancer, thus allowing a test sample to be classified as indicative of a certain cancer type.

[0058] In some aspects, the foregoing score may be associated with a methylation sequencing pipeline in which biological samples are collected and bisulfite conversion is implemented to prepare cfDNA. Subsequent high-throughput sequencing, data preprocessing, and methylation calling may be conducted to identify methylated and unmethylated CpG sites. Differential methylation analysis may pinpoint cancer-associated regions, and feature selection processes may extract relevant CpG sites. A trained machine learning model may be configured to analyze the relevant features and generate a score (e.g., a cancer score) that represents the likelihood of disease presence. A thresholding process may be employed to categorize samples into minimal residual disease (MRD) positive or MRD-negative categories. This integrated pipeline may help support clinical decisions by providing a quantitative MRD assessment based on methylation data.

[0059] As disclosed herein, network communication module 170 may be used to facilitate communications between a user device, one or more databases, and any other suitable system or device through a wired or wireless network connection. Any communication protocol / device may be used, including, without limitation, a modem, an Ethernet connection, a network card (wireless or wired), an infrared communication device, a wireless communication device, and / or a chipset (such as a Bluetooth™ device, an 802.11 device, a WiFi device, a WiMax device, cellular communication facilities, etc.), a near-field communication (NFC), a Zigbee communication, a radio frequency (RF) or radio-frequency identification (RFID) communication, a PLC protocol, a 3G / 4G / 5G / LTE based communication, and / or the like. For example, a user device having a user interface platform forprocessing / analyzing tumor fraction data may communicate with another user device with the same platform, a regular user device without the same platform (e.g., a regular smartphone), a remote server, a physical device of a remote loT local network, a wearable device, a user device communicably connected to a remote server, and etc.

[0060] The functional modules described herein are provided by way of example. It will be understood that different functional modules may be combined to create different utilities. It will also be understood that additional functional modules or sub-modules may be created to implement a certain utility.

[0061] Referring now to FIG. 3, an exemplary workflow 300 is provided for establishing cancer and non-cancer samples to be utilized for synthetic sample mixture generation. Aspects of the exemplary workflow 300 may be performed in accordance with some or all components described in FIG. 2A and 2B.

[0062] In the process of creating synthetic data to augment the training of disease classifiers, e.g., cancer classifiers such as NSCLC classifiers, the first step 305 is the selection and characterization of the diluent. The diluent is the non-cancer background sample used to dilute the cancer sample to create one or more synthetic mixtures that mimic a desired tumor fraction. In an aspect, the selection of the diluent may be guided by specific criteria to ensure that the resulting synthetic mixtures accurately represent the desired cancer signal range. For instance, one criteria may be the scores associated with the non-cancer sample. Preferably, in an aspect, the diluent sample should have received high non-cancer scores (or conversely low cancer scores) in prior classifier runs. This ensures that the diluent samples are representative of the non-cancer background and are suitable for creating mixtures that maintain the desired balance of cancer and non-cancer components. In anaspect, the non-cancer scores may be selected only if they exceed a predetermined threshold to be eligible for diluent consideration. For instance, samples with scores greater than the 90thpercentile of non-cancer scores may be preferred. In an aspect, in a situation where multiple samples are available for a particular subject, the highest scored non-cancer sample may be chosen for use. Another criteria may be the sample source. In some aspects, diluent samples may be derived from biological samples, such as plasma samples, but the selection process may also consider other factors, such as the source of the sample (e.g., from different participants or patients). The choice of source may impact the diversity and representativeness of the synthetic data.

[0063] At step 310, the process of generating synthetic data may additionally involve the selection of a spike-in cancer sample. The spike-in sample represents the cancer signal that will be added to the diluent to create one or more synthetic mixtures. Similar to the selection of the diluent, the selection of the spike-in sample may be guided by certain criteria to ensure that the resulting synthetic mixture accurately represents the desired cancer signal. For instance, the spike-in sample may be limited to a specific cancer type, e.g., NSCLC. This ensures that the spike-in sample contains the relevant cancer signal for the target application. More particularly, by restricting the spike-in samples to a specific cancer type, such as NSCLC, the synthetic data generated will be enriched with the methylation patterns and other characteristics unique to NSCLC. This ensures that the cancer signal extracted is specific to NSCLC. Another criteria may be the scores associated with the cancer sample. Preferably, in an aspect, the spike-in sample should have received high cancer scores in prior classifier runs. These “high score” samples are likely to exhibit methylation patterns and other biomarkers that may be indicative oflate-stage cancers, which may present a high cancer signal. Additionally, in some aspects, the spike-in sample may be identified from samples for which the cancer status has been confirmed, for example, through follow-up diagnostic workup. This ensures that the spike-in sample has a significant, reliable, and well-defined cancer signal, thereby indicating that it may be suitable for creating synthetic mixtures. Additionally or alternatively, in an aspect, the non-cancer score for the sample may need to fall below a predetermined threshold to be eligible for spike-in consideration. More particularly, establishing a threshold for non-cancer scores ensures that the selected spike-in sample has a minimal non-cancer signal. For instance, samples with non-cancer scores lower than the 10thpercentile of non-cancer scores may be preferred. In an aspect, in a situation where multiple samples are available for a particular subject, the lowest scored non-cancer sample may be chosen for use. Another criteria may be the sample source. In some aspects, the spike-in samples may be derived from plasma samples, or other biological samples, but the selection process may also consider other factors, such as the source of the sample (e.g., from different participants or patients). The choice of source may impact the diversity and representativeness of the synthetic data.

[0064] At step 315, the spike-in sample may be matched with one or more selected diluent samples. The goal in matching the samples may be to pair specific diluent and spike-in samples that, when mixed, will create one or more realistic synthetic mixtures, each having a specific targeted tumor fraction. To achieve this, several matching criteria may be considered and compared between samples, including one or more of: the type of assay that was performed on the sample to obtain the sample data, the sex, age, smoking status, race, medical history, etc., of the subject from whom the sample was derived, or one or more other factors. Anideal mixture may include a spike-in sample having a very low non-cancer score (e.g., under 10thpercentile) and a diluent sample having a very high non-cancer score (e.g., above 90thpercentile) with a majority of the aforementioned criteria matching between the two samples (e.g., both samples may be from the same subject). However, in practice, an ideal match between diluent and spike-in samples may not always be achievable. Therefore, a hierarchy of matching requirements may be defined. For instance, the highest priority may be given to matching by the assay type utilized to obtain the sample data and / or the sex of the subject from whom the sample was derived. Stated differently, in one aspect, each of the paired samples may have been processed through the same assay and may both be associated with subjects of the same sex (e.g., the spike-in and diluent sample must both be either male or female). More particularly, differences in assay type may have significant effects on the resultant sample data so it is important that the diluent and spike-in samples are associated with the same assay. These differences may include the breadth and / or depth of sequencing (e.g., whole genome vs. targeted panels, larger vs. smaller targeted panels, etc.), the chemicals used to prepare the samples, etc. Secondary criteria, such as age and smoking status, may be rank-ordered based on their importance and / or based on the type of cancer being examined. These characteristics may impact the cancer signal and its variability, so matching them where possible may improve the relevance of the synthetic mixtures and may mimic “real data” as closely as possible. For example, it may be advantageous to avoid matching a diluent sample from a female subject with a spike-in sample from a prostate cancer positive subject. As another example, it may be further advantageous to ensure that smoking status matches between the diluent sample and the spike-in sample for certain cancers (e.g., NSCLC) due to the correlationbetween smoking status and the prevalence of a specific cancer type. Other criteria, such as race, may or may not have a rank order, depending on the disease state being classified. In other aspects, only the cancer or non-cancer signal scores and assay type of the diluent and spike-in samples may be considered when selecting and matching the samples; or, only the cancer or non-cancer signal scores may be considered, and no further criteria may be considered for matching the diluent and spike-in samples at step 315.

[0065] Although step 305 is depicted as occurring before 310, step 310 may be performed before step 305, or steps 305 and 310 may be performed in parallel. In some aspects, steps 305 and 310 may be performed in parallel with step 315. For example, if multiple suitable diluent and spike-in samples are available, then the individual diluent sample may be selected based on the availability of a suitable matching spike-in sample. In some aspects, if multiple suitable diluent and spike-in samples are available, then the diluent and spike-in samples that best match with one another, e.g., share the most criteria in common, may be selected for mixing.

[0066] In regards to steps 305 - 315, in one aspect, the selection of the diluent and the cancer sample may occur dynamically, e.g., without explicit or additional user involvement. For instance, the computer system may be configured to dynamically identify (e.g., via analysis of metadata attached to the data files of each sample) a set of potential spike-in and diluent samples that may be utilized for synthetic mixture creation. The computer system may be further configured to match two or more samples together based upon a determination of matching compatibility (e.g., as dictated by digital instructions that delineate hierarchical matching requirements). Additionally or alternatively, the selection of the diluent and the spikein samples, and the matching thereof, may occur manually. For instance, a user maymanually select a specific cancer sample and a specific non-cancer sample based on information available to the user about each sample. The user may additionally pair the two samples based upon their determination of matching compatibility (e.g., as informed by the available information for each sample).

[0067] At step 320, a “p” value may be computed based on the designated targeted tumor fraction, e.g., targeted tumor methylation fraction (targeted TMeF), and the other known metrics in the cancer and non-cancer data sets. In the context of this application, the p-value refers to the fraction of DNA fragments from the spikein (cancer) sample that may be combined with the diluent (non-cancer) sample to generate a mixed, synthetic sample. Accordingly, calculation of the p-value is necessary to identify the proportion of cancer and non-cancer fragments to use to create the synthetic mixture, as further described herein.

[0068] An exemplary illustration of this combination process is depicted in FIG. 4. More particularly, FIG. 4 depicts a cancer plasma sample 405 having a first binary coverage (Ci) and a first tumor methylation fraction (TMeFi), a non-cancer plasma sample 410 having a second binary coverage (C2) and second tumor methylation fraction (TMeF2), and a synthetic mixture 415 having a targeted tumor methylation fraction (targeted TMeF). To create the synthetic mixture 415, a “p” fraction of cancer fragments from cancer sample 405 may be combined with a (1 - “p”) fraction of non-cancer fragments from non-cancer sample 410 to generate the synthetic mixture 415.

[0069] The mixture coverage may be computed by the following Equation One:Mixture Coverage = pCi + (1-p)C2

[0070] The tumor methylation fraction of the mixture may be computed by the following Equation Two:Mixture TMeF = [pCi * TMeFi] I [pCi + (1-p)C2]

[0071] Given the foregoing equations, and having known values for Ci, C2, and TMeFi (e.g., based on previously computed data for cancer sample 405 and non-cancer sample 410), a user may designate virtually any targeted TMeF they want to be attributed to the synthetic mixture. Inputting the targeted TMeF into Equation Two may thereby enable calculation of the fraction of cancer fragments (p) from any given cancer sample that need to be combined with the fraction of noncancer fragments (1-p) from a paired non-cancer sample to generate the synthetic mixture with the targeted TMeF. Because, as discussed herein, TMeF serves as a representative or corollary value for disease projection, a classifier training system or user may precisely specify different stages of progression of the resultant synthetic mixture.

[0072] As previously mentioned, in the context of this application, the synthetic “mixture” is a computational mixture of data values, a first portion of which comes from data obtained for the spike-in sample and a second portion of which comes from data obtained for the diluent sample. All of this accumulated data is contained in separate data files (e.g., stored locally on a computer server, stored remotely on another server or storage location, etc.) that may be accessed upon sample matching and mixing. For instance, a first “p” value may be represented in computer code that represents the percentage of cancer fragments to be taken from a cancer sample (e.g., the first p-value may be 0.925011350328812, which may correspond to approximately 92.5% of cancer fragments in the cancer sample). A second “p” value may also be present in the computer code that represents thepercentage of non-cancer fragments to be taken from a non-cancer sample (e.g., the second p-value may be 0.0749886496711882, which may correspond to approximately 7.5% of non-cancer fragments in the non-cancer sample). Knowing these p-values, the system may therefore access and utilize the designated proportion of cancer and non-cancer data in the relevant system files.

[0073] In an aspect, the computational mixing of data associated with different samples may generate synthetic samples that have different tumor fractions than the original real cancer sample. For example, a cancer sample may contain 10% tumor fraction (e.g., the cancer sample in this instance is composed of 9 parts non-cancer and 1 part cancer, mathematically represented as 0.9NC + 0.1C, where NC stands for ‘non-cancer1and C stands for ‘cancer’). Upon dilution of the cancer sample with 50% non-cancer DNA fragments, the resultant tumor fraction associated with the synthetic sample may be 5% (e.g., mathematically represented as 0.5(0.9NC + 0.1C) + 0.5*1 NC = 0.95NC + 0.05C). Through the foregoing process, multiple synthetic samples may be generated that each contain a different tumor fraction, thereby increasing the diversity of data that a machine learning classifier may be trained on, as further described herein.

[0074] In an aspect, the selection of the data for each specific DNA fragment from the cancer and non-cancer samples may either be random or more deterministic. For instance, with respect to the former, given a designated p-value for a cancer sample (e.g., 0.1 , which corresponds to 10%), the system may pull data associated with a randomly selected 10% of DNA fragments in the cancer sample. Given the random nature of the selection process, the specific characteristics of the synthetic mixture may change for each run, despite utilizing a consistent p-value across all runs. For instance, in a first run, the randomly selected 10% of DNAfragments may contain a majority of DNA fragments that exhibit the cancer signal, whereas in a second run, the randomly selected 10% of DNA fragments may contain a majority of DNA fragments that do not exhibit the cancer signal. With respect to the latter case, in an aspect, a user may manually designate which types of DNA fragments should be included in the proportional pool. For instance, a user may specify that each of selected DNA fragments from the cancer sample must exhibit a cancer signal for consideration to be included in the pool of DNA fragments that will be used to generate the synthetic mixture.

[0075] Referring now to FIG. 5, a table is provided that illustrates how identifying metrics for synthetic mixtures may be organized and stored in a data table. More particularly, each row in table 500 may represent a single synthetic mixture, and each column may provide identifying characteristics about that mixture. For instance, with respect to Row 1 , the cancer sample utilized in the synthetic mixture may be selected from a first subject having participant ID “ABC123,” and the non-cancer sample utilized in the synthetic mixture may be selected from a second subject having participant ID “X1 Y2Z3”. Each of these samples may be associated with the same cohort or study (e.g., “ccga”), have been processed through the same assay type and may have the same primary cancer tissue of origin indication (e.g., “lung”). Row 1 further reveals that the first subject may be a 64 year old female with stage 3 cancer. In this exemplary table, two synthetic mixtures were generated using the same cancer sample. For instance, with respect to Row 2, the cancer sample from the first subject remains the same as in the mixture represented in Row 1 , but the non-cancer sample is pulled from a third subject, i.e., having participant ID “A1 B2C3”). It is also important to note that the column designations included in table500 are exemplary and not exhaustive. More particularly, table 500 may includeadditional, or alternative, columns that provide more information about each sample and / or subject (e.g., smoking status, race, etc.).

[0076] Referring now to FIG. 6, a diagram is provided that further illustrates how data from a single cancer sample 605 may be utilized to generate a diverse array of synthetic mixture data. For instance, DNA fragments from a cancer sample 605 may be mixed with DNA fragments from a variety of different non-cancer samples 610, 615 in proportions that are dictated by various targeted TMeF values 620. For instance, six synthetic mixtures may be generated between the combination of the cancer sample 605 and the non-cancer sample 610 based on the six designated targeted TMeF values (e.g., 5e'5, 1 e-4, 2e~4, 5e'4, 1 e-3, and 2e'3). Because each of these synthetic mixtures has a different targeted TMeF value, the proportion of cancer sample that will be utilized in the mixture is different (as signified by the equations described above), despite the fact that the same samples are utilized. Furthermore, in an aspect, even if the same six targeted TMeF values were utilized to generate synthetic mixtures between the cancer sample 605 and one or more different non-cancer samples, the mixture characteristics would be different because different DNA fragments would be chosen from the cancer sample and non-cancer sample in each run, thereby producing different iterations of unique articles of synthetic data. Through this process, a multitude of unique synthetic mixtures may be generated that contribute to the volume and diversity of the training data set that a classifier may be trained on.

[0077] Table 1 below presents a subset of exemplary values for synthetic mixtures that may be formed, for instance, from the combination of cancer sample 605 and a plurality of non-cancer samples 610. As can be observed from Table 1 , the spike-in sample remains the same, while different diluent samples are utilized tocreate each synthetic mixture. The p-values represent the proportion of the real cancer sample that is utilized to create the mixture, and each p-value associated with a mixing pair is different and dependent on the targeted TMeF. For instance, the combination of spike-in ABC123 and diluent X1 Y2Z3 may generate p1 of 0.6 for a first targeted TMeF (e.g., 5e'5) and p2 of 0.4 for a second targeted TMeF (e.g., 1e-4). In an aspect, in practice, each synthetic sample may be labeled with the same participant ID or path ID as the cancer sample.Table 1

[0078] In an aspect, the mixing process may be initiated by system 210 responsive to detection of one or more predetermined events. For instance, in an aspect, system 210 may be configured to initiate the mixing process upon detection that a training data set has a data set volume below a predetermined threshold. For example, upon receiving a command to apply the training data set to a machine learning classifier, system 210 may perform a check to identify whether the training data set contains a predetermined volume of data. If it does not, then system 210 may dynamically generate additional data (e.g., system 210 may generate random synthetic samples of varying tumor fraction values, etc.) to increase the volume of training data. In another aspect, system 210 may be configured to initiate the mixingprocess upon detection that a population of samples (e.g., samples having a specific tumor fraction value or range of tumor fraction values) are underrepresented in the training data set. For instance, system 210 may identify (e.g., by analyzing metadata associated with each data sample) that the training data set may have an abundance of real sample data corresponding to non-cancer or the early and / or late stages of cancer but contains little or no amount of real data corresponding to the intermediate stages of cancer signal (e.g., in the range designated by box 6 in FIG. 1). In this circumstance, system 210 may be configured to generate synthetic data samples that have tumor fraction values corresponding to these intermediate stages. In another aspect, the mixing process may be initiated upon detection of an explicit user command to mix two or more designated samples together. In some aspects, the need for synthetic training data may be identified during or after training, in addition to or instead of prior to training. In other instances, the overall volume of training data or volume of training data having a tumor fraction falling within a certain range may be identified as low, and an indication may be generated. A user may then be able to initiate the generation of synthetic data samples in response to the indication. In still other examples, it may be assumed that training data sets may have little or no amount of real data corresponding to the intermediate ranges of cancer signal (e.g., in the range designated by box 6 in FIG. 1), and synthetic data may automatically be generated without assessing the volume of data.

[0079] Referring now to FIG. 7, a process of combining real cancer data and a plurality of real non-cancer data sets based on a designated targeted tumor fraction is disclosed. In an aspect, the original data set may be partitioned into distinct folds, with one fold designated as the validation set and the remaining folds designated as the training set. The training set may be employed to instruct theclassifier, while the validation set may assesse the classifier’s performance on unseen data. It is important to ensure that samples from the same subjects are not distributed across different folds in order to prevent data leakage (e.g., where a sample used for classifier training is also used for validation). In an aspect, a synthetic sample may be generated by combining a sample associated with a cancer-positive subject (Subject A) with a sample associated with a cancer-negative subject (Subject B). The synthetic sample, along with other samples from Subjects A and B, may be utilized to train the classifier. To mitigate the data leakage issue, when any samples from Subject A are allocated to a particular fold, the system may ensure that remaining samples from Subject A and samples from Subject B are assigned to the same fold, thereby preventing data leakage. Accordingly, referring to diagram 700 in FIG. 7, both real and synthetic samples associated with cancer participant Subject A 705 were placed in validation fold 710. In view of this occurrence, any sample from non-cancer participant Subject B 720, used to form synthetic samples associated with Subject A 705, cannot be placed in any other training fold, e.g., fold 715, but rather, will be placed in validation fold 710 with the samples from cancer participant Subject A 705.

[0080] In an aspect, the generated synthetic data may be utilized as part of a training data set that is used to train a machine-learning classifier to identify genomic patterns in a biological sample (e.g., methylation patterns, DNA mutation patterns, etc.) that may be indicative of a disease (e.g., cancer). More particularly, the training data set may contain a combination of real sample data and synthetic data. For instance, in an exemplary use case, the real sample data may include early and late stage NSCLC cancer sample data, whereas the synthetic data may represent intermediate stage NSCLC data (e.g., the type of data that may be conventionallydifficult to obtain). The inclusion of synthetic data in the training data set may help to expand the volume and diversity of sample types in the training data, which may correspondingly improve the ability of a classifier to identify indications of a disease in a sample. The improved classifier performance that may be achieved through this process may help better inform adjuvant treatment decisions. For instance, the classifier may be better able to identify MRD levels in a sample after surgery and / or treatment, a use case in which it may be particularly beneficial to be able to identify cancer signals earlier. The ability to identify MRD levels may aid a healthcare professional in determining whether additional or alternative treatment is needed and, if so, what type of treatment that may be.

[0081] In an aspect, the proportional constitution of the training data set may be varied as desired. More particularly, the training data set may be composed of specific proportions of real data from biological samples and synthetic data from samples mixed according to the embodiments described herein. For instance, in the conventional use-case, the training data set is composed of 100% real data from biological samples. To address potential deficiencies in classifier performance related to the use of only real data in the training set, and thus the potential dearth or under-representation of data from certain sample types, the training set may be modified to include some proportion of synthetic data. For example, a training data set containing 1 .25X training data may be comprised of X real NSCLC data (i.e. , 100% of the real NSCLC sample data) + 0.25X synthetic NSCLC data. The classifier may be trained on different proportions of real and synthetic data (e.g., 1 ,5X, 4X, 10X, etc.) to identify which proportional mixture generates optimal performance results for the particular use case.

[0082] Synthetic data, while being valuable in addressing data scarcity issues, may not fully capture the complexity of real-world cases (e.g., patient profiles, disease stages, treatment responses, etc.) and may introduce biases or inaccuracies into the training data set. Accordingly, in an aspect, the real sample data may be more heavily weighted and emphasized in the training data set over the synthetic data. This weighting process may ensure that real data, which may be considered more reliable and representative of actual sample signals, has a stronger influence on the classifier’s training compared to synthetic data. In an aspect, the weighting mechanism may assign weights to each data point in the training data set based on one or more factors. For instance, a primary factor influencing the weighting mechanism may be the source of the data point (e.g., real sample data may be assigned higher weights to emphasize its authenticity and clinical relevance whereas synthetic data may be assigned lower weights to acknowledge potential limitations of its generated nature). Another factor considered in the weighting decision may be the confidence in the accuracy and representativeness of each data point (e.g., real data may generally be considered to be more accurate, and real data with a verified cancer diagnosis or data points with high confidence in their accuracy may receive higher weights). In an aspect, the actual numerical values of the weight discussed above may be determined through various methods, such as manual assignment, statistical analysis, or one or more machine learning algorithms. In an aspect, the weighting mechanism may not be static and may be subject to iterative adjustments. As more real data becomes available, or as the model’s performance is assessed on real-world cases, the weighting scheme may need to be fine-tuned to optimize model performance.

[0083] Referring now to FIG. 8, a graph is provided that illustrates the performance improvement of a classifier that is trained on a mix of real and synthetic sample data. Each data point on the graph represents the detection rate per percentage of NSCLC cancer samples used in the training data set. Each sample percentage after 100% represents the amount of synthetic sample data utilized. For instance, all x-axis values up to 100 indicate that only real NSCLC sample data was utilized (e.g., x-axis value of 25 indicates that 25% of the available real samples were utilized, etc.). For x-axis values over 100, the training data set contains some proportional mix of real to synthetic data, with the proportion of synthetic data increasing from left to right. For instance, an x-axis value of 150 indicates that the training data set was composed of a 2:1 ratio of real-to-synthetic data (e.g., 100% of the real NSCLC data and 50% of the real NSCLC data as synthetic sample data). In an aspect, box 800 encompasses the detection rate values achieved when only real data is utilized to train the classifier. As can be observed, once additional synthetic data is included in the training set, the detection rate improves, e.g., from -0.77 at 100% real data usage to -0.79 at 175% data usage (i.e., 100% real data and 75% synthetic data). In an aspect, the detection rate may reach a threshold level of improvement, beyond which additional synthetic data may no longer improve the detection rate and, in some situations, may actually lower it. For instance, the detection rate appears to reach a maximum of 0.79 when the ratio of real to synthetic data is approximately 2:1 (i.e., x-axis value of 150) or 4:3 (i.e., x-axis value of 175).

[0084] Referring collectively to FIG. 9A and FIG. 9B, the improvement in model performance by using a synthetically-expanded training data set may be observed in the graphs of 9A and 9B. In an aspect, each of the graphs is focused on a specific subtype of NSCLC (e.g., graph 9A depicts detection rate results for lungadenocarcinoma, and graph 9B depicts detection rate results for squamous cell lung cancer) and includes data points that represent the detection rate per percentage of NSCLC cancer samples used in the training data set for the respective subtype. Upon examination of these graphs, similar conclusions may be drawn as with the data set portrayed by the graph of FIG.8. More particularly, the detection rate of each subtype classifier exhibits noticeable improvement after synthetic data is utilized to expand the training data set. This improvement eventually saturates in both classifiers (e.g., around 150 in both of graphs 9A and 9B).

[0085] Referring now to FIGS. 10A-10D, the improvement in model performance across different cancer stages by using a synthetically-expanded training data set may further be observed. In an aspect, the graphs of 10A-10D display the detection rate per percentage of NSCLC cancer samples used in the training data set at each stage of cancer progression. The graph of 10A depicts cancer detection rates for cancer samples from subjects having stage I cancer, the graph of 10B depicts cancer detection rates for cancer samples from subjects having stage II cancer, the graph of 10C depicts cancer detection rates for cancer samples from subjects having stage III cancer, and the graph of 10D depicts cancer detection rates for cancer samples from subjects having stage IV cancer. Across all stages, I- IV, it can be observed that the utilization of additional synthetic mixture data in the training data set improves the detection rate of the classifier compared to those classifiers just trained on the real data. This improvement is more apparent during stage 1 of the cancer, where the detection rate increases from approximately 0.23 to 0.27 once synthetic samples begin to be utilized. Accordingly, embodiments of this disclosure may be particularly useful for improving the detection of early stage diseases.

[0086] Similar improvements to the detection rate can be observed in theNSCLC subtypes across the different cancer stages, as illustrated in the graphs inFIGS. 11A - 11 D. In an aspect, the graphs of 11 A and 11 B display the detection rate per percentage of NSCLC cancer samples used in the training data set at each stage of cancer progression (stages l-IV) for the lung adenocarcinoma subtype. Across stages I, III, and IV (stage II being the exception in which the detection rate remains stagnant upon the introduction of synthetic data to the training data set), it can be observed that the utilization of additional synthetic mixture data in the training data set improves the detection rate of the classifier compared to those classifiers just trained on real data. In another aspect, the graphs of 11C and 11 D display the detection rate per percentage of NSCLC cancer samples used in the training data set at each stage of cancer progression (stages l-IV) for the squamous cell lung cancer subtype. Across stages I and II, it can clearly be observed that the utilization of additional synthetic mixture data in the training data set improves the detection rate of the classifier compared to those classifiers just trained on real data. Across stages III and IV, however, the addition of synthetic mixture data to the training data set produces only marginal improvements, if any, to detection rate.

[0087] Referring now to FIG. 12, graphs 1205, 1210 are provided that illustrate the improvement in the 50% endpoint Limit of Detection (LODso) in two different NSCLC subtypes, adenocarcinoma (associated with graph 1205) and squamous cell lung cancer (associated with graph 1210). In an aspect, each data point in the graphs 1205, 1210 represents the lowest level tumor methylation fraction that could be detected per percentage of NSCLC cancer samples used in the training data set. Similar to the graphs depicted in FIGs. 9A - 12, each percentage of NSCLC cancer samples utilized after 100% indicates that at least some syntheticdata was utilized in the training data set. It can be observed that, for each subtype, the LOD is improved for classifiers trained on at least some proportion of synthetic mixture training data.

[0088] There may be various additional applications for the concepts described herein. For instance, in an aspect, by creating large synthetic cohorts of data, researchers may evaluate the sensitivity and specificity of diagnostic tests and biomarkers at specific LOD thresholds. This may aid in improving the accuracy of cancer detection. In another aspect, further feature optimization may be conducted to identify the optimal number of relevant features to utilize in the training process to maximize classifier performance. Those meaningful features, which may not be identified when the real samples are limited, may be helpful in detecting the cancer when TMeF is low. In another aspect, further research may be conducted into the impact that the synthetic mixtures may have on the sequencing depth in real samples. For instance, the generation and inclusion of synthetic data in the training data set may preclude the conventional need to sequencing the real samples at higher depth. In another aspect, more synthetic data may be gathered within a cancer patient so that researchers may obtain a more accurate estimation of LOD within a subject, which may correspondingly lead to a more accurate estimation of LOD of each cancer type.

[0089] Referring now to FIG. 13, an exemplary workflow for generating synthetic data is disclosed. The exemplary workflow may be performed by components of computer system 210 (shown in FIG. 2B).

[0090] At step 1305, computer system 210 may receive an indication to generate a synthetic data set. In an aspect, the synthetic data set may be a combination of at least a portion of a first data set associated with a real cancersample and at least a portion of a second data set associated with a real non-cancer sample. The first and second data sets may contain DNA data associated with the cancer and non-cancer samples, which may have previously been derived from the performance of a DNA sequencing process on each sample type. Although DNA data is referenced herein, any suitable nucleic acid data may be used.

[0091] In an aspect, the indication may take on one or more different forms.For instance, in one aspect, the indication may be an explicit user command, provided to a component of computer system 210, to generate a synthetic data set. In another aspect, the indication may be resultant from a dynamic, computer-realized determination. For instance, computer system 210 may identify that a training data set contains a volume of training data that is below a threshold volume and may therefore generate an instruction to generate synthetic data to expand the volume of the training set. In another similar aspect, computer system 210 may identify that a training data set lacks a diversity of training data and may generate an instruction to generate a variety of different types of synthetic data to expand the volume diversity.

[0092] In some aspects, the indication may specify a target tumor fraction that the synthetic data set should represent. For instance, an explicit user command to generate synthetic data may contain a designation of the exact tumor fraction that the data composition of the synthetic data set will exhibit. In another aspect, the tumor fraction of the synthetic data set may be dynamically determined by computer system 210. For instance, responsive to identifying that a training data set lacks samples having a particular type of tumor fraction, system 210 may be configured to generate synthetic data sets that exhibit the tumor fraction that is lacking in the training data.

[0093] At step 1310, computer system 210 may identify which data sets from a real cancer sample and a real non-cancer sample to utilize to construct the synthetic data set. More particularly, computer system 210 may have access to data associated with a plurality of real cancer samples and real non-cancer samples and may be configured to identify which of these data sets are compatible to generate a synthetic data set. In an aspect, the selection of both the cancer and non-cancer samples may be guided by certain criteria. One criteria may be the classifier scores previously assigned to each sample. For instance, for selection of the non-cancer sample, only those samples having a score greater than a predetermined threshold non-cancer score (e.g., 90thpercentile non-cancer score, etc.) may be eligible for synthetic data generation. Similarly, for selection of the cancer sample, only those samples having a score lower than a predetermined threshold non-cancer score (e.g., 10thpercentile non-cancer score, etc.) may be eligible for synthetic data generation.

[0094] In an aspect, computer system 210 may consider a hierarchical list of other matching criteria when selecting the data sets associated with the cancer and non-cancer sample. For instance, computer system 210 may select those eligible cancer and non-cancer samples (e.g., those samples having the requisite non- cancer scores) that have both been processed through the same assay and may both be associated with subjects of the same sex. If computer system 210 determines that no sample pairs match the foregoing criteria (e.g., no real cancer sample shares the same processing assay and subject sex as another non-cancer sample, etc.), then computer system 210 may examine whether the eligible cancer and non-cancer samples share one or more secondary criteria, such as age andsmoking status, to determine whether a match can be made, or a cancer and a noncancer sample may simply be chosen, even if they do not meet the criteria.

[0095] At step 1315, computer system 210 may determine a first proportion of a first data set (i.e. , associated with a selected real cancer sample) and a second proportion of a second data set (i.e., associated with a selected real non-cancer sample) to utilize in the synthetic mixture. In an aspect, the first and second proportions of the first and second data sets may be dictated by the designated targeted tumor fraction associated with the synthetic mixture. More particularly, a p- value, which is representative of the fraction of DNA fragments from the real cancer sample that may be combined with the real non-cancer sample to generate the synthetic mixture, may be algorithmically related to the designated tumor fraction, as previously discussed above in connection with FIGS. 4 and 6. The identification of the p-value may correspondingly enable computer system 210 to designate the first proportion of the first data set and the second proportion of the second data set to utilize in the generation of the synthetic data set.

[0096] At step 1320, upon determining the proportions with which the first data set and the second data set should be mixed, computer system 210 may select a first subset of the first data set that corresponds to the first mixing proportion and a second subset of the second data set that corresponds to the second mixing proportion. In an aspect, the selection of the subset in each of the first and second data sets may be substantially random. For instance, if the targeted tumor fraction dictates that 10% of the first data set should be mixed with 90% of the second data set, then computer system 210 may randomly select any 10% of the first data set to mix with any 10% of the second data set. Alternatively, in another aspect, the selection of the subsets may be more intentional. For example, assuming the samemixing proportion as delineated in the previous example (10% cancer and 90% noncancer), computer system 210 may select only the 10% of data in the first data set that is associated with a positive cancer signal.

[0097] At step 1325, computer system 210 may combine the first subset of the first data set with the second subset of the second data set to form a synthetic data mixture. This combination process may be performed multiple times to generate a plurality of synthetic data mixtures having varying tumor fraction representations. These articles of synthetic data may be utilized to expand the volume and / or diversity of a training data set, which may correspondingly be used to train a machine-learning classifier. In an aspect, the ratio of real data to synthetic data in the training data set may be based on preconfigured settings. In an aspect, real data may be weighted more heavily than synthetic data in the training data set in certain situations, e.g., in order to avoid or reduce the introduction of inaccuracies or biases stemming from the synthetic data.Additional Considerations / . TMeF Range

[0098] In an aspect, during the generation of synthetic data for training machine learning classifiers, it may be desirable to exclude synthetic samples with TMeFs significantly below the LoD to maintain empirical specificity levels. More particularly, it was observed that including synthetic samples with TMeFs far below the LoD reduced empirical specificity by more than 0.5%, which may be significant for maintaining reliable classifier performance. Specifically, synthetic cancer samples with very low TMeFs resemble non-cancer samples in their characteristics. Including such samples in the positive class during training may introduce ambiguity, complicating the classification task and threshold-setting process for the machinelearning model. To address the foregoing, a lower bound on the TMeFs of synthetic cancer samples may be set to ensure that the synthetic data remains sufficiently distinct from non-cancer samples. This calibration ensures that synthetic data complements real data effectively, maintaining the model’s ability to achieve its target specificity.

[0099] Further to the foregoing, Table 2 below represents an experimental setup for evaluating NSCLC classifier performance by varying the TMeF ranges of synthetic cancer samples added to the training data and adjusting the ratio of synthetic to real cancer samples. Table 2 categorizes models based on whether the synthetic data has unrestricted TMeFs, TMeFs below the LoD, TMeFs near the LoD, or TMeFs above the LoD, thereby enabling analysis of how these factors impact classifier accuracy and specificity.Table 2

[0100] Referring now to FIG. 14, graph 1400 illustrates specificity levels across the different cross-validation models presented above in Table 2. The models were trained with a target specificity of 97%. The dashed line across the graph at a specificity of 0.970 indicates the target specificity. Examination of graph 1500 reveals a clear drop in specificity when synthetic samples with TMeFs far below the LoD are included. Under both training sample sizes, the models that only include synthetic samples significantly below LoD have empirical specificity levels that are greater than 0.5% lower than expected. Under both training sample sizes, models with synthetic data that included TMeFs near the LoD, or TMeFs above the LoD, tended to have higher specificities closer to the target specificity. / / . Sample Weighting

[0101] In some aspects, an issue was identified during classifier training when integrating synthetic samples. In particular, in one implementation a sample weighting mechanism assigned weights based on a parameter intended to distribute representation across samples uniformly. In this disclosure, this parameter will be referred to as a “shuffler key.” It was observed that when synthetic samples were introduced, this approach caused multiple participants to share the same shuffler key, leading to unintended consequences, such as the under-representation of noncancer participants used as diluents. This misrepresentation occurred because the shuffler key failed to account for participant-specific weighting, resulting in incorrect weight distribution, especially when non-cancer samples with low cancer signals were intentionally selected for synthetic mixtures as spike-ins. In an aspect, to address this issue, a sample weighting strategy may implement two corrective measures. The first is that only real samples may be used during a thresholding step to eliminate the influence of synthetic samples in determining binary cutoff values forcancer scores. Second, an individual sample’s weight may be calculated based on the number of samples included from a given participant. The samples may be weighted as 1 / (the number of samples for a given participant), such that the total weight assigned to a participant’s samples sums to one. This change in the weighting of samples promotes equal representation of all non-cancer participants in the thresholding process, regardless of the number of samples each contributed. As a result, the actual classifier specificity may properly represent the targeted specificity in the study population.

[0102] Further to the foregoing, and referring now to FIG. 15, diagram 1500 presents an exemplary scenario that illustrates how the inclusion of synthetic mixtures affects the weighting of real samples in the thresholding step of classifier training, leading to under-representation of certain non-cancer samples. In diagram 1500, synthetic data may be generated by mixing a cancer sample (e.g., spike-in, sample 01) with non-cancer samples (diluents, samples 02-07). Without synthetic mixtures, each non-cancer sample in the thresholding process would receive a weight of 1 , because only real samples are included. However, with the addition of synthetic mixtures, the same non-cancer samples share their weight with synthetic counterparts, thereby reducing their weight to 1 / 6 in the thresholding step. This under-representation occurs because the synthetic data generation process relies on specific non-cancer samples with correct predictions and low cancer signals to act as diluents. These selected non-cancer samples are repeatedly used across multiple synthetic mixtures, inadvertently skewing the weights assigned to them during the thresholding step. As a result, the classifier’s ability to maintain accurate specificity is comprised, as the representation of non-cancer samples in thresholding no longer reflects the intended population distribution.

[0103] Diagram 1600 in FIG. 16 further highlights how including synthetic mixtures may affect the weighting of non-cancer samples during the specificity determination phase of machine learning classifier training. More particularly, the issue may arise because all non-cancer samples used as diluents (combined with cancer samples to create synthetic mixtures) share the same groupJD (i.e. , shuffler key). Specifically, when inverse sample weights are calculated based on the shuffler key, the weight assigned to each non-cancer sample is 1 / (m+n+j), where m, n, and j represent the number of samples from different non-cancer participants (e.g., participants B, C, and D). This method aggregates all diluents into a single group, leading to an underrepresentation of individual non-cancer participants in the overall weighting process. For instance, if participant B has m samples, the weight for each of those samples is reduced proportionally because the total pool of samples includes contributions from participants C and D as well. In an aspect, in an alternative approach, weights may be assigned based on the participant ID. Here, the weight for samples from each participant is calculated as 1 / m, 1 / n, or 1 / y, such that each participant’s contributions are equally represented regardless of the number of other participants in the group. This method avoids diluting the influence of any single participant’s samples during the specificity determination process.

[0104] Referring collectively to FIGS. 17 and 18, graphs 1700 in FIG. 17 and 1800 in FIG. 18 graphically illustrate the differences in the weighting approaches described above. More particularly, graph 1700 illustrates the empirical specificity of a cancer classifier trained using different data configurations, specifically highlighting the impact of including synthetic samples on specificity. Graph 1700 compares three models: the base model, a candidate model trained with real samples, and a candidate model trained with synthetic samples. The horizontal dashed line indicatesthe target specificity of 95%, which the classifiers aim to achieve. For example, if target specificity is 95%, then the final cutoff of cancer score would make 95% of non-cancer samples have a cancer score of less than or equal to the binary cutoff. Graph 1700 shows that the base and candidate models trained with real samples generally achieve or exceed the target specificity across both v1 and v2 assays. However, the candidate model trained with synthetic samples exhibits a slight deviation from the target specificity in some cases. This discrepancy was not expected, because the synthetic samples were not included in the thresholding step, where the binary cutoff for cancer scores is determined to meet the target specificity. Graph 1800 illustrates the improvement in the empirical specificity of a classifier after implementing updated, and more accurate, sample weighting strategies during the thresholding step. More particularly, after updating the weighted mechanism to ensure that each participant is equally represented, the model’s specificity becomes relatively more closely aligned with the target. This adjustment illustrates that the integration of synthetic data does not disrupt the balance of real sample representation, and graph 1800 demonstrates that the updated approach effectively resolves the specificity discrepancies introduced by synthetic data, enabling the classifier to meet its performance objectives.

[0105] Referring now to FIG. 19, graph 1900 further illustrates how the implementation of a new weighting method addresses the underrepresentation issue for non-cancer samples used as diluents during the specificity determination process in disease classifier training. More particularly, graph 1900 compares sample weights calculated based on two strategies: using the groupJD (shuffler key) approach versus using the participant D approach. Examination of graph 1900 reveals that, when using the groupJD approach, weights are distributed across allsamples within the same group, which can result in lower weights for diluents. This occurs because multiple non-cancer samples from different participants are pooled under the same group identifier. The participantJD approach addresses this issue by assigning weights based on the individual participant, thereby giving higher weights to their samples compared to the groupJD approach.

[0106] In general, any process discussed in this disclosure that is understood to be computer-implementable may be performed by one or more processors of a computer system, such as system environment 210, as described above. A process or process step performed by one or more processors may also be referred to as an operation. The one or more processors may be configured to perform such processes by having access to instructions (e.g., software or computer-readable code) that, when executed by the one or more processors, cause the one or more processors to perform the processes. The instructions may be stored in a memory of the computer server. A processor may be a central processing unit (CPU), a graphics processing unit (GPU), or any suitable types of processing unit.

[0107] A computer system, such as system environment 210, may include one or more computing devices. If the one or more processors of the computer system are implemented as a plurality of processors, the plurality of processors may be included in a single computing device or distributed among a plurality of computing devices. If a system environment comprises a plurality of computing devices, the memory of the computer system may include the respective memory of each computing device of the plurality of computing devices.

[0108] FIG. 20 is a simplified functional block diagram of a computer system2000 that may be configured as a computing device for executing the processesdescribed herein, according to exemplary embodiments of the present disclosure. FIG. 20 is a simplified functional block diagram of a computer that may be configured according to exemplary embodiments of the present disclosure. In various embodiments, any of the systems herein may be an assembly of hardware including, for example, a data communication interface 2020 for packet data communication. The platform also may include a central processing unit (“CPU”) 2002, in the form of one or more processors, for executing program instructions. The platform may include an internal communication bus 2008, and a storage unit 2006 (such as ROM, HDD, SDD, etc.) that may store data on a computer readable medium 2022, although the system 2000 may receive programming and data via network communications via electronic network 2025 (e.g., voice, video, audio, images, or any other data over the electronic network 2025). The system 2000 may also have a memory 2004 (such as RAM) storing instructions 2024 for executing techniques presented herein, although the instructions 2024 may be stored temporarily or permanently within other modules of system 2000 (e.g., processor 2002 and / or computer readable medium 2022). The system 2000 also may include input and output ports 2012 and / or a display 2010 to connect with input and output devices such as keyboards, mice, touchscreens, monitors, displays, etc. The various system functions may be implemented in a distributed fashion on a number of similar platforms, to distribute the processing load. Alternatively, the systems may be implemented by appropriate programming of one computer hardware platform.

[0109] In this disclosure, the term “based on” means “based at least in part on.” The singular forms “a,” “an,” and “the” include plural referents unless the context dictates otherwise. The term “exemplary” is used in the sense of “example” rather than “ideal.” The terms “comprises,” “comprising,” “includes,” “including,” orother variations thereof, are intended to cover a non-exclusive inclusion such that a process, method, or product that comprises a list of elements does not necessarily include only those elements, but may include other elements not expressly listed or inherent to such a process, method, article, or apparatus. Relative terms, such as “about,” “approximately,” “substantially,” and “generally,” are used to indicate a possible variation of ±10% of a stated or understood value. In addition, the term “between” used in describing ranges of values is intended to include the minimum and maximum values described herein. The use of the term “or” in the claims and specification is used to mean “and / or” unless explicitly indicated to refer to alternatives only if the alternatives are mutually exclusive, although the disclosure supports a definition that refers to only alternatives and “and / or.” As used herein “another” may mean at least a second or more.

[0110] As used herein, the term “user” generally encompasses any person or entity, such as a researcher and / or a care provider (e.g., a doctor, etc.), that may desire information, resolution of an issue, or engage in any other type of interaction with a provider of the systems and methods described herein (e.g., via an application interface resident on their electronic device, etc.). The term “electronic application” or “application” may be used interchangeably with other terms like “program,” or the like, and generally encompasses software that is configured to interact with, modify, override, supplement, or operate in conjunction with other software.

[0111] Program aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of executable code and / or associated data that is carried on or embodied in a type of machine-readable medium. “Storage” type media include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as varioussemiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer of the mobile communication network into the computer platform of a server and / or from a server to the mobile device. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links, or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.

[0112] Furthermore, while some embodiments described herein include some but not other features included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the invention, and form different embodiments, as would be understood by those skilled in the art. For example, in the following claims, any of the claimed embodiments can be used in any combination.

[0113] Thus, while certain embodiments have been described, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the invention, and it is intended to claim all such changes and modifications as falling within the scope of the invention. For example,functionality may be added or deleted from the block diagrams and operations may be interchanged among functional blocks. Steps may be added or deleted to methods described within the scope of the present invention.

[0114] The above disclosed subject matter is to be considered illustrative, and not restrictive, and the appended claims are intended to cover all such modifications, enhancements, and other implementations, which fall within the true spirit and scope of the present disclosure. Thus, to the maximum extent allowed by law, the scope of the present disclosure is to be determined by the broadest permissible interpretation of the following claims and their equivalents, and shall not be restricted or limited by the foregoing detailed description. While various implementations of the disclosure have been described, it will be apparent to those of ordinary skill in the art that many more implementations are possible within the scope of the disclosure. Accordingly, the disclosure is not to be restricted except in light of the attached claims and their equivalents.

Claims

WHAT IS CLAIMED IS:

1. A computer-implemented method, the computer-implemented method comprising: receiving, at a computing device, an indication to generate a synthetic data set having characteristics associated with a designated tumor methylation fraction; identifying, based on the received indication and using a processor of the computing device, a first data set associated with a real cancer sample and a second data set associated with a real non-cancer sample; determining, based on the designated tumor methylation fraction and using the processor, a first proportion of the first data set to combine with a second proportion of the second data set; selecting, based on the determining and using the processor, a first subset of the first data set that corresponds to the first proportion and a second subset of the second data set that corresponds to the second proportion; and generating, using the processor, at least one synthetic data sample satisfying the designated tumor methylation fraction in the synthetic data set by combining the first subset of the first data set with the second subset of the second data set.

2. The computer-implemented method of claim 1 , wherein the real cancer sample is a non-small cell lung cancer (NSCLC) sample.

3. The computer-implemented method of claim 1 , wherein the real cancer sample is one of: a tissue sample, a urine sample, and a blood sample.

4. The computer-implemented method of claim 1 , wherein the identifying the first data set and the second data set comprises: determining, using the processor and via analyzing metadata associated with a data set pool containing the first data set and the second data set, that the first data set and the second data set satisfy one or more selection criteria; and selecting, responsive to the determining, the first data set and the second data set.

5. The computer-implemented method of claim 4, wherein the selection criteria are selected based on a disease type associated with the real cancer sample.

6. The computer implemented method of claim 5, wherein the disease type is selected from a group comprising: breast cancer, lung cancer, prostate cancer, colorectal cancer, renal cancer, uterine cancer, pancreatic cancer, cancer of the esophagus, a lymphoma, head / neck cancer, ovarian cancer, a hepatobiliary cancer, a melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, and gastric cancer.

7. The computer-implemented method of claim 4, wherein the one or more selection criteria for the first data set comprise a non-cancer score below a first predetermined threshold, and wherein the one or more selection criteria for the second data set comprise the non-cancer score above a second predetermined threshold.

8. The computer-implemented method of claim 4, wherein the determining that the first data set and the second data set satisfy the one or more selection criteria comprises: identifying that the real cancer sample and the real non-cancer sample share at least one selection criteria selected from the group of: assay type, subject sex, subject age, subject smoking status, subject race, and subject medical history.

9. The computer-implemented method of claim 1 , further comprising: assembling, using the processor, a training data set that contains a mix of (i) a real sample data set comprising the first data set and the second data set and (ii) the synthetic data set.

10. The computer-implemented method of claim 9, wherein a ratio of the synthetic data set to the real sample data set in the mix is less than 4 to 1 .11 . The computer-implemented method of claim 9, further comprising: weighting the real sample data set more than the synthetic data set in the training data set.

12. The computer-implemented method of claim 9, further comprising: repeating the steps of the method to generate a plurality of synthetic data samples in the synthetic data set by combining portions of the first data set with portions of the second data set.

13. The computer-implemented method of claim 9, further comprising training, using the training data set, a machine-learning classifier to identify methylation patterns in a biological sample.

14. The computer-implemented method of claim 13, wherein the machinelearning classifier is further configured to generate, based on the identified methylation patterns, a disease state score.

15. The computer-implemented method of claim 14, further comprising determining an empirical sensitivity for the machine-learning classifier by: weighting each synthetic data sample in the training data set such that each contribution of the second portion of the second data set are equally represented regardless of the number contributions in the group.

16. The computer-implemented method of claim 14, further comprising: receiving a test sample associated with a test subject; applying the test sample to the trained machine-learning classifier; receiving, from the trained machine-learning classifier, the disease state score; and categorizing, upon comparing the disease state score to a predetermined threshold, the test sample as being minimal residual disease (MRD) positive or MRD negative.

17. The computer-implemented method of claim 9, wherein the receiving, at the computing device, the indication to generate the synthetic data set is in responseto identifying that a real sample data set in a training data set contains less than a predetermined volume of real data corresponding to a predetermined cancer stage.

18. A system, the system comprising: one or more processors; one or more computer readable media storing instructions that are executable by the one or more processors to perform operations to: receive, at a computing device associated with the system, an indication to generate a synthetic data set having characteristics associated with a designated tumor methylation fraction; identify, based on the received indication, a first data set associated with a real cancer sample and a second data set associated with a real noncancer sample; determine, based on the designated tumor methylation fraction, a first proportion of the first data set to combine with a second proportion of the second data set; select, based on the determining, a first subset of the first data set that corresponds to the first proportion and a second subset of the second data set that corresponds to the second proportion; and generate, using the one or more processors, at least one synthetic data sample satisfying the designated tumor methylation fraction in the synthetic data set by combining the first subset of the first data set with the second subset of the second data set.

19. The system of claim 18, wherein the real cancer sample is a non-small cell lung cancer (NSCLC) sample.

20. The system of claim 18, wherein the real cancer sample is one of: a tissue sample, a urine sample, and a blood sample.21 . The system of claim 18, wherein the operations to identify the first data set and the second data set further comprise operations to: determine, via analyzing metadata associated with a data set pool containing the first data set and the second data set, that the first data set and the second data set satisfy one or more selection criteria; and select, responsive to the determining, the first data set and the second data set.

22. The system of claim 21 , wherein the selection criteria are selected based on a disease type associated with the real cancer sample.

23. The system of claim 22, wherein the disease type is selected from a group comprising: breast cancer, lung cancer, prostate cancer, colorectal cancer, renal cancer, uterine cancer, pancreatic cancer, cancer of the esophagus, a lymphoma, head / neck cancer, ovarian cancer, a hepatobiliary cancer, a melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, and gastric cancer.

24. The system of claim 21 , wherein the one or more selection criteria for the first data set comprise a non-cancer score below a first predetermined threshold and wherein the one or more selection criteria for selecting the second data set comprise the non-cancer score above a second predetermined threshold.

25. The system of claim 21 , wherein the operations to determine that the first data set and the second data set satisfy the one or more selection criteria comprise: identify that the real cancer sample and the real non-cancer sample share at least one selection criteria selected from the group of: assay type, subject sex, subject age, subject smoking status, subject race, and subject medical history.

26. The system of claim 18, wherein the operations further comprise operations to: assemble a training data set that contains a mix of (i) a real sample data set comprising the first data set and the second data set and (ii) the synthetic data set.

27. The system of claim 26, wherein a ratio of the synthetic data set to the real sample data set in the mix is less than 4 to 1 .

28. The system of claim 26, wherein the operations further comprise operations to: weight the real sample data set more than the synthetic data set in the training data set.

29. The system of claim 26, wherein the operations further comprise operations to: repeat the steps in the instructions to generate a plurality of synthetic data samples in the synthetic data set by combining portions of the first data set with portions of the second data set.

30. The system of claim 26, wherein the operations further comprise operations to: train, using the training set, a machine-learning classifier to identify methylation patterns in a biological sample.31 . The system of claim 30, wherein the machine-learning classifier is further configured to generate, based on the identified methylation patterns, a disease state score.

32. The system of claim 31 , wherein the operations further comprise operations to determine an empirical sensitivity for the machine-learning classifier by: weighting each synthetic data sample in the training data set such that each contribution of the second portion of the second data set are equally represented regardless of the number contributions in the group.

33. The system of claim 31 , wherein the operations further comprise operations to: receive a test sample associated with a test subject;apply the test sample to the trained machine-learning classifier; receive, from the trained machine-learning classifier, the disease state score; and categorize, upon comparing the disease state score to a predetermined threshold, the test sample as being minimal residual disease (MRD) positive or MRD negative.

34. The system of claim 26, wherein the operations to receive the indication to generate the synthetic data set is in response to operations that identify that a real sample data set in a training data set contains less than a predetermined volume of real data corresponding to a predetermined cancer stage.

35. A non-transitory computer-readable medium storing computer-executable instructions which, when executed by a system, cause the system to perform operations comprising: receiving, at a computing device, an indication to generate a synthetic data set having characteristics associated with a designated tumor methylation fraction; identifying, based on the received indication and using a processor of the computing device, a first data set associated with a real cancer sample and a second data set associated with a real non-cancer sample; determining, based on the designated tumor methylation fraction and using the processor, a first proportion of the first data set to combine with a second proportion of the second data set;selecting, based on the determining and using the processor, a first subset of the first data set that corresponds to the first proportion and a second subset of the second data set that corresponds to the second proportion; and generating, using the processor, at least one synthetic data sample satisfying the designated tumor methylation fraction in the synthetic data set by combining the first subset of the first data set with the second subset of the second data set.

Citation Information

Patent Citations

  • Tumor fraction estimation using methylation variants

    US20230272486A1

  • Cancer classification with synthetic spiked-in training samples

    WO2021202424A1