Generative Adversarial Networks for Urinary Biomarkers
Patent Information
- Application Number
- JP2024532725
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-11-30
- Filing Date
- 2022-11-23
- Publication Date
- 2025-09-17
AI Technical Summary
Machine learning algorithms face challenges when dealing with small, imbalanced datasets, leading to biased performance towards the majority class and poor results in applications like credit card fraud detection, spam detection, and medical diagnostics.
A Generative Adversarial Network (GAN) is employed to create synthetic features for biomedical datasets, specifically using a Conditional Tabular Generative Adversarial Network (CTGAN) to balance imbalanced data sets by generating synthetic features for machine learning systems.
The GAN-based approach enhances the performance of machine learning models by reducing class imbalance, improving accuracy and robustness in classifying medical samples, particularly in kidney transplant rejection prediction.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] cross reference This application claims priority to U.S. Provisional Application No. 63 / 284,590, filed November 30, 2021, the disclosure of which is incorporated herein by reference in its entirety.
[0002] The present invention relates generally to a methodology for balancing imbalanced biological data sets. [Background technology]
[0003] Several state-of-the-art artificial intelligence applications have a challenging and long-standing unsolved problem of dealing with small imbalanced datasets in their implementations. The class imbalance problem arises when there is an uneven number of samples for all classes present in a dataset, which can result in machine learning algorithms favoring bias towards the majority class while producing poor performance for the minority class. This is a common problem that affects many real-world applications such as credit card fraud detection, spam detection, cham prediction, medical diagnosis, dense object detection, among others. There is a pressing need for techniques that can address the biases introduced into machine learning systems trained with small imbalanced datasets. Summary of the Invention
[0004] In the following discussion, certain articles and methods are described to provide background and introduction. Nothing contained herein should be construed as an "admission" of prior art. Applicants expressly reserve the right, if necessary, to demonstrate that the articles and methods referenced herein do not constitute prior art under any provision of applicable law.
[0005] Disclosed herein is a system and use of Generative Adversarial Network (GAN) based data augmentation methods to create synthetic features, particularly in scenarios involving small and imbalanced biomedical datasets for machine learning systems, such as complex multivariate analyses of biomarkers from urine samples.
[0006] In some aspects, the disclosure provides a system configured to balance imbalanced datasets obtained from biological samples, the system comprising: one or more computer subsystems; and one or more components executed by the one or more computer subsystems, the one or more components including a generative adversarial network trained on a first training set including data corresponding to amounts of cell-free DNA (cfDNA) biomarkers from subjects with organ dysfunction designated as a first training input; and a second training set including data corresponding to amounts of cell-free DNA (cfDNA) biomarkers from subjects without organ dysfunction designated as a second training input, the first and second datasets being imbalanced; and the one or more computer subsystems configured to generate a set of synthetic features for the first dataset and / or the second dataset by inputting a portion of the data from the first training input and the second training input into the generative adversarial network.
[0007] In some cases, the generative adversarial network is configured as a conditional generative adversarial network, a vanilla generative adversarial network, a table generative adversarial network, or a tabular generative adversarial network. In some cases, the generative adversarial network is further trained with an additional training set including data corresponding to the amount of methylated cfDNA (m-cfDNA) biomarkers from subjects with organ dysfunction, specified as additional training inputs, and an additional training set including data corresponding to the amount of methylated cfDNA (m-cfDNA) biomarkers from subjects without organ dysfunction, specified as additional training inputs.
[0008] In some instances, the generative adversarial network is further trained with an additional training set including data corresponding to the amount of an inflammatory biomarker from subjects with organ dysfunction, specified as additional training inputs, and an additional training set including data corresponding to the amount of an inflammatory biomarker from subjects without organ dysfunction, specified as additional training inputs. The inflammatory biomarker can be a member of the chemokine (CXC motif) ligand family, such as CXC motif chemokine ligand 1 (CXCL1), CXC motif chemokine ligand 2 (CXCL2), CXC motif chemokine ligand 5 (CXCL5), CXC motif chemokine ligand 9 (CXCL9) (MIG), or CXC motif chemokine ligand 10 (CXCL10) (IP-10).
[0009] In some cases, the generative adversarial network is further trained with an additional training set including data corresponding to the amount of an apoptotic biomarker from subjects with organ dysfunction, specified as an additional training input, and an additional training set including data corresponding to the amount of an apoptotic biomarker from subjects without organ dysfunction, specified as an additional training input, in some cases, the apoptotic biomarker is clusterin.
[0010] In some cases, the generative adversarial network is further trained with an additional training set including data corresponding to the amount of a protein from subjects with organ dysfunction, specified as additional training inputs, and an additional training set including data corresponding to the amount of a protein from subjects without organ dysfunction, specified as additional training inputs. In some cases, the protein is albumin, but the protein could also be total protein.
[0011] In some aspects, the one or more computer subsystems are further configured to determine one or more characteristics of the synthetic features for the first data set and / or the second data set. In other aspects, the one or more computer subsystems are further configured to train a machine learning model using the simulated images. Such a machine learning model may be trained on the first data input, the second data input, or any number of data inputs. In some cases, the machine learning model is trained on the first data input and the second data input, rather than on the set of synthetic features. In some cases, the machine learning model is CTGAN, SMOTE, SVM-SMOTE, ADASYN.
[0012] In some cases, the biological sample is urine, but the biological sample may also be blood, bronchiolar lavage, or another suitable bodily fluid. In some cases, the organ is an allograft and the injury is caused by rejection of the allograft by the subject. In some cases, the organ is a kidney, pancreas, heart, lung, or liver. In some cases, the organ is a kidney. In some cases, the injury is chronic kidney injury (CKI) or acute kidney injury (AKI). In some cases, the injury is caused by a viral infection suffered by the subject, such as a viral infection caused by Sars-CoV-2, CMV, or BKV. In some cases, the injury is a cancer that harms an organ, such as bladder cancer or kidney cancer. In some cases, the subject is a human.
[0013] In some aspects, the present disclosure provides a system configured to analyze a dataset obtained from a biological sample, the system comprising one or more computer subsystems and one or more components executed by the one or more computer subsystems, the one or more components including a generative adversarial network trained with a training set corresponding to an amount of cfDNA from a subject, the one or more computer subsystems configured to generate a synthetic dataset from the biological sample by inputting a subset of the training data into the generative adversarial network. In some cases, at least one subset of the training data is annotated with a biological state, such as an acute rejection biological state, a chronic kidney injury (CKI) biological state, an acute kidney injury (AKI), a COVID-19 biological state, or a healthy or stable biological state. In some cases, the cfDNA is from a urine sample. In other embodiments, the cfDNA is from a blood sample or plasma sample, although various bodily fluids, such as saliva or bronchiolar lavage, are suitable.
[0014] In some cases, the generative adversarial network is further trained with an additional training set including data corresponding to the amount of a methylated cfDNA (m-cfDNA) biomarker from the subject, and an additional training set including data corresponding to the amount of an inflammatory biomarker from the subject, such as a member of the chemokine (CXC motif) ligand family, e.g., CXC motif chemokine ligand 1 (CXCL1), CXC motif chemokine ligand 2 (CXCL2), CXC motif chemokine ligand 5 (CXCL5), CXC motif chemokine ligand 9 (CXCL9) (MIG), or CXC motif chemokine ligand 10 (CXCL10) (IP-10). In some cases, the generative adversarial network is further trained with an additional training set including data corresponding to the amount of an apoptotic biomarker from the subject, such as clusterin.
[0015] In some cases, the generative adversarial network is further trained on an additional training set that includes data corresponding to the amount of a protein, such as albumin or total protein. In some cases, the subject is a human.
[0016] In some aspects, the present disclosure provides a non-transitory computer-readable medium storing program instructions executable on one or more computer systems for executing a computer-implemented method for generating a simulated image of a specimen, the computer-implemented method including one or more computer subsystems and one or more components executed by the one or more computer subsystems, the one or more components including a generative adversarial network trained with a training set corresponding to an amount of cfDNA from a subject, the one or more computer subsystems configured to generate a synthetic dataset from the biological sample by inputting a subset of the training data into the generative adversarial network.
[0017] In some aspects, the disclosure provides a non-transitory computer-readable medium storing program instructions executable on one or more computer systems for executing a computer-implemented method for generating a simulated image of a specimen, the computer-implemented method including one or more computer subsystems and one or more components executed by the one or more computer subsystems, the one or more components including a generative adversarial network trained with a training set including data corresponding to amounts of cell-free DNA (cfDNA) biomarkers from subjects with organ dysfunction designated as a first training input, and a second training set including data corresponding to amounts of cell-free DNA (cfDNA) biomarkers from subjects without organ dysfunction designated as a second training input, wherein the first and second datasets are unbalanced, and the one or more computer subsystems are configured to generate a set of synthetic features for the first dataset and / or the second dataset by inputting a portion of the data from the first training input and the second training input into the generative adversarial network. [Brief description of the drawings]
[0018] The foregoing and other features and advantages of the present invention will be more fully understood from the following detailed description of illustrative embodiments taken in conjunction with the accompanying drawings. [Figure 1] The conventional oversampling method (SMOTE) is shown. [Diagram 2] We present strategies for expanding the training dataset using different data augmentation methods. [Diagram 3] We present strategies for training different generative adversarial networks (GANs), including incorporating external data (i.e., synthetic samples or synthetic features or external data) into them and then training different algorithms. [Figure 4A]collectively show a comparison between a range of time points and example biomarkers, which are measured based on their distributions generated by a Conditional Tabular Generative Adversarial Network (CTGAN) using biological original samples (i.e., features of the biological original samples) and synthetic samples (i.e., synthetic features). [Figure 4B] collectively show a comparison between a range of time points and example biomarkers, which are measured based on their distributions generated by a Conditional Tabular Generative Adversarial Network (CTGAN) using biological original samples (i.e., features of the biological original samples) and synthetic samples (i.e., synthetic features). [Figure 4C] collectively show a comparison between a range of time points and example biomarkers, which are measured based on their distributions generated by a Conditional Tabular Generative Adversarial Network (CTGAN) using biological original samples (i.e., features of the biological original samples) and synthetic samples (i.e., synthetic features). [Figure 4D] collectively show a comparison between a range of time points and example biomarkers, which are measured based on their distributions generated by a Conditional Tabular Generative Adversarial Network (CTGAN) using biological original samples (i.e., features of the biological original samples) and synthetic samples (i.e., synthetic features). [Figure 4E]collectively show a comparison between a range of time points and example biomarkers, which are measured based on their distributions generated by a Conditional Tabular Generative Adversarial Network (CTGAN) using biological original samples (i.e., features of the biological original samples) and synthetic samples (i.e., synthetic features). [Figure 4F] collectively show a comparison between a range of time points and example biomarkers, which are measured based on their distributions generated by a Conditional Tabular Generative Adversarial Network (CTGAN) using biological original samples (i.e., features of the biological original samples) and synthetic samples (i.e., synthetic features). [Figure 4G] collectively show a comparison between a range of time points and example biomarkers, which are measured based on their distributions generated by a Conditional Tabular Generative Adversarial Network (CTGAN) using biological original samples (i.e., features of the biological original samples) and synthetic samples (i.e., synthetic features). [Figure 4H] collectively show a comparison between a range of time points and example biomarkers, which are measured based on their distributions generated by a Conditional Tabular Generative Adversarial Network (CTGAN) using biological original samples (i.e., features of the biological original samples) and synthetic samples (i.e., synthetic features). [Figure 5A]collectively show a comparison between a range of time points and example biomarkers, which are measured based on the first two principal components generated by a Conditional Tabular Generative Adversarial Network (CTGAN) using biological original samples (i.e., features of the biological original samples) and synthetic samples (i.e., synthetic features). [Figure 5B] collectively show a comparison between a range of time points and example biomarkers, which are measured based on the first two principal components generated by a Conditional Tabular Generative Adversarial Network (CTGAN) using biological original samples (i.e., features of the biological original samples) and synthetic samples (i.e., synthetic features). [Figure 5C] collectively show a comparison between a range of time points and example biomarkers, which are measured based on the first two principal components generated by a Conditional Tabular Generative Adversarial Network (CTGAN) using biological original samples (i.e., features of the biological original samples) and synthetic samples (i.e., synthetic features). [Figure 5D] collectively show a comparison between a range of time points and example biomarkers, which are measured based on the first two principal components generated by a Conditional Tabular Generative Adversarial Network (CTGAN) using biological original samples (i.e., features of the biological original samples) and synthetic samples (i.e., synthetic features). [Figure 5E]collectively show a comparison between a range of time points and example biomarkers, which are measured based on the first two principal components generated by a Conditional Tabular Generative Adversarial Network (CTGAN) using biological original samples (i.e., features of the biological original samples) and synthetic samples (i.e., synthetic features). [Figure 5F] collectively show a comparison between a range of time points and example biomarkers, which are measured based on the first two principal components generated by a Conditional Tabular Generative Adversarial Network (CTGAN) using biological original samples (i.e., features of the biological original samples) and synthetic samples (i.e., synthetic features). [Figure 5G] collectively show a comparison between a range of time points and example biomarkers, which are measured based on the first two principal components generated by a Conditional Tabular Generative Adversarial Network (CTGAN) using biological original samples (i.e., features of the biological original samples) and synthetic samples (i.e., synthetic features). [Figure 5H] collectively show a comparison between a range of time points and example biomarkers, which are measured based on the first two principal components generated by a Conditional Tabular Generative Adversarial Network (CTGAN) using biological original samples (i.e., features of the biological original samples) and synthetic samples (i.e., synthetic features). [Figure 6A]Collectively, we present a result analysis of the performance of machine learning algorithms on training samples + synthetic samples augmented by different oversampling techniques. [Figure 6B] Collectively, we present a result analysis of the performance of machine learning algorithms on training samples + synthetic samples augmented by different oversampling techniques. [Figure 7] The table lists the results of training on the original data, the generated samples of SMOTE, the generated samples of ADASYN, the generated samples of SVMSMOTE, and the generated samples of CTGAN for the Random Forest, XGBoost, and LightGBM algorithms. The figure shows the feasibility of using different strategies to augment samples in a synthetic manner in a manner that generally reproduces the ROC-AUC obtained with the original data. [Figure 8A] Collectively, we show the performance of the random forest models oversampled by CTGAN and baseline (Figure 8A), the random forest models oversampled by SVMSMOTE and SMOTE (Figure 8B), and the random forest model oversampled by ADASYN (Figure 8C) on the kidney transplant rejection dataset using synthetic urine samples. [Figure 8B] Collectively, we show the performance of the random forest models oversampled by CTGAN and baseline (Figure 8A), the random forest models oversampled by SVMSMOTE and SMOTE (Figure 8B), and the random forest model oversampled by ADASYN (Figure 8C) on the kidney transplant rejection dataset using synthetic urine samples. [Figure 8C]Collectively, we show the performance of the random forest models oversampled by CTGAN and baseline (Figure 8A), the random forest models oversampled by SVMSMOTE and SMOTE (Figure 8B), and the random forest model oversampled by ADASYN (Figure 8C) on the kidney transplant rejection dataset using synthetic urine samples. [Figure 9] 9 shows non-parametric results of a random forest-based rejection score using the SMOTE synthetic data generation method to provide Q-scores. The axes in FIG. 9 represent SMOTE-generated Q-scores (Y-axis) against SMOTE phenotypes (X-axis). [Figure 10] The non-parametric results of the random forest-based rejection score using the method of generating original data (i.e., biological data) to provide Q-scores are shown. The axis in Figure 10 represents the Q-score of the original data against the original phenotype (Y-axis). [Figure 11] Figure 11 shows non-parametric results of a random forest-based rejection score using the GAN synthetic data generation method to provide a Q-score. The axes in Figure 11 represent the GAN-generated Q-score (Y-axis) against the GAN phenotype (X-axis). [Figure 12] 12 shows non-parametric results of a random forest-based rejection score using the ADASYN synthetic data generation method to provide a Q-score. The axes in FIG. 12 represent ADASYN-generated Q-scores (Y-axis) against ADASYN phenotypes (X-axis). [Figure 13] Figure 13 shows non-parametric results of a random forest-based rejection score using the SVM synthetic data generation method to provide a Q-score. The axes in Figure 13 represent SVM-generated Q-scores (Y-axis) versus phenotype (X-axis).
[0019] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0020] In medical diagnostic applications, information about healthy patients is much richer than information about diseased patients. Thus, machine learning algorithms tend to misclassify some unhealthy patients as healthy. Furthermore, acquiring biological data is difficult and expensive, since generating training samples in the biomedical field requires experts and a series of long-term experiments. If synthetic data can be used to complement and improve real data, more valuable applications can be realized with less current data in different domains. Creating synthetic bioinformatics data is a challenging task, since the synthetic data must preserve the underlying biological effects.
[0021] For example, kidney diseases are known to be largely multifactorial, with complex and overlapping clinical phenotypes and morphologies, which often result in delayed diagnosis and chronic progression. Despite advances in computational power and the evolution of machine learning-based methods, the underlying biological complexity of various kidney diseases and their progression to kidney transplant rejection continues to make early diagnosis and intervention a challenge, especially in resource-poor areas. Currently, existing academic and applied research focuses on leveraging such methods to better understand the division and function of multiple organs, and machine learning methods have contributed to more accurate and timely predictions, and better understanding of histological pathology to some extent. However, such methods are limited to the fields of transplantation and rejection monitoring due to insufficient data availability, and therefore have not yet penetrated into standard medical practice and diagnostic procedures. With the help of artificial intelligence (AI), it is possible to perform large-scale health screening for potential kidney diseases as well as targeted biomarker and drug discovery, thus enabling clinicians to treat patients in a more targeted manner.
[0022] Furthermore, artificial intelligence-assisted diagnostic applications can help unravel the various etiologies of kidney diseases for more accurate phenotyping or outcome prediction, thus reducing the chances of misdiagnosis. The generalization of machine learning models typically depends on the quality of the dataset, since a good dataset allows the machine learning classifier to efficiently capture the underlying characteristics. As a result, the machine learning classifier is likely to be more robust in effectively generalizing the underlying characteristics to unseen data. To yield a good dataset, the data should be a good representation of the true distribution and cover as many cases as possible with a reasonably large number of samples. However, collecting biomedical data usually requires the involvement of expert physicians, leading to high collection costs and data. Therefore, it is not always possible to access data from a larger number of patients. Therefore, creating synthetic datasets is beneficial when machine learning algorithms attempt to learn the underlying characteristics of data from small imbalanced datasets.
[0023] Described herein is a generative adversarial network (GAN) system that introduces synthetic data (i.e., data augmentation) into a biological dataset, for example, by generating synthetic data in tabular format that reduces class imbalance when the number of samples is uneven for all classes present in the dataset. The systems and processes described herein add external synthetic training data to a training set derived from biological samples to improve the performance of machine learning algorithms and significantly reduce or eliminate bias generated from uneven numbers of samples. In some aspects, the system of the present disclosure describes the addition of external synthetic data to a kidney transplant rejection dataset that is trained primarily on six biomarker features, along with a time feature (days post-transplant: 0 days (day of surgery), -1 days (day before surgery), +1 days (24 hours post-surgery), etc.) representing the number of days since organ transplantation (e.g., kidney transplant, pancreas transplant, double kidney + pancreas transplant) to predict early kidney transplant failure.
[0024] In some aspects, the present disclosure provides systems using different GAN architectures and the effectiveness of synthetic data generated by GAN-based methods for machine learning algorithms and processes for utilizing the same. In some aspects, the present disclosure describes a comparison of the distribution of the first two principal components and the cumulative sum per feature in a dataset including only original data collected from biological samples against a synthetic training set to which synthetic biomarker data (i.e., external data) has been added. In additional aspects, the present disclosure describes the ROC-AUC, sensitivity, and specificity scores obtained by a machine learning classifier trained with additional synthetic data versus a classifier trained only on the original data. In further aspects, the present disclosure describes the performance of a machine learning classifier on a dataset augmented with one or more GAN architectures described herein, including but not limited to the Conditional Tabula GAN (CTGAN) architecture, the statistical oversampling SMOTE architecture, the ADASYN architecture, and the SVMSMOTE architecture.
[0025] This disclosure demonstrates with experimental results that systems and processes utilizing GAN-based data augmentation achieve significantly higher accuracy in accurately classifying medical samples when compared to traditional statistical oversampling methods. The use of such GAN-based data augmentation approaches for medical tabula rasa data provides a new generation of artificial intelligence applications in the medical field.
[0026] Generative Adversarial Networks for Biomarker Analysis The presence or absence of a combination of biomarkers in a sample may reflect the state of a subject's organ. Identification of biomarkers typically involves the use of biochemical assays to identify the "amount" or "level" of a biomarker in a sample. Many assays exist in the art that can be used to detect biomarkers in biological samples (e.g., urine or blood), such as gene or protein array analysis, or metabolite analysis. The use of biochemical assays in this context may require probing for functional changes in genes and proteins, the need for a priori knowledge of their function (e.g., antibody detection), and extensive assay development and optimization.
[0027] In many diseases (e.g., allograft rejection or organ damage), the presence of observable functional biomarkers often occurs late in the disease state. For example, the presence of serum creatinine (sCR), a biomarker commonly used to screen for kidney allograft rejection, is only detected as a late marker of allograft rejection. Thus, preventative measures for allograft rejection or kidney damage may not be effective if developed solely in conjunction with the detection of late markers of rejection, such as serum creatinine.
[0028] Contributions to understanding the individual biomarkers expressed in allograft rejection, particularly in renal, lung and cardiac allograft rejection, have been made by systematic evaluation of gene expression data and "omics" studies. See, for example, Sigdel TK, Bestard O, Tran TQ, et al. A Computational Gene Expression Score for Predicting Immune Injury in Renal Allografts. PLoS One. 2015;10(9):e0138133, published 14 September 2015, vdoi:10.1371 / journal.pone.0138133. See also Sigdel, Tara, et al., "Assessment of 19 Genes and Validation of CRM Gene Panel for Quantitative Transcriptional Analysis of Molecular Rejection and Inflammation in Archival Kidney Transplant Biopsies." Frontiers in Medicine, vol. 6, 2019, doi: 10.3389 / fmed.2019.00213. See also Sigdel, Tara K., et al., "A Urinary Common Rejection Module (UCRM) Score for Non-Invasive Kidney Transplant Monitoring." PLOS ONE, vol. 14, no. 7, 2019, doi: 10.1371 / joumal.pone.0220052.See also Khatri, Purvesh, et al. "A Common Rejection Module (CRM) for Acute Rejection across Multiple Organs Identifies Novel Therapeutics for Organ Transplantation." Journal of Experimental Medicine, vol. 210, no. 11, 2013, pp. 2205-2221., doi:10.1084 / jem.20122709.
[0029] Other studies have considered donor derived cell-free DNA (dd-cfDNA) as a potential surrogate biomarker for allograft injury, first in blood and then in urine samples. dd-cfDNA continuously flows into the circulation from the moment the transplanted organ is transplanted. One rationale for monitoring dd-cfDNA in transplantation is that cellular damage to the allograft that occurs up until or during the rejection sequence releases DNA into the recipient's circulation, thereby elevating dd-cfDNA levels. Thus, due to continuous cellular turnover, strategies to measure donor derived cell-free DNA (dd-cfDNA) levels as a potential surrogate biomarker for allograft injury have been investigated as a potential surrogate biomarker for transplant injury (see, for example, Sarwal and Sigdel, WO2014 / 145232). However, such applications are limited by the technology available for capturing dd-cfDNA.
[0030] For example, some methods for dd-cfDNA capture / detection required either gender mismatching between donor and recipient or prior genotyping of donor and recipient. This allows for quantification of dd-cfDNA by PCR amplification of genes found on the Y chromosome, such as the SRY gene. Snyder et al. described a universal approach to dd-cfDNA assessment that does not require gender mismatching (see TMSnyder, KK. Valantine, SR Quake., Universal noninvasive detection of solid organ transplant rejection. Proc Natl Acad Sci, 108 (2011), pp. 6229-6234). Snyder used genome-wide sequencing of plasma cfDNA in heart transplant recipients to evaluate for SNPs known to be homozygous with sequences that differ between donors and recipients and to calculate the percentage of dd-cfDNA relative to total cfDNA. The study found that at some frequency dd-cfDNA levels are elevated prior to a pathological diagnosis of rejection. However, this approach requires DNA from a donor, which is often impractical and particularly difficult if the transplant was performed several years ago.
[0031] Improvements in these technologies have necessitated the use of targeted next-generation sequencing (NGS) technologies to quantify dd-cfDNA without the need for prior genotyping of donors and recipients. These NGS assays include AlloSure® (CareDx, Brisbane, CA) and Prospera® (Natera, San Carlos, CA). AlloSure® has been analytically validated in a Clinical Laboratory Improvement Amendments (CLIA) setting. Prospera® (Natera, San Carlos, CA) was adapted for use in kidney transplantation from an approach developed for non-invasive prenatal testing (NIPT). Yet, both approaches are often impractical as they require NGS sequencing of samples, making these products costly for continuous monitoring.
[0032] Sarwal et al. have investigated the use of various samples, including urine, as a non-invasive source of other informative biomarkers for monitoring different types of solid organ transplants (see, e.g., U.S. Pat. Nos. 10,982,272, 10,995,368, 11,124,824, and U.S. Patent Application Nos. 17 / 376,919, 17 / 498,489). Sarwal recognized that Alu elements are the most abundant transposable elements in the human genome, containing over one million copies distributed throughout the human genome. Recognizing the abundance of Alu repeats, Sarwal created a ratio of Alu repeats in urine samples from transplant patients to the number of Alu repeats in urine samples from normal populations. This ratio could be used as a surrogate for damage, but was not sufficiently informative by itself.
[0033] Further studies have been initiated to explore potential combinations of biomarkers as surrogates for allograft failure. For example, QSant™ utilizes a composite score of various biomarkers with different biochemical characteristics, i.e., proteins, metabolites, and nucleic acids. (See Yang, Sarwal, et al., A urine score for noninvasive accurate diagnosis and prediction of kidney transplant rejection. Science Translational Medicine, 18 Mar 2020, Vol. 12, Issue 535). Yang et al. showed that a urinary composite score of six biomarkers, consisting of an inflammatory biomarker (e.g., CXCL-10, also known as IP-10), an apoptotic biomarker (e.g., clusterin), a cfDNA biomarker, a DNA methylation biomarker, a creatinine biomarker, and a total protein, allows the diagnosis of acute rejection (AR) with an area under the receiver operating characteristic curve of 0.99 and an accuracy of 96%. Remarkably, QSant™ (formerly known as QiSant™) predicts acute rejection before a rise in the stand-alone serum creatinine test is observed, allowing for earlier detection of rejection than is currently possible with current standard of care tests.
[0034] However, analysis of data obtained in such studies can be difficult in part because many biological datasets available for these studies come from imbalanced datasets, which are datasets with an uneven number of samples for all classes present in the dataset. This can cause machine learning algorithms to perform poorly for minority classes while favoring a bias towards the majority class. The present disclosure explores scenarios in which synthetic data can be used to supplement and improve the real data obtained in such studies to reduce class imbalance and achieve more valuable applications in different domains with less existing data.
[0035] Generative Adversarial Networks Creating synthetic datasets is beneficial when machine learning algorithms attempt to learn underlying characteristics of data from small imbalanced datasets. Machine learning algorithms find and apply patterns in the data. Multivariate machine learning and linear and non-linear fitting algorithms can also be applied to biomarker search. Machine learning is generally supervised or unsupervised. In supervised learning, the most common data is labeled to tell the machine exactly what patterns to look for. For example, samples from patients with a known diagnosis of acute rejection are labeled as "acute rejection". Samples from "normal" patients are labeled as "stable". The algorithm then starts looking for patterns that are clearly different between "normal" and "acute rejection". In unsupervised learning, the data has no labels. The machine algorithm looks for any patterns it can find. This may be interesting, for example, if all samples analyzed are from subjects who have received allogeneic transplants. Machine algorithms can be used, for example, to detect specific markers of a wide range of allografts.
[0036] The generalization of a machine learning model depends on the quality of the dataset, since a good dataset allows the machine learning classifier to capture the underlying characteristics well. As a result, the machine learning classifier will be more robust in effectively generalizing the underlying characteristics to unseen data. To yield a good dataset, the data should generally be a good representation of the true distribution and should cover as many cases as possible with a reasonably large number of samples. However, collecting biomedical data usually requires the involvement of expert physicians, leading to high collection costs and data. Therefore, it is not always possible to access data of more patients. Another reason for creating synthetic data is to avoid using original data to train machine models for privacy reasons. For example, medical samples consisting of sensitive personal information about patients, such as weight, height, and date of birth, should be strictly protected for privacy reasons, since working directly with such information may compromise its security. The present disclosure addresses these challenges by a) generating synthetic (i.e., synthetic datasets that augment inputs in biological samples by providing external training features to the original data, and b) training a machine learning model on the generated synthetic dataset without training on the original data.
[0037] There has been an explosion of biomarker discovery efforts using genomics, proteomics, and metabolomics, but these techniques also focus on characterizing biomarkers present in the biological original sample. Biological samples can particularly benefit from synthetic data augmentation techniques, in part due to the challenges of obtaining sufficient amounts of original sample or preserving the integrity of all biomarkers in the biological original sample to feature in machine learning models. This disclosure demonstrates the utility of synthetic data augmentation techniques in biological samples, and in a specific embodiment of a renal transplant rejection dataset consisting of six biomarkers: cell-free DNA (cfDNA), methylated cell-free DNA (m-cfDNA), at least one inflammatory marker(s), at least one apoptotic marker(s), total protein, and creatinine, to predict early renal transplant failure. The biological role of these biomarkers to assess renal injury and acute rejection in patients has demonstrated efficacy in supporting important patient management decisions, with a turnaround time of less than three days. See, for example, U.S. Patent Nos. 10,982,272 and 10,995,368. Also see A urine score for noninvasive accurate diagnosis and prediction of kidney transplant rejection, Science Translational Medicine 18 Mar 2020: Vol.12, Issue 535, eaba2501. Following kidney transplantation, it is essential to monitor subjects for evidence of rejection to reduce the risk of graft loss. In this disclosure, it has been shown that the performance of machine learning algorithms improves when algorithms trained on datasets obtained from subjects' urine samples (i.e., true training data) are combined with synthetic data generated by GAN-based data augmentation methods.
[0038] Generative Adversarial Networks for the Analysis of Urinary Biomarkers In one aspect, the present disclosure provides a synthetic data augmentation approach for medical tabula data that improves the analysis of a combination of biomarkers that can be used to accurately monitor the integrity of solid organ allografts after transplantation. The present disclosure describes such an analysis in a kidney transplant rejection dataset consisting of six biomarkers: cell-free DNA (cfDNA), methylated cell-free DNA (m-cfDNA), CXCL10, clusterin, total protein, and creatinine, to predict early kidney transplant failure.
[0039] Kidney disease is a significant medical and public health burden worldwide, with both AKI and CKD resulting in high morbidity and mortality, and contributing to enormous healthcare costs. Due to the high heterogeneity in disease manifestations, progression, and treatment response, the present disclosure explores leveraging novel big data and AI methods to solve the challenges associated with addressing these complex diseases and disease-related injuries. The present disclosure considers generative adversarial networks (GANs), first introduced in 2014 by Goodfellow et al., to significantly improve upon foundational approaches to provide new opportunities to solve data scarcity problems and help powerful machine learning applications overcome the barriers of small biological sample sizes, particularly sample sizes with uneven distribution.
[0040] GANs provide a strategy to train generative models that automatically discover and learn patterns based on deep neural networks consisting of a generator network and a discriminator network. The role of the generator is to generate new, plausible examples from the problem domain, and the role of the discriminator is to classify the examples as either true (from the domain) or false (e.g., synthetic or generated). The two neural networks learn simultaneously from the training data in an adversarial zero-sum game fashion, where the loss of one neural network is the gain of the other.
[0041] In this disclosure, we show that GAN-based data augmentation methods can be applied to generate quality synthetic samples that resemble the original distribution of the real-world data they are provided with. But more importantly, the above studies and related academic research show that GAN-based data generation methods can reproduce the biological complexity found in various types of genetic, proteomic, and cell-type data that are often analyzed in diagnostic and therapeutic studies. This disclosure shows that the systems, processes, and methods disclosed herein can be successfully applied to biological data in various medical fields. This shows that GAN-based generative models can be a useful tool to generate synthetic biomarker data of biological samples for more robust analysis.
[0042] To address the problem of small sample size, some oversampling methods have been proposed in previous studies. The present disclosure provides the use of GAN-based synthetic data techniques, which may be a more effective strategy than previous oversampling methods to overcome the problem of imbalanced datasets. In some aspects, the present disclosure considers and implements oversampling methods, including random oversampling in its analysis. Figure 1 shows a conventional oversampling method (SMOTE). As shown in Figure 1, input data (majority class samples are larger circles and minority class samples are smaller circles) is processed using the SMOTE methodology (minority oversampling) for synthetic data calculation, and then synthetic data is generated.
[0043] In some embodiments, the present disclosure contemplates the use of Synthetic Minority Oversampling Technique (SMOTE), Borderline SMOTE, Borderline Oversampling with SVM, and Adaptive Synthetic Sampling (ADASYN), as well as other suitable methodologies, for analyzing biomarkers in biological samples (e.g., blood or urine).
[0044] In some aspects, an example oversampling method contemplated in this disclosure involves randomly replicating training examples of the minority class (i.e., random oversampling).
[0045] In some aspects, example oversampling methods contemplated in this disclosure include the Synthetic Minority Oversampling Technique (SMOTE), which works by selecting close examples in the feature space, drawing lines between the samples in the feature space, and drawing new samples as points along the lines.
[0046] In yet another aspect, an exemplary oversampling method contemplated in this disclosure includes a novel minority oversampling technique that considers a k-nearest neighbor classification model to generate only minority synthetic samples near the boundary. The SMOTE-SVM oversampling method is an extension of SMOTE, which fits a support vector machine algorithm to the dataset and generates synthetic samples using a decision boundary defined by the support vectors.
[0047] In another aspect, exemplary oversampling methods contemplated in this disclosure include an adaptive synthetic sampling approach, which utilizes a weighted distribution of the minority class to generate synthetic samples that are inversely proportional to the density of examples in the minority class.
[0048] In another aspect, the present disclosure contemplates Majority Weighted Minority Oversampling TEchnique (MWMOTE), a method that aims to generate more selected synthetic minority class samples by assigning weights based on Euclidean distance from the nearest majority class instance.
[0049] Other methods have been developed to meet the demands of datasets, and this disclosure considers a suitable method to replace imbalanced learning for machine learning algorithms by rebalancing class distributions for imbalanced datasets.
[0050] Other definitions For the purposes of interpreting this specification the following definitions will apply and, where appropriate, terms used in the singular will also include the plural and vice versa.
[0051] sample The term "biological sample" or "sample" as used herein refers to a mixture of cells, tissues, and fluids obtained or derived from an individual that contains cells and / or other molecular entities to be characterized and / or identified, for example, based on physical, biochemical, chemical, and / or physiological characteristics. In one embodiment, the sample is a liquid (i.e., a biological fluid), such as urine, blood, serum, plasma, saliva, sputum, etc. In other embodiments, the sample is a histological section, such as a solid tissue section from a biopsy.
[0052] Subject The subject can be any human or animal (collectively, "individual") that has received an allograft. For example, the subject can be a human, a non-human primate such as a chimpanzee, as well as other ape and monkey species, livestock such as cows, horses, sheep, goats, pigs, farm animals such as rabbits, dogs, cats, and laboratory animals including rats, mice, guinea pigs, and the like. The subject can be of any age. The subject can be, for example, a geriatric, adult, adolescent, pre-adolescent, child, infant, or toddler. In certain cases, the subject is a pediatric recipient of an allograft.
[0053] A "subject," also referred to as an "individual," may be a "patient." A "patient" refers to a subject under the care of a treating physician. In one embodiment, the patient is suffering from renal damage or impairment. In another embodiment, the patient is suffering from renal disease or failure. In another embodiment, the patient has undergone a renal transplant and is experiencing renal transplant rejection. In yet another embodiment, the patient has been diagnosed with renal damage, renal disease, or renal transplant rejection, but is not receiving any treatment to address the diagnosis.
[0054] probe "Hybridization", "probe hybridization", "cfDNA probe hybridization", or "Alu probe hybridization" refers to a reaction in which one or more polynucleotides react to form a complex stabilized through hydrogen bonds between the bases of the nucleotide residues. Hydrogen bonds can occur by Watson Crick base pairing, Hoogstein binding, or in any other sequence-specific manner. The complex may include two strands forming a double-stranded structure, three or more strands forming a multistranded complex, a single self-hybridizing strand, or any combination of these. A hybridization reaction may constitute a step in a more extensive process, such as pairing with a cfDNA sequence (e.g., probe hybridization to an Alu region of cfDNA), initiating PCR, or cleaving a polynucleotide with an enzyme. A sequence that can hybridize to a given sequence is called the "complement" of the given sequence.
[0055] The terms "polynucleotide", "nucleotide", "nucleotide sequence", "nucleic acid" and "oligonucleotide" are used interchangeably. They refer to polymeric forms of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or their analogs. Polynucleotides may have any three-dimensional structure and may perform any function, known or unknown. Non-limiting examples of polynucleotides include coding or non-coding regions of genes or gene fragments, loci defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, small interfering RNA (siRNA), small hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. The term also encompasses nucleic acid-like structures with synthetic backbones. See, for example, Eckstein, 1991; Baserga et al., 1992; Milligan, 1993; WO97 / 03211; WO96 / 39154; Mata, 1997; Strauss-Soukup, 1997; and Samstag, 1996. A polynucleotide may contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure may be imparted before or after self-assembly of the polymer. The sequence of nucleotides may be interrupted by non-nucleotide moieties. A polynucleotide may be further modified after polymerization, such as by conjugation with a labeling moiety.
[0056] As used herein, the term "genomic locus" or "locus" (plural loci) is the specific location of a gene or DNA sequence on a chromosome. A "gene" is a stretch of DNA or RNA that encodes a polypeptide or an RNA chain that is taken to be the molecular unit of heredity in an organism because of the functional role it has in that organism. For the purposes of the present invention, a gene may be considered to have regions that regulate the production of a gene product, whether or not such regulatory sequences flank the coding and / or transcribed sequence. Thus, genes include, but are not necessarily limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, origins of replication, matrix attachment sites, and locus control regions.
[0057] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to amino acid polymers of any length. The polymers may be linear or branched, may comprise modified amino acids, and may be interrupted by non-amino acids. These terms also encompass amino acid polymers that have been modified by, for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as conjugation with a labeling component.
[0058] As used herein, the term "amino acid" includes natural and / or unnatural or synthetic amino acids, including glycine and both the D or L optical isomers, as well as amino acid analogs and peptidomimetics.
[0059] As used herein, the term metabolite refers to an intermediate or end product of metabolism. The term metabolite is typically used for small molecules, but can also include amino acids, vitamins, nucleotides, antioxidants, organic acids, and vitamins.
[0060] As used herein, the term "domain" or "protein domain" refers to a portion of a protein sequence that can exist and function independently of the rest of the protein chain.
[0061] As used herein, the terms "disorder" or "disease" and "injury" or "damage" are used interchangeably and refer to any change in the state of the body, or of one of its organs and / or tissues, that impedes or interferes with the ability of organ and / or tissue function (e.g., causing organ dysfunction) and / or causes symptoms, such as discomfort, dysfunction, suffering, or death, in an afflicted subject.
[0062] A subject "at risk" of developing renal injury, renal disease, or renal graft rejection may or may not have detectable disease or symptoms, and may or may not exhibit detectable disease or symptoms of disease prior to the treatment methods described herein. "At risk" means that the subject has one or more risk factors described herein and known in the art that are measurable parameters that correlate with the development of renal injury, renal disease, or renal graft rejection. A subject with one or more of these risk factors has a higher probability of developing renal injury, renal disease, or renal graft rejection than a subject without any of these risk factor(s).
[0063] The term "condition" is used herein to refer to the identification or classification of a medical or pathological condition, disease, or diagnosis. For example, "condition" may refer to a healthy state of a subject, a stable state of a subject who has received an allograft, or may refer to the identification of a disease. The disease may be renal injury, renal disease (e.g., CKI or AKI), or renal graft rejection. "Diagnosis" may also refer to the classification of the severity of renal injury, renal disease, or renal graft rejection. The diagnosis of renal injury, renal disease, or renal graft rejection may be made according to any protocol used by a person skilled in the art (e.g., a nephrologist).
[0064] The term "companion diagnostic" is used herein to refer to a method that aids in making a clinical decision regarding the presence, extent, or other nature of a particular type of symptom or condition of renal injury, renal disease, or renal transplant rejection. For example, a companion diagnostic for renal injury, renal disease, or renal transplant rejection can include measuring cell-free DNA fragment size.
[0065] The term "prognosis" is used herein to refer to the prediction of the likelihood of development and / or recurrence of a disorder, such as renal injury, renal disease, or renal graft rejection, during treatment with an allograft. The predictive method of the present invention can be used clinically to make treatment decisions by selecting the most appropriate treatment modality for any particular patient. The predictive method of the present invention is a useful tool in predicting and / or aiding in the diagnosis of whether a patient is likely to develop renal injury, renal disease, or renal graft rejection, whether they will have a recurrence of renal injury, renal disease, or renal graft rejection, and / or whether symptoms of renal injury, renal disease, or renal graft rejection are worsening.
[0066] "Treating" and "treatment" refer to clinical intervention in an attempt to change the natural course of an individual, and can occur before, during, or after a clinical diagnosis or prognosis. Desirable effects of treatment include preventing the occurrence or recurrence of renal injury, renal disease, or renal graft rejection, or a condition or symptom thereof, alleviating a condition or symptom of renal injury, renal disease, or renal graft rejection, reducing any direct or indirect pathological consequences of renal injury, renal disease, or renal graft rejection, reducing the rate of progression or severity of renal injury, renal disease, or renal graft rejection, and / or improving or alleviating renal injury, renal disease, or renal graft rejection. In some embodiments, the methods and compositions of the present invention are used on patient subpopulations identified as at risk for developing renal injury, renal disease, or renal graft rejection. In some cases, the methods and compositions of the present invention are useful in an attempt to delay the onset of renal injury, renal disease, or renal graft rejection. The beneficial or desired clinical outcomes are known or can be easily obtained by those skilled in the art. For example, the beneficial or desired clinical outcomes can include, but are not limited to, one or more of the following: monitoring kidney damage, detecting kidney damage, identifying the type of kidney damage, assisting kidney transplant physicians in deciding whether to send transplant patients for biopsy, and helping kidney transplant physicians make decisions for clinical management and therapeutic intervention purposes.
[0067] The term "wild type" as used herein is a term of the art understood by those skilled in the art and refers to the typical form of a naturally occurring organism, strain, gene, or characteristic as distinguished from mutant or variant forms. The term "variant" as used herein should be interpreted to mean exhibiting a quality having a pattern that deviates from that which occurs in nature. The terms "orthologue" (also referred to herein as "ortholog") and "homologue" (also referred to herein as "homolog") are well known in the art. By further guidance, a "homolog" of a protein as used herein is a protein of the same species that performs the same or similar function as the protein to which it is a homolog. Homolog proteins may, but need not, be structurally related, and are only partially structurally related. An "orthologue" of a protein as used herein is a protein of a different species that performs the same or similar function as the protein that is the ortholog of the orthologous protein, and is not necessarily structurally related, or is only partially structurally related. Homologs and orthologs may be identified by homology modeling (see, for example, Greer, Science vol. 228 (1985) 1055, and Blundell et al. Eur J Biochem vol 172 (1988), 513, or "structural BLAST" (Dey F, Cliff Zhang Q, Petrey D, Honig B. Toward a "structural BLAST": using structural relationships to infer function. Protein Sci. 2013 April; 22(4): 359-66. doi: 10.1002 / pro.2225). EXAMPLES
[0068] Example 1. Generative adversarial network to generate synthetic biomarker data for urine samples Data collection The study included 379 independent biopsy-matched urine samples obtained with informed consent from 309 pediatric recipients (3–18 years) and adult recipients (18–76 years) of renal allografts transplanted at three different transplant centers (University of California, San Francisco (UCSF), San Francisco, USA), Stanford University (Palo Alto, CA), and the Instituto Nacional de Ciencias Medicas y Nutricion Salvador Zubiran (Mexico City, Mexico).
[0069] Of the 379 samples, acute kidney allograft rejection (AR) was confirmed in 243 samples by paired biopsy readout, and no rejection or stable (STA) phenotype was confirmed in 136 samples. Urine samples were collected from these patients on days 1-1539 post-transplant. Custom generated ELISAs for m-cfDNA, CXCL10, and clusterin concentrations were used for these biomarkers. cfDNA was detected using probes as described in Sarwal et al. (see, e.g., U.S. Patent Nos. 10,982,272, 10,995,368, 11,124,824, and U.S. Patent Application Nos. 17 / 376,919, 17 / 498,489). Both DNA assays used SuperSignal ELISA for luminescent detection. Analyte concentrations from 379 independent biological samples with corresponding biomarker data (cfDNA, m-cfDNA, CXCL10, clusterin, creatinine, and total protein) were measured.
[0070] Synthetic urine samples were generated from learned distributions of urinary analyte concentrations based on true biological samples with corresponding biomarker data (cfDNA, m-cfDNA, CXCL10, clusterin, creatinine, and total protein).
[0071] The original data was randomly split into a 70% training set and a 30% test set, with 174 biopsy-confirmed acute renal allograft rejection (AR) phenotypes and 91 no rejection (NR) or stable (STA) phenotypes in the training set. There were 69 biopsy-confirmed acute renal allograft rejection (AR) phenotypes and 45 no rejection or stable (STA) phenotypes in the test set. The following scheme was used to expand the training dataset using different data augmentation methods.
[0072] Figure 2 is a schematic diagram of various GANS strategies utilized on the aforementioned dataset to test the process for expanding the dataset using different data augmentation methods. As shown in Figure 2, an original sample set of 295 inputs (174 acute rejection (AR) and 91 no rejection (NR)) was used for training.
[0073] We then used the Synthetic Minority Oversampling Technique (SMOTE) as a statistical technique to increase the number of cases in the dataset in a balanced way. The module worked by generating new instances from the existing minority cases (NR=91) provided as input. This implementation of SMOTE did not change the number of majority cases. Furthermore, the new synthetic data was not simply a copy of the existing minority cases. Instead, the algorithm took samples of the feature space for each target class and its nearest neighbors, creating balanced samples with AR=174 and NR=174, totaling n=348.
[0074] In parallel, we used the methodology of the Adaptive Synthetic Sampling Approach for Imbalanced Learning (ADASYN) to generate the synthetic data points required to balance the dataset. The main difference between SMOTE and ADASYN is the difference in the generation of synthetic sample points for minority data points. In ADASYN, we use a density distribution r xWe considered the weighting factor, which determines the number of synthetic samples generated for a particular point, whereas in SMOTE, the weighting factor is uniform across all minority points. This strategy produced balanced samples with AR=174 and NR=172, totaling n=346, as shown in Figure 2.
[0075] In parallel, we developed CTGAN, a deep learning-based synthetic data generator ensemble for single-table data. CTGAN (from “Conditional Tabular Generative Adversarial Network”) was completed by constructing a synthetic data table using GAN. GAN is a pair of neural networks that creates the first row of synthetic data, and the second row, called the discriminator, tries to determine whether it is true or not. Finally, the generator was able to generate synthetic data that the discriminator cannot distinguish from the true data. This strategy produced balanced samples with AR=784 and NR=784, totaling n=1565, as shown in Figure 2.
[0076] Example 2: Creating a machine learning classifier using various GANs Synthetic urine samples were generated from learning distributions of urinary analyte concentrations based on real biological samples with corresponding biomarker data (cfDNA, m-cfDNA, CXCL10, clusterin, creatinine, and total protein). Figure 3 shows the strategy for training different generative adversarial networks (GANs) incorporating external data (i.e., synthetic samples or synthetic features or external data) within them, followed by training different algorithms as outlined in this example.
[0077] To develop and train the different GANs, we split the data into training and testing sets using a random 70 / 30 split, respectively, and ran four different GANs: Conditional Tabula Generative Adversarial Network (CTGAN), Vanilla GAN, Tabula GAN (TGAN), and Table GAN. See Figure 3. "Training Different GANs".
[0078] A logarithmic transformation was applied to the data to transform the skewed distribution of the aforementioned biomarkers and help reduce the range of values the generator must generate. A model was then trained using both identified and unidentified target variables to generate high-quality synthetic minority samples.
[0079] TGAN is a tabular data synthesizer that uses LSTMs to generate synthetic data column-by-column, where each column depends on previously generated columns. When generating columns, TGAN's attention mechanism focuses on previous columns that are highly related to the current column.
[0080] Table GAN uses convolutional networks in both the generator and the classifier. When the tabular data contains a label column, a prediction loss is added to the generator to explicitly improve the correlation between the label column and other columns.
[0081] Vanilla GAN uses a minimax algorithm, includes a classifier and a generator with four dense layers in its architecture, optimizes a binary cross-entropy loss function, and calculates the logarithmic loss of the predicted probabilities of both the generator and the classifier.
[0082] TabulaGAN is a GAN-based data augmentation method to address challenges in the tabula data generation task that previous statistical and deep neural network methods could not address, such as non-Gaussian, multimodal distributions, and imbalanced discrete sequences.
[0083] 4A through 4H collectively show a comparison between a range of time points and exemplary biomarkers measured based on their distributions generated by a Conditional Tabular Generative Adversarial Network (CTGAN) using biological original samples (i.e., biological original sample features) and synthetic samples (i.e., synthetic features). FIG. 4A shows a comparison between the original samples and the synthetic samples (i.e., synthetic features) based on a feature-by-feature cumulative sum of six biological features generated by the CTGAN over a period of time after transplantation. FIGS. 4B through 4G show a comparison between the original samples and the synthetic samples (i.e., synthetic features) based on each of the individual biological features used in an exemplary test, i.e., the QSant™ diagnostic test for allograft rejection. Figure 4B shows the performance of the creatinine biomarker, Figure 4C shows the performance of the total protein biomarker, Figure 4D shows the performance of an exemplary inflammation biomarker, Figure 4E shows the performance of an exemplary clusterin biomarker, Figure 4F shows the performance of an exemplary cfDNA biomarker, and Figure 4H shows the distribution of true and false phenotypes.
[0084] 5A through 5H collectively show a comparison between a range of time points and exemplary biomarkers measured based on the first two principal components generated by a Conditional Tabular Generative Adversarial Network (CTGAN) using biological original samples (i.e., biological original sample features) and synthetic samples (i.e., synthetic features). FIGS. 5B through 5G show a comparison between the original samples and synthetic samples (i.e., synthetic features) based on each of the individual biological features used in an exemplary test, i.e., the QSant™ diagnostic test for allograft rejection. FIG. 5B shows the performance of the creatinine biomarker, FIG. 5C shows the performance of the total protein biomarker, FIG. 5D shows the performance of an exemplary inflammation biomarker, FIG. 5E shows the performance of an exemplary clusterin biomarker, and FIG. 5F shows the performance of an exemplary cfDNA biomarker. FIG. 5H shows the phenotype. 6A-6B collectively show the results analysis of the performance of machine learning algorithms on training samples plus synthetic samples augmented by different oversampling techniques.
[0085] It was observed that training CTGAN, where no class labels were generated, provided realistic synthetic data of biomarker values (with high sensitivity and specificity) compared to other GAN architectures based on distribution, feature-wise cumulative sum, and the first two principal components. A machine learning classifier was then built on the training set merged with the synthetic samples, and the performance of the classifier oversampled by CTGAN was compared to traditional oversampling methods such as SMOTE, SVM SMOTE, ADASYN, and a baseline on non-oversampled data.
[0086] But more importantly, this data suggests that a variety of different methods can be used to generate synthetic data that closely mimics the capabilities of biological data. Based on the foregoing data, various systems can be configured to balance imbalanced datasets obtained from biological samples, which can be trained using CTGAN, vanilla GAN, TGAN, and table GAN strategies to generate synthetic data. Such synthetic data can then be used to train various ML algorithms, including CTGAN, SMOTE, SVM-SMOTE, and ADASYN machine learning algorithms.
[0087] The present disclosure contemplates that such strategies may be used with biological samples obtained from urine as described in the examples, but also from blood, serum, plasma, bronchoalveolar lavage fluid, or another suitable source of biological material.
[0088] Example 3. Synthetic urine samples generated using Conditional Tabula Generative Adversarial Network (CTGAN) In this example, a conditional tabula generative adversarial network (CTGAN) with Wasserstein loss (W-loss) and gradient penalty was used as an example GAN architecture to generate the final synthetic urine samples. In contrast to min-max standardization used by previous models to manage complex distributions, CTGAN introduced new techniques such as conditional generators and training by sampling to manage imbalanced discrete sequences and mode-specific standardization. The training process of traditional GANs was a minimax game using binary cross-entropy loss (Bce loss). However, training GANs with Bce loss was prone to mode collapse and vanishing gradient problems, especially when the generated examples are significantly different from the true examples. Mode collapse occurs when a generator generates examples from a single class from the entire training dataset, such as handwritten digit 1, and learns to fool the classifier by collapsing into a single mode or the entire distribution of possible handwritten digits. A real-world dataset may have many modes associated with each possible class in the dataset, such as digits in a dataset of handwritten digits.
[0089] To solve this mode collapse and vanishing gradient descent, the present disclosure uses CTGAN to apply a Wasserstein loss (W-loss) function with a gradient penalty regularization term along with a critic network / classifier that seeks to maximize the distance between the true and fake distributions, approximating the Earth Mover's Distance (EM-Distance), i.e., the amount of effort it takes to make the generated distribution equal to the true distribution. W-loss can be expressed as:
[0090]
number
[0091]
number
[0092]
number
[0093] Thus, the generator provides useful feedback from the critic and prevents mode collapse in the vanishing gradient problem. In other words, the 1-Lipschitz continuity condition helps the training of GANs to remain more stable by ensuring that the W loss function is not only continuous and differentiable at every single value. The W loss with the 1-Lipschitz continuity condition can be expressed as follows:
[0094]
number
[0095] Example 4. Results analysis of the performance of machine learning algorithms on training samples plus synthetic samples augmented by different oversampling techniques The disclosed experiments aimed to achieve the following analyses: i) To understand whether GAN-based data augmentation methods can be utilized to generate high-quality synthetic urine samples. ii) to understand whether such methods may outperform traditional oversampling methods for improving biomarker data quality; and iii) To conclude whether GANs can provide an opportunity to improve the performance of supervised machine learning classifiers on small imbalanced datasets to predict the attenuation of kidney transplant rejection.
[0096] We implemented Table GAN, Vanilla GAN, TGAN, and CTGAN models and tested their performance in constructing high-quality synthetic data. The results showed that the disclosed GAN method performed best in generating synthetic data that closely matched the biopsy data. The CTGAN model outperformed other architectures in generating synthetic data. Therefore, we selected the CTGAN model for further analysis.
[0097] CTGAN analysis From the 265 samples in our training set, we used CTGAN to generate 1300 synthetic urine samples for further training samples. We then implemented machine learning classifiers such as Random Forest, Xgboost, and LightGBM classifiers to determine whether at least the disclosed machine learning classifiers could benefit from adding additional synthetic training data to the real training set.
[0098] We compared the performance of the oversampled classifiers with conditional TabulaGAN, SMOTE, SVMSMOTE, and ADASYN, including non-oversampled data as baseline data. We trained all classifiers with hyperparameters selected based on a comprehensive hyperparameter grid search conducted on a new training set consisting of 30% synthetic samples and 70% original samples. These classifiers were then tested on a 30% test set of the original dataset (n=114), and the performance of the classifiers was measured based on roc-auc, sensitivity, and specificity metrics. Based on the results from Figure 7, the machine learning classifiers performed well when high-quality synthetic training data was added to augment the biological data, where the GAN-based data augmentation method, in particular, helped all three machine learning classifiers to predict acute rejection in kidney transplants more accurately than other oversampling techniques.
[0099] We also analyzed the feature importance of the random forest classifier after training, and the feature importance confirmed that important biomarkers from a biological perspective still appear to contribute most in the algorithm. Thus, the potential use of this technique to create synthetic data in scenarios with small imbalanced datasets provides a useful solution for machine learning applications in the biomedical field.
[0100] Table 1 tabulates the performance of the machine learning algorithms on the kidney transplant rejection dataset disclosed in Example 1 using synthetic urine samples.
[0101] [Table 1]
[0102] Table 2 lists the performance of machine learning algorithms trained on different GAN architectures.
[0103] [Table 2]
[0104] Table 3 lists the performance of machine learning algorithms trained on different GAN architectures.
[0105] [Table 3]
[0106] Our experiments demonstrated the potential use of generative adversarial network-based data augmentation methods to create synthetic urine samples in scenarios involving small and imbalanced biomedical datasets for machine learning systems. By comparing GAN-based data augmentation methods with traditional statistical sampling techniques, we confirmed that GAN-based techniques can model the complex distribution of tabular data for more robust results of machine learning algorithms.
[0107] 8A-8C, 9, 10, 11 show the non-parametric results of the random forest-based kidney rejection score using different synthetic data generation methods (0=stable, 1=acute kidney rejection). 8A-8C collectively show the performance of the random forest models oversampled by CTGAN and baseline (FIG. 8A), the performance of the random forest models oversampled by SVMSMOTE and SMOTE (FIG. 8B), and the performance of the random forest model oversampled by ADASYN (FIG. 8C) on the kidney transplant rejection dataset using synthetic urine samples. FIG. 9 shows the non-parametric results of the random forest-based rejection score using the SMOTE synthetic data generation method to provide Q-scores. The axes in FIG. 9 represent SMOTE-generated Q-scores (Y-axis) against SMOTE phenotypes (X-axis). FIG. 10 shows the non-parametric results of the random forest-based rejection score using the original data (i.e., biological data) generation method to provide Q-scores. The axes in FIG. 10 represent the Q-scores of the original data against the original phenotype (Y-axis). FIG. 11 shows non-parametric results of a random forest-based rejection score using the GAN synthetic data generation method to provide a Q-score. The axes in FIG. 11 represent the GAN-generated Q-scores (Y-axis) against the GAN phenotype (X-axis). FIG. 12 shows non-parametric results of a random forest-based rejection score using the ADASYN synthetic data generation method to provide a Q-score. The axes in FIG. 12 represent the ADASYN-generated Q-scores (Y-axis) against the ADASYN phenotype (X-axis). FIG. 13 shows non-parametric results of a random forest-based rejection score using the SVM synthetic data generation method to provide a Q-score. The axes in FIG. 13 represent the SVM-generated Q-scores (Y-axis) against the phenotype (X-axis).
[0108] Although the present invention may be satisfied by embodiments in many different forms, as will be described in detail in connection with the preferred embodiment of the present invention, it is understood that the present invention is not intended to be limited to the specific embodiment illustrated and described herein. Numerous modifications are possible for those skilled in the art without departing from the spirit of the present invention. The scope of the present invention is determined by the appended claims and their equivalents. The abstract and title of the invention should not be construed as limiting the scope of the present invention, but rather their purpose is to enable the appropriate authorities and the general public to quickly determine the general nature of the invention. In the following claims, unless the term "means" is used, none of the features or elements recited therein should be construed as a means-plus-function limitation pursuant to 35 USC § 112 (35 USC § 112 ¶ 6).
Claims
1. 1. A system configured to balance an imbalanced dataset obtained from a biological sample, comprising: one or more computer subsystems; and one or more components executed by the one or more computer subsystems, the one or more components comprising: a first training set including data corresponding to amounts of cell-free DNA (cfDNA) biomarkers from subjects with organ dysfunction, designated as a first training input; a second training set including data corresponding to amounts of cell-free DNA (cfDNA) biomarkers from subjects without organ dysfunction designated as a second training input; and the first and second datasets are imbalanced, and the one or more computer subsystems are configured to generate a set of synthetic features for the first dataset and / or the second dataset by inputting portions of the data from the first training input and the second training input into the generative adversarial network.
2. The system of claim 1 , wherein the generative adversarial network is configured as a conditional generative adversarial network.
3. The system of claim 1 , wherein the generative adversarial network is configured as a vanilla generative adversarial network.
4. The system of claim 1 , wherein the generative adversarial network is configured as a table generative adversarial network.
5. The system of claim 1 , wherein the generative adversarial network is configured as a tabular generative adversarial network.
6. The generative adversarial network is further configured to: provide an additional training set, which includes data corresponding to the amount of a methylated cfDNA (m-cfDNA) biomarker from a subject with organ dysfunction, as an additional training input; and an additional training set including data corresponding to the amount of methylated cfDNA (m-cfDNA) biomarkers from subjects without organ dysfunction, specified as additional training input.
7. The generative adversarial network is configured to: The system of claim 1 , further trained with an additional training set including data corresponding to amounts of inflammatory biomarkers from subjects without organ damage, specified as additional training input.
8. The system of claim 7, wherein the inflammatory biomarker is a member of the chemokine (CXC motif) ligand family.
9. 9. The system of claim 8, wherein the member of the chemokine (C-X-C motif) ligand family is C-X-C motif chemokine ligand 1 (CXCL1), C-X-C motif chemokine ligand 2 (CXCL2), C-X-C motif chemokine ligand 5 (CXCL5), C-X-C motif chemokine ligand 9 (CXCL9) (MIG), or C-X-C motif chemokine ligand 10 (CXCL10) (IP-10).
10. The generative adversarial network is further configured to: The system of claim 1 , further trained with an additional training set including data corresponding to the amount of apoptotic biomarkers from subjects without organ damage, specified as additional training input.
11. The system of claim 10 , wherein the apoptosis biomarker is clusterin.
12. The generative adversarial network is configured to: and an additional training set including data corresponding to protein abundances from subjects without organ damage, specified as additional training input.
13. The system of claim 12 , wherein the protein is albumin.
14. The system of claim 12 , wherein the protein is a total protein.
15. The system of claim 1 , wherein the one or more computer subsystems are further configured to determine one or more characteristics of the composite feature for the first data set and / or the second data set.
16. The system of any one of claims 1 to 15, wherein the one or more computer subsystems are further configured to train a machine learning model using the simulated images.
17. The system of claim 16 , wherein the machine learning model is trained on the first data input and the second data input.
18. 20. The system of claim 17, wherein the machine learning model is trained on the first data input and the second data input, rather than on the set of synthetic features.
19. The system of claim 16 , wherein the machine learning model is a CTGAN.
20. The system of claim 16 , wherein the machine learning model is SMOTE.
21. The system of claim 16 , wherein the machine learning model is SVM-SMOTE.
22. The system of claim 16 , wherein the machine learning model is ADASYN.
23. The system of claim 1 , wherein the biological sample is urine.
24. The system of claim 1 , wherein the biological sample is blood.
25. The system of claim 1 , wherein the organ is an allograft and the damage is caused by a rejection reaction of the allograft by the subject.
26. The system of claim 1 , wherein the organ is a kidney, a pancreas, a heart, a lung, or a liver.
27. 27. The system of claim 26, wherein the organ is a kidney.
28. 27. The system of claim 26, wherein the disorder is chronic kidney injury (CKI) or acute kidney injury (AKI).
29. The system of claim 1 , wherein the disorder is caused by a viral infection suffered by the subject.
30. The system of claim 1, wherein the viral infection is caused by SARS-CoV-2, CMV, or BKV.
31. The system of claim 1 , wherein the disorder is a cancer that harms the organ.
32. The system of claim 1 , wherein the subject is a human.
33. 1. A system configured to analyze a dataset obtained from a biological sample, comprising: one or more computer subsystems; and one or more components executed by the one or more computer subsystems, the one or more components including a generative adversarial network trained with a training set corresponding to an amount of cfDNA from a subject; and the one or more computer subsystems configured to generate a synthetic dataset from the biological sample by inputting a subset of the training data to the generative adversarial network.
34. 34. The system of claim 33, wherein the subset of the training data is annotated with biological states.
35. 34. The system of claim 33, wherein at least a subset of the training data is annotated with a biological state of acute rejection.
36. 34. The system of claim 33, wherein at least a subset of the training data is annotated with a biological state of chronic kidney injury (CKI) or acute kidney injury (AKI).
37. 34. The system of claim 33, wherein at least a subset of the training data is annotated with a biological state of COVID-19.
38. 34. The system of claim 33, wherein at least a subset of the training data is annotated with a healthy or stable biological state.
39. 34. The system of claim 33, wherein the cfDNA is from a urine sample.
40. 34. The system of claim 33, wherein the cfDNA is from a blood sample or a plasma sample.
41. 34. The system of claim 33, wherein the generative adversarial network is further trained with an additional training set including data corresponding to the amount of a methylated cfDNA (m-cfDNA) biomarker from a subject.
42. 34. The system of claim 33, wherein the generative adversarial network is further trained with an additional training set including data corresponding to amounts of inflammatory biomarkers from a subject.
43. 43. The system of claim 42, wherein the inflammatory biomarker is a member of the chemokine (CXC motif) ligand family.
44. 44. The system of claim 43, wherein the member of the chemokine (C-X-C motif) ligand family is C-X-C motif chemokine ligand 1 (CXCL1), C-X-C motif chemokine ligand 2 (CXCL2), C-X-C motif chemokine ligand 5 (CXCL5), C-X-C motif chemokine ligand 9 (CXCL9) (MIG), or C-X-C motif chemokine ligand 10 (CXCL10) (IP-10).
45. 34. The system of claim 33, wherein the generative adversarial network is further trained with an additional training set including data corresponding to amounts of apoptotic biomarkers from a subject.
46. 46. The system of claim 45, wherein the apoptosis biomarker is clusterin.
47. 34. The system of claim 33, wherein the generative adversarial network is further trained with an additional training set including data corresponding to protein abundances.
48. 48. The system of claim 47, wherein the protein is albumin.
49. 48. The system of claim 47, wherein the protein is total protein.
50. 34. The system of claim 33, wherein the subject is a human.
51. 1. A non-transitory computer-readable medium storing program instructions executable on one or more computer systems for performing a computer-implemented method for generating a simulated image of a specimen, the computer-implemented method comprising:
1. A non-transitory computer-readable medium, comprising: one or more computer subsystems; and one or more components executed by the one or more computer subsystems, the one or more components including a generative adversarial network trained with a training set corresponding to an amount of cfDNA from a subject, the one or more computer subsystems configured to generate a synthetic dataset from the biological sample by inputting a subset of the training data into the generative adversarial network.
52. 1. A non-transitory computer-readable medium storing program instructions executable on one or more computer systems for performing a computer-implemented method for generating a simulated image of a specimen, the computer-implemented method comprising: one or more computer subsystems; and one or more components executed by the one or more computer subsystems, the one or more components including a generative adversarial network trained with a training set including data corresponding to amounts of cell-free DNA (cfDNA) biomarkers from subjects, the first training set including data corresponding to amounts of cell-free DNA (cfDNA) biomarkers from subjects with organ dysfunction designated as a first training input, and a second training set including data corresponding to amounts of cell-free DNA (cfDNA) biomarkers from subjects without organ dysfunction designated as a second training input; the first and second datasets are imbalanced, and the one or more computer subsystems are configured to generate a set of synthetic features for the first dataset and / or the second dataset by inputting portions of the data from the first training input and the second training input into the generative adversarial network.