A method for determining disease risk combining downsampling of class-imbalanced sets with survival analysis

Downsampling class-imbalanced datasets in survival analysis improves sensitivity and specificity by balancing the majority and minority classes, addressing the imbalanced prediction issues in disease risk assessment, particularly in low-prevalence diseases.

JP7680950B2Active Publication Date: 2025-05-21SOMALOGIC OPERATING CO INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021530139
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-12-21
Filing Date
2019-11-21
Publication Date
2025-05-21
Estimated Expiration
2039-11-21

AI Technical Summary

Technical Problem

Existing methods for identifying disease-related biomarkers face challenges with high-dimensional data, leading to imbalanced sensitivity and specificity in survival analysis due to class imbalance, particularly in low-prevalence diseases, where the majority class dominates feature selection, often misclassifying the minority class.

Method used

Downsampling the majority class in class-imbalanced datasets to balance sensitivity and specificity by generating a survival model through cross-validation, using techniques like Cox proportional hazards analysis and elastic net penalization.

Benefits of technology

Improves the AUC, sensitivity, and specificity of survival models by equally considering diagnosed and undiagnosed individuals, enhancing the predictive accuracy of disease risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007680950000008
    Figure 0007680950000008
  • Figure 0007680950000009
    Figure 0007680950000009
  • Figure 0007680950000010
    Figure 0007680950000010
Patent Text Reader

Abstract

1. A method for downsampling a class-imbalanced set using survival analysis, comprising: obtaining a class-imbalanced dataset, the class-imbalanced dataset including biological data from a plurality of subjects, the biological data for each subject including an observation, a time value, and a plurality of clinical measurements, the biological data being classified as part of a majority data class or a minority data class, the majority data class including more observations than the minority data class; downsampling the class-imbalanced dataset, the downsampling producing a majority data class including an equal or substantially equal number of observations as the minority data class; and performing cross-validation on the downsampled dataset using survival analysis to generate a survival model, the observations either including an event or not including an event at a particular time value.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 773,028, filed November 29, 2018, and U.S. Provisional Patent Application No. 62 / 783,733, filed December 21, 2018, the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates generally to the field of disease risk assessment, and more specifically to systems and methods for processing electronic data to assess disease risk. [Background technology]

[0003] Methods for identifying biomarkers associated with the risk of various disease-related conditions or events, such as cardiovascular events, diabetes diagnosis, and various cancer types, have improved, primarily due to the discovery of high-throughput techniques such as gene sequencing, transcriptomics, proteomics, and metabolomics. However, these technologies also complicate the problem by generating high-dimensional data that represent complex biological processes that can make the extraction of meaningful biomarker signatures difficult.

[0004] When the main goal is to correctly identify individuals who experience a disease-related condition or event within a specified time period, an analysis that typically only uses a classification approach can be enhanced by framing it as a special type of classification problem that incorporates a survival modeling approach in combination with a classification tool. However, survival analysis can suffer from an imbalance in the number of patients who experience a disease-related condition or event and those who do not. Predictive classifiers are known to generally perform poorly on imbalanced data, since the models are trained to be accurate "as often as possible." This effect arises because a larger majority class drives the features selected for the model. The majority class is still accurately predicted, while the minority class may be frequently misclassified. However, the sensitivity and specificity become imbalanced, with one maximized over the other, relying on groups with a larger number of observations. In health outcome modeling, it is common for disease prevalence in a cohort to be low, forming a minority class. In such situations, specificity is maximized at the expense of sensitivity. This becomes problematic when the goal is to identify as many individuals as possible who are at risk for developing a condition or event.

[0005] Thus, there continues to be a need for alternative methods for improved methods for identifying molecular signatures or biomarkers for particular diseases or conditions. The present disclosure meets such a need by providing methods for improved biomarker discovery. Summary of the Invention

[0006] According to some aspects of the present disclosure, the disclosed systems and methods relate to downsampling of majority classes, i.e., classes with more observations, of class-imbalanced datasets that include time values, to improve sensitivity and specificity in survival analysis. The purpose of downsampling is to "bias" the classifier to equally consider diagnosed and undiagnosed individuals in order to balance the sensitivity and specificity of the model.

[0007] In one embodiment, the method includes obtaining a class imbalance dataset, the class imbalance dataset being and performing cross-validation on the downsampled dataset using survival analysis to generate a survival model. Disclosed is a method for generating a survival model comprising: obtaining a class-imbalanced dataset including biological data from a plurality of subjects, the biological data for each subject including an observation, a time value, and a plurality of clinical measurements, the biological data being classified as part of a majority data class or a minority data class, the majority data class including more observations than the minority data class; down-sampling the class-imbalanced dataset to generate a downsampled dataset, the down-sampling generating a majority data class including an equal or substantially equal number of observations as the minority data class; and performing cross-validation on the downsampled dataset using survival analysis to generate a survival model, the observations including an event or no event at a particular time value.

[0008] According to aspects of the present disclosure, the area under the curve (AUC), sensitivity, specificity, and / or C-index of a survival model is closer to 1 than the AUC, sensitivity, specificity, and / or C-index of a survival model where the class-imbalanced dataset was not downsampled prior to survival analysis.

[0009] In other examples, the class imbalanced dataset is a survival dataset and / or the event is a disease, disorder, or condition of interest. In further examples, the survival analysis is selected from the group consisting of Cox proportional hazards analysis, random forest analysis, accelerated failure time analysis, and any combination thereof, and includes fitting machine learning techniques such as penalized regression techniques. The method may further include elastic net penalization.

[0010] In other embodiments, the cross-validation is at least 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, 15-fold, 16-fold, 17-fold, 18-fold, 19-fold, or 20-fold cross-validation. In other embodiments, the survival model comprises 5 to 1000 features, each feature selected from the group consisting of a protein measurement, a clinical factor, and combinations thereof. The clinical factor is selected from the group consisting of age, weight, blood pressure, height, BMI, cholesterol, sex, and combinations thereof.

[0011] In further embodiments, the clinical measurements are selected from proteomic measurements, genomic measurements, transcriptomic measurements, metabolomic measurements, and combinations thereof.Furthermore, the cross-validation is selected from k-fold cross-validation, generalized Monte Carlo cross-validation, and leave-p-out cross-validation or bootstrapping techniques.

[0012] According to aspects of the present disclosure, the majority data class is 95% of the class imbalanced dataset and the minority data class is 5% of the class imbalanced dataset, or the majority data class is 90% of the class imbalanced dataset and the minority data class is 10% of the class imbalanced dataset, or the majority data class is 85% of the class imbalanced dataset and the minority data class is 15% of the class imbalanced dataset, or the majority data class is 80% of the class imbalanced dataset and the minority data class is 20% of the class imbalanced dataset, or the majority data class is 75% of the class imbalanced dataset and the minority data class is 25% of the class imbalanced dataset, or the majority data class is 70% of the class imbalanced dataset and the minority data class is 30% of the class imbalanced dataset, or the majority data class is 65% of the class imbalanced dataset and the minority data class is 35% of the class imbalanced dataset, or the majority data class is 60% of the class imbalanced dataset and the minority data class is 40% of the class imbalanced dataset.

[0013] In another embodiment, a method includes downsampling a class imbalanced dataset. and performing cross-validation on the downsampled dataset using survival analysis to generate a survival model; the observations either include an event or no event at a particular time value; the class-imbalanced dataset includes biological data from a plurality of subjects, the biological data for each subject including an observation, a time value, and measurements of a plurality of proteins, the biological data being classified as part of a majority data class or a minority data class, the majority data class including more observations than the minority data class.

[0014] According to aspects of the present disclosure, the AUC, sensitivity, specificity, and / or C-index of a survival model is closer to 1 than the AUC, sensitivity, specificity, and / or C-index of a survival model where the class-imbalanced dataset was not downsampled prior to survival analysis.

[0015] In the examples of the present disclosure, the AUC is calculated based on determining whether a subject has an event by a particular time point.

[0016] Also disclosed is a computer-implemented method for determining risk of disease, comprising obtaining a class-imbalanced dataset, the class-imbalanced dataset including biological data from a plurality of subjects, the biological data for each subject including an observation, a time value, and a plurality of clinical measurements, the biological data being classified as part of a majority data class or a minority data class, the majority data class including more observations than the minority data class; downsampling the class-imbalanced dataset to generate a downsampled dataset, the downsampling generating a majority data class including an equal or substantially equal number of observations as the minority data class; and performing cross-validation on the downsampled dataset using survival analysis to generate a survival model; the observations either include an event or do not include an event at a particular time value; and the downsampling and cross-validation steps are computed using a computer system.

[0017] According to aspects of the present disclosure, the AUC, sensitivity, specificity, and / or C-index of a survival model is closer to 1 than the AUC, sensitivity, specificity, and / or C-index of a survival model where the class-imbalanced dataset was not downsampled prior to survival analysis.

[0018] Also disclosed is a computer readable program storage device tangibly embodying a program of instructions executable by the computer to perform the method steps of a method for determining risk of disease comprising obtaining a class-imbalanced dataset, the class-imbalanced dataset including biological data from a plurality of subjects, the biological data for each subject including an observation, a time value, and a plurality of clinical measurements, the biological data being classified as part of a majority data class or a minority data class, the majority data class including more observations than the minority data class; downsampling the class-imbalanced dataset to generate a downsampled dataset, the downsampling generating a majority data class including an equal or substantially equal number of observations as the minority data class; and performing cross-validation on the downsampled dataset using survival analysis to generate a survival model; the observations including an event or not including an event at a particular time value.

[0019] According to an embodiment of the present disclosure, the AUC, sensitivity, specificity, and / or C-index of a survival model are measured using a 1:1 ratio (AUC), 1:2 ratio (S1), and 1:3 ratio (S2), where the class-imbalanced dataset was not downsampled prior to survival analysis. UC, sensitivity, specificity, and / or C-index of survival models are closer to 1.

[0020] Also disclosed is a computing system for determining risk of disease, the computing system including a memory for storing programmed instructions and a processor configured to execute the programmed instructions to perform operations including obtaining a class-imbalanced dataset, the class-imbalanced dataset including biological data from a plurality of subjects, the biological data for each subject including an observation, a time value, and a plurality of clinical measurements, the biological data being classified as part of a majority data class or a minority data class, the majority data class including more observations than the minority data class; down-sampling the class-imbalanced dataset to generate a down-sampled dataset, the down-sampling generating a majority data class including an equal or substantially equal number of observations as the minority data class; and performing cross-validation on the down-sampled dataset using survival analysis to generate a survival model, the observation including an event or not including an event at a particular time value.

[0021] According to aspects of the present disclosure, the AUC, sensitivity, specificity, and / or C-index of a survival model is closer to 1 than the AUC, sensitivity, specificity, and / or C-index of a survival model where the class-imbalanced dataset was not downsampled prior to survival analysis.

[0022] Also disclosed is a non-transitory computer-readable medium having stored thereon instructions executable by a processor to perform the operations of obtaining a class-imbalanced dataset, the class-imbalanced dataset including biological data from a plurality of subjects, the biological data for each subject including an observation, a time value, and a plurality of clinical measurements, the biological data being classified as part of a majority data class or a minority data class, the majority data class including more observations than the minority data class; down-sampling the class-imbalanced dataset to generate a down-sampled dataset, the down-sampling generating a majority data class including an equal or substantially equal number of observations as the minority data class; and performing cross-validation on the down-sampled dataset using survival analysis to generate a survival model, the observations including an event or not including an event at a particular time value.

[0023] According to aspects of the present disclosure, the AUC, sensitivity, specificity, and / or C-index of a survival model is closer to 1 than the AUC, sensitivity, specificity, and / or C-index of a survival model where the class-imbalanced dataset was not downsampled prior to survival analysis.

[0024] Also disclosed is a computer-implemented method for determining risk of disease, comprising: receiving, on a computer, a class-imbalanced dataset, the class-imbalanced dataset including biological data from a plurality of subjects, the biological data for each subject including an observation, a time value, and a plurality of clinical measurements, the biological data being classified as part of a majority data class or a minority data class, the majority data class including more observations than the minority data class; down-sampling, on a computer, the class-imbalanced dataset to generate a down-sampled dataset, the down-sampling generating a majority data class including an equal or substantially equal number of observations as the minority data class; and performing, on a computer, cross-validation on the down-sampled dataset using survival analysis to generate a survival model, the observations including an event or no event at a particular time value.

[0025] According to aspects of the present disclosure, the AUC, sensitivity, specificity, and / or C-index of a survival model is closer to 1 than the AUC, sensitivity, specificity, and / or C-index of a survival model where the class-imbalanced dataset was not downsampled prior to survival analysis. [Brief description of the drawings]

[0026] [Figure 1] 1 illustrates an example of a networked computing environment in which methods, systems, and other aspects of the present disclosure may be implemented. [Diagram 2] FIG. 1 illustrates a high-level architecture diagram of a disease risk analysis platform for acquiring and processing clinical data in accordance with the present disclosure. [Diagram 3] Kaplan-Meier survival curves for myocardial infarction (MI) in the HUNT3 CHD subcohort. [Figure 4-1]Kaplan-Meier survival curves for MI in the test set, stratified by predicted event. For each method, the test set is split into high-risk and average-risk individuals using the threshold identified by cross-validation. Kaplan-Meier curves are then calculated for both groups. The logistic regression model results in everyone being predicted to be at low risk, hence there is only one survival curve. [Figure 4-2] Continued from Figure 4-1. [Figure 5-1] Kaplan-Meier survival curves for MI in the test set are shown, using the downsampled Cox elastic net model to predict MI ≤ 4 years. Various thresholds for classifying individuals as high risk were investigated. [Figure 5-2] Continued from Figure 5-1. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0027] Unless otherwise noted, technical terms are used according to conventional usage. Definitions of common terms in molecular biology can be found in Benjamin Lewin, Genes V, published by Oxford University Press, 1994 (ISBN 0-19-854287-9), Kendrew et al. (eds.), The Encyclopedia of Molecular Biology, published by Blackwell Science Ltd., 1994 (ISBN 0-632-02182-9), and Robert A. Meyers (ed.), Molecular Biology and Biotechnology: a Comprehensive Desk Reference, published by VCH Publishers, Inc., 1995 (ISBN 1-56081-569-8). Unless otherwise explained, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this disclosure belongs. The singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. "Comprising A or B" means including A, or B, or A and B. It should be further understood that all base or amino acid sizes and all molecular weight or molecular mass values ​​given for nucleic acids or polypeptides are approximate and are provided for illustration purposes.

[0028] Additionally, ranges provided herein are understood to be shorthand notations for all values ​​within that range. For example, the range of 1 to 50 is understood to include any number, combination of numbers, or subranges (as well as fractions thereof) from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50. Any concentration range, percentage range, ratio range, or integer range, unless otherwise indicated, includes any integer within the recited range, and, where appropriate, fractions thereof (such as tenths and hundredths of an integer). Also, any numerical range recited herein for any physical characteristic, such as polymer subunits, size or thickness, should be understood to include every integer in the recited range, unless otherwise indicated. As used herein, "about" or "consisting essentially of" means ±20% of the recited range, value, or structure, unless otherwise indicated. As used herein, the terms "include" and "comprise" are open ended and are used synonymously.

[0029] Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, suitable methods and materials are described below. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including explanations of terms, will control. In addition, the materials, methods, and examples are illustrative only and are not intended to be limiting.

[0030] As used herein, "SOMAmers" or slow off-rate modified aptamers refer to aptamers with improved off-rate properties. SOMAmers can be generated using the improved SELEX method described in U.S. Patent No. 7,947,447, entitled "Method for Generating Aptamers with Improved Off-Rates."

[0031] The terms "biological sample", "sample" and "test sample" are used interchangeably herein and refer to any material, biological fluid, tissue or cell obtained from or otherwise derived from an individual. This includes blood (including whole blood, white blood cells, peripheral blood mononuclear cells, buffy coat, plasma and serum), sputum, tears, mucus, nasal washings, nasal aspirates, exhaled breath, urine, semen, saliva, peritoneal washings, ascites, cyst fluid, cerebrospinal fluid, amniotic fluid, glandular fluid, lymph, nipple aspirate, bronchial aspirate (e.g., bronchoalveolar lavage), bronchial scrapings, synovial fluid, joint aspirate, organ secretions, cells, cell extracts and cerebrospinal fluid. This also includes all the experimentally separated fractions mentioned above. For example, blood samples can be fractionated into serum, plasma or into fractions containing specific types of blood cells, such as red blood cells or white blood cells (leukocytes). In some embodiments, the sample may be a combination of samples from an individual, such as a combination of tissue and liquid samples. The term "biological sample" also includes materials containing homogenized solid material, such as from a fecal sample, a tissue sample, or a tissue biopsy. The term "biological sample" also includes materials derived from tissue culture or cell culture. Any suitable method for obtaining a biological sample may be used, and exemplary methods include, for example, phlebotomy, swabs (e.g., buccal swabs), and fine needle aspiration cytology procedures. Exemplary tissues that can be aspirated include lymph nodes, lung, lung lavage, BAL (bronchoalveolar lavage), thyroid, breast, pancreas, and liver. Samples may also be collected, for example, by microdissection (e.g., laser capture microdissection (LCM) or laser microdissection (LMD)), bladder washing, smear (e.g., PAP smear), or ductal lavage. A "biological sample" obtained or derived from an individual includes any such sample that has been processed in any suitable manner after being obtained from the individual.

[0032] As used herein, "biological data" refers to any data derived from a biological sample, including, but not limited to, proteomic data collected using aptamers specific to protein targets, optionally in multiplexed aptamer-based assays.

[0033] As used herein, "clinical factors" refer to physiological attributes that may be associated with an increased risk of a medical condition or event. Clinical factors include age, weight, blood pressure, height, BMI, cohort, and other factors. These include, but are not limited to, cholesterol, and gender.

[0034] As used herein, "class imbalance" refers to a property of a dataset that describes when a set of data is classified into two or more classes, the two or more classes having a substantially unequal number of observations.

[0035] As used herein, "cross-validation" refers to any model building and validation technique for evaluating the performance of a model on the data used to build the model and methods in which the results of statistical analyses are generalized to independent data sets, including, but not limited to, k-fold cross-validation, Monte Carlo cross-validation, and leave-p-out cross-validation (where p can range from 1 to the total number of samples minus 1).

[0036] As used herein, "downsampling" refers to subsetting a class' data with more observations, i.e., the majority data class, to reduce class imbalance.

[0037] As used herein, "equivalent" or "substantially equivalent" refers to a difference between the compared classes where the difference in the number of observations is less than 10%.

[0038] As used herein, a "feature" refers to a measurable trait or characteristic of an object in a dataset. Features include, but are not limited to, protein measurements and clinical factors.

[0039] As used herein, "majority data class" refers to the class with a greater number of observations in a class-imbalanced dataset having two classes.

[0040] As used herein, a "minority data class" refers to the class with a smaller number of observations in a class-imbalanced dataset having two classes.

[0041] As used herein, "survival analysis" refers to any modeling of time-to-event data. Methods of survival analysis can be used for any time-to-event outcome, such as time to MI, onset of diabetes, onset of various forms of cancer, etc. Survival analysis includes, but is not limited to, Cox proportional hazards analysis, random forest analysis, and accelerated failure time analysis.

[0042] As used herein, a "survival dataset" refers to any dataset that includes both a time value and an event status value that indicates whether an event of interest occurred during the time period that the subject was observed.

[0043] In survival analysis, class imbalance poses a major problem, where the number of disease (or event)-free individuals exceeds the number of disease-affected individuals within a given time frame. This imbalance can lead to inaccurate predictions of risk for individuals at high risk of the disease. Downsampling mitigates this problem by balancing the number of individuals in the minority and majority classes, thus improving the detection and selection of features associated with individuals in the minority class and their putative influence on the risk of developing a disease or event.

[0044] One context in which downsampling of class-imbalanced datasets for survival analysis was demonstrated to improve AUC was in protein biomarkers generated by the SOMAscan® proteomics assay, which was used to identify circulating protein biomarkers associated with risk of cardiovascular events in patients with stable coronary heart disease (CHD). The resulting models provide superior capabilities to existing clinical risk tools and have broad applicability and generalizability across composite cardiovascular endpoints.

[0045] This disclosure describes a targeted model for predicting secondary MI among patients with stable CHD. Proteomic data was used to identify patients likely to experience secondary MI within 4 years of blood draw in patients with stable CHD. In addition to the proteomic signal, the data contains information on whether specific cardiovascular events occurred during the observation period and the length of time to either a) the event, or b) the end of the study due to other factors. These time-to-event data make the problem highly suitable for survival analysis approaches.

[0046] If the primary goal is to correctly identify individuals who will experience an MI event within 4 years, the analysis can be reframed as a classification problem. In this case, if the event occurred prior to 4 years, the individual is in the “positive” class, and if the individual remained in the study beyond the 4-year time frame without MI, the individual is labeled as the “negative” class. By incorporating time to MI in the deployment of the classifier, survival analysis tools improve the predictive accuracy of the model (compared to standard classification models) because the survival model “uses all the information”. This reframing also allows standard classification metrics such as AUC and confusion matrices to be used to evaluate model performance. Although this method of evaluating survival models is not a traditional approach, event-specific classification offers many advantages to the clinical setting. Labeling patients as “positive” or “negative” is more easily understood among a broad audience (compared to hazard ratios or probabilities, for example). By improving this understanding of prognostic testing, clinicians can provide more precise and targeted medical management. However, like standard classification modeling, this approach to survival analysis can suffer from an imbalance between patients who experience an event and those who do not.

[0047] For example, only 8.1% of individuals in the subcohort analyzed in Example 1 will develop a secondary MI within 4 years, while more than eight times as many participants (66.9%) will survive 4 years or more without an event. The purpose of downsampling is to "bias" the classifier to give equal consideration to diagnosed and undiagnosed individuals in order to balance the sensitivity and specificity of the model. Although resampling techniques have been applied in various machine learning techniques, class imbalance is an unexplored topic in machine learning using survival modeling techniques.

[0048] In Example 1, downsampling is combined with a Cox proportional hazards elastic net regression model to evaluate the prediction of MI events within 4 years of first blood draw.

[0049] As is evident from Example 1, the performance of survival analyses, such as the Cox proportional hazards elastic net model (i.e., the "Coxnet" model), can be improved by downsampling the data during modeling. The present disclosure effectively shows that the downsampled Coxnet model outperforms the standard Coxnet model, the downsampled elastic net logistic regression model, and the standard elastic net logistic regression model.

[0050] In addition to downsampling, there are other methods for dealing with class imbalance that can also be incorporated into survival models. For example, case weighting, simple oversampling, or more complex oversampling techniques such as synthetic minority oversampling techniques (SMOTE) can be considered in traditional survival analysis, as well as in extended machine learning methods such as survival random forests.

[0051] Although Example 1 details the combination of downsampling with survival analysis in the context of predicting MI events within a specified time frame, the methods disclosed herein can be applied to any prediction of risk of a medical condition or disease-related event within a selected time frame.

[0052] FIG. 1 is a block diagram of a networked computing environment 100 for processing electronic data to determine risk of disease, for example, by downsampling class imbalance data, according to an aspect of the disclosure. As shown in FIG. 1, the networked computing environment 100 may include a disease risk analysis platform 102, including a server system 104 and an electronic database 106. The server system 104 may store and execute software modules, algorithms, or other subsystems of the disease risk analysis platform 102 for use over an electronic network 108, such as the Internet. A user may access the disease risk analysis platform 102 over the electronic network 108 by a user device 110, such as a computing device. The user device 110 may enable a user to display a web browser for accessing the disease risk analysis platform 102 hosted by the server system 104 over the electronic network 108. The user device 110 may be any type of device for accessing a web page, such as a personal computing device, a mobile computing device, or the like. A source device 112 may provide and / or receive data to the disease risk analysis platform 102 over the electronic network 108. The source device 112 may be any type of device for accessing web pages, such as a personal computing device, a mobile computing device, or the like.

[0053] FIG. 1 is presented as an example only. Other examples are possible and may differ from the networked computing environment 100 of FIG. 1. Also, the number and arrangement of devices and networks shown in the networked computing environment 100 are presented as an example. In practice, there may be additional devices, fewer devices and / or networks, different devices and / or networks, or devices and / or networks arranged differently than those shown in the networked computing environment 100. Furthermore, two or more devices shown in FIG. 1 may be implemented within a single device, and a single device shown in FIG. 1 may be implemented as multiple distributed devices. Additionally or alternatively, one or more user devices and / or a server system of the networked computing environment 100 may perform one or more functions of the server system 104 and / or the disease risk analysis platform 102.

[0054] FIG. 2 illustrates an example computer architecture 200 for processing electronic data to determine risk of disease. Specifically, FIG. 2 illustrates an example computer architecture 200 configured to combine downsampling of class-imbalanced sets with survival analysis, according to one or more embodiments of the present disclosure. As illustrated in the computer architecture 200 of FIG. 2, the server system 104 of the disease risk analysis platform 102 may include a data acquisition module 212, a downsampling module 214, and a cross-validation module 216. The disease risk analysis platform 102 may further include one or more databases or data stores, whether accessed locally or remotely. For example, as illustrated in FIG. 2, the disease risk analysis platform 102 may include a class-imbalanced dataset 206 including majority class data 202 and minority class data 204. The ... 8 and survival model 210. It should be understood that one or more of the data acquisition module 212, the downsampling module 214, the cross-validation module 216, the class imbalanced dataset 206, the downsampled dataset 208, and the survival model 210 may have some or all of its functionality and content stored or executed locally, remotely, or both locally and remotely, and that its functionality may be combined or distributed with other components of the platform.

[0055] In one embodiment of the exemplary computer architecture 200, the data acquisition module 212 may receive a class imbalanced dataset 206 including majority class data 202 and minority class data 204 from the user device 110 or the source device 112. The class imbalanced dataset 206 may be processed by a downsampling module 214 to generate a downsampled dataset 208. The downsampled dataset 208 may be processed by a cross-validation module 216 to generate a survival model 210. The survival model 210 may then be transmitted to the user device 100 and / or the source device 112 via the electronic network 108.

[0056] When programmable logic is used, such logic can be executed on a commercially available processing platform or a dedicated device. Those skilled in the art will appreciate that embodiments of the disclosed subject matter can be practiced with a variety of computer system configurations, including multi-core multi-processor systems, minicomputers, mainframe computers, computers linked or clustered with distributed functions, as well as pervasive or small computers that can be embedded within virtually any device.

[0057] For example, at least one processor device and memory may be used to implement the above-described embodiments. The processor device may be a single processor, multiple processors, or a combination thereof. The processor device may include one or more processor "cores."

[0058] Various embodiments of the present disclosure may be implemented using a processor device, as described in the examples of Figures 1 and 2 above. After reading this description, it will be apparent to one of ordinary skill in the art how to implement embodiments of the present disclosure using other computer systems and / or computer architectures. Although operations may be described as sequential processes, some of the operations may in fact be performed in parallel, simultaneously, and / or in a distributed environment, and with program code stored locally or remotely for access by a single or multi-processor machine. Additionally, in some embodiments, the order of operations may be rearranged without departing from the spirit of the disclosed subject matter.

[0059] It should be understood that the disease risk analysis platform 102 and / or any device used to access the disease risk analysis platform 102, such as the user device 110 or the source device 112, may include a central processing unit (CPU). Such a CPU may be any type of processor device, including, for example, any type of special purpose or general purpose microprocessor device. As will be appreciated by those skilled in the art, the CPU may also be a single processor of a multi-core / multi-processor system, a system operating alone or in a cluster of computing devices, in a cluster or in a server farm. The CPU may be connected to a data infrastructure, for example, a bus, a message queue, a network, or a multi-core message passing scheme.

[0060] It should be further understood that the disease risk analysis platform 102 and / or any device used to access the disease risk analysis platform 102, such as the user device 110 or the source device 112, may also include a main memory, such as random access memory (RAM), and may also include a secondary memory. The secondary memory, such as read only memory (ROM), may be, for example, a hard disk drive or a removable storage drive. Such a removable storage drive may include, for example, a floppy disk drive, a magnetic tape drive, an optical disk drive, a flash memory, or the like. The removable storage drive in this example reads from and / or writes to a removable storage unit in a well-known manner. The removable storage unit may include a floppy disk, a magnetic tape, an optical disk, etc., that is read and written by the removable storage drive. As will be appreciated by those skilled in the art, the removable storage unit generally includes a computer usable storage medium having computer software and / or data stored thereon.

[0061] In alternative embodiments, the secondary memory may include other similar means that allow computer programs or other instructions to be loaded into the device. Examples of such means may include program cartridges and cartridge interfaces (such as those found in video game machines), removable memory chips (such as EPROMs, or PROMs) and associated sockets, and other removable storage units and interfaces that allow software and data to be transferred from the removable storage units to the device.

[0062] It should further be appreciated that the disease risk analysis platform 102 and / or any device used to access the disease risk analysis platform 102, e.g., the user device 110 or the source device 112, may also include a communications interface ("COM"). The communications interface allows software and data to be transferred between the device and external devices. The communications interface may include a modem, a network interface (such as an Ethernet card), a COM port, a PCMCIA slot and card, or the like. The software and data transferred via the communications interface may be in the form of signals, which may be electrical, electromagnetic, optical, or other signals that can be received by the communications interface. These signals may be provided to the communications interface via the communications path of the device, which may be implemented using, for example, wire or cable, fiber optics, a phone line, a cellular phone link, an RF link, or other communications channel.

[0063] The hardware elements, operating systems, and programming languages ​​of such devices are conventional in nature and are presumed to be fully familiar to those skilled in the art. The device used to access the disease risk analysis platform may also include input and output ports for connecting to input and output devices such as a keyboard, mouse, touch screen, monitor, display, etc. Of course, the functions of the various servers may be implemented in a distributed manner on many similar platforms to distribute the processing load. Alternatively, the server may be implemented by appropriate programming of one computer hardware platform.

[0064] The systems, apparatus, devices, and methods disclosed herein are described in detail by way of example and with reference to the figures. The examples discussed herein are merely examples and are presented to supplement the explanation of the apparatus, devices, systems, and methods described herein. Features or components shown in the figures or described below may be used in any manner that is not specifically limited. Unless otherwise specified as essential in the drawings, no particular implementation of any of the apparatus, devices, systems, or methods should be considered essential. For ease of reading and clarity, certain components, modules, or methods may be described only with respect to certain figures. In this disclosure, any identification of a particular technique, arrangement, etc. is either related to the particular example presented, or is merely a general description of such technique, arrangement, etc. Identification of specific details or examples is not intended and should not be construed as essential or limiting unless specifically so specified. Failure to specifically describe a combination or subcombination of components should not be understood as indicating that any combination or subcombination is not possible. It will be understood that changes to the disclosed and described examples, arrangements, configurations, components, elements, apparatus, devices, systems, methods, etc. can be made and may be desirable for particular applications. Additionally, for any described method, whether or not the method is described in conjunction with a flow diagram, unless otherwise specified or required by context, it should be understood that any explicit or implicit ordering of steps performed in performing the method does not imply that those steps must be performed in the order presented, but instead may be performed in another order or in parallel.

[0065] Throughout this disclosure, references to components or modules generally refer to items that can be logically grouped together to perform a function or group of related functions. The components and modules can be implemented in software, hardware, or a combination of software and hardware. The term "software" is used broadly to include executable code, e.g., machine-executable or machine-interpretable instructions, as well as data structures, data stores, and computational instructions stored in any suitable electronic format, including firmware, and embedded software. The terms "information" and "data" are used broadly to include a wide variety of electronic information, e.g., executable code; content, e.g., text, video data, and audio data; and various codes or flags. The terms "information," "data," and "content" may be used interchangeably where the context permits. EXAMPLES

[0066] The following examples are presented to more fully illustrate some embodiments of the present invention. However, they should not be interpreted as limiting the broad scope of the present invention in any way. Those skilled in the art can easily adopt the principles underlying this discovery to design various mixtures without departing from the spirit of the present invention.

[0067] Example 1 This example provides an explanation of downsampling combined with a Cox proportional hazards elastic net regression model to evaluate the prediction of myocardial infarction (MI) events within 4 years of first blood draw, as can be executed within the exemplary data risk analysis platform in Figure 2.

[0068] The objectives of this example are at least twofold: 1) to select and identify features that are predictive of both the minority and majority classes, and 2) to derive effect sizes estimated such that the risk of the minority class is well predicted. In contrast, we examined the predictive ability of a logistic regression elastic net model, with and without downsampling, and a Cox elastic net model without downsampling.

[0069] Materials and Methods - Dataset The samples used in the analysis were from a subcohort of the HUNT3 study, a prospective Norwegian randomized controlled trial. This was a cohort study and included blood samples collected from study participants and follow-up health information. The CHD subcohort has been described previously (Peter Ganz, et al. Development and validation of a protein-based risk score for cardiovascular outcomes among patients with stable coronary heart disease. Jama, 315(23):2532-2541, 2016), and inclusion criteria included evidence of existing but stable CHD via a history of MI >6 months prior, stenosis, inducible ischemia, or previous coronary revascularization. Plasma samples were assayed using the SOMAscan® Assay (SomaLogic, Inc; Boulder, CO, USA), which measures relative protein abundance using Slow Off-rate Modified Aptamer (SOMAmer®) reagents. The V4 assay measures 5,220 protein analytes and is an established platform for protein biomarker discovery.

[0070] In the subcohort, 8.1% of patients experienced a secondary MI within 4 years (Table 1). Kaplan-Meier survival curves for MI in the CHD subcohort are shown in Figure 3. Kaplan-Meier curves are an empirical nonparametric method for examining how the probability of event-free (i.e., MI-free) status changes over time. In the CHD subcohort of the HUNT3 dataset, the event-free probability of MI declines over time. Table 1 shows the incidence and demographic information of MI in the CHD subcohort. [Table 1]

[0071] Materials and Methods - Cox Elastic Net Model Survival data are characterized by outcomes that are the time to an event, which can cover a wide range of topics such as MI events, death from cancer, re-hospitalization due to disease, or mechanical part failure. The nature of time-dependent data is that events are not observed in some individuals if they occur outside the study period. These individuals are "censored." However, this can occur for multiple reasons (e.g. death from causes unrelated to MI, withdrawal of the individual from the study, occurrence of MI after the end of the study framework). There are multiple types of censoring, but the data include individuals who are right-censored, meaning that for patients without an MI event, it is assumed that they occurred after the last observation.

[0072] Survival data are characterized by the survival function S(.), which is the probability of being free of an event at time t, calculated as:

number

number

[0073] Because the number of features in the dataset is greater than the sample size, elastic net penalization can be incorporated into the model and is a form of penalized regression that combines the least absolute shrinkage and selection operator (i.e., lasso) with ridge regression or Tikhonov regularization. This tool performs feature selection via a lasso routine, leaving correlated features together in the model such that p is greater than n. In standard regression models, the effect of a feature, β, is typically I , and the predictor X' i β, where, in elastic net regularization, the estimated feature effect is calculated as:

number

[0074] Survival analyses were combined with elastic net penalties by using the Cox elastic net model implemented via the glmnet package available in CRAN-R. The Cox elastic net model merges the standard Cox proportional hazards model with elastic net penalties, allowing classifiers to be expanded using survival methods while still providing the benefits of penalized regression.

[0075] To mitigate class imbalance, we combined the Cox proportional hazards elastic net model with a downsampling technique. This approach allowed us to identify the features that best predicted whether an individual was at high risk of experiencing an MI event within 4 years, using a hazard ratio threshold identified by cross-validation to compute a "high risk" classifier. Furthermore, this technique estimated the effects of features in a way that allowed features that accurately predicted individuals at high risk to have different "weights" (i.e., beta estimates) than if they were derived using the full cohort.

[0076] For comparison, we performed two elastic net logistic regression models (with and without downsampling, which can be implemented via the caret package in R) and a Cox elastic net model that does not incorporate the downsampling method. Models were compared using AUC, sensitivity, specificity, and C-Index, as appropriate.

[0077] Analyses were performed using R version 3.4.4 with RStudio Server version 1.1.453.

[0078] Materials and Methods - Data Subsetting The dataset was split into a training set (80% of the data) and a test set (20%). The training set was used to build the model and the final model was evaluated on the test set. The threshold for prediction on the test set of the Cox elastic net model was calculated as the average of the thresholds generated per split during cross-validation. Univariate filtering was performed using the training set before implementing the penalized regression model. Student's t-tests were calculated for each analyte to evaluate whether the mean values ​​were statistically significantly different between individuals who experienced and did not experience an MI event in the study framework. For consistency in showing the utility of the method, the top 100 analytes (ranked by false discovery rate values) are included throughout the model deployment.

[0079] result Results from the downsampled Cox elastic net model were compared to two logistic regression elastic net models (downsampled and not downsampled) and a Cox elastic net model that did not use downsampling. For ease of notation, the Cox elastic net model is referred to as the "Coxnet" model and the elastic net logistic regression model is referred to as the "LRnet" model. Downsampled models are appended with "DS" (e.g., the Cox elastic net model that implements downsampling is "DS-Coxnet").

[0080] Across models, five-fold cross-validation with five iterations was used to select the best model within each model type using the training set. The best model was selected via maximum AUC. Feature selection, estimated effects, and classification thresholds were allowed to vary between models. Following cross-validation, the predictive ability of the top models in each category was evaluated on the test dataset.

[0081] During model development, a Coxnet model was built using the original data, but optimized for classification using the AUC metric at 4 years. This means that a standard survival model was built, but a binary 4-year mark classifier (positive / negative for MI before 4 years) was used to calculate the AUC and optimize the model. The 4-year outcome was used to develop a logistic regression model, which was also optimized using the AUC. The C-Index was used for the purpose of comparing models using standard survival model metrics. was calculated for survival models.

[0082] Model Results and Comparison Cross-validation results show that both Coxnet models significantly outperformed the standard LRnet model (see Table 2). This result is expected because survival analysis methods use time-to-event information as part of feature selection and model development. A more compelling result is that the DS-Coxnet model outperformed both the DS-LRnet and standard Coxnet models across all classification metrics (AUC, sensitivity, specificity). Furthermore, the DS-Coxnet model had a higher C-Index than the standard Coxnet model, indicating that the downsampled model better predicts the order of time to MI. [Table 2]

[0083] Following model optimization by cross-validation, the predictive ability of the top models was evaluated on the test set. This included consideration of sensitivity and specificity based on correctly predicting individuals as “high risk” of experiencing MI by the 4-year mark. Performance metrics for all models on the test set are shown in Table 3. The DS-Coxnet model is the only model that performs better than “random chance” with an AUC of 0.63. Furthermore, the DS-Coxnet model has the highest sensitivity and specificity compared to both the DS-LRnet and standard Coxnet models (unsurprisingly, the LRnet model performs as poorly on the test dataset as it does on the training dataset). [Table 3]

[0084] To further demonstrate the benefit of the downsampled survival model approach, for each model, Kaplan-Meier curves were generated on the test set and stratified by whether individuals were predicted as high risk or not, using the model-specific thresholds identified by cross-validation (see Figure 4). In this comparison, the thresholds for the standard and DS-Coxnet models were calculated as the average thresholds across cross-validation iterations. This method of visual inspection shows that the thresholds for the DS-Coxnet model are used to very clearly separate the high- and average-risk groups; this separation is not clearly defined for the other models.

[0085] The combined evidence from the figures and model evaluation metrics (Table 3) makes a compelling case that a downsampled survival model approach is beneficial for identifying individuals at high risk of MI within 4 years.

[0086] Investigating thresholds for downsampled Coxnet models The threshold used to predict the test set using the DS-Coxnet model was the average over all thresholds from the cross-validation iterations. This threshold led to higher sensitivity and specificity than the other models, but the values ​​were still significantly unbalanced. An important consideration is whether the sensitivity / specificity tradeoff can be further balanced by manipulating the prediction threshold.

[0087] As with classification models, the threshold can be adjusted to find a value that maximizes sensitivity, maximizes specificity, or minimizes the difference between sensitivity and specificity on the test set. Table 4 shows performance metrics for various thresholds on the test set, and Figure 5 plots the Kaplan-Meier curves for each. As shown in Table 4, varying the prediction threshold results in sensitivities above 60% without decreasing the AUC. However, the Kaplan-Meier curves (Figure 5) show the widest separation between high-risk and average-risk individuals using the average threshold. [Table 4]

[0088] Although the sensitivity and specificity remain relatively lower than typically desired (i.e., 70% or higher), this result may be due to the fact that the test set only had 13 subjects with an MI event 4 years prior, limiting the deployment of the model. However, the analysis shows that the thresholds used to classify the level of risk in the survival model can be adjusted in the same way as in the classification model.

[0089] It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the present disclosure being indicated by the following claims.

Claims

1. 1. A computer-implemented method for identifying biomarkers for a disease or condition, comprising: a) obtaining a class-imbalanced dataset, the class-imbalanced dataset comprising biological data from a plurality of subjects, the biological data for each subject comprising an observation, a time value, and a plurality of clinical measurements, the biological data being classified as part of a majority data class or a minority data class, the majority data class comprising more observations than the minority data class; b) down-sampling the class-imbalanced dataset by subsetting data of the majority data class to reduce class imbalance between a number of observations of the majority data class and a number of observations of the minority data class to generate a down-sampled dataset, wherein the down-sampling generates the majority data class that includes a number of observations equivalent to the minority data class; and c) performing a survival analysis by training a Cox proportional hazards model using cross-validation on the downsampled data set to generate an elastic net penalty that identifies features that classify observations between the minority data class and the majority data class, and a survival model, the elastic net penalty being combined with the survival analysis; The observation either includes an event or does not include an event at a particular time value; and the AUC, sensitivity, specificity, and / or C-index of the survival model is closer to 1 than the AUC, sensitivity, specificity, and / or C-index of the survival model where the class-imbalanced dataset was not downsampled prior to the survival analysis; The method.

2. The method of claim 1 , wherein the class imbalanced dataset is a survival dataset.

3. The method of claim 1 or 2, wherein the event is a disease, disorder, or condition of the subject.

4. The method according to any one of claims 1 to 3, wherein the cross-validation is a 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, 15-fold, 16-fold, 17-fold, 18-fold, 19-fold, or 20-fold cross-validation.

5. The method of any one of claims 1 to 4, wherein the survival model comprises between 5 and 1000 features, each feature selected from the group consisting of a protein measurement, a clinical factor, and a combination thereof.

6. 6. The method of claim 5, wherein the clinical factors are selected from the group consisting of age, weight, blood pressure, height, BMI, cholesterol, sex, and combinations thereof.

7. The method of any one of claims 1 to 6, wherein the clinical measurements are selected from proteomic measurements, genomic measurements, transcriptomic measurements, metabolomic measurements, or combinations thereof.

8. The method of any one of claims 1 to 7, wherein the cross-validation is selected from k-fold cross-validation, Monte Carlo cross-validation, and leave-N-out cross-validation.

9. The method of any one of claims 1 to 8, wherein the majority data class is 95% of the class imbalanced dataset and the minority data class is 5% of the class imbalanced dataset.

10. The method of any one of claims 1 to 8, wherein the majority data class is 90% of the class imbalanced dataset and the minority data class is 10% of the class imbalanced dataset.

11. The method of any one of claims 1 to 8, wherein the majority data class is 85% of the class imbalanced dataset and the minority data class is 15% of the class imbalanced dataset.

12. The method of any one of claims 1 to 8, wherein the majority data class is 80% of the class imbalanced dataset and the minority data class is 20% of the class imbalanced dataset.

13. The method of any one of claims 1 to 8, wherein the majority data class is 75% of the class imbalanced dataset and the minority data class is 25% of the class imbalanced dataset.

14. The method of any one of claims 1 to 8, wherein the majority data class is 70% of the class imbalanced dataset and the minority data class is 30% of the class imbalanced dataset.

15. The method of any one of claims 1 to 8, wherein the majority data class is 65% of the class imbalanced dataset and the minority data class is 35% of the class imbalanced dataset.

16. 9. Any of claims 1 to 8, wherein the majority data class is 60% of the class imbalanced data set and the minority data class is 40% of the class imbalanced data set.

13. The method according to claim 1.

17. 1. A computer-implemented method for identifying biomarkers for a disease or condition, comprising: a) downsampling the class-imbalanced dataset by subsetting data of a majority data class to reduce class imbalance between a number of observations of the majority data class and a number of observations of a minority data class to generate a downsampled dataset, wherein the downsampling generates the majority data class that includes a number of observations equivalent to the minority data class; and b) performing a survival analysis by training a Cox proportional hazards model using cross-validation on the downsampled data set to generate an elastic net penalty that identifies features that classify observations between the minority data class and the majority data class, and a survival model, the elastic net penalty being combined with the survival analysis; the observation either includes an event or does not include an event at a particular time value; the class-imbalanced dataset comprises biological data from a plurality of subjects, the biological data for each subject comprising an observation, a time value, and a plurality of clinical measurements, the biological data being classified as part of the majority data class or the minority data class, the majority data class comprising more observations than the minority data class; and the AUC, sensitivity, specificity, and / or C-index of the survival model is closer to 1 than the AUC, sensitivity, specificity, and / or C-index of the survival model where the class-imbalanced dataset was not downsampled prior to the survival analysis; The method.

18. 20. The method of claim 17, wherein the AUC is calculated based on a determination of whether a subject has an event by a particular time point.

19. A computer-implemented method of the method of any one of claims 1 to 16, comprising: The method, wherein steps b) and c) are calculated using a computer system.

20. The method of claim 19 , wherein the class imbalance dataset in step a) is received by a computer system.

21. 1. A computer readable program storage device, comprising: Said apparatus having stored thereon a program of instructions for carrying out each of the method steps of the method according to any one of claims 1 to 16.

22. 1. A computing system for identifying biomarkers for a disease or condition, comprising: a memory for storing programmed instructions; and a processor configured to execute the programmed instructions to perform operations, The system, wherein the operations are for carrying out a method according to any one of claims 1 to 16.

23. A non-transitory computer-readable medium, comprising: Stores instructions executable by a processor to perform operations; The non-transitory computer readable medium, wherein the operations are for performing the method of any one of claims 1 to 16.

Citation Information

Patent Citations

  • Compositions and methods for detecting predisposition to cardiovascular disease

    WO2017214397A1

  • Methods and systems for detecting usual interstitial pneumonia

    WO2018048960A1

  • Proadm as marker indicating an adverse event

    WO2018141840A1