A system and method for predicting insulin resistance or pancreatic beta-cell function, and a computer-readable medium thereof.
A system using machine learning models with age, sex, and BMI features effectively predicts insulin resistance and beta-cell function, addressing the limitations of existing methods by achieving high accuracy and clinical significance in mortality prediction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- TAICHUNG VETERANS GENERAL HOSPITAL
- Filing Date
- 2024-03-29
- Publication Date
- 2026-05-28
AI Technical Summary
Existing clinical methods for predicting insulin resistance and pancreatic beta-cell function, such as HOMA-IR and HOMA-β, are not widely used due to the necessity of fasting plasma insulin levels, which are not regularly monitored, and existing machine learning models primarily focus on diabetes mellitus rather than early diagnosis in non-diabetic populations.
A system and method utilizing a database, feature extraction, and machine learning model building to predict insulin resistance and pancreatic beta-cell function, incorporating features like age, sex, and body mass index, with algorithms like XGBoost, Random Forest, and Deep Neural Networks for accurate prediction.
The method achieves high accuracy in predicting insulin resistance and beta-cell function, demonstrated by ROC curves exceeding 0.8, with significant clinical significance in predicting cardiovascular and all-cause mortality, providing early indicators for diabetes management.
Smart Images

Figure 0007867035000026 
Figure 0007867035000027 
Figure 0007867035000028
Abstract
Description
Technical Field
[0001] The present disclosure is related to medical monitoring applications, and more specifically, to the prediction of insulin resistance and / or pancreatic beta-cell function.
Background Art
[0002] Insulin resistance (IR) is a state in which the effects of insulin are reduced in cells of skeletal muscle, liver, and adipose tissue. IR occurs when insulin reaches a predetermined concentration along with insufficient glucose response. There are numerous underlying factors that may contribute to the development of IR, such as obesity, stress, drugs (e.g., steroids), pregnancy, insulin antibodies, and genetic defects in the insulin signaling pathway. IR is also associated with clinical diseases such as diabetes (DM), coronary artery disease, metabolic syndrome, polycystic ovary syndrome, and non-alcoholic fatty liver disease.
[0003] IR is well known as an indicator for the early diagnosis of DM. One of the first signs of metabolic syndrome is the onset of IR. The decrease in the effect of insulin (i.e., IR) causes overwork of beta cells that secrete more insulin as a compensatory mechanism to maintain the level of plasma glucose. As metabolic syndrome worsens, the number of beta cells decreases and the level of plasma glucose begins to rise. Eventually, the elevated plasma glucose level meets the diagnostic definition of DM. It is noteworthy that the decline in beta-cell function (decrease in insulin secretion) begins as early as 12 years before receiving a DM diagnosis and continues throughout the disease course. Numerous clinical studies have suggested that the earlier DM is treated, the better the DM outcome, and that early control of DM has a great advantage in reducing the legacy effect (i.e., metabolic memory) related to glucose metabolism disorders. Therefore, the identification of IR (i.e., the discovery at the earliest stage of DM) or the confirmation of beta-cell function is important in the clinical spectrum of DM.
[0004] Static tests used in clinical practice include the homeostasis model assessment of IR (HOMA-IR) and the homeostasis model assessment of β-cell function (HOMA-β). HOMA estimates the degree of β-cell deficiency and insulin sensitivity using equations relating fasting plasma insulin concentration and fasting plasma blood glucose levels. However, while fasting plasma insulin levels are not regularly monitored, they are necessary for calculating HOMA-IR and HOMA-β. Therefore, despite their significant clinical importance, IR and β-cell deficiency are not widely used as indicators in clinical practice.
[0005] Recently, machine learning has been used to address this situation. However, most studies are based on predictive models in the field of diabetes mellitus (DM), rather than on indicators for early diagnosis of DM in non-diabetic patients (i.e., impaired IR and β-cell function).
[0006] From the above perspective, there is an unmet need in this technology field for using artificial intelligence to predict IR and / or β-cell function in non-diabetic populations. [Overview of the Initiative]
[0007] To address the above issues, this disclosure provides a system for predicting the insulin resistance and / or pancreatic β-cell function of a subject, comprising: a database configured to provide a dataset; a feature extraction module configured to collect and process features from the dataset to generate a feature set relating to a subject, wherein the feature set includes the subject's age, sex, race, and body mass index; and a model building and optimization module configured to build a machine learning model based on the feature set for predicting the subject's insulin resistance and / or pancreatic β-cell function.
[0008] The Disclosure further provides a method for predicting insulin resistance and / or pancreatic β-cell function of a subject of interest, comprising the steps of: configuring a database to provide a dataset; configuring a feature extraction module to collect and process features from the dataset to generate a feature set relating to the subject, the feature set including the subject's age, sex, race, and body mass index; and configuring a model building and optimization module to build a machine learning model based on the feature set for predicting insulin resistance and / or pancreatic β-cell function of the subject.
[0009] The Disclosure further provides a computer-readable medium storing computer-executable code, which, when executed, allows the computer-executable code to implement the methods of the Disclosure.
[0010] The above and other objects of the present invention will become undoubtedly apparent to those skilled in the art upon reading the following detailed description of preferred embodiments, which are illustrated with various diagrams and drawings. [Brief explanation of the drawing]
[0011] The patent or application file includes at least one color drawing. Copies of the patent or published patent gazette, including the color drawing, are available from the United States Patent and Trademark Office upon request and payment of the necessary fees.
[0012] Figure 1 is a schematic diagram showing an example of the structure of a system and method for predicting insulin resistance and / or pancreatic β-cell function according to an embodiment of the present disclosure.
[0013] Figure 2A is a curve graph showing receiver operating characteristic (ROC) curves for XGboost, random forest, logistic regression, and deep neural network (DNN) for a training group to predict insulin resistance (IR) in patients with chronic kidney disease (CKD) according to an embodiment of the present disclosure.
[0014] Figure 2B is a curve graph showing receiver operating characteristic (ROC) curves for an internal validation group for predicting insulin resistance (IR) in patients with chronic kidney disease (CKD) according to an embodiment of this disclosure, using XGboost, random forest, logistic regression, and deep neural network (DNN).
[0015] Figure 2C is a curve graph showing receiver operating characteristic (ROC) curves for XGboost, random forest, logistic regression, and deep neural network (DNN) for a test group used to predict insulin resistance (IR) in patients with chronic kidney disease (CKD) according to an embodiment of this disclosure.
[0016] Figure 3A is a bar graph illustrating the detailed feature importance of the XGboost prediction model according to the embodiment of this disclosure.
[0017] Figure 3B is a bar graph showing the detailed feature importance of the random forest prediction model according to the embodiment of this disclosure.
[0018] Figure 4A is a heat graph illustrating the positive or negative influence of features on predicting insulin resistance using SHAP (SHapley Additive exPlanations) values of the XGboost prediction model according to the embodiment of this disclosure.
[0019] Figure 4B is a heat graph illustrating the positive or negative influence of features on predicting insulin resistance using SHAP values from a random forest prediction model according to the embodiment of this disclosure.
[0020] Figures 5A and 5B are Kaplan-Meier survival curves of CV mortality based on the presence or absence of insulin resistance, showing the clinical significance (cardiovascular (CV) mortality and all-cause mortality) predicted by the XGboost prediction model derived from the database of Taiwan Biobank according to an embodiment of the present disclosure.
[0021] Figures 5C and 5D are Kaplan-Meier survival curves of all-cause mortality based on the presence or absence of insulin resistance, showing the clinical significance (cardiovascular (CV) mortality and all-cause mortality) predicted by the XGboost prediction model derived from the database of Taiwan Biobank according to an embodiment of the present disclosure.
[0022] Figure 6A is a bar graph showing the personalized IR risk and the description of the feature set according to an embodiment of the present disclosure.
[0023] Figure 6B is a bar graph showing the personalized β-cell insufficiency risk and the description of the feature set according to an embodiment of the present disclosure.
Mode for Carrying Out the Invention
[0024] The following embodiments are provided to explain the present disclosure in detail. Those skilled in the art can easily understand the advantages and effects of the present disclosure by reading the present disclosure, and can also introduce or apply it to other different embodiments. Therefore, any element or method included in the scope of the present disclosure disclosed herein can be combined with all the elements or methods disclosed in all the embodiments of the present disclosure.
[0025] The relationships, structures, steps, factors, and other features shown in the drawings attached to this disclosure are used only for explaining embodiments so that those skilled in the art can read and understand the disclosure from them, and are not intended to limit the scope of the disclosure. Any changes, modifications, or adjustments to the above features that do not affect the designed purpose and effects of this disclosure should be included within the technical scope of this disclosure.
[0026] As used herein, when an object is described as "comprising", "including", or "having" a certain limitation, unless otherwise specified, it may further include other modules, elements, members, structures, parts, devices, systems, steps, connections, factors, features, etc., and should not exclude others.
[0027] As used herein, terms representing order such as "first", "second", "third", "fourth", etc. are shown for convenience only for describing or distinguishing certain limitations such as modules, sets, databases, elements, members, structures, parts, devices, systems, steps, connections, factors, or features, and are not intended to limit the scope of this disclosure or the spatial order of these limitations. Furthermore, unless otherwise specified, singular expressions such as "a", "an", "the", etc. are also relevant in the plural case, and expressions such as "or", "and / or" may be interchangeable with each other.
[0028] As used herein, the terms "subject", "participant", "individual", and "patient" may be interchangeable with each other.
[0029] As used herein, the terms "fasting blood glucose level" and "fasting plasma glucose level" may be interchangeable with each other.
[0030] As used herein, the terms “decline in β-cell function,” “decline of β-cell function,” and “β-cell deficiency” may be interchangeable.
[0031] As used herein, the terms "variable," "factor," and "feature" may be interchangeable.
[0032] As used herein, the terms “comprise,” “comprising,” “include,” “including,” “have,” “having,” “contain,” and “containing,” or any other variations thereof, are intended to mean non-exclusive inclusion. For example, a module, process, or method having a set of elements may not necessarily contain only these elements, but may also contain other elements that are not expressly enumerated or that are specific to such module, process, or method.
[0033] As used herein, the terms “Homeostasis Model Assessment of IR (HOMA-IR)” and “Homeostasis Model Assessment of β-cell Function (HOMA-β)” are used in relation to the assessment of IR and β-cell function. HOMA-IR and HOMA-β are the most preferred methods for estimating declines in IR and β-cell function. Specifically, variables for calculating HOMA-IR include age, sex, race, year of assessment, BMI, and smoking status. Elevated IR (i.e., HOMA-IR) is associated with a higher IR. Furthermore, β-cell function is strongly associated with the development of type 2 diabetes mellitus (DM). This association is statistically independent of impaired glucose tolerance, obesity, and body fat distribution. High HOMA-IR and HOMA-β values are independently associated with a higher risk of developing prediabetes. HOMA-IR employs the following formula to indicate insulin resistance: Fasting plasma insulin level (μU / ml) × Fasting plasma blood glucose level (mg / dL) / 405. HOMA-β employs the following formula to evaluate β-cell function: (20 × fasting plasma insulin level (μU / ml)) / (fasting plasma blood glucose level (mg / dL) - 63). For adults, the cutoff value indicating IR is a HOMA-IR value of 2.5, and the cutoff value indicating decreased β-cell function is a HOMA-β value of 66%. In some embodiments, HOMA-IR and HOMA-β values are calculated for subjects in a third database, i.e., an integrated database of the first and second databases, and patients or participants are divided into non-IR group, IR group, non-β-cell dysfunction group, and β-cell dysfunction group according to their corresponding HOMA-IR value (≤ or > 2.5) and HOMA-β value (< 66%).
[0034] In at least one embodiment of this disclosure, the database includes a first database and a second database. The first and second databases are derived from a first population of subjects and a second population of subjects, respectively. In some embodiments, the race of the first population differs from that of the second population. For example, the first population may be of non-Asian race and the second population may be of Asian race, but this disclosure is not limited thereto. method
[0035] Figure 1 shows a system and method for predicting insulin resistance and / or pancreatic β-cell function. The steps of the procedure are indicated by arrows and are explained below. Specifically, the term "ML" shown in Figure 1 is an abbreviation for Machine Learning. Research protocols and participants in this disclosure
[0036] In some embodiments of this disclosure, a first dataset derived from a first database (e.g., the National Health and Nutrition Examination Survey (NHANES) database), a second dataset derived from a second database (e.g., the MAJOR (MJ) survey database), and a fourth dataset derived from a fourth database (e.g., the Taiwan Biobank (TWB) database) are used.
[0037] In some embodiments, the first database is the NHANES database, a U.S. health-related program. Health surveys within this program are regularly conducted by the National Center for Health Statistics (NCHS) of the Centers for Disease Control and Prevention (CDC). The first data is freely available for research purposes. The NCHS Research Ethics Review Board approved the research in this disclosure, and all participants or their representatives submitted informed consent. Examinations included physical measurements, health and nutrition questionnaires, and clinical tests. Participants completed the questionnaires during home interviews. This disclosure analyzes NHANES participants from January 1, 1999, to December 31, 2012. Participants under 18 years of age, those with incomplete clinical data, and those with diabetes mellitus were excluded from the analysis.
[0038] In some embodiments, the second database is the MJ Survey Database, the most detailed and accurate medical database in Taiwan (Republic of China). The Taiwan MJ Cohort resource is an ongoing, dynamic prospective study of women and men participating in a large-scale health checkup program conducted by the MJ Health Management Institution (www.mjclinic.com.tw) in Taiwan. The MJ Group in Taiwan is the largest health management group in Asia. Details of the Taiwan MJ Cohort study's target population and data collection methodology have been reported elsewhere. Since 1994, approximately 600,000 individuals from Taiwan have been enrolled in this MJ Cohort. Participants complete all standard protocols, including medical history questionnaires, multi-phase blood tests, and clinical tests such as pulmonary function tests, electrocardiograms, urinalysis, and stool tests. These procedures comply with the International Organization for Standardization 9001 standard for quality management. After the initial examination, all participants were encouraged to participate annually. All data is updated by the MJ Health Research Foundation.
[0039] In some embodiments, step S1 represents the database integration module integrating the first database and the second database into a third database, that is, the third database is the result of integrating the first database and the second database.
[0040] In some embodiments, the fourth database is the Taiwan Biobank (TWB) database. The TWB project was initiated by Taiwan's Ministry of Health and Welfare to collect national genetic and clinical data of the Taiwanese people. Taiwan has the world's most comprehensive health database, covering up to 99% of its population. TWB is a rich database for biomedical research. The TWB database currently includes 174,077 participants, 84,276,467 questionnaires, 72,076,086 height / weight data, and 79,314,928 clinical data. For all participants, biomarkers and genetic data are generated from urine and blood samples. This disclosure analyzes the TWB population from December 10, 2008 to November 30, 2018.
[0041] In some embodiments, a fourth database of the present disclosure, such as the TWB database, is used for external validation to determine whether the application of a diagnostic algorithm can predict the mortality rate of individuals who are not diabetic, based on IR or β-cell function estimated from the third database.
[0042] In some embodiments, participants are collected from a first, second, and fourth database, for example, the NHANES database (1999–2012), the MJ database (2008–2017), and the TWB database (October 10, 2008–November 30, 2018), respectively. Baseline variables include age, sex, body mass index (BMI) (weight in kg divided by the square of height in m), HOMA-IR (except for the TWB database), total cholesterol (TC) (mg / dL), high-density lipoprotein (HDL) (mg / dL), triglyceride (mg / dL), FPG (mg / dL), and glycated hemoglobin (HbA1c) (%). DM is defined according to the American Diabetes Association (ADA) guidelines for DM.
[0043] In some embodiments, the third feature set of the third database of this disclosure is used for machine learning to enable the maximum utilization of IR or β-cell function. The third feature set is well known in the art as factors related to IR or β-cell function. The process of building machine learning models in this disclosure
[0044] In some embodiments, step S2 means that 55% of the patient population in the third database of the present disclosure are randomly selected as the training group for model building, 15% as the internal validation group for hyperparameter optimization, and 30% as the independent test group.
[0045] In some embodiments, 78.6% (55 / 70) of the patient population in the third database of this disclosure (i.e., 55% in the test group and 15% in the internal validation group) are excluded from the predictive model for training the machine learning model of this disclosure to prevent overfitting. This ratio (78.6%) was generally lower than that for other predictive models (80%). The ratio for internal validation of performance (21.4%) (15 / 70) was generally higher than that for other predictive models (20%), indicating that the algorithm of this disclosure uses a larger number of patients for validating model performance.
[0046] In some embodiments, the Synthetic Minority Oversampling Technique (SMOTE) preprocessing algorithm is used to balance the sample between the target and non-target populations within the training group. In some embodiments, step S3 means that the third set of features of this disclosure used to predict IR (HOMA-IR > 2.5) are readily available variables, including, but not limited to, age, sex, race, body mass index (BMI), fasting plasma glucose or fasting blood glucose (fg / FBG / FPG), glycated hemoglobin (HbA1c), triglyceride (tg), total cholesterol (TC / chol), and / or high-density lipoprotein (HDL) cholesterol (HDL cholesterol / hdlc). In some embodiments, step S3 means that the third set of features of this disclosure used to predict β-cell function (HOMA-β value < 66%) are readily available variables, including, but not limited to, age, sex, race, body mass index (BMI), fasting plasma glucose or fasting blood glucose (fg / FBG / FPG), triglyceride levels (tg), total cholesterol (TC / chol), high-density lipoprotein (HDL) cholesterol (HDL cholesterol / hdlc), glutamate oxaloacetate transaminase (GOT), glutamate pyruvate transaminase (GPT), waist circumference (WC), total bilirubin (TB), albumin (alb), systolic blood pressure (sbp), diastolic blood pressure (dbp), estimated glomerular filtration rate (eGFR), and / or creatinine (CRE).
[0047] In some embodiments, a deep neural network (DNN) was used to first estimate the method to choose. Other conventional methods of machine learning algorithms, such as the Random Forest (RFs) algorithm, the eXtreme Gradient Boosting (XGboost) algorithm, and the logistic regression algorithm, were also used. Their accuracy was compared with that of the logistic regression model. After the models were trained, they were tested on a different group. For the training group, the ROC (Receiver Operating Characteristic) curves of the various algorithms were compared. The classification performance of the various classifiers was compared using both the ROC curve and the Area Under the Curve (AUC). A target AUC (>0.80) suggested that the models of this disclosure were sufficient to predict IR and / or β-cell function. In the DNN, the overall structure of the deep neural network is designed as follows: 9 input layers → 18 intermediate hidden layers → 27 intermediate hidden layers → 36 intermediate hidden layers → one-dimensional output layer.
[0048] In some embodiments, a binary outcome of insulin resistance or β-cell function is set as the output layer. To avoid overfitting in deep learning model training, the inventors added a dropout layer between the hidden layers. The dropout rate was set to 0.2. As the activation function, a Scaled Exponential Linear Unit is used in the hidden layers and a hard sigmoid function is used in the output layer. In the hyperparameter tuning process of XGboost and RFs, grid search is used to identify the optimal values of the parameter combinations of possible values.
[0049] In some embodiments, the Gini index is used to calculate feature importance, i.e., the importance of each variable in the third feature set of this disclosure. For model description, SHAP values (SHapley Additive exPlanations) are used to explain how various machine learning models function. Finally, since IR and β-cell function are reported for predicting mortality in individuals without diabetes, the IR prediction algorithm and β-cell function prediction algorithm of this disclosure also consider predictive values for mortality.
[0050] In some embodiments, step S4 means that once excellent diagnostic performance is achieved for the above model, the best model is applied to a fourth database (e.g., the TWB database) for external validation, and its potential clinical significance (e.g., cardiovascular disease (CV) mortality and all-cause mortality) is determined. Clinical significance of IR and β-cell function: CV mortality and all-cause mortality
[0051] The clinical significance of IR and β-cell function is analyzed using the above-mentioned trained model, which demonstrated the best predictive performance. Mortality information was derived through data linkage with active follow-up and death certification. In the International Classification of Diseases (ICD) 9th or 10th edition, all-cause mortality is coded as ICD-9 0001-E999 and ICD-10 A00-Y98. CV mortality is coded under CV disease (ICD-9 390-456 and ICD-10 I00-I99), and includes coronary heart disease (ICD-9 140-414 and ICD-10 I20-I25), stroke (ICD-9 430-438 and ICD-10 I60-I69), ischemic stroke (ICD-9 434 and ICD-10 I63), and hemorrhagic stroke (ICD-9 4330-432 and ICD-10 I60-I62).
[0052] In some embodiments, the fourth database of the Disclosure is configured to provide a fourth dataset of the Disclosure, which is used for external validation and evaluation of clinical significance. In some embodiments, cardiovascular mortality and all-cause mortality are included in clinical significance. Statistical analysis in this disclosure
[0053] NHANES is a multifaceted and complex study. To represent sample-weighted data, it is necessary to calculate weighted data according to the analytical guidelines (US NHANES: Analytical Guidelines, 2011-2014 and 2015-2016, available online). However, when building the machine learning and deep learning models in this disclosure, the original unweighted data obtained from NHANES and the MJ database is used. There are two reasons why weighted data is not used in this disclosure. First, weighted data is typically used to estimate national incidence or prevalence. In training the above models, national prevalence estimates are not required; only the relationship between IR or β-cell function and the third set of features in this disclosure between individuals is needed. Second, it is necessary to integrate NHANES and the MJ database, but the MJ database does not provide corresponding weights. Therefore, unweighted data was also used for NHANES. For unweighted data in the NHANES, MJ, and TWB databases, continuous variables are reported as mean ± standard deviation (SD), and categorical data are reported as numerical values (proportions). Differences in clinical variables between IR or β-cell function states are assessed using chi-square tests for categorical variables and independent t-tests for continuous variables. Univariate and multivariate logistic regression using restricted cubic spline methods is employed to identify nonlinear relationships between selected features and IR or β-cell function, which may be useful for comparing the associated patterns of machine learning methods and conventional statistical methods.
[0054] In some embodiments, the feature extraction module of the Disclosure collects and analyzes a dataset of the Disclosure by determining at least one of the following: mean, standard deviation of the mean, numerical value, proportion, coefficient of variance, and the slope and coefficient of determination of a linear regression, in order to build a machine learning model of the Disclosure.
[0055] In some embodiments, a fourth dataset derived from the TWB database is applied as input to the above algorithm, and the output is the predicted probability of IR or β-cell function for each participant. A predicted probability greater than 0.5 is defined as the cutoff point for belonging to the IR group or the β-cell deficiency group. Kaplan-Meier survival curves using the Log-Rank test and a proportional HR model for survival analysis trials are used to compare CV mortality or all-cause mortality among the predicted IR group, non-IR group, β-cell deficiency group, and non-β-cell deficiency group. All reported p-values are two-sided p-values and are considered significant if p < 0.05. Deep learning algorithms and other ML (including XGBoost, RFs, and DNNs) are performed using Keras (version 2.4.0), TensorFlow (version 1.10.0), and Python (version 3.6.5). Univariate and multivariate analyses of CV mortality and all-cause mortality are also performed. The inventors also compared the algorithm's predictive performance with the Framingham score using the C index and misclassification statistics. All statistical analyses were performed using SAS for Windows (version 9.4, SAS, Inc., Cary, New Connecticut). This disclosure has been approved by the Ethics Committee of Taichung Veterans General Hospital (IRB number: CE20023A). All procedures were carried out in accordance with relevant guidelines and regulations. result Participant selection from the NHANES database, MJ database, and TWB database.
[0056] Initially, this disclosure includes 71,916 participants from the NHANES database (1999–2012) and 14,359 participants from the MJ database (2008–2017). After exclusions, the final analysis included 25,737 participants (14,211 from NHANES and 12,526 from the MJ database). The excluded participants were 26,373 from NHANES and 72 from the MJ database (due to being under 18 years of age), 28,939 from NHANES and 657 from the MJ database (due to incomplete clinical data), and 2,393 from NHANES and 1,104 from the MJ database (due to a history of diabetes mellitus). From all participants in the third database of this disclosure (i.e., the integrated database of the NHANES database and the MJ database), 14,705 participants were randomly selected for the training group and 4,018 participants for the internal validation group. In the training group, models were trained using algorithms, namely XGBoost, RFs, logistic regression, and DNN, and their AUCs were finally compared. The trial group for model evaluation included a total of 8,014 participants. For external validation and evaluation of clinical significance, participants from the TWB database will be further collected from December 10, 2008 to November 30, 2018. Baseline characteristics of participants derived from the Caucasian database (NHANES) and the Asian database (MJ, Taiwan).
[0057] The first and second databases in this disclosure (NHANES vs MJ) include two racial populations, namely the Caucasian population and the Asian population, respectively. However, the age (44.84 years vs 45.06 years) and sex ratios (48.29% vs 50.44%) are similar. In addition, fasting plasma glucose levels or fasting blood glucose levels (91.54 mg / dL vs 100.07 mg / dL), glycated hemoglobin levels (5.32% vs 5.2%), and total cholesterol levels (198.45 mg / dL vs 199.13 mg / dL) were also similar. However, insulin resistance was significantly higher in patients from the NHANES population than in patients from the MJ population (2.84 vs 1.92). Furthermore, individuals derived from NHANES showed higher plasma insulin levels (12.1 mIU / L vs 7.67 mIU / L) and BMI (27.86 kg / m²). 2 vs 23.78 kg / m 2 ), and the triglyceride levels (126.21 mg / dL vs 116.59 mg / dL) were high. Baseline characteristics of participants with or without diabetes mellitus (DM), based on HOMA-IR (≤2.5 or >2.5) and / or HOMA-β (<66%).
[0058] Regarding the three groups described in this disclosure—the training group, the internal validation group, and the trial group—no statistically significant differences were observed among these three groups. Table 1 shows relevant detailed information, based on the presence or absence of insulin resistance, obtained from all participants in the first, second, and third databases. Participants from NHANES had higher mean HOMA-IR (2.84 vs 1.92), higher fasting plasma insulin levels (12.1 mIU / L vs 7.67 mIU / L), and higher BMI (27.86 kg / m²). 2 vs 23.78 kg / m 2Except for higher triglyceride levels (126.21 mg / dL vs 116.59 mg / dL), lower FPG levels (91.54 mg / dL vs 100.07 mg / dL), and lower HDL cholesterol levels (53.88 mg / dL vs 57.77 mg / dL), the baseline characteristics of participants from the NHANES database and those from the MJ database were similar. In the integrated NHANES and MJ databases (i.e., the third database of this disclosure), the IR group (IR>2.5) was significantly (p<0.0001), older (46.06±16.58 vs 45.28±15.34), had a higher proportion of males (54.3% vs 46.5%), higher HOMA-IR (4.53±2.94 vs 1.47±0.53), higher plasma insulin levels (17.84±10.88 vs 6.17±2.17), higher BMI (30.05±6.05 vs 24±4.03), higher HbA1c (5.43±0.41 vs 5.21±0.41), and higher FPG (99.08±10.62 vs (93.92±9.78), HDL cholesterol levels were low (49.19±12.97 vs 59.2±15.43), TC was high (201.35±39.97 vs 197.16±38.17), and triglyceride levels were high (155.51±112.06 vs 104.25±78.04).
[0059] Table 2 shows relevant detailed information, based on pancreatic β-cell function, obtained from all participants from the first, second, and third databases. (Table 1) This disclosure includes baseline data, based on the presence or absence of insulin resistance, obtained from all participants in the first, second, and third databases of this disclosure, namely the NHANES database, the MJ database, and the integrated NHANES and MJ databases. JPEG0007867035000001.jpg233158JPEG0007867035000002.jpg233157JPEG0007867035000003.jpg68169 (Table 2) Baseline data according to pancreatic β-cell function obtained from all participants in the first, second, and third databases of this disclosure, namely the NHANES database, the MJ database, and the integrated NHANES database and MJ database. JPEG0007867035000004.jpg233158JPEG0007867035000005.jpg235158JPEG0007867035000006.jpg164169Baseline data from the TWB database for external validation and evaluation of clinical significance
[0060] In some embodiments, a fourth dataset derived from a fourth database of this disclosure, such as the TWB database, is fed to the XGBoost algorithm and divided into a predicted non-IR group (n=86,283) and a predicted IR group (8,774). Participants predicted to be IR were significantly (p<0.0001), younger (47.51±10.81 years vs 49.42±10.86 years), more male (48.80% vs 33.39%), and had a higher BMI (30.27±3.54 kg / m²). 2 vs 23.32±2.97kg / m 2 ), HbA1c was high (5.83±0.32% vs 5.57±0.33%), FPG was high (99.35±9.44 mg / dL vs 90.98±6.99 mg / dL), HDL cholesterol was low (43.39±8.59 mg / dL vs 56.32±13.25 mg / dL), TC was high (200.22±35.75 mg / dL vs 195.44±34.98 mg / dL), and triglyceride levels were high (201.95±150.76 mg / dL vs 100.54±67.6 mg / dL). Predicted values from machine learning models
[0061] Table 3 shows the results for the internal validation and test groups when using XGboost, RFs, logistic regression, and DNN as algorithms. Figure 2A shows the ROC curve for the training group. Figure 2B shows that in the internal validation group, the AUC of the ROC was greater than 0.8 for all algorithms (the largest being XGboost at 0.87). According to Table 3, in the internal validation group, XGboost also had the highest accuracy (0.80), sensitivity (0.63), specificity (0.89), positive predictive value (PPV) (0.74), and negative predictive value (NPV) (0.8). Figure 2C shows that in the test group as well, the AUC of the ROC was greater than 0.80 for all algorithms (the largest being XGboost at 0.88). Furthermore, as shown in Table 3, XGboost had the highest accuracy (0.81), sensitivity (0.64), and NPV (0.83) in the test group, while RFs had the highest specificity (0.90) and PPV (0.75) in the test group. (Table 3) Area under the receiver operating characteristic (ROC) curve (AUC), accuracy, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) for various models. JPEG0007867035000007.jpg233158
[0062] Tables 4 and 5 show the AUC of ROC curves generated by machine learning models for predicting IR and β-cell function of this disclosure, using each of the feature sets of this disclosure. This indicates the predictive performance.
[0063] In some embodiments, the machine learning models of this disclosure for predicting IR and β-cell function may be employed in classification algorithms and additional algorithms, but this disclosure is not limited thereto. In some embodiments, the classification algorithm may be, but is not limited to, XGBoost, RFs, logistic regression, and / or DNN. In some embodiments, the additional algorithm may be, but is not limited to, linear regression, polynomial regression, support vector regression, decision tree regression, random forest regression, ridge regression, lasso regression, elastic net regression, logistic regression, gradient boosting, and / or K-approximation regression.
[0064] In some embodiments, the true values of a subject's HOMA-IR and / or HOMA-β may be determined by a classification algorithm, a predicted probability of the subject's IR or β-cell function, and additional algorithms, the analysis algorithms of the Disclosure may be, but are not limited to, XGBoost, RFs, logistic regression, and / or DNN, and the additional algorithms of the Disclosure may be, but are not limited to, linear regression, polynomial regression, support vector regression, decision tree regression, random forest regression, ridge regression, lasso regression, elastic net regression, logistic regression, gradient boosting, and / or K-approximation regression.
[0065] In some embodiments, the model building and optimization module of the Disclosure builds a machine learning model of the Disclosure based on the feature set of the Disclosure by classifying non-diabetic patients into insulin-resistant, non-insulin-resistant, β-cell deficiency, and non-β-cell deficiency groups to predict insulin resistance and pancreatic β-cell function in non-diabetic patients, and generating corresponding predicted values and classification performance values.
[0066] In some embodiments, the classification performance value is the area under the curve of the receiver operating characteristic curve of the machine learning model of the present disclosure. (Table 4) Area under the receiver operating characteristic (ROC) curve (AUC) of a machine learning model used to predict IR (N / A) data is not available. JPEG0007867035000008.jpg233160JPEG0007867035000009.jpg235160JPEG0007867035000010.jpg233155J PEG0007867035000011.jpg234155JPEG0007867035000012.jpg237155JPEG0007867035000013.jpg93169 (Table 5) Area under the receiver operating characteristic (ROC) curve (AUC) of a machine learning model used to predict β-cell function ("N / A" indicates that data is not available). JPEG0007867035000014.jpg233155JPEG0007867035000015.jpg233155JPEG0007867035000016.jpg226155JPEG0007867035000017.jpg226155JPEG0007867035000018.jpg227155JPEG0007867035000019.jpg151169 Relative importance of parameters in XGBoost and Random Forests (RFs)
[0067] Figures 3A and 3B show the relative importance of each feature in the feature set of this disclosure, examined for the XGboost predictive model and the RFs predictive model, respectively. The importance of all features was similar across both the XGboost and RFs algorithms. BMI had the greatest impact on IR (0.43 for the XGboost algorithm and 0.47 for the RFs algorithm), and its impact was significantly greater than that of other features. Other variables influencing IR, in descending order of importance, were FPG (0.12 for the XGboost algorithm and 0.13 for the RFs algorithm), HDL (0.11 for the XGboost algorithm and 0.10 for the RFs algorithm), and triglyceride levels (0.097 for the XGboost algorithm and 0.13 for the RFs algorithm).
[0068] Referring to Figures 4A and 4B, the relative contribution of features to the algorithm (i.e., SHAP value) in the XGboost and RFs algorithms showed a positive or negative correlation with IR (Figure 4A shows the case for XGboost, and Figure 4B shows the case for RFs). In both algorithms, all features were clearly divided into those with a positive correlation to IR and those with a negative correlation, and similar patterns were obtained in both algorithms. Some features, such as BMI, FPG, triglyceride levels, sex, and HbA1c, positively influence IR. Other features, such as race, age, HDL cholesterol levels, and total cholesterol levels, negatively influence IR.
[0069] Specifically, SHAP (SHapley Additive exPlanations) is a method that links the contribution of each feature to predictions made by a machine learning model. In calculating the SHAP value, the contribution of each feature among all features is averaged. In other words, the SHAP value of each feature evaluates the significance of that feature in the machine learning model of some embodiments of this disclosure. As shown in Figures 4A and 4B, a positive SHAP value is associated with an increase in IR prediction, and a negative SHAP value is associated with a decrease in IR prediction. Features with high values are shown in red, and features with low values are shown in blue. Figure 4A shows nine bands, and Figure 4B also shows nine bands, with each feature band representing the distribution of all subjects in this application. Taking the first band in Figure 4A as an example, the higher the BMI, the higher the SHAP value, which suggests that the higher the BMI, the higher the probability that the machine learning model will predict IR. Furthermore, taking the sixth band in Figure 4A as an example, the SHAP value increases as the HDL cholesterol level decreases, suggesting that the lower the HDL cholesterol level, the higher the probability that the machine learning model will predict IR. Additionally, the feature value for Asians is 1, while the feature value for non-Asians is 0. The influence of features on IR obtained from machine learning algorithms and restricted cubic spline analysis
[0070] An increasing trend is observed for BMI, FPG, triglyceride levels, and HbA1c (SHAP values increased with increasing features). A decreasing trend is observed for HDL, age, and total cholesterol levels (SHAP values decreased with increasing features). All patterns of SHAP values obtained through machine learning in this disclosure are similar to the associated patterns.
[0071] An XGboost algorithm for predicting cardiovascular mortality and all-cause mortality, externally validated by the fourth database (Taiwan Biobank) of this disclosure.
[0072] Because the XGboost model provides the best prediction of IR, the inventors applied the XGboost algorithm to the fourth dataset of this disclosure derived from the TWB database. Here, participants from the TWB database are divided into a predicted HOMA-IR(+)(IR) group and a predicted HOMA-IR(-)(non-IR) group.
[0073] Figures 5A, 5B, 5C, and 5D show that the Kaplan-Meier survival curves (Figures 5A and 5B for CV mortality, and Figures 5C and 5D for all-cause mortality) indicate a significant difference between the predicted IR group and the non-predicted IR group (p<0.0001 for CV mortality, p=0.0006 for all-cause mortality). Therefore, predicting IR using the XGboost model has clear clinical significance, particularly for predicting the worsening of CV mortality and all-cause mortality. Table 6 shows the potentially relevant factors based on multivariate analysis. (Table 6) Univariate and multivariate analyses of factors associated with cardiovascular mortality and all-cause mortality in the TWB database. JPEG0007867035000020.jpg232169JPEG0007867035000021.jpg62169
[0074] Table 7 shows that when comparing HOMA-IR > 2.5 with ≤ 2.5, the hazard ratios (HR) for all-cause mortality and cardiovascular mortality are clearly higher in both univariate and multivariate analyses. (Table 7) Proportional hazards model for survival time analysis JPEG0007867035000022.jpg213169
[0075] When comparing CV mortality with the Framingham score (shown in Table 8), XGboost's predictive performance was inferior (0.547 vs 0.643, concordance 0.557). However, in Table 9, under condition 1 (cutoff value: 10% CV risk over 10 years estimated by the Framingham score), an increase in mortality risk was observed in the group with HOMAIR > 2.5 (XGboost) + ≤ 10% (Framingham score) even when the risk prediction of the Framingham score was ≤ 10% (all-cause mortality: 1.556 (95% confidence interval (CI): 1.186~20.41), p=0.0014) (CV mortality: 3.374 (95% CI: 2.051~5.552, p<0.0001)). Similarly, in the group under Condition 2 (cutoff value: 10-year CV risk estimated by the Framingham score is 20%), HOMAIR > 2.5 (XGboost) + & ≤ 20% (Framingham score), an increase in mortality risk was observed even when the risk prediction of the Framingham score was ≤ 20% (all-cause mortality: 1.433 (95% CI: 1.141~1.8), p=0.002) (CV mortality: 2.655 (95% CI: 1.752~4.023, p<0.0001)). (Table 8) The C index for comparing the prediction algorithm (XGboost) and the Framingham score regarding CV mortality. JPEG0007867035000023.jpg101169 (Table 9) Misclassification statistics of XGboost and Framingham scores for CV mortality JPEG0007867035000024.jpg233156JPEG0007867035000025.jpg233158
[0076] Figures 6A and 6B show the results and feature sets for personalized IR risk (Figure 6A) and personalized β-cell deficiency risk (Figure 6B). Specifically, wc is an abbreviation for waist circumference, tg for triglyceride level, got for glutamate oxaloacetate transaminase level, tb for total bilirubin level, gpt for glutamate pyruvate transaminase level, alb for albumin level, chol for cholesterol level, dpb for diastolic blood pressure, cre for creatinine level, hdlc for high-density lipoprotein cholesterol level, fg for fasting blood glucose level, sbp for systolic blood pressure, MDRD for Modification of Diet in Renal Disease formula, and EGFR for estimated glomerular filtration rate. Patients and healthcare professionals can adjust treatment and personal lifestyles based on the above results.
[0077] In some embodiments, the model building and optimization module of the present disclosure builds a machine learning model based on a feature set of a third dataset of the present disclosure to predict pancreatic β-cell function in non-diabetic patients, the feature set of which includes age, sex, race, body mass index (BMI), fasting blood glucose / plasma blood glucose (fg / FBG / FPG), triglyceride levels (tg), total cholesterol (TC / chol), high-density lipoprotein (HDL) cholesterol (HDL cholesterol / hdlc), glutamate oxaloacetate transaminase (got), glutamate pyruvate transaminase (gpt), waist circumference (wc), total bilirubin (tb), albumin (alb), systolic blood pressure (sbp), diastolic blood pressure (dbp), estimated glomerular filtration rate (eGFR), and / or creatinine (cre).
[0078] In some embodiments, the machine learning of the present disclosure using a third dataset derived from a third database (e.g., an integrated database of the NHANES and MJ databases) can be applied to find a good model for predicting IR and β-cell function in a non-diabetic population with high diagnostic performance (AUC of 0.88) based on a feature set of the present disclosure that can be easily obtained from patients.
[0079] In some embodiments, the predictive IR using this cross-racial algorithm was associated with clinical significance validated using the TWB database, specifically with worsening of cardiovascular mortality and all-cause mortality. Furthermore, although its predictive performance for cardiovascular mortality risk was lower (smaller C index) compared to the Framingham score,
[0080] Even in patients with a HOMA-IR > 2.5 (indicating a low predicted cardiovascular risk based on Framingham scores), there is a risk of high cardiovascular mortality and all-cause mortality. In particular, for the non-DM population in this disclosure, Framingham does not fully predict cardiovascular mortality and all-cause mortality. This also reflects the importance of this disclosure.
[0081] Numerous machine learning studies on diabetes mellitus (DM) have been published to date. However, findings from the UKPDS, ADVANCE, VADT, and ACCORD studies suggest that early glucose control is crucial for achieving favorable vascular outcomes. Based on the DM spectrum, it can be inferred that identifying IR and β-cell function is extremely important. However, calculating HOMA-IR and HOMA-β requires insulin levels from patients that are not regularly collected in clinical practice. Therefore, it is important to devise predictive models for IR and β-cell function (without using actually measured insulin levels).
[0082] The predictive machine learning model disclosed herein has two notable strengths: superior predictive performance (high AUC) and generalizability (it uses only easily obtainable features and is applicable across racial groups).
[0083] The significant impact of fasting plasma insulin levels on HOMA-IR also suggests the importance of this disclosure. Firstly, this disclosure does not use fasting plasma insulin level data in the calculation of HOMA-IR, despite it having the greatest impact on IR. Secondly, HOMA-IR values differ significantly between different racial populations (e.g., Asian and Caucasian) (2.84 (2.77-2.91) in the NHANES population (Caucasian), and 1.92±1.26 in the MJ database (Asian)). This difference in IR is mainly due to fasting plasma insulin levels (12.1±10.35 mIU / L in NHANES, compared to 7.67±4.76 in the MJ database), rather than FPG levels (91.54 mg / dL in NHANES, compared to 100.07±8.3 in the MJ database). This disclosure aims to train a machine learning model that can be used in general across various ethnic groups to predict intra-radicular risk (IR) without using fasting plasma insulin levels.
[0084] Those skilled in the art will readily see that many modifications and alterations can be made to the methods and apparatus while maintaining the teachings of the present invention. Therefore, the above disclosure should be limited only by the limits and boundaries of the appended claims.
Claims
1. A system for predicting insulin resistance or pancreatic β-cell function in subjects requiring a study, A database configured to provide a dataset, A feature extraction module configured to collect and process features from the dataset in order to generate a feature set relating to the subject, wherein the feature set comprises the subject's age, sex, race, and body mass index. The module includes a machine learning model built based on the feature set to predict the insulin resistance or pancreatic β-cell function of the subject, and a module configured to build and optimize the model to classify the subject into an insulin-resistant group and a non-insulin-resistant group, or a β-cell deficiency group and a non-β-cell deficiency group, The aforementioned subjects are non-diabetic patients, according to the system.
2. The system according to claim 1, wherein the feature set further comprises the subject's fasting blood glucose level and / or glycated hemoglobin level.
3. The system according to claim 2, wherein the feature set further comprises the subject's total cholesterol level and / or high-density lipoprotein cholesterol level.
4. The system according to claim 1, wherein the feature set further comprises the triglyceride levels of the subject.
5. The system according to claim 4, wherein the feature set further comprises the subject's total cholesterol level and / or high-density lipoprotein cholesterol level.
6. The system according to claim 5, wherein the feature set further comprises the subject's fasting blood glucose level and at least one selected from the group consisting of glutamate oxaloacetate transaminase level, glutamate pyruvate transaminase level, total bilirubin level, and albumin level.
7. The system according to claim 1, wherein the feature set further comprises the total bilirubin level and / or albumin level of the subject.
8. The system according to claim 7, wherein the feature set further comprises the subject's fasting blood glucose level and at least one selected from the group consisting of total cholesterol level and high-density lipoprotein cholesterol level.
9. The system according to claim 1, wherein the machine learning model is trained with the features labeled with outcomes related to the impaired insulin resistance and / or β-cell function of the subject.
10. The aforementioned database is A first database derived from the first population of the aforementioned subjects, The second database comprises a second population of the subjects, The system according to claim 1, wherein the race of the first population is different from the race of the second population.
Citation Information
Patent Citations
CN115188493A
JP2006304833A
JP2017205161A
WO2023239869A1