Risk prediction model for future disease progression in early stage of diabetes and construction method thereof

By using the DDRTree algorithm to reduce and visualize the metabolic phenotype of prediabetes individuals, a two-dimensional tree structure is constructed, which solves the problem of difficult to identify individual risk differences in prediabetes individuals and predict future disease progress in the prior art, and achieves more accurate risk assessment and personalized management.

CN120148854APending Publication Date: 2025-06-13SHANGHAI INST FOR ENDOCRINE & METABOLIC DISEASES +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510223018.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify risk differences among individuals in pre-diabetic stages, and it is difficult to effectively predict individual future disease progression into diabetes and complications.

Method used

The discriminant dimensionality reduction tree (DDRTree) algorithm was used to reduce and visualize 12 consecutive metabolic clinical phenotype variables of individuals in prediabetes, and a two-dimensional tree structure was constructed, and the individual was divided into four different subtypes, and the individual's coordinates in the two-dimensional tree were positioned to construct a risk prediction model for future disease progression in prediabetes.

Benefits of technology

Through the application of the DDRTree algorithm, metabolic heterogeneity and future risk of disease progression in individuals in pre-diabetes can be more accurately captured, and personalized prediction and management strategies can be provided to help early intervention and effectively manage diabetes and complications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148854A_ABST
    Figure CN120148854A_ABST
Patent Text Reader

Abstract

The invention discloses a risk prediction model for future disease progression of prediabetic patients and a construction method thereof, and the method comprises the steps: collecting clinical index data of the prediabetic patients in a training queue and a verification queue, including gender, age, WHR, BMI, HOMA-IR, HOMA-B, ALT, AST, GGT, FPG, PBG, HbA1c, TG and HDL-C; a DDRTree algorithm is used to establish a two-dimensional tree structure of the heterogeneity of the early stage of diabetes in the training queue; dividing the prediabetic population into different subtypes, obtaining the distribution coordinates of the prediabetic individuals in the two-dimensional tree, evaluating the risk that the prediabetic progression is T2DM, CKD and CVD diseases, and constructing a risk prediction model of the future disease progression of the prediabetic; model accuracy verification is carried out in the verification queue, the risk of future disease progression in the early stage of diabetes is accurately evaluated, and early intervention on diabetes and complications is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of disease prediction models, and particularly relates to a risk prediction model for future disease progression in pre-diabetes and a construction method thereof. Background Art

[0002] Pre-diabetes represents an intermediate stage of glucose dysregulation and may precede type 2 diabetes (T2DM). In the China Cardiovascular Metabolic and Malignancy Cohort with a 5-year follow-up, the cumulative incidence of pre-diabetes, jointly defined by impaired fasting glucose tolerance, impaired glucose tolerance, and elevated HbA1c, progressing to diabetes was 21.3%. In another randomized clinical trial of diabetes prevention, during the 30-year follow-up of the China Daqing Study, the cumulative diabetes incidence among pre-diabetes participants in the control group reached 95.9%. Therefore, early identification of high-risk pre-diabetes individuals, effective screening, and risk stratification management are crucial for preventing disease progression to T2DM and its complications.

[0003] Individuals with pre-diabetes show significant differences in insulin sensitivity, insulin secretion, blood glucose levels, and lipid metabolism, and these factors affect the metabolic state, pathophysiological mechanisms, and the occurrence of various complications such as progression from pre-diabetes to T2DM, cardiovascular disease (CVD), and chronic kidney disease (CKD). Measuring blood glucose alone is insufficient to capture the risk differences among individuals, and clinical indicators such as body fat, insulin resistance, blood lipids, blood pressure, and liver enzymes can facilitate the early identification of precise subsets of pre-diabetes.

[0004] The discriminant dimensionality reduction tree (DDRTree) is a computational method for visualizing and analyzing high-dimensional biological data. It integrates the principles of dimensionality reduction and decision tree algorithms. Recent research has applied the DDRTree algorithm to distinguish the heterogeneity between type 2 diabetes and type 1 diabetes, achieving clear two-dimensional visualization of complex and continuous phenotypic variations among individuals. Considering the differences in the pathophysiology of prediabetes, applying this method can provide new insights into the heterogeneity of prediabetes. Compared with traditional data-driven disease classification methods such as k-means, the DDRtree method has no clear classification boundaries and can better reflect the complexity of the classification of the prediabetes population and the continuity of future disease progression. Wagner and his team used the K-means clustering method to classify prediabetes into different groups to more accurately describe the future diabetes and complication risks of each group. The K-means algorithm is widely used in disease classification research due to its simplicity and applicability to large datasets. It can define clear clustering categories and performs better when the data structure is relatively simple and the disease type boundaries are clear. However, when dealing with continuous phenotypic variables and the heterogeneity of diseases with complex structures, the DDRTree algorithm performs more excellently. It does not simply define several specific disease subtypes and regard different individuals within a certain subtype as homogeneous. Instead, it combines continuous phenotypes to personalized predict the future disease progression of specific individual patients, can capture the dynamic trajectory of disease development, and thus has greater clinical utility. The clustering results of DDRTree show an easy-to-visualize and easier-to-understand two-dimensional tree structure. Phenotypic variables and related disease risks are continuously distributed along the entire two-dimensional tree. The coordinates of individuals on the two-dimensional tree can be used to predict the future multiple disease progressions of prediabetes individuals. Summary of the Invention

[0005] Aiming at the above technical problems, the main object of the present invention is to provide a risk prediction model for the future disease progression of prediabetes, aiming to clarify: 1) how the tree structure distinguishes the metabolic parameters between prediabetes individuals; 2) how the heterogeneity of specific metabolic phenotypes predicts the progression of prediabetes individuals to diabetes and its complications. Prediabetes is the intermediate stage of diabetes development and has multi-level heterogeneity. By accurately assessing the risk of future disease progression of prediabetes, early intervention, effective management, and targeted prevention of diabetes and its complications can be achieved.

[0006] Another object of the present invention is to provide a method for constructing a risk prediction model for future disease progression of prediabetes. By applying the DDRTree algorithm, heterogeneous data of prediabetes from 55,777 participants in the China Cardiometabolic Disease and Cancer Cohort (4C) study are used. Based on 12 continuously measured metabolically clinically available phenotypes, the prediabetes population is divided into four different subtypes, the coordinates of individuals in the two-dimensional tree are located, a risk prediction model for future disease progression of prediabetes is constructed, and the risks of future disease progression of prediabetes to T2DM, CKD, and CVD are evaluated through it.

[0007] The above object of the present invention is achieved by the following technical solutions:

[0008] In the first aspect of the present invention, a method for constructing a risk prediction model for future disease progression of prediabetes is provided, including the following steps:

[0009] S1: Collect clinical index data of prediabetes patients in the training cohort and the validation cohort, including the gender, age of the patients, and 12 clinical phenotype variables, namely waist-to-hip ratio (WHR), body mass index (BMI), insulin resistance index (HOMA-IR), pancreatic islet β-cell function index (HOMA-B), alanine aminotransferase (ALT), aspartate aminotransferase (AST), γ-glutamyl transferase (GGT), fasting plasma glucose (FPG), 2-hour postprandial blood glucose (PBG), glycated hemoglobin (HbA1c), triglyceride (TG), and high-density lipoprotein cholesterol (HDL-C);

[0010] S2: Use the DDRTree algorithm to perform dimensionality reduction and visualization on the clinical index data of prediabetes patients in the training cohort in step S1, and construct a two-dimensional tree structure of prediabetes heterogeneity;

[0011] S3: Based on the two-dimensional tree structure in step S2, divide the prediabetes population into different subtypes, obtain the distribution coordinates of prediabetes individuals in the two-dimensional tree, evaluate the risks of prediabetes progression to T2DM, CKD, and CVD diseases, and construct a risk prediction model for future disease progression of prediabetes;

[0012] S4: Verify the model accuracy of the risk prediction model for future disease progression of prediabetes in step S3 in the validation cohort;

[0013] Wherein, the prediabetes is defined as: among participants without diabetes, FPG is 5.6 mmol / L to 6.9 mmol / L, or PBG is 7.8 mmol / L to 11.0 mmol / L, or HbA1c is 5.7% to 6.4%.

[0014] Preferably, in step S1, the training cohort is from the China Cardiometabolic Disease and Cancer Cohort (4C), including 55,777 participants with prediabetes, and clinical variable data of the prediabetes population is collected; the validation cohort is a community-based prospective cohort established in Songnan area, Shanghai, China from June to August 2009, including 4,012 individuals aged ≥40 years old, followed up until November 2021.

[0015] Preferably, in step S1, clinical index data of prediabetes patients are collected for data screening and variable extraction, and 12 clinical phenotype variables related to the disease risks of T2DM, CKD, and CVD are obtained, including WHR, BMI, HOMA-IR, HOMA-B, ALT, AST, GGT, FPG, PBG, HbA1c, TG, and HDL-C.

[0016] Preferably, in step S2, a linear regression model is established in the training cohort, with gender and age as independent variables and 12 clinical phenotype variables as dependent variables respectively, to obtain the residual matrix of the 12 clinical phenotype variables. The DDRTree algorithm is used to reduce the dimension and visualize the residual matrix, and the prediabetes population is displayed as a two-dimensional tree structure, showing phenotypic changes and individual distributions in the tree structure, so as to reflect the heterogeneity of prediabetes.

[0017] Preferably, in step S3, the prediabetes subtype characterized by hyperglycemia, insulin resistance, obesity, elevated triglycerides, and elevated liver enzymes has a high risk of progressing to T2DM.

[0018] Preferably, in step S3, the prediabetes subtype characterized by obesity, insulin resistance, hyperglycemia, and dyslipidemia has a high risk of progressing to CKD.

[0019] Preferably, in step S3, the subtypes characterized by hyperglycemia, insulin resistance, obesity, elevated triglycerides, and elevated liver enzymes and the subtypes characterized by obesity, insulin resistance, hyperglycemia, and dyslipidemia have a high risk of progressing to CVD, and the CVD subtype distributions are very different.

[0020] Preferably, in step S3, the progression of prediabetes to CVD diseases includes stroke, myocardial infarction, and heart failure subtypes.

[0021] In the second aspect of the present invention, a risk prediction system for the progression of prediabetes to T2DM, CKD, and CVD diseases is provided, including:

[0022] The data collection module is used to collect heterogeneous phenotypic data of patients with prediabetes, including gender, age, and 12 clinical phenotypic variables, namely WHR, BMI, HOMA-IR, HOMA-B, ALT, AST, GGT, FPG, PBG, HbA1c, TG, and HDL-C;

[0023] Data preprocessing module, used to remove missing values ​​and outliers, where outliers are values ​​outside the range of mean ± 5 standard deviations (SD);

[0024] The data analysis module is used to construct a two-dimensional tree structure using the DDRTree algorithm for the 12 clinical phenotype variables in the training cohort;

[0025] The prediction model building module is used to divide the prediabetes population into different subtypes, use the two-dimensional coordinates of the prediabetes individuals in the tree structure as independent variables, establish a Cox regression model, evaluate the risk of prediabetes progressing to T2DM, CKD and CVD diseases, and build a risk prediction model for future disease progression in prediabetes;

[0026] The verification module is used to verify the accuracy of the risk prediction model for future disease progression in prediabetes by using a verification cohort.

[0027] Preferably, the data acquisition module collects clinical indicator data of prediabetes patients in the training cohort and the validation cohort; the data preprocessing module performs data screening and feature extraction to obtain 12 clinical phenotypic variables corresponding to the risk of prediabetes progressing to T2DM, CKD and CVD diseases, including WHR, BMI, HOMA-IR, HOMA-B, ALT, AST, GGT, FPG, PBG, HbA1c, TG and HDL-C.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] 1. The DDRTree algorithm is applied to the pre-diabetes population in a national prospective cohort study for the first time in the present invention. A sufficient sample size can construct a robust two-dimensional tree structure, which performs excellently in exploring disease progression and understanding disease heterogeneity, making it a valuable tool for capturing the complex dynamics of disease evolution. This analysis reduces the dimensionality of multi-dimensional continuous clinical variables in the pre-diabetes population to describe disease heterogeneity, which helps to more personalized guide pre-diabetes care. The prevention strategy based on DDRTree will likely bring benefits to high-risk populations of diabetes, CKD, CVD and their subtypes. In a large-scale national pre-diabetes prospective cohort, the DDRTree algorithm can capture the complex interactions of disease risk factors and track the complex dynamic changes of the disease, so as to realize the risk prediction of pre-diabetes progression and provide guidance for the personalized care of pre-diabetes.

[0030] 2. In the risk prediction model of future disease progression of pre-diabetes in the present invention, the risk of T2DM in Group 4 characterized by hyperglycemia, insulin resistance, obesity, elevated triglycerides and liver enzymes is the highest, while the risk of CKD in Group 3 characterized by obesity, insulin resistance, hyperglycemia and dyslipidemia is the highest. The risk of CVD in Group 3 and Group 4 is relatively high, and the distribution of CVD subtypes varies greatly. These findings are well verified in the external validation cohort SN_2009 - 2021. In addition, a user-friendly online tool is also developed to evaluate the future disease risk of individuals in the pre-diabetes population, with strong applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 Phenotypic characteristics of 55,777 pre-diabetes participants in the training cohort in the embodiment; among them, A: Distribution of 12 clinical phenotypic variables (WHR, BMI, HOMA-IR, HOMA-B, ALT, AST, GGT, FPG, PBG, HbA1c, TG and HDL-C) on the pre-diabetes two-dimensional tree, and each point in the figure represents an individual; B: Correlation coefficients (95% confidence interval) between the contents of 12 phenotypic variables and the two-dimensional coordinates, red: correlation coefficient between the content and the first-dimensional coordinate; green: correlation coefficient between the content and the second-dimensional coordinate; C: Spatial autocorrelation (Moran index) of 12 phenotypes on the two-dimensional tree, and a higher Moran index indicates that the phenotype has higher spatial autocorrelation; D: Tree structure of the two-dimensional tree and pre-diabetes subtypes (subtypes 1 to 4).

[0032] Figure 2For predicting the probabilities of developing T2D and CKD in pre-diabetic patients in the training cohort of the embodiment within the next 5 years, a COX risk model constructed with the coordinates of an individual on a two-dimensional tree as independent variables is used as the disease prediction model; where A: the probability of developing T2DM within 5 years; B: the probability of developing CKD within 5 years; C: the HR (95% confidence interval) of the occurrence probability on the two-dimensional coordinate, red: the HR value of the occurrence probability on the first-dimensional coordinate; green: the HR value of the occurrence probability on the second-dimensional coordinate; D: the spatial autocorrelation of the occurrence probabilities of T2DM and CKD on the two-dimensional tree.

[0033] Figure 3 For predicting the occurrence probabilities of 5-year CVD and its subtypes in pre-diabetic participants in the training cohort of the embodiment (n = 55,777), a COX risk model constructed with the coordinates of an individual on a two-dimensional tree as independent variables is used as the disease prediction model; where A: the probabilities of developing CVD, stroke, myocardial infarction, and heart failure within 5 years; B: the HR (95% confidence interval) of the occurrence probability on the two-dimensional coordinate, red: the HR value of the occurrence probability on the first-dimensional coordinate; green: the HR value of the occurrence probability on the second-dimensional coordinate; C: the spatial autocorrelation of the occurrence probabilities of CVD and its subtypes on the two-dimensional tree.

[0034] Figure 4 For predicting the occurrence probabilities of 10-year T2DM, CVD, and CKD in pre-diabetic participants in the validation cohort SN_2009 study of the embodiment (n = 1,651), a method of mapping first and then predicting is adopted. First, calculate the two-dimensional coordinates of an individual, map them to the two-dimensional tree, obtain the coordinates of the nearest neighboring points on the two-dimensional tree, and construct a COX risk model with the neighboring point coordinates as independent variables to predict the occurrence probabilities of T2DM and CVD within 10 years, and construct a Logistic regression model to predict the occurrence probability of CKD; A: the occurrence probability of T2DM within 10 years; B: the occurrence probability of CVD within 10 years; C: the occurrence probability of CKD; D: the HR (95% confidence interval) of the occurrence probability on the two-dimensional coordinate, red: the HR value of the occurrence probability on the first-dimensional coordinate; green: the HR value of the occurrence probability on the second-dimensional coordinate; E: the spatial autocorrelation of the occurrence probabilities of T2DM, CVD, and CKD on the two-dimensional tree.

[0035] Figure 5 For the inclusion and exclusion process of 55,777 pre-diabetic participants in the training cohort of the embodiment.

[0036] Figure 6 For the inclusion and exclusion process of 1,651 pre-diabetic participants in the validation cohort SN_2009 study of the embodiment. Detailed implementation manners

[0037] To more fully understand and demonstrate the technical solutions, objectives, and advantages of the present invention, the technical effects produced by the present invention will be further described in detail and completely below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. It should be noted that for those of ordinary skill in the art, other embodiments obtained without departing from the concept of the present invention all fall within the protection scope of the present invention.

[0038] In the following embodiments, the 4C study is a prospective cohort study of Chinese adults aged 40 and above, and the research subjects are older pre-diabetes patients. Pre-diabetes is an intermediate stage in the development of diabetes and has multi-faceted heterogeneity. Emphasis is placed on conducting precise risk assessments in order to achieve early intervention, effective management, and targeted prevention of diabetes and its complications.

[0039] In the following embodiments, the DDRTree algorithm is used to study the pre-diabetes heterogeneity of 55,777 participants from the China Cardiometabolic Disease and Cancer Cohort (4C) study. According to 12 clinical phenotype variables including WHR, BMI, HOMA-IR, HOMA-B, ALT, AST, GGT, FPG, PBG, HbA1c, TG, and HDL-C, different risks of pre-diabetes progressing to T2DM, CKD, and CVD (including stroke, myocardial infarction, and heart failure subtypes) are evaluated. Group 4, characterized by hyperglycemia, insulin resistance, obesity, elevated triglycerides, and elevated liver enzymes, has the highest risk of T2DM, while Group 3, characterized by obesity, insulin resistance, hyperglycemia, and dyslipidemia, has the highest risk of CKD. Groups 3 and 4 have higher risks of CVD, and there are significant differences in the distribution of CVD subtypes. These findings are well validated in an external validation cohort study with a follow-up of more than 10 years starting from SN_2009 - 2021. The complex dynamic changes in the progression of pre-diabetes are demonstrated through the DDRTree algorithm, and the importance of personalized management strategies in pre-diabetes care is emphasized.

[0040] Example 1

[0041] 1. Materials and Methods

[0042] Cohort and variable definitions. Two community-dwelling prospective cohort datasets were used for analysis:

[0043] (1) The 4C study - a national, multi-center, prospective, population-based study of Chinese adults;

[0044] (2) The SN study - a prospective cohort study established in the Songnan area of Shanghai, China.

[0045] 2. The 4C Study

[0046] The 4C study aimed to examine the relationship between glycemic parameters and clinical outcomes, including diabetes, cardiovascular diseases, chronic kidney disease, and mortality. From 2011 to 2012, a total of 193,846 adults aged 40 and above were recruited from 20 communities in different geographical regions of China. At the baseline visit, a comprehensive set of questionnaires, clinical measurements, and laboratory tests were conducted. Body weight, height, waist circumference, and hip circumference were measured by trained research nurses according to a standard protocol. Three blood pressure measurements obtained after 5 minutes of sitting were averaged using an automated electronic device (Omron HEM-752FUZZY model, Dalian, China). All participants underwent an oral glucose tolerance test (OGTT), and blood glucose was obtained at 0 and 2 hours during the test. Under strict quality control procedures, plasma glucose concentration was analyzed using the glucose oxidase or hexokinase method within two hours after blood sample collection. A glycated hemoglobin collection system (Bio-Rad Laboratories, California, USA) was used to collect capillary fingertip blood samples, which were transported to the certified central laboratory of Ruijin Hospital at 2 - 8°C. This clinical laboratory has been double-certified by the National Glycohemoglobin Standardization Program of the United States and the College of American Pathologists (CAP) laboratory.

[0047] HbA1c was measured using high-performance liquid chromatography (VARIANT II system, Bio-Rad Laboratories, California, USA). Serum insulin, ALT, AST, GGT, HDL-C, and TG were detected using an automated analyzer (ARCHITECT ci16200 analyzer, Abbott Laboratories, Illinois, USA) in the central laboratory. The glomerular filtration rate eGFR was calculated using the Chronic Kidney Disease Epidemiology Collaboration (CKD-EPI) equation, and insulin resistance was estimated using HOMA-IR, with the formula: fasting insulin (μIU / mL) × fasting blood glucose (mg / dL) / 405. The function of β-cells was evaluated by HOMA-B, with the formula: (360 × fasting insulin [μIU / mL]) / (fasting blood glucose [mg / dL] - 63).

[0048] During the period from 2014 to 2016, all participants were invited to participate in a face-to-face follow-up. Trained staff asked about medical history using the same standard questionnaire as at the baseline visit. Anthropometric measurements and blood pressure were measured for OGTT, and blood samples were obtained using the same protocol as at the baseline examination.

[0049] In the 4C study, 64,650 participants were diagnosed with prediabetes based on OGTT and HbA1c measurements. Among them, 7,416 participants with missing biochemical test and physical measurement data, and 1,457 participants with outliers in clustering variables were excluded. Finally, 55,777 individuals diagnosed with prediabetes were included in the analysis ( Figure 5) To explore the phenotypic heterogeneity of prediabetes, a forward stepwise logistic regression model was used to select phenotypes from common clinical indicators such as blood glucose, blood lipids, blood pressure, and liver enzymes. Finally, 12 clinical indicators, namely WHR, BMI, HOMA-IR, HOMA-B, ALT, AST, GGT, FPG, PBG, HbA1c, TG, and HDL-C, were identified as key biomarkers related to the future disease progression of prediabetes patients and could be used to distinguish the disease heterogeneity of prediabetes.

[0050] The 4C study has been approved by the Medical Ethics Committee of Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, and all participants have signed written informed consent forms.

[0051] 3. SN Cohort

[0052] The SN_2009 study is a community-based prospective cohort established in Songnan, Shanghai, China from June to August 2009, including 4,012 individuals aged ≥40 years, followed up until November 2021. All participants underwent an OGTT, and blood samples were collected at 0 and 2 hours. FPG and PBG were measured using the glucose oxidase method on an automated analyzer (ADVIA-1650 Chemistry System, Germany). Plasma and serum samples were collected and immediately stored in Eppendorf tubes at -80°C. ALT, AST, GGT, HDL-C, and TG were detected using an automated analyzer (ADVIA VR 1650 Chemistry System, Germany); serum insulin was measured using electrochemiluminescence immunoassay (Roche Diagnostics, Basel, Switzerland), and HbA1c was determined using a fully automated high-performance liquid chromatography analyzer (Bio-Rad, Hercules, CA). Based on OGTT and HbA1c tests, 1,651 adults with prediabetes were identified, and complete baseline information was obtained from this cohort ( Figure 6 ), with a median follow-up time of 12.4 (12.3 - 12.4) years.

[0053] Definition of prediabetes: Prediabetes was defined according to the 2010 American Diabetes Association criteria, that is, in participants without diabetes, FPG was 5.6 mmol / L to 6.9 mmol / L, or PBG was 7.8 mmol / L to 11.0 mmol / L, or HbA1c was 5.7% to 6.4%.

[0054] 4. Outcome Definition

[0055] T2DM: In the 4C study, diabetes events were defined as FPG ≥ 7.0 mmol / L, 2-hour post-load plasma glucose ≥ 11.1 mmol / L, or HbA1c ≥ 6.5%, or physician-diagnosed diabetes. In the SN study, new-onset diabetes was defined as physician-diagnosed diabetes based on the National Health Insurance System.

[0056] CVD: Information on vital status and clinical outcomes was collected from the local death and disease registries of the National Disease Surveillance Points System and the National Health Insurance System. Major CVD events included first non-fatal myocardial infarction, non-fatal stroke, hospitalization or treatment for heart failure during follow-up, and cardiovascular death.

[0057] CKD: Included new-onset renal failure requiring dialysis or alternative treatment, death due to renal causes, decline in eGFR (eGFR < 60 mL / min / 1.73m 2 ) at follow-up or a certain decline in eGFR category (from eGFR ≥ 90 mL / min / 1.73m 2 at baseline to 60 - 89 mL / min / 1.73m 2 ) at follow-up), accompanied by a 25% or greater decrease in eGFR from baseline.

[0058] 5. Statistical Analysis

[0059] The baseline characteristics of the study participants were presented as the mean ± standard deviation or median (interquartile range) for continuous variables and the number (percentage) for categorical variables. Missing values or outliers with a difference from the mean of more than 5 SD were excluded first. A linear regression model was established with gender and age as independent variables and 12 clinical phenotype variables as dependent variables respectively to obtain a residual matrix of 12 variables. The DDRTree algorithm in the Monocle R package was used to reduce the dimension of the phenotypic residual matrix based on age and gender to obtain a visualized two-dimensional tree structure of pre-diabetes heterogeneity. This tree structure is a continuous display of multidimensional clinical indicators in two dimensions, with each point representing an individual, the branches reflecting the subtype characteristics of the population, and the points on the branches reflecting individual characteristics. The clinical variable characteristics and future disease risks of the pre-diabetes population can be intuitively displayed on this two-dimensional tree.

[0060] To construct a predictive model for future disease risks, based on a two-dimensional tree structure, using the two-dimensional coordinates of an individual in the tree structure as independent variables, a Cox regression model was established to evaluate the risks of pre-diabetes progressing to diseases such as T2DM, CKD, and CVD. To observe the clinical variable characteristics of the pre-diabetes population and the heterogeneity of future disease risks, the clinical variable characteristics and the predicted disease risks were respectively displayed on the two-dimensional tree. To quantitatively evaluate the distribution trends of clinical variable characteristics and future disease risks on the two-dimensional tree, a linear regression model was used to calculate the correlation coefficients (95% confidence intervals) between 12 phenotypic variables, disease risks (T2DM, CKD, CVD and their subtypes) and the two-dimensional coordinates. To quantitatively evaluate the clustering of clinical variable characteristics and future disease risks on the two-dimensional tree, global and local spatial autocorrelation coefficients (Moran's index) were introduced for spatial autocorrelation degree analysis. A positive Moran's index indicates that the variables are clustered together, while a negative value indicates that the variables are not clustered. All statistical analyses were performed using R software (version 4.3.3).

[0061] 6. Results

[0062] The baseline characteristics of the 4C study population are shown in Table 1. The 12 variables used to determine the heterogeneity of pre-diabetes were (mean [SD] or median [IQR]): WHR 0.88 (0.06), BMI 24.42 kg / m 2 (3.28 kg / m 2 )、HOMA-IR 1.67 (1.18 - 2.34), HOMA-B 67.42 (47.86 - 94.51), ALT 14.00 U / L (10.80 - 20.00 U / L), AST 21.78 U / L (7.39 U / L), GGT 20.00 U / L (14.00 - 30.00 U / L), FPG 5.57 mmol / L (0.52 mmol / L), PBG 7.12 mmol / L (1.68 mmol / L), HbA1c 5.81% (0.33%), TG 1.31 mmol / L (0.94 mmol / L - 1.87 mmol / L), HDL-C 1.35 mmol / L (0.35 mmol / L).

[0063] Table 1: Baseline characteristics of the study population

[0064] Variables 4C Study (N = 55,777) SN Study (N = 1,651) Males 18,762(33.64%) 608(36.8%) Age (year) 56.26±8.50 58.32±9.07 WHR 0.88±0.06 0.89±0.06 <![CDATA[BMI (kg / m 2 )]]> 24.42±3.28 24.90±3.39 FPG (mmol / L) 5.57±0.52 4.99±0.53 PBG (mmol / L) 7.12±1.68 6.98±1.65 HbA1c (%) 5.81±0.33 5.98±0.25 HOMA-IR 1.67(1.18-2.34) 1.45(0.93-2.14) HOMA-B 67.42(47.86-94.51) 88.64(58.00-138.00) TG (mmol / L) 1.31(0.94-1.87) 1.32(0.96-1.89) HDL-C (mmol / L) 1.35±0.35 1.38±0.31 ALT (U / L) 14.00(10.80-20.00) 17.00(13.00-25.00) AST (U / L) 21.78±7.39 22.11±6.79 GGT (U / L) 20.00(14.00-30.00) 21.63(15.93-31.91)

[0065] Note: Data are presented as mean ± standard deviation, number (percentage), or median (interquartile range); ALT: alanine aminotransferase; AST: aspartate aminotransferase; WHR: waist-hip ratio; BMI: body mass index; FPG: fasting plasma glucose; GGT: gamma-glutamyl transferase; HDL-C: high-density lipoprotein cholesterol; HOMA-B: homeostasis model assessment of beta-cell function; HOMA-IR: homeostasis model assessment of insulin resistance; PBG: post-load plasma glucose; TG: triglyceride.

[0066] 6.1, Phenotypic characteristics and gradients of the tree structure

[0067] Display the data characteristics of 12 independent clinics on a two-dimensional tree ( Figure 1 A), and use the correlation coefficient to evaluate the distribution trend of clinical variables on the dimensional tree ( Figure 1 B), and use the spatial autocorrelation coefficient to evaluate the aggregation of clinical variables on the two-dimensional tree ( Figure 1 C). WHR, BMI, HOMA-IR, and HOMA-B showed a consistent gradient distribution pattern in both dimensions of the tree, with the highest values in the lower left branch and the lowest values in the upper right branch. The liver enzyme levels, including ALT, AST, and GGT, had significant gradients in both dimensions of the tree, with the lowest in the lower right branch and the highest in the upper left branch. In addition, blood glucose characteristics such as FPG and PBG showed a gradient decline along dimension 1, with higher concentrations aggregated on the left side, while HbA1c showed an opposite pattern in dimension 2. TG and HDL-C showed completely opposite gradients in both dimensions, with HDL-C reaching the highest values in the upper half and the right half of the tree, while TG was similar to the blood glucose characteristics and had higher levels in the left half of the tree. The global spatial autocorrelation coefficient (Moran's index) describes the strength of the spatial correlation of the two-dimensional tree ( Figure 1 C). Among the 12 variables analyzed, ALT (r = 0.72), AST (r = 0.72), and HOMA-IR (r = 0.64) had the greatest impact on the tree coordinates, followed by BMI and HOMA-B, and the impacts of HDL-C, GGT, TG, and WHR were relatively small. Blood glucose characteristics, including FPG, PBG, and HbA1c, showed the smallest changes throughout the tree structure. Given the definition of prediabetes, the change in baseline blood glucose levels in this population was relatively small, thus limiting its potential for discrimination.

[0068] According to the clustering results, each branch within the tree was grouped for subsequent description ( Figure 1D). Individuals in Group 1 had the healthiest metabolism, with the lowest obesity index, insulin resistance, blood glucose, and lipid levels. At the same time, Group 2 had higher levels of obesity index, insulin resistance, blood glucose, and lipid traits, but the lowest liver enzyme levels. Patients in Group 3 and Group 4 both had unhealthy metabolic characteristics, characterized by high TG, high blood glucose levels, obesity, and insulin resistance. The difference was that: Group 3 included patients with the highest HbA1c and the lowest HDL-C, while Group 4 included patients with elevated liver enzymes.

[0069] 6.2, Relationship between Heterogeneity of Prediabetic Patients and Clinical Outcomes

[0070] During a 5-year follow-up, 3,615 cases of T2DM, 1,020 cases of CKD, and 933 cases of CVD (582 cases of stroke, 142 cases of myocardial infarction, and 70 cases of heart failure) were observed. Figure 2 and Figure 3 Visualizing the heterogeneity of clinical outcomes of the progression of prediabetes in four groups, the probabilities of T2DM, CKD, and CVD showed differences throughout the DDRtree. The risk of T2DM peaked in the upper left quadrant of the tree (Group 4), with an HR (95% CI) of 0.62 (0.60 - 0.65) in Dimension 1 and an HR (1.09 - 1.20) of 1.14 in Dimension 2 ( Figure 2 A and 2C). The high probability of T2DM was associated with higher baseline blood glucose characteristics, liver enzyme levels, insulin resistance, and obesity. The risk of CKD was highest in the third group, with an HR (95% CI) of 0.84 (0.78 - 0.91) in Dimension 1 and an HR (95% CI) of 0.69 (0.62 - 0.76) in Dimension 2. The risk of developing CKD was jointly driven by obesity, insulin resistance, blood glucose characteristics, and dyslipidemia.

[0071] The overall cardiovascular outcomes increased in the left part of the tree (Group 3 and Group 4), with an HR (95% CI) of 0.91 (0.84 - 0.98) in Dimension 1 and an HR (95% CI) of 0.99 (0.90 - 1.09) in Dimension 2 ( Figure 3 ), corresponding to the area with an unhealthy metabolic profile. Then, a detailed assessment was conducted for cardiovascular disease subtypes (including myocardial infarction, stroke, and heart failure), and their probabilities were shown on the two-dimensional tree ( Figure 3 A). Stroke accounted for 57% of the overall CVD, and its risk distribution was similar, with higher risks in Group 3 and Group 4 (0.83 (0.75 - 0.92) in the first dimension and 0.98 (0.86 - 1.12) in the second dimension). However, for other CVD subtypes, namely myocardial infarction (highest risk in Group 3) and heart failure (highest risk in Group 4), the significant distributions in the second dimension were different, being 0.74 (0.56 - 0.98) and 1.10 (0.74 - 1.62) respectively ( Figure 3B).

[0072] 6.3 External validation of pre-diabetes heterogeneity

[0073] For external validation of the method and results, the two-dimensional tree established in the 4C study was used as the reference tree, and individuals in the external dataset were mapped into the reference tree. The mapping function consisted of two steps as follows.

[0074] Step 1: Use a generalized additive model to predict the two-dimensional coordinate dimensions (x-axis coordinate and y-axis coordinate) of each individual (Table 2).

[0075] Step 2: Calculate the Euclidean distance between the coordinates of the external individual and the coordinates of all individuals on the reference tree, and map the coordinate to the nearest point on the reference tree.

[0076] The SN_2009 cohort was used as the external validation cohort ( Figure 6 ). Disease prediction results for T2DM, CVD, and CKD were obtained in SN_2009 ( Figure 4 ). The results of the validation cohort were highly consistent with those in 4C. The occurrence probability of T2DM was the highest in group 4, and the HR (95% CI) on the first dimension reached 0.34 (0.22 - 0.54). The distributions of CVD and CKD were also similar to those in the 4C cohort, and the lowest risks were concentrated in group 1. In groups 3 and 4, individuals with higher risks tended to be distributed in the lower part and left side of the two-dimensional tree.

[0077] 6.4 Visualization validation of pre-diabetes progression risk

[0078] To better assist pre-diabetes patients in disease risk prediction and help clinicians in the diagnosis of pre-diabetes subtypes, an online application (https: / / www.rjh.com.cn / 2018RJPortal / 4c / preDM-ddrtree / ) was developed. By inputting an individual's gender, age, and the risks of 12 clinical indicators on the website, the coordinate position of the individual on the two-dimensional tree was obtained, and the risks of developing T2DM, CKD, and CVD within 5 years were predicted.

[0079] 6.5 Evaluation of the effectiveness of the disease prediction model

[0080] The area under the receiver operating characteristic curve (ROC AUC) was used to evaluate the performance of the disease risk prediction models constructed in 4C and SN_2009. The prediction models showed good performance. In the 4C cohort, the ROC AUC values of the prediction models for T2DM, CKD, and CVD were 0.632, 0.706, and 0.716, respectively; for the prediction models of CVD subtypes, the ROC AUC values ranged from 0.698 to 0.705. In the validation cohort SN_2009, the ROC AUC values of the prediction models for T2DM, CKD, and CVD were between 0.700 and 0.814, and the models showed good robustness.

[0081] 7. Discussion and Conclusions

[0082] The DDRTree algorithm was used to clarify the heterogeneity of prediabetes in a national, population-based prospective Chinese cohort and was validated in another independent cohort. The innovative application of DDRTree identified four distinct metabolic branches based on 12 clinical phenotype variables, which could effectively represent the differences in the metabolic profiles and emphasized the different risks of T2DM, CKD, and CVD events; the fourth group characterized by hyperglycemia, insulin resistance, obesity, elevated triglycerides, and elevated liver enzymes had the highest risk of T2DM, while the third group characterized by obesity, insulin resistance, hyperglycemia, and dyslipidemia had the highest risk of CKD. The third and fourth groups had higher CVD risks and greater differences in the distribution of CVD subtypes. These findings were well validated in the external independent cohort SN_2009 - 2021.

[0083] Table 2: Generalized additive models predicting dimensions 1 and 2 of DDRTree, fitted with cubic regression splines

[0084]

[0085] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for constructing a risk prediction model for future disease progression in prediabetes, characterized in that: The following steps are involved: S1: Clinical indicator data of patients with prediabetes in the training cohort and validation cohort were collected, including the patients' gender, age, and 12 clinical phenotypic variables, namely WHR, BMI, HOMA-IR, HOMA-B, ALT, AST, GGT, FPG, PBG, HbA1c, TG, and HDL-C; S2: Using the DDRTree algorithm, the clinical indicator data of prediabetes patients in the training cohort described in step S1 are reduced in dimension and visualized to construct a two-dimensional tree structure of prediabetes heterogeneity; S3: Based on the two-dimensional tree structure described in step S2, the prediabetes population is divided into different subtypes, the distribution coordinates of the prediabetes individuals in the two-dimensional tree are obtained, the risk of prediabetes progressing to T2DM, CKD and CVD diseases is assessed, and a risk prediction model for future disease progression of prediabetes is constructed; S4: Validate the accuracy of the risk prediction model for future disease progression in prediabetes described in step S3 in the validation cohort; Prediabetes was defined as: in participants without diabetes, FPG was 5.6 mmol / L to 6.9 mmol / L, or PBG was 7.8 mmol / L to 11.0 mmol / L, or HbA1c was 5.7% to 6.4%.

2. The method for constructing a risk prediction model for future disease progression in prediabetes according to claim 1, characterized in that: In step S1, the training cohort is from the Chinese Cardiometabolic Disease and Cancer Cohort (4C), including 55,777 participants with prediabetes, and clinical variable data of the prediabetic population are collected; the validation cohort is a community-dwelling prospective cohort established in Songnan District, Shanghai, China from June to August 2009, including 4012 individuals aged ≥40 years, and followed up until November 2021.

3. The method for constructing a risk prediction model for future disease progression in prediabetes according to claim 1, characterized in that: In step S1, clinical indicator data of patients with prediabetes were collected for data screening and variable extraction to obtain 12 clinical phenotypic variables related to the risk of T2DM, CKD, and CVD diseases, including WHR, BMI, HOMA-IR, HOMA-B, ALT, AST, GGT, FPG, PBG, HbA1c, TG, and HDL-C.

4. The method for constructing a risk prediction model for future disease progression in prediabetes according to claim 1, characterized in that: In step S2, a linear regression model is established in the training cohort, with gender and age as independent variables and 12 clinical phenotypic variables as dependent variables, to obtain the residual matrix of the 12 clinical phenotypic variables, and the residual matrix is ​​reduced and visualized by the DDRTree algorithm, and the prediabetes population is displayed as a two-dimensional tree structure, in which the phenotypic changes and individual distribution are displayed, thereby reflecting the heterogeneity of prediabetes.

5. The method for constructing a risk prediction model for future disease progression in prediabetes according to claim 1, characterized in that: In step S3, the prediabetes subtype characterized by hyperglycemia, insulin resistance, obesity, elevated triglycerides, and liver enzymes is at high risk of progressing to T2DM.

6. The method for constructing a risk prediction model for future disease progression in prediabetes according to claim 1, characterized in that: In step S3, the prediabetes subtype characterized by obesity, insulin resistance, hyperglycemia, and dyslipidemia has a high risk of progression to CKD.

7. The method for constructing a risk prediction model for future disease progression in prediabetes according to claim 1, characterized in that: In step S3, the subtype characterized by hyperglycemia, insulin resistance, obesity, triglycerides, and elevated liver enzymes and the subtype characterized by obesity, insulin resistance, hyperglycemia, and dyslipidemia progressed to high CVD risk, and the distribution of CVD subtypes was highly different.

8. The method for constructing a risk prediction model for future disease progression in prediabetes according to claim 1, characterized in that: In step S3, the prediabetes progresses to CVD diseases including stroke, myocardial infarction and heart failure subtypes.

9. A risk prediction system for the progression of prediabetes to T2DM, CKD and CVD diseases, characterized in that: include: The data collection module is used to collect heterogeneous phenotypic data of patients with prediabetes, including age, gender, and 12 clinical phenotypic variables, namely WHR, BMI, HOMA-IR, HOMA-B, ALT, AST, GGT, FPG, PBG, HbA1c, TG, and HDL-C; Data preprocessing module, used to remove missing values ​​and outliers, where outliers are values ​​outside the range of ±5 standard deviations from the mean; The data analysis module is used to construct a two-dimensional tree structure using the DDRTree algorithm for the 12 clinical phenotype variables in the training cohort; The prediction model building module is used to divide the prediabetes population into different subtypes, use the two-dimensional coordinates of the prediabetes individuals in the tree structure as independent variables, establish a Cox regression model, evaluate the risk of prediabetes progressing to T2DM, CKD and CVD diseases, and build a risk prediction model for future disease progression in prediabetes; A verification module is used to verify the accuracy of the risk prediction model for future disease progression of prediabetes based on a verification cohort.

10. The risk prediction system for prediabetes to progress to T2DM, CKD and CVD diseases according to claim 9, characterized in that: The data acquisition module collects clinical indicator data of prediabetes patients in the training cohort and the validation cohort; the data preprocessing module performs data screening and feature extraction to obtain 12 clinical phenotypic variables corresponding to the risk of prediabetes progressing to T2DM, CKD and CVD diseases, including WHR, BMI, HOMA-IR, HOMA-B, ALT, AST, GGT, FPG, PBG, HbA1c, TG and HDL-C.

Citation Information

Cited By

  • Typing model training method and device applied to knee osteoarthritis and typing method

    CN122158152A