Risk prediction model for future disease progression of obese people and construction method thereof

By using the DDRTree method and reverse graph embedding technology, a risk prediction model for future disease progress in obese people was constructed, the problem of obesity heterogeneity treatment was solved, and the implementation of personalized treatment plans and the accuracy of disease risk prediction was achieved.

CN120164616APending Publication Date: 2025-06-17SHANGHAI INST FOR ENDOCRINE & METABOLIC DISEASES +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510223024.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the heterogeneity of obesity and cannot accurately identify high-risk subgroups, making it difficult to implement personalized treatment plans.

Method used

A discriminant dimensionality reduction tree (DDRTree) method was used to combine the reverse graph embedding technology, and a nonlinear two-dimensional tree structure was constructed using ten clinical indicators from the 4C study to construct a risk prediction model for future disease progress in obese people.

Benefits of technology

A more refined classification of obese people is achieved, disease risks of different subtypes are identified, personalized treatment strategies are provided, and the accuracy of disease progression risk prediction is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164616A_ABST
    Figure CN120164616A_ABST
Patent Text Reader

Abstract

The invention discloses a risk prediction model for future disease progression of obese people and a construction method thereof, and the method comprises the steps: collecting clinical index data, including gender, age, HbA1c, HOMA-IR, WHR, SBP, DBP, TC, HDL-C, TG, ALT and CREAT, of the obese people in a training queue and a verification queue, and building a two-dimensional tree structure of the heterogeneity of the obese people in the training queue by using a DDRTree algorithm; the method comprises the following steps: dividing obese people into different subtypes, obtaining distribution coordinates of individuals in a two-dimensional tree, evaluating the risk that the obese people develop into T2DM, CKD and CVD diseases, constructing a risk prediction model for future disease progression of the obese people, performing model accuracy verification in a verification queue, and accurately evaluating the risk of future disease progression of the obese people. And early intervention on obesity-related diseases is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of disease prediction models, and particularly relates to a risk prediction model for future disease progression in obese populations and a method for constructing the same. Background Art

[0002] With the rapid development of the global economy and significant changes in dietary patterns and lifestyles, obesity has become a global health problem. As a major common risk factor for common chronic diseases such as type 2 diabetes (T2DM), cardiovascular disease (CVD), and chronic kidney disease (CKD), obesity has brought a huge socioeconomic and health burden.

[0003] Obesity is a heterogeneous and multifactorial disease, with individuals showing different metabolic characteristics and susceptibilities to clinical endpoints. This heterogeneity emphasizes the need for precision medicine approaches to accurately classify obesity subtypes, develop models for the diverse characteristics of obesity, and formulate targeted intervention measures. Although the existing concept of "metabolically healthy obesity (MHO)" stratifies individuals into high-risk and low-risk groups based on body mass index (BMI), there is currently no universally accepted definition, and the cut-off values of classification indicators vary widely in different studies. The crude dichotomous risk factor classification cannot effectively handle the observable continuous risk gradient in classification factors and the differences during disease progression. Therefore, there is an urgent need to develop a more refined classification system to reflect the heterogeneity of obesity, identify high-risk subgroups, and achieve individualized treatment regimens.

[0004] Discriminative dimensionality reduction tree (DDRTree), as a dimensionality reduction method mainly used in computational biology, especially in single-cell RNA sequencing data analysis, has been applied to the precise subtype classification of diabetes. It can effectively capture the complex branching structure of high-dimensional data and identify different diabetes subtypes, paving the way for more personalized and effective treatment strategies. Given the core differences in pathophysiology of obesity, applying this method may provide new insights into the heterogeneity of obesity. Summary of the Invention

[0005] The main objective of the present invention is to provide a method for constructing a risk prediction model for future disease progression in obese populations. By using the DDRTree method in combination with reverse graph embedding technology, a non-linear two-dimensional tree structure is constructed using ten clinical indicators of obese individuals from the 4C study, resulting in a risk prediction model for future disease progression in obese populations, which is further validated in an independent cohort in China.

[0006] Another objective of the present invention is to provide a risk prediction model for future disease progression in obese populations.

[0007] The above objectives of the present invention are achieved through the following technical solutions:

[0008] In the first aspect of the present invention, a method for constructing a risk prediction model for future disease progression in obese individuals is provided, comprising the following steps:

[0009] S1: Collect clinical index data of obese individuals in the training cohort and the validation cohort, including gender, age, and 10 clinical phenotype variables, namely glycated hemoglobin (HbA1c), homeostatic model assessment-insulin resistance (HOMA-IR), waist-to-hip ratio (WHR), systolic blood pressure (SBP), diastolic blood pressure (DBP), total cholesterol (TC), high-density lipoprotein cholesterol (HDL-C), triglyceride (TG), alanine aminotransferase (ALT), and creatinine (CREAT);

[0010] S2: Use the DDRTree algorithm to perform dimensionality reduction and visualization on the clinical index data of obese individuals in the training cohort in step S1, and construct a two-dimensional tree structure of the heterogeneity of obese individuals;

[0011] S3: Based on the two-dimensional tree structure in step S2, divide obese individuals into different subtypes, obtain the distribution coordinates of individual obese individuals in the two-dimensional tree, evaluate the risks of obese individuals progressing to diseases such as T2DM, CKD, and CVD, and construct a risk prediction model for future disease progression in obese individuals;

[0012] S4: Verify the model accuracy of the risk prediction model for future disease progression of obese individuals in step S3 in the validation cohort; wherein the obese individuals are defined as: BMI≥28.0 kg / m 2 。

[0013] Preferably, in step S1, the training cohort is from the China Cardiometabolic Disease and Cancer Cohort (4C), including 193,846 obese individual participants with BMI≥28.0 kg / m 2 and collect the clinical variable data of obese individuals; the validation cohort is a community-based prospective cohort established in Karamay, Xinjiang Uygur Autonomous Region, China from March to June 2011, including 9,834 individuals aged ≥40 years old, followed up until November 2021.

[0014] Preferably, in step S1, collect the clinical index data of obese individuals for data screening and variable extraction, and obtain 10 clinical phenotype variables related to the risks of T2DM, CKD, and CVD diseases, including HbA1c, HOMA-IR, WHR, SBP, DBP, TC, HDL-C, TG, ALT, and CREAT.

[0015] Preferably, in step S2, a linear regression model is established in the training cohort, with gender and age as independent variables and 10 clinical phenotype variables as dependent variables respectively, to obtain the residual matrix of the 10 clinical phenotype variables. The DDRTree algorithm is used to reduce the dimension and visualize the residual matrix, and the obese population is displayed as a two-dimensional tree structure, and the phenotypic changes and individual distributions are shown in the tree structure, so as to reflect the heterogeneity of the obese population.

[0016] Preferably, in step S3, the obese population subtype characterized by hyperglycemia and insulin resistance has a high risk of progressing to T2DM.

[0017] Preferably, in step S3, the obese population subtype characterized by hyperglycemia, insulin resistance, hypertension and dyslipidemia has a high risk of progressing to CKD.

[0018] Preferably, in step S3, the subtype characterized by hypertension and dyslipidemia has a high risk of progressing to CVD.

[0019] In the second aspect of the present invention, a risk prediction system for the progression of the obese population to T2DM, CKD and CVD diseases is provided, including:

[0020] A data acquisition module for collecting heterogeneous data of the obese population, including 10 clinical phenotype variables, namely HbA1c, HOMA-IR, WHR, SBP, DBP, TC, HDL-C, TG, ALT and CREAT;

[0021] A data preprocessing module for removing missing values and outliers, where the outliers are values outside the range of the mean ± 5 standard deviations (SD);

[0022] A data analysis module for constructing a two-dimensional tree structure for 10 clinical phenotype variables in the training cohort using the DDRTree algorithm;

[0023] A prediction model construction module for classifying the obese population into different subtypes, using the two-dimensional coordinates of the obese population individuals in the tree structure as independent variables, establishing a Cox regression model, evaluating the risks of the obese population progressing to T2DM, CKD and CVD diseases, and constructing a risk prediction model for the future disease progression of the obese population;

[0024] A verification module for verifying the model accuracy of the risk prediction model for the future disease progression of the obese population by using a validation cohort.

[0025] Preferably, the data acquisition module collects clinical index data of obese people in the training cohort and the validation cohort; the data preprocessing module performs data screening and feature extraction to obtain 10 clinical phenotype variables corresponding to the risks of obese people progressing to diseases such as T2DM, CKD, and CVD, including HbA1c, HOMA-IR, WHR, SBP, DBP, TC, HDL-C, TG, ALT, and CREAT.

[0026] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0027] 1. The present invention first applies the DDRTree algorithm to obese people in a national prospective cohort study. A sufficient sample size can construct a robust two-dimensional tree structure, which performs excellently in exploring disease progression and understanding disease heterogeneity, making it a valuable tool for capturing the complex dynamics of disease evolution. This analysis reduces multi-dimensional continuous clinical variables for obese people to describe disease heterogeneity, which helps to more personalized guide the care of obese people. The prevention strategy based on DDRTree may bring benefits to high-risk populations of future T2DM, CKD, and CVD. In a large-scale national prospective cohort of obese people, the DDRTree algorithm can capture the complex interactions of disease risk factors, track the complex dynamic changes of diseases, so as to realize the risk prediction of the progression of obese people and provide guidance for the personalized care of obese people.

[0028] 2. In the risk prediction model for the future disease progression of obese people in the present invention, the subtype of obese people characterized by hyperglycemia and insulin resistance has a high risk of progressing to T2DM; the subtype of obese people characterized by hyperglycemia, insulin resistance, hypertension, and dyslipidemia has a high risk of progressing to CKD; the subtype characterized by hypertension and dyslipidemia has a high risk of progressing to CVD. These findings have been verified in the external validation cohort XJ_2011 and have strong applicability. Description of the Drawings

[0029] Figure 1 It is a conceptual framework diagram for constructing a risk prediction model for the future disease progression of obese people in the embodiment; A) Heterogeneity of obesity and phenotypes and their relationship with the risks of T2DM, CVD, and CKD, combining (1) phenotype data of 18,733 obese participants to establish obese subgroups, (2) an external dataset of 1,653 obese participants from XJ_2011 followed up for 10.5 years; B) Five groups with different sub-phenotypes are determined according to ten obesity-related phenotype variables including HbA1c, HOMA-IR, WHR, SBP, DBP, TC, HDL-C, TG, ALT, and CREAT, and the sub-phenotypes are further distinguished according to the probability of obesity-related progression results.

[0030] Figure 2Visual representation of the phenotypic characteristics of 18,733 obese participants in the training cohort of the example; A) Distribution of ten obesity-related phenotypic variables (HbA1c, HOMA-IR, WHR, SBP, DBP, TC, HDL-C, TG, ALT, CREAT) on the obesity two-dimensional tree, where each point in the figure represents an individual; B) Correlation coefficients (95% confidence intervals) between the contents of the ten phenotypic variables and the two-dimensional coordinates, red: correlation coefficient between the content and the first-dimensional coordinate; green: correlation coefficient between the content and the second-dimensional coordinate; C) Spatial autocorrelation (Moran's index) of the ten phenotypes on the two-dimensional tree, and a higher Moran's index indicates a higher spatial autocorrelation of the phenotype; The P-values of all values are less than 0.0001.

[0031] Figure 3 Visualization of the heterogeneity of future disease progression in the training cohort of the example; A) Probability of developing T2DM within 5 years; B) Probability of developing CVD within 5 years; C) Probability of developing CKD within 5 years; For all outcomes (A - C), a COX risk model constructed with the coordinates of the individuals on the two-dimensional tree as independent variables was used as the disease prediction model; D) HR (95% confidence interval) of the probability of occurrence of each outcome on the two-dimensional coordinates, red: HR value of the probability of occurrence on the first-dimensional coordinate; green: HR value of the probability of occurrence on the second-dimensional coordinate; E) Spatial autocorrelation of the probabilities of occurrence of the three outcomes.

[0032] Figure 4 Visualization of the heterogeneity of disease progression in the validation cohort XJ_2011 study in the example. The obese individuals (n = 1,653) in the XJ_2011 study were located on the 4C (reference) tree using a mapping function, and each point represents an individual in the XJ_2011 study; A) Probability of onset of T2DM; B) Probability of developing CVD within 5 years; C) Probability of onset of CKD. The probability of the CVD outcome (B) was generated by a Cox proportional hazards model constructed based on the DDRTree dimensions. For the T2DM and CKD outcomes (A and C), the probabilities were generated by a Logistic regression model constructed based on the DDRTree dimensions; D) HR (95% confidence interval) and OR (95% confidence interval) of the probability of occurrence on the two-dimensional coordinates; E) Spatial autocorrelation of the three outcomes; The Moran's I statistic is shown on the x-axis, and higher values indicate stronger spatial autocorrelation of the phenotype; The P of all values is less than 0.0001, and all predictions are from models containing DDRTree dimensions.

[0033] Figure 5 Flowchart of participant selection in the 4C study of the example (related to Table 1).

[0034] Figure 6Flowchart for selecting test participants in the XJ study of the embodiment (related to Table 1).

[0035] Figure 7 Local Moran's index distribution pattern of phenotypes and outcome events of 4C obese participants in the embodiment.

[0036] Figure 8 Visualization of phenotypic characteristics of obese participants (n = 1653) at baseline in the XJ study of the embodiment. Detailed implementation mode

[0037] In order to more fully understand and demonstrate the technical solutions, objectives, and advantages of the present invention, the technical effects produced by the present invention will be further described in detail and completely below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. It should be noted that for those of ordinary skill in the art, other embodiments obtained without departing from the concept of the present invention all belong to the protection scope of the present invention.

[0038] Embodiment 1

[0039] 1. Method

[0040] 1.1 Research participants

[0041] This study used data from two cohorts:

[0042] (1) China Cardiometabolic Disease and Cancer Cohort (4C) Study;

[0043] (2) Xinjiang Uygur Autonomous Region (XJ) Study, the conceptual framework is as Figure 1 shown.

[0044] 1.2 4C Study

[0045] Using data from the 4C study, which recruited 193,846 adults over 40 years old from 20 communities in different geographical regions of China between 2011 and 2012. A comprehensive set of questionnaires, clinical measurements, and laboratory tests were conducted at the baseline visit. Obese participants with a BMI ≥ 28.0 kg / m 2 (according to the definition of the Obesity Working Group) were included. Exclusion criteria included: obese participants with cardiovascular disease, cancer, chronic kidney, liver, or pancreatic diseases at baseline, as well as those with missing data on obesity-related phenotypic variables or outliers. Finally, 18,733 obese participants were included in the analysis, with a median follow-up period of 3.3 years. The selection process of the study sample is as Figure 5 shown.

[0046] All study participants had their blood pressure, weight, height, waist circumference, and hip circumference measured according to the standard protocol during outpatient visits. BMI was calculated as weight (in kilograms) divided by the square of height (in meters). Systolic and diastolic blood pressure were measured three times in the non-dominant arm, and the average value was taken for analysis after the subject had been sitting quietly for at least 5 minutes. An automatic electronic device (Omron HEM-752FUZZY model, Dalian, China) was used to average the three blood pressure measurements obtained after sitting quietly for 5 minutes. All participants underwent an oral glucose tolerance test, and blood glucose was obtained at 0 hours and 2 hours during the test. Under strict quality control procedures, plasma glucose concentration was analyzed using the glucose oxidase or hexokinase method within two hours after blood sample collection. A glycated hemoglobin collection system (Bio-Rad Laboratories, California, USA) was used to collect capillary fingertip blood samples, which were transported to the certified central laboratory of Ruijin Hospital at 2-8°C. This clinical laboratory has been double-certified by the National Glycohemoglobin Standardization Program of the United States and the College of American Pathologists (CAP) laboratory. Serum insulin, CREAT, ALT, TC, HDL-C, and TG were detected using an automatic analyzer (ARCHITECT ci16200 analyzer, Abbott Laboratories, Illinois, USA) in the central laboratory. HOMA-IR was used to estimate insulin resistance, with the formula: fasting insulin (μIU / mL) × fasting blood glucose (mg / dL) / 405.

[0047] 1.3 XJ Cohort

[0048] The XJ_2011 study was a community-based prospective cohort that included 9,834 individuals aged ≥40 years. It was established in the Karamay area of Xinjiang Uygur Autonomous Region, China, from March to June 2011 and followed up until November 2021. HbA1c was measured using high-performance liquid chromatography with a VARIANTII hemoglobin detection system (Bio-Rad Laboratories). Serum insulin, CREAT, ALT, TC, HDL-C, and TG were measured using a fully automated biochemical analyzer (ARCHITECT ci16200 analyzer, Abbott Laboratories, Illinois, USA). 1,653 adults with a BMI ≥28.0 kg / m 2 and complete baseline information were identified from this cohort ( Figure 6 ), with a median follow-up time of 10.5 years (interquartile range: 10.3 - 10.6).

[0049] Each study was approved by the Medical Ethics Committee of Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, and all study participants provided written informed consent.

[0050] 1.4 Endpoint Definition

[0051] T2DM: In the 4C study, diabetes events were defined as FPG ≥ 7.0 mmol / L, 2-hour post-load plasma glucose ≥ 11.1 mmol / L, or HbA1c ≥ 6.5%, or physician-diagnosed diabetes. In the XJ_2011 study, incident diabetes was defined as physician-diagnosed diabetes based on the National Health Insurance System.

[0052] CKD: eGFR was calculated using the Chronic Kidney Disease Epidemiology Collaboration (CKD-EPI) equation. CKD included new-onset renal failure requiring dialysis or alternative treatment, death due to renal causes, decline in eGFR (eGFR < 60 mL / min / 1.73m 2 ) or a certain decline in eGFR category (from eGFR ≥ 90 mL / min / 1.73m 2 at baseline to 60 - 89 mL / min / 1.73m 2 ) at follow-up, accompanied by a 25% or greater decrease in eGFR compared to baseline.

[0053] CVD: Information on vital status and clinical outcomes was collected from the local death and disease registries of the National Disease Surveillance Points System and the National Health Insurance System. Major CVD events included first non-fatal myocardial infarction, non-fatal stroke, hospitalization or treatment for heart failure during follow-up, and cardiovascular death.

[0054] II. Statistical Analysis

[0055] Obesity subgroups with different subphenotypes were constructed in the 4C study and validated in XJ_2011 as follows:

[0056] 2.1 Multidimensional Phenotype Data Analysis of DDRTree in the 4C Study

[0057] The DDRTree algorithm in the Monocle R package was used with ten clinically available phenotypes, including WHR, HOMA-IR, HbA1c, HDL-C, TG, TC, ALT, CREAT, SBP, and DBP. Outliers beyond 5 standard deviations from the mean were excluded, and the data were rank-normalized. In addition, each phenotype was adjusted for age and sex by linear regression analysis. Subsequently, the residual matrix of these ten phenotypes was processed by the DDRTree algorithm to reduce the multidimensional obesity-related phenotype data to a low-dimensional space, which was visualized as a tree structure. The ends of the branches represented different phenotypes, while the center of the tree represented the mixed phenotypes.

[0058] 2.2 Evaluation of Long-Term Risk Differences between Subgroups

[0059] To obtain the individual probabilities of various obesity-related outcomes for each study participant, the two-dimensional coordinates derived from DDRTree were used as independent variables to construct a Cox proportional hazards model to predict the event probabilities for each individual. In all models, individuals without follow-up data or those who had experienced the relevant events at baseline were excluded. These event probabilities were then superimposed on the dendrogram to visualize the heterogeneity of obesity progression.

[0060] 2.3 Statistical evaluation of the phenotypic or outcome distribution in the tree structure

[0061] Although most phenotypes and outcomes were clearly different in the tree structure, this difference could be quantified by two methods. First, a linear regression model was fitted between the two dimensions of the tree and each phenotype, and the regression coefficients and their 95% confidence intervals were plotted. For obesity-related outcomes, the hazard ratio HR was derived from the Cox proportional hazards model, where HR indicates how the probability of obesity outcomes (probability of T2DM / CVD / CKD) changes in the two-dimensional space (x-axis and y-axis). Second, global and local Moran's I statistics were used for spatial autocorrelation analysis. In this analysis, the position of each individual (using its spatial coordinates) was considered, and its relationship with the relevant phenotypic and event probability values was evaluated. A positive value of Moran's I indicates that similar values are clustered together, while a negative value indicates that the variable is not clustered. For the estimation of local Moran's I, 1000 neighbors were considered, and the individual-level Moran's I was calculated and visualized as a continuous variable.

[0062] To map obese individuals onto the 4C reference tree, a "mapping function" was constructed using 10 obesity-related phenotypes (HbA1c, HOMA-IR, WHR, TC, HDL-C, TG, ALT, CREAT, SBP, and DBP), while considering age and gender, to assign each individual a position in the 4C tree.

[0063] 2.4 Online application of obesity-related outcomes and development of prediction models

[0064] Using the "mapping function" described above, an application was further developed that can determine the position of obese individuals on the tree based on clinically available variables (age, gender, WHR, HOMA-IR, HbA1c, HDL-C, TG, TC, ALT, CREAT, SBP, and DBP) and provide the predicted probabilities of each obesity-related outcome over the next 5 years.

[0065] III. Results

[0066] The key baseline characteristics of the participants in the two datasets are shown in Table 1. The average BMI in the 4C study was 30.37 kg / m 2, similar to the average BMI of 30.50 kg / m in the XJ_2011 study 2 Similar

[0067] Table 1: Phenotypic characteristics of the study population

[0068]

[0069]

[0070] Data are presented as mean ± SD, number (percentage), or median (interquartile range). Alanine aminotransferase, ALT; Body mass index, BMI; Diastolic blood pressure, DBP; Glycated hemoglobin, HbA1c; Homeostasis model assessment of insulin resistance, HOMA-IR; High-density lipoprotein cholesterol, HDL-c; Systolic blood pressure, SBP; Waist-to-hip ratio, WHR.

[0071] 3.1 Baseline obesity phenotypic heterogeneity

[0072] Based on ten clinically available variables, including WHR, SBP, DBP, HbA1c, HOMA-IR, CREAT, ALT, TC, HDL-C, and TG, the DDRTree algorithm was used to divide the 18,733 participants in the 4C study into different obesity subgroups, simplifying the complex phenotypic data into a two-dimensional tree structure. As Figure 2 shown in A, the data of ten obesity-related characteristics were superimposed on the tree to visualize and characterize phenotypic changes. Magenta indicates high values, yellow indicates lower values, and extreme values appear at the distal ends of the branches, while intermediate values appear at the proximal ends.

[0073] According to the distal part of the two-dimensional visualization, five groups of phenotypes were identified ( Figure 1 B). Individuals in the upper right tree (P1) are metabolically healthy, with higher HDL-C and lower blood glucose, blood lipid, and blood pressure levels. In contrast, the tree in the lower left part (P5) includes high blood glucose, central obesity, high TG and ALT levels, and low HDL-C. The tree in the lower right part (P2) includes participants with lower HDL-C, while individuals in the upper left tree (P4) have higher blood pressure, cholesterol, and creatinine levels, as well as moderate levels of HbA1c, TG, and ALT. Meanwhile, individuals in the upper part of the tree (P3) have higher HDL-C and slightly higher blood pressure and cholesterol levels, being in a transitional intermediate state from metabolically healthy (P1, P2) to unhealthy (P4, P5).

[0074] In addition to the visual gradient, two metrics were used to evaluate the distribution pattern of phenotypes in the tree. Figure 2B shows the regression relationship between each phenotype and the two dimensions of the tree. The 2D visual representation can be interpreted along the horizontal axis (dimension 1) and the vertical axis (dimension 2). Dimension 1 is negatively correlated with HbA1c, WHR, HOMA-IR, TC, TG, ALT, CREAT, SBP, and DBP, while positively correlated with HDL-C, indicating that the health status of obese subjects improves to the right along the horizontal axis. In addition, Figure 2 C plots the global Moran’s I value for each phenotype, representing the strength of the spatial correlation of the phenotype on the tree. HDL-C, TC, and HOMA-IR contribute the most to phenotypic heterogeneity, followed by DBP, SBP, TG, then HbA1c and ALT, and CREAT or WHR shows the least variation on the tree.

[0075] 3.2 Heterogeneity of obesity-related outcomes

[0076] Next, investigate how phenotypic heterogeneity at baseline translates into heterogeneity of obesity-related progression outcomes. Using the Cox proportional hazards model, with the coordinates of each individual in the two-dimensional space as covariates, evaluate the event probability that each obese individual progresses to a specific outcome within the next 5 years ( Figure 3 A, B, C), and obtain the HR and its 95% CI to show how the risk of obesity outcomes varies in the two-dimensional space ( Figure 3 D). In the 4C study, the probability distribution of T2DM (HR (95% CI)) is 0.54 (0.49 - 0.59) in dimension 1 and 0.84 (0.74 - 0.94) in dimension 2 ( Figure 3 D), resulting in the peak of the T2DM risk in the lower left part of the tree (P5; Figure 3 A), corresponding to the region of hyperglycemia and insulin resistance at baseline. While the HR (95% CI) of the probability distribution of CVD is 0.74 (0.66 - 0.84) in dimension 1 and 1.26 (1.08 - 1.47) in dimension 2 ( Figure 3 B), the upper left part of the tree contains individuals with hypertension and dyslipidemia at baseline, where the CVD incidence risk is the highest (P4; Figure 3 D). The risk of CKD decreases significantly along the x-axis (HR dim1 = 0.79; 95% CI, 0.71 - 0.88), while it varies little along the y-axis (HR dim2 = 1.01; 95% CI, 0.87 - 1.16; Figure 3 D), with the highest risk in the left part of the tree (P4 and P5) and the lowest on the right. The strength of the global spatial correlation, as shown by Moran’s I ( Figure 3 E), indicates that each outcome has a strong and significant spatial correlation on the tree (P < 0.001). Local Moran’s I ( Figure 7)It reveals that the individuals within the tree have a strong correlation with their neighbors at each outcome.

[0077] 3.3 External validation of obesity heterogeneity

[0078] For external validation of the method and results, the tree derived from the 4C study was used as the reference tree, and an external individual was projected onto the reference tree using a mapping function that consisted of two steps. In the first step, two generalized additive models were used to predict dimension 1 (x - coordinate) and dimension 2 (y - coordinate) for each individual. These two models had excellent performance, with adjusted R 2 values of 0.877 and 0.837 respectively (Table 2). In the second step, a distance - estimation algorithm was used to assign the individuals to the reference tree.

[0079] For external validation, the data of the XJ_2011 cohort ( Figure 8 ) was used. The T2DM, CVD, and CKD outcomes in the XJ_2011 cohort were obtained ( Figure 4 ) As expected, the probability superposition of T2DM on the tree revealed a pattern consistent with that observed in the 4C study. The probability distributions of CVD and CKD also followed a similar pattern: the higher risks were concentrated in the lower part of the tree. The risks in the upper - left branch were slightly lower, which could be attributed to the smaller number of participants with hypertension at baseline in the XJ_2011 cohort (Table 1).

[0080] 4 Visualizing individual risks with their phenotypes

[0081] To help clinicians intuitively understand individual profiles and the risks of related disease complications, an application (http: / / 118.25.189.214:30044 / obesity - ddrtree / ) was developed. This tool can place obesity within the obesity continuum (tree - like structure) and predict their risks of developing T2DM, CVD, and CKD within 5 years. These predictions were made using age, gender, and ten other continuous phenotypic measurements used to define the tree. Based on these variables of continuous measurements, the estimated area under the receiver operating characteristic curve (AUC) of the three outcome models (T2DM, CVD, and CKD) in the 4C study ranged from 0.66 to 0.72. External validation showed that the AUC values were between 0.68 and 0.81, and the calibration of each model was good.

[0082] IV. Supplementary methods

[0083] 4.1 Development of the mapping function for external validation

[0084] To map obese individuals onto the 4C reference tree, a "mapping function" was constructed. This function uses 10 obesity-related phenotypes (HbA1c, HOMA-IR, WHR, TC, HDL-C, TG, ALT, CREAT, SBP, and DBP), while taking age and gender into account, to assign each individual a position in the 4C tree. The mapping function consists of two components:

[0085] (1) Two generalized additive models that use cubic regression splines to fit smooth terms and predict two DDRTree dimensions (Table 2) with 10 phenotypes, age, and gender as covariates.

[0086] (2) A distance estimation algorithm that calculates the Euclidean distance between points in a two-dimensional space. When all 10 phenotypes, as well as age and gender, are provided, the trained spline model predicts the two dimensions of an individual, thus determining their provisional position on the tree. In the second step, the distance estimation algorithm calculates the distances between the individual's provisional position and all positions on the 4C tree. The mapping function then reassigns the individual to the position on the reference tree that is closest to the provisional point, a process that identifies the individual among 18,733 participants who is closest to the newly mapped individual.

[0087] Finally, the predicted probabilities of phenotypes and obesity-related outcomes in the XJ_2011 cohort were overlaid onto this tree to further evaluate their distribution on the tree and compare it with the distribution in the 4C cohort. The XJ_2011 data used a method similar to that of the 4C study to evaluate CVD risk and probability by constructing a Cox proportional hazards model and used a Logistic regression model to evaluate the risks and probabilities of T2DM and CKD.

[0088] The performance of the Cox proportional hazards model constructed based on DDRTree dimensions for T2DM, CVD, and CKD in 4C_2011 and XJ_2011 was evaluated by ROCAUC. ROCAUC is an indicator that measures the discrimination ability of a model, and a higher value indicates better discrimination ability of the model.

[0089] Table 2: Generalized additive models for predicting DDRTree dimension 1 and dimension 2, fitted with cubic regression splines

[0090]

[0091] In summary, the present invention adopts a novel method to perform dimensionality reduction on the health record data of 18,733 obese individuals by using a series of clinically relevant phenotypic data available at baseline in a national, population-based prospective cohort. The resulting tree structure provides a two-dimensional representation of the increasingly recognized complex phenotypic variations in obesity, enabling the overlay of clinically relevant phenotypes and their probabilities of predicting the occurrence of complications, and validating obesity-related outcomes using another independent dataset. The present invention demonstrates how the application of the DDRTree method can intuitively show the transformation of phenotypic heterogeneity into differences in the deterioration of blood glucose and the incidence of microvascular and macrovascular diseases during the obesity disease process.

[0092] Obesity is a multifaceted disease, and BMI alone is insufficient to reflect the risk differences among individuals. DDRTree is a novel data dimensionality reduction technique that can identify five different metabolic phenotypes based on ten clinically relevant continuous phenotypic variables, summarizing the heterogeneity of obesity progression. For example, the sub-phenotype characterized by hyperglycemia and insulin resistance (P5) significantly increases the risk of developing T2DM and CKD. Another sub-phenotype, characterized by hypertension and dyslipidemia (P4), significantly increases the risk of developing T2DM and CKD. In contrast, a "metabolically healthy" group (P1) was identified, which is characterized by healthy metabolic characteristics, including good blood glucose control, less visceral fat, low levels of insulin resistance, TG, and blood pressure, and this group has a lower risk of obesity-related diseases, which also partially conforms to the new definition of MHO proposed by Schulze MB et al.

[0093] Different from traditional obesity classification methods (such as MHO or clustering methods), the DDRtree method provides a more detailed and comprehensive representation of the complexity and continuity of obesity progression, and this method has been applied to study the heterogeneity of diabetes. By combining routinely measured clinical indicators and considering the effects of age and gender, the model shows the potential for wide application in the general population. In addition, a user-friendly web-based tool has been developed to provide a convenient method for clinicians to guide personalized management strategies for obesity, highlighting the importance of precision medicine in the effective management of obesity.

[0094] This study has several major advantages, which are as follows:

[0095] First, the discovery cohort has a large sample of obese individuals nationwide. This large sample size provides a solid foundation for understanding obesity-related health risks. Implementing centralized measurements ensures a high degree of accuracy and consistency of the dataset, further enhancing the reliability of the research results. The research results of the discovery cohort have been strictly verified in an independent cohort, emphasizing the reliability and reproducibility of the results.

[0096] Second, the incidence of CVD and CKD in the validation cohort has been tracked for over 10 years, providing valuable insights into these long-term outcomes.

[0097] Finally, the new application of the DDRTree method successfully depicts the heterogeneity of obesity, reveals different metabolic phenotypes and their associated risks, and incorporating individual continuous phenotypes into clinical practice makes a significant contribution to the precision management of obesity.

[0098] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for constructing a risk prediction model for future disease progression in obese people, characterized in that: The following steps are involved: S1: Clinical indicator data of obese people in the training cohort and validation cohort were collected, including gender, age and 10 clinical phenotypic variables, namely HbA1c, HOMA-IR, WHR, SBP, DBP, TC, HDL-C, TG, ALT and CREAT; S2: Using the DDRTree algorithm, the clinical indicator data of the obese population in the training cohort described in step S1 are reduced in dimension and visualized to construct a two-dimensional tree structure of the heterogeneity of the obese population; S3: Based on the two-dimensional tree structure described in step S2, the obese population is divided into different subtypes, the distribution coordinates of the obese individuals in the two-dimensional tree are obtained, the risk of the obese population progressing to T2DM, CKD and CVD diseases is assessed, and a risk prediction model for future disease progression in the obese population is constructed; S4: Verify the accuracy of the risk prediction model for future disease progression in the obese population described in step S3 in the validation cohort, where the obese population is defined as: BMI ≥ 28.0 kg / m 2 .

2. The method for constructing a risk prediction model for future disease progression in obese people according to claim 1, characterized in that: In step S1, the training cohort was from the Chinese Cardiometabolic Disease and Cancer Cohort (4C), including 193,846 subjects with BMI ≥ 28.0 kg / m 2 Obese participants were recruited to collect clinical variable data of obese people; the validation cohort was a community-dwelling prospective cohort established in Karamay, Xinjiang Uygur Autonomous Region, China from March to June 2011, including 9834 individuals aged ≥40 years, and followed up until November 2021.

3. The method for constructing a risk prediction model for future disease progression in obese people according to claim 1, characterized in that: In step S1, clinical indicator data of obese people were collected for data screening and variable extraction to obtain 10 clinical phenotypic variables related to the risk of T2DM, CKD and CVD diseases, including HbA1c, HOMA-IR, WHR, SBP, DBP, TC, HDL-C, TG, ALT and CREAT.

4. The method for constructing a risk prediction model for future disease progression in obese people according to claim 1, characterized in that: In step S2, a linear regression model is established in the training cohort, with gender and age as independent variables and 10 clinical phenotypic variables as dependent variables, to obtain the residual matrix of the 10 clinical phenotypic variables, and the residual matrix is ​​reduced and visualized by the DDRTree algorithm, and the obese population is displayed as a two-dimensional tree structure, in which the phenotypic changes and individual distribution are displayed, thereby reflecting the heterogeneity of the obese population.

5. The method for constructing a risk prediction model for future disease progression in obese people according to claim 1, characterized in that: In step S3, the subtype of obese people characterized by hyperglycemia and insulin resistance is at high risk of developing T2DM.

6. The method for constructing a risk prediction model for future disease progression in obese people according to claim 1, characterized in that: In step S3, the subtype of obese people characterized by hyperglycemia, insulin resistance, hypertension, and dyslipidemia are at high risk of progressing to CKD.

7. The method for constructing a risk prediction model for future disease progression in obese people according to claim 1, characterized in that: In step S3, the subtype characterized by hypertension and dyslipidemia progresses to high CVD risk.

8. A risk prediction system for obese people to develop T2DM, CKD and CVD diseases, characterized in that: include: The data collection module is used to collect heterogeneous data of obese people, including 10 clinical phenotypic variables, namely HbA1c, HOMA-IR, WHR, SBP, DBP, TC, HDL-C, TG, ALT and CREAT; Data preprocessing module, used to remove missing values ​​and outliers, where outliers are values ​​outside the range of ±5 standard deviations from the mean; The data analysis module is used to construct a two-dimensional tree structure using the DDRTree algorithm for the 10 clinical phenotype variables in the training cohort; The prediction model building module is used to divide obese people into different subtypes, use the two-dimensional coordinates of obese individuals in the tree structure as independent variables, establish a Cox regression model, evaluate the risk of obese people progressing to T2DM, CKD and CVD diseases, and build a risk prediction model for future disease progression in obese people; The verification module is used to verify the accuracy of the risk prediction model for future disease progression in the obese population by using a verification cohort.

9. The risk prediction system for obese people to develop T2DM, CKD and CVD diseases according to claim 8, characterized in that: The data acquisition module collects clinical indicator data of obese people in the training cohort and the validation cohort; the data preprocessing module performs data screening and feature extraction to obtain 10 clinical phenotypic variables corresponding to the risk of obese people developing T2DM, CKD and CVD diseases, including HbA1c, HOMA-IR, WHR, SBP, DBP, TC, HDL-C, TG, ALT and CREAT.

Citation Information

Cited By

  • Typing model training method and device applied to knee osteoarthritis and typing method

    CN122158152A