CVD (chemical vapor deposition) prediction model for IBD (infectious bursal disease) population, application and early screening kit

By combining cardiovascular metabolism and plasma proteomics of IBD patients, a machine learning model was constructed, which solved the problem of insufficient prediction of traditional models in the IBD population, and achieved high-accuracy early prediction and risk identification of cardiovascular diseases in IBD patients.

CN121122684APending Publication Date: 2025-12-12GUANGDONG GENERAL HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510999087.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing cardiovascular disease prediction models are not accurate enough in patients with inflammatory bowel disease (IBD). Traditional models underestimate their cardiovascular disease risk and lack precise assessment tools specifically for IBD populations, making it impossible to effectively identify high-risk individuals and conduct early intervention.

Method used

We developed a cardiovascular disease prediction model based on cardiac metabolism and plasma proteomics in patients with inflammatory bowel disease. The model incorporates 10 cardiovascular metabolic risk factors and 5 CVD-related proteins, and is constructed using machine learning methods. It is then applied to a cardiovascular disease early screening kit, and the prediction accuracy is improved through Cox regression analysis and 5-fold cross-validation.

Benefits of technology

It enables ultra-early warning of cardiovascular disease in IBD patients, with higher prediction accuracy than traditional models. It can provide accurate predictions up to 10 years before the onset of the disease, revealing key proteomic pathways and providing a reference for clinical applications and prevention strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121122684A_ABST
    Figure CN121122684A_ABST
Patent Text Reader

Abstract

The invention provides a CVD (Chemical Vapor Deposition) prediction model aiming at IBD (Infectious Bursal Disease) population, and the CVD prediction model comprises 10 cardiovascular metabolism risk factors (CMFs) which comprise physical indexes and blood biochemical indexes; the physical indexes comprise systolic pressure [SBP], diastolic pressure [DBP], body mass index [BMI] and waistline, and the blood biochemical indexes comprise high density lipoprotein cholesterol [HDL-C], low density lipoprotein cholesterol [LDL-C], total cholesterol, glycosylated hemoglobin [HbA1c] rise, blood sugar level and triglyceride-glucose index [TyG]; the five CVD related proteins comprise IGF2R, LGALS4, NUP50, PIGR and VSIG2, and the five CVD related proteins comprise the IGF2R, LGALS4, NUP50, PIGR and VSIG2. The cardiovascular disease prediction model provided by the invention can provide ultra-early warning within 10 years before attack, is high in prediction accuracy, and is superior to the conventional model in short-term and long-term prediction aspects; meanwhile, a key proteome pathway is disclosed, and reference can be provided for clinical application and prevention strategies. The invention also provides application based on the cardiovascular disease prediction model and an early screening kit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a cardiovascular disease prediction model based on cardiac metabolism and plasma proteomics in patients with inflammatory bowel disease, its application, and an early screening kit. Background Technology

[0002] Inflammatory bowel disease (IBD) is a group of chronic inflammatory diseases of the gastrointestinal tract with unknown causes, including ulcerative colitis (UC) and Crohn's disease (CD). Due to its unknown etiology and lack of effective treatment, it is gradually becoming a global public health problem. IBD is widely recognized as a multisystem disease not limited to the digestive system, and previous studies have shown that patients with IBD have a more than 10% increased risk of cardiovascular disease (CVD) compared to the general population. Cardiovascular disease is a leading cause of death worldwide, and timely identification of high-risk individuals for CVD among IBD patients and implementation of preventive measures are crucial to improving the public health burden.

[0003] However, early identification of CVD events in IBD patients remains an unresolved clinical problem. Currently, most CVD risk scoring and assessment focuses only on type 2 diabetes or the general population, and CVD prediction tools specifically for IBD patients are still lacking.

[0004] Traditional cardiovascular disease risk prediction models, such as the Systemic Coronary Risk Assessment 2 (SCORE2), often underestimate the cardiovascular disease risk in IBD patients and are prone to classification errors when extrapolating traditional cardiovascular disease models to the IBD population. Therefore, there is an urgent need to develop tools that can more accurately predict the cardiovascular disease risk in the IBD population, enabling targeted lifestyle modifications, timely drug treatment, and effective reduction of cardiovascular disease incidence. This will provide more personalized and precise medical services for IBD patients and improve their long-term health.

[0005] Intra-intestinal dysplasia (IBD), characterized by chronic intestinal inflammation and immune dysregulation, can induce systemic metabolic disorders. The accumulation of cardiovascular metabolic risk factors (CMFs), including dyslipidemia and insulin resistance, can significantly increase the risk of cardiovascular disease. Furthermore, advanced technologies such as plasma proteomics and metabolomics analysis have enabled the identification of novel biomarkers and dysregulated pathways associated with the pathogenesis of CVD in IBD. The polygenic risk score (PRS) for CVD plays a crucial role in quantifying genetic susceptibility, particularly in younger IBD patients where the predictive power of traditional risk factors is relatively weak. Therefore, the multidimensional integration of CMFs, genomic biomarkers, and genetic susceptibility has emerged as a promising approach for achieving precise risk stratification in the IBD population.

[0006] This invention aims to combine multi-omics data with cardiovascular metabolic risk factors to develop a machine learning-based CVD event prediction tool for IBD patients, providing new insights for future clinical practice and health management of IBD patients. Summary of the Invention

[0007] To address the shortcomings of existing technologies, the main objective of this invention is to provide a cardiovascular disease prediction model and its application based on cardiac metabolism and plasma proteomics in patients with inflammatory bowel disease.

[0008] To achieve the aforementioned main objectives, on the one hand, the present invention provides a cardiovascular disease prediction model based on cardiac metabolism and plasma proteomics in patients with inflammatory bowel disease, the cardiovascular disease prediction model comprising:

[0009] Ten cardiovascular metabolic risk factors (CMFs) were identified, including physical indicators and blood biochemical indicators. Physical indicators included systolic blood pressure [SBP], diastolic blood pressure [DBP], body mass index [BMI], and waist circumference. Blood biochemical indicators included high-density lipoprotein cholesterol [HDL-C], low-density lipoprotein cholesterol [LDL-C], total cholesterol, elevated glycated hemoglobin [HbA1c], blood glucose levels, and triglyceride-glucose index [TyG].

[0010] Five CVD-related proteins were identified, including IGF2R, LGALS4, NUP50, PIGR, and VSIG2.

[0011] According to another specific embodiment of the present invention, the method for constructing a cardiovascular disease prediction model includes the following steps:

[0012] A. Using prospective cohort study data from biobanks (e.g., the UK Biobank), we analyzed a large number (e.g., tens of thousands) of adult individuals who had no cardiovascular disease (CVD) at baseline but had a history of inflammatory bowel disease (IBD) before baseline, and randomly assigned them to a training set (85%) and a test set (15%).

[0013] B. Identify CVD-related proteins and metabolites using Cox regression analysis;

[0014] C. Five-fold cross-validation was used to perform machine learning on the training set, and the accuracy of the cardiovascular disease prediction model for CVD was evaluated on the test set.

[0015] According to another specific embodiment of the present invention, the cardiovascular disease prediction model further includes a CVD-related metabolite, namely creatinine.

[0016] On the other hand, the present invention provides an application of the above-mentioned cardiovascular disease prediction model based on cardiac metabolism and plasma proteomics of patients with inflammatory bowel disease, which is applied to the preparation of cardiovascular disease early screening kits.

[0017] In another aspect, the present invention provides an early screening kit for cardiovascular disease prediction, the early screening kit comprising detection reagents for detecting the concentrations of the following five proteins: IGF2R, LGALS4, NUP50, PIGR and VSIG2.

[0018] Patients with inflammatory bowel disease (IBD) have a significantly higher risk of cardiovascular disease (CVD) than the general population. This invention integrates the proteomics and CMFs of IBD patients to construct a novel, accurate, and robust CVD prediction model, which is of great significance for identifying high-risk individuals and early intervention.

[0019] This invention utilizes machine learning algorithms to develop a proteomics-based model for the early identification of cardiovascular diseases, and compares its predictive accuracy with existing CVD models. The joint model of proteins and CMFs demonstrates excellent predictive performance for CVD events throughout the follow-up period. This invention offers the following advantages:

[0020] The cardiovascular disease prediction model of this invention can provide ultra-early warning within 10 years before the onset of the disease. Its prediction accuracy is high, and it is superior to previous models in both short-term and long-term prediction. At the same time, it reveals key proteomic pathways, which can provide a reference for clinical applications and prevention strategies.

[0021] The present invention will now be described in further detail with reference to the accompanying drawings. Attached Figure Description

[0022] Figure 1 It is a research and design flowchart;

[0023] First, data were extracted from 5,248 participants with pre-baseline inflammatory bowel disease (IBD) and a median follow-up of 13.76 years from UK biobank participants. This included disease data for IBD and cardiovascular disease as defined by ICD-10 codes, over 2,700 plasma proteomics and over 250 plasma metabolomics data, 10 cardiovascular metabolic risk factor indicators, and predictors of multigene risk scores for cardiovascular disease. Next, feature selection was performed using the results of a Cox proportional hazards model. The population was divided into training and test sets at an 85% to 15% ratio based on different recruitment centers, using four machine learning methods.

[0024] The model was established and evaluated through 10 replicates of internal 5-fold cross-validation, and its predictive power was quantified by the average AUC of the 10 replicates. A series of sensitivity analyses further evaluated the model's predictive power for new CVD events within 10 years, including replicate analyses after dividing the training and test sets by different proportions and excluding individuals diagnosed with CVD within one year of baseline, and comparisons with the existing model SCORE2. Furthermore, the relationship between modifiable clinical phenotypes and cardiovascular metabolic risk factors and plasma proteins was explored.

[0025] Figure 2 shows the prediction accuracy of all models for new CVD events, where:

[0026] Figure 2A The prediction accuracy of all models for new CVD events throughout the follow-up period is shown in the bar chart, which displays the area under the receiver operating characteristic (AUC) curves for different CVD prediction models on test sets that are geographically different from the training set.

[0027] Figure 2B The prediction accuracy of all models for new CVD events within 10 years is shown. The bar chart shows the area under the receiver operating characteristic (AUC) curves of different CVD prediction models on test sets that are geographically different from the training set.

[0028] Figure 3 This is a SHAP visualization of the joint model of 5 proteins and 10 cardiovascular metabolic risk factors in the Extra Trees algorithm. The width of the horizontal axis can be interpreted as the degree of contribution to CVD prediction; the wider the range, the greater the contribution. The colors represent the magnitude of the predictor variables, encoded as a gradient from blue (low) to red (high), as shown by the color bar on the right. The x-axis direction represents the probability of having CVD (right) or being healthy (left).

[0029] Abbreviations: BMI, Body Mass Index; CVD, Cardiovascular Disease; CMFs, Cardiovascular Metabolic Risk Factors; DBP, Diastolic Blood Pressure; HDL-C, High-Density Lipoprotein; HbA1c, Glycated Hemoglobin; SBP, Systolic Blood Pressure; SHAP, Shapley Additive Explanation; TyG, Triglyceride-Glucose Index.

[0030] Figure 4 shows the association between modifiable clinical phenotypes and the parameter levels of cardiovascular disease prediction models, where:

[0031] Figure 4AThis study demonstrates the association between modifiable clinical phenotype and the levels of 10 cardiovascular metabolic risk factors. The heatmap shows the association between modifiable clinical phenotype and the levels of 10 cardiovascular metabolic risk factors, including lack of exercise, smoking, alcohol consumption (≥3 times per week), poor diet, poor sleep, and sedentary lifestyle. The analysis used a linear regression model, adjusted for age, sex, and race, with the levels of the 10 cardiovascular metabolic risk factors as the outcome variable and modifiable clinical phenotype as the explanatory variable. The β coefficient represents the effect size, with red indicating a positive correlation and blue indicating a negative correlation. Statistical significance was determined using false discovery rate (FDR) correction (*P<0.05, **P<0.01, ***P<0.001, ****P<0.0001).

[0032] Abbreviations: BMI, Body Mass Index; DBP, Diastolic Blood Pressure; HDL-C, High-Density Lipoprotein; HbA1c, Glycated Hemoglobin; SBP, Systolic Blood Pressure; TyG, Triglyceride-Glucose Index;

[0033] Figure 4B The association between modifiable clinical phenotype and five protein levels is shown in the heatmap. Modifiable clinical phenotypes include obesity, lack of exercise, smoking, alcohol consumption (≥3 times per week), poor diet, poor sleep, and sedentary lifestyle. The analysis used a linear regression model, adjusted for age, sex, and race, with the five protein levels as outcome variables and modifiable clinical phenotype as an explanatory variable. The β coefficient represents the effect size, with red indicating a positive correlation and blue indicating a negative correlation. Statistical significance was determined using false discovery rate (FDR) correction (*P<0.05, **P<0.01, ***P<0.001, ****P<0.0001). Detailed Implementation

[0034] Example 1

[0035] Patients with inflammatory bowel disease (IBD) have a significantly higher risk of cardiovascular disease (CVD) than the general population. Reliable tools for identifying CVD in IBD patients, especially long-term predictive tools, remain to be determined.

[0036] In this study, feature selection and model development were performed on IBD patients from the UK Biobank. Participants were randomly assigned to a training set (85%) and a test set (15%) based on their recruitment center. Cox regression analysis was used to identify CVD-related proteins and metabolites. This invention was developed and internally validated using four machine learning algorithms, with 10 replicates of 5-fold cross-validation and independent validation in cohorts geographically different from the training set. Furthermore, 10-year CVD risk was assessed, and the predictive accuracy of the model was compared with existing CVD models. The performance of each model was evaluated using the area under the receiver operating characteristic (AUC) and 95% confidence intervals.

[0037] During a median follow-up of 13.76 years, 715 CVD events occurred among 5248 IBD patients. Ten cardiovascular metabolic risk factors (CMFs), polygenic risk scores (PRS), five CVD-related proteins, and one CVD-related metabolite were used to construct the model. For CVD events throughout the follow-up period, the combined protein and CMF model demonstrated excellent predictive performance, with an AUC (95% CI) of 0.88 (0.86–0.89) on the training set and an AUC (95% CI) of 0.79 (0.79–0.80) on the geographically differentiated test set. Both the protein model (AUC: 0.76 [95% CI: 0.75–0.76]) and the CMF model (AUC: 0.72 [95% CI: 0.72–0.72]) outperformed other single-omics models. Notably, the protein model performed exceptionally well in predicting the 10-year risk of CVD (AUC: 0.87 [95% CI: 0.85–0.90]). In the UK biobank population, both the protein model and the combined model outperformed traditional CVD models (AUC 0.73, P < 0.05).

[0038] This study is the first to integrate proteomics and CMFs of IBD patients to construct a novel, accurate and robust CVD prediction model, which is of great significance for identifying high-risk individuals and early intervention.

[0039] This embodiment develops a cardiovascular disease prediction model based on cardiac metabolism and plasma proteomics in patients with inflammatory bowel disease, which can predict cardiovascular disease up to 10 years in advance in the general population.

[0040] I. Methods

[0041] 1. Study population

[0042] This study used data extracted from the UK Biobank. The UK Biobank (UKB) is a global biomedical database containing over 500,000 UK participants with a baseline age range of 37 to 73 years. Participants were recruited from 22 centers in the UK between 2006 and 2010. More details about the UK Biobank can be found in other studies.

[0043] This study included patients with a pre-baseline history of IBD and excluded those diagnosed with CVD at baseline, resulting in a total of 5248 participants. Participants were randomly assigned to the training and test sets by the UKB recruitment center, with a ratio of approximately 85%:15%. The training set included participants from Manchester, Oxford, Cardiff, Glasgow, Edinburgh, Reading, Bury, Newcastle, Leeds, Nottingham, Sheffield, Liverpool, Middlesbrough, Hounslow, Croydon, Birmingham, Swansea, and Wrexham, while the test set included participants from Stockport (pilot), Stock, Bristol, and Butts (Supplementary Table 1). Figure 1 The research design is explained.

[0044] 2. Definitions of IBD and CVD

[0045] Individuals diagnosed with IBD are identified using the International Classification of Diseases 10 (ICD-10), coded K50 and K51. New-onset CVD is defined as the first occurrence of ischemic heart disease (ICD-10 codes I20-25) or cerebrovascular disease (ICD-10 codes I60-69), determined by linking it to hospital event statistics (HES) or the national mortality index.

[0046] Self-reported cases were excluded from this study to ensure the reliability of IBD and CVD diagnoses. Detailed information on ICD-10 codes can be found in Supplementary Table 2.

[0047] 3. Plasma proteome

[0048] The UKB Pharmaceutical Proteomics Project (UKB-PPP) consortium provided a wealth of proteomics data from blood samples. Most of these samples were collected during participants' first visits to the UK assessment centre between 2007 and 2010, with others coming from consortium members and participants in the COVID-19 repeat imaging study.

[0049] Blood samples were collected using EDTA tubes, centrifuged at 4°C for 10 minutes to obtain plasma, and stored at -80°C. They were then transported to Olink Analytical Services in Sweden using dry ice for analysis. TMProteomics analysis was performed using Explore 3072 Proximity Extended Analysis (PEA). Following rigorous quality control measures, this study quantified 2,736 unique proteins from 54,219 participants, distributed across eight proteomes: Cardiometabolism, Cardiometabolism II, Inflammation, Inflammation II, Neurology, Neurology II, Oncology, and Oncology II. Protein levels were converted to normalized protein expression (NPX) values ​​on a log2 scale.

[0050] 4. Plasma metabolome

[0051] EDTA plasma samples were collected at baseline and tested between June 2019 and April 2020. Quantitative analysis of the samples was performed using a high-throughput NMR-based metabolic biomarker analysis platform developed by Nightingale Health Ltd. Metabolomics encompassed absolute levels (mmol / L) of 251 metabolites from approximately 280,000 UK Biobank participants, including amino acids, glycolytic metabolites, ketone bodies, lipids, lipoproteins, and transfatty acids. Further details on metabolomics quantification can be found at:

[0052] https: / / biobank.ndph.ox.ac.uk / showcase / ukb / docs / NMR_companion_ phase2.pdf .

[0053] 5. Definition of PRS for CMFs and CVD

[0054] Ten CMFs were selected as predictive features for this study because they have a clear association with cardiovascular disease risk. The selected features included physical parameters (systolic blood pressure [SBP], diastolic blood pressure [DBP], body mass index [BMI], waist circumference) and blood biochemical parameters (high-density lipoprotein cholesterol [HDL-C], low-density lipoprotein cholesterol [LDL-C], total cholesterol, elevated glycated hemoglobin [HbA1c], blood glucose levels, and triglyceride-glucose index [TyG]). Detailed definitions of CMFs are provided in Supplementary Table 3. The PRS (Field ID 26223) for CVD was calculated based on external GWAS data from Genomics PLC under UKB Project 9659.

[0055] 6. Statistical Analysis

[0056] Participant baseline characteristics were described as means or percentages. In descriptive analyses, chi-square tests were used for categorical variables and t-tests were used for continuous variables to compare differences between groups (CVD vs. non-CVD). Participant follow-up began on the date of their first visit to the UKB assessment center (FieldID53) and continued until the earliest date of death, first diagnosis, or review (October 30, 2022).

[0057] A Cox proportional hazards regression model was run on the training set to estimate the association between the NPX values ​​of each plasma protein or metabolite and new-onset CVD, and HR values, 95% confidence intervals (CI), and p-values ​​were reported. The model was adjusted for age, sex, race, BMI, smoking, alcohol consumption, chronic obstructive pulmonary disease (COPD), and diabetes. We used the MICE package in R, applying chain equations and predicted mean matching, combined with regression and nearest neighbor models to resolve missing values ​​for covariates. Bonferroni correction was used to assess significant associations (P < 0.05), taking into account the number of proteins detected (n = 2736) and metabolites (n = 251).

[0058] This invention selects mature and efficient machine learning algorithms such as Light Gradient Boosting Machine (LGBM), eXtreme Gradient Boosting (XGBoost), Random Forest (RF), and Extra Trees as benchmark methods. Ten CMFs, CVD PRS, and proteins and metabolites that remain significant after Bonferroni correction are individually or jointly input into the FLAML Automated Machine Learning (AutoML) model. Participants are divided into training and test sets based on recruitment, and all models are built using the same AutoML framework described above.

[0059] To further reduce overfitting, the four machine learning algorithms performed internal 10-fold 5-fold cross-validation on the training set, with a maximum iteration count limited to 100 and early termination. Hyperparameter tuning was performed using the built-in AutoML feature to determine the optimal hyperparameters for maximum performance. Automatic class weight balancing in scikit-learn's "Balanced" mode was used to address class imbalance, and the computed weights were incorporated into automatic training (via the `sample_weight` parameter).

[0060] Finally, five machine learning models were developed based on 10 CMFs (Continuous Metabolomics Functions) for CVD, proteomics, metabolomics, and PRS (Progressive Reproductive Stimulus). Model performance was primarily evaluated using receiver operating characteristic (ROC) area under the curve (AUC) analysis to assess the predictive accuracy of these models for new-onset CVD in the test set. The training and testing performance of each model was measured using the average of 10 crossovers of the bootstrap method (1000 iterations per run). The Mann-Whitney U test was used to determine whether there were statistically significant differences in AUC values ​​between models. Furthermore, the algorithm with the highest AUC on the test set, geographically different from the training set, was used to rank predicted features based on the mean of the Shapley additive explanations (SHAP) from 10 repeated evaluations. Shapley additive explanations are a metric for feature importance, assessing their contribution to model performance, and then visualized using SHAP plots. The Shap plots visually explain the impact of each feature in the model through the magnitude of its values ​​(encoded with color gradients) and the direction of the trend on the horizontal axis (the probability of developing CVD).

[0061] To evaluate the generalization and predictive accuracy of the above model, the following analyses were performed: (1) Participants were randomly assigned to a 75% training set and a 25% test set based on the recruitment center; (2) Individuals who developed CVD within the first year after baseline were excluded from the analysis; and (3) CVD events were replicated over 10 years. In the corresponding analyses, individuals diagnosed with CVD within 10 years after baseline were considered the case group, and those diagnosed with CVD after 10 years were considered healthy individuals. To further evaluate the superiority of the model, the existing SCORE2 model was applied to the UKB population to develop the SCORE2 model for comparison with our model. The SCORE2 model consists of clinical characteristics such as age, sex, smoking status, history of diabetes, systolic blood pressure, total cholesterol (TC), and HDL-C.

[0062] Furthermore, this invention performed linear regression analysis to assess the cross-sectional association between CVD-related protein and CMF levels and modifiable clinical phenotypes. CVD-related protein and CMF levels were used as outcome variables, while phenotypes, including obesity, physical activity, smoking status, frequency of alcohol consumption, dietary habits, sedentary time, and sleep duration, were used as explanatory variables. All these analyses were adjusted for age, sex, and ethnicity. The mediating role of CVD-related plasma proteins and CMFs in the association between IBD and CVD was further assessed using the “mediation” package in R.

[0063] All statistical analyses and visualizations were performed using R 4.2.1 software and Python 3. A p-value < 0.05 was considered statistically significant (two-sided).

[0064] II. Results

[0065] 1. Baseline characteristics

[0066] During a median follow-up of 13.76 years (IQR 13.12–14.45), 715 CVD events occurred among 5248 IBD participants. At baseline, compared with the control group, IBD patients with future CVD events had higher mean age, BMI, blood pressure, waist circumference, and levels of certain blood biochemistry parameters (including glycated hemoglobin, blood glucose, and TyG), were more likely to be smokers, and consumed alcohol almost daily. The case group had a lower mean for high-density lipoprotein (HDL) (P<0.05). Detailed baseline characteristics are shown in Table 1.

[0067] Table 1 Baseline characteristics of the study population

[0068]

[0069]

[0070] For continuous variables, the Student's t-test is used; for discrete variables, the Pearson chi-square test is used to assess differences between groups. Abbreviations: CVD, cardiovascular disease; TyG, triglycerides-blood glucose.

[0071] 2. Plasma proteins and metabolites related to cardiovascular disease

[0072] Among 2,736 plasma proteomic biomarkers and 251 metabolites, after adjustment for multiple covariates, 256 proteins and 16 proteins were significantly associated with increased CVD risk (Supplementary Tables 4-5). After Bonferroni correction, five proteins—IGF2R (HR 7.84, P = 1.68 × 10⁻²), LGALS4 (HR 1.98, P = 2.53 × 10⁻²), NUP50 (HR 2.06, P = 2.70 × 10⁻²), PIGR (HR 3.10, P = 4.47 × 10⁻²), and VSIG2 (HR 1.91, P = 4.70 × 10⁻²)—and one metabolite—creatinine (HR 326.85, P = 1.83 × 10⁻⁵)—remained significant.

[0073] 3. Predictive accuracy of new CVD events during the follow-up period

[0074] For new CVD events during the follow-up period, in a test set geographically different from the training set, the AUC ranged from 0.76 to 0.78 for the 5 CVD-related protein models and from 0.70 to 0.72 for the 10 CMFs models. Figure 2A (See Supplementary Tables 6-7). Among all four machine learning algorithms, the panel of 5 CVD-related proteins and 10 CMFs showed fairly stable predictive performance compared to the PRS model for CVD (AUC: 0.53-0.55) or the metabolomics model (AUC: 0.59-0.60) (Mann-Whitney U test: P < 0.05). Therefore, this invention further constructed a joint model of 5 CVD-related proteins and 10 CMFs. The AUC of the joint model ranged from 0.73 to 0.79, with the AUC of Extra trees reaching 0.79 and 95% CI of 0.79-0.81. After dividing the population into training / test sets at a ratio of 75% / 25% (Supplementary Tables 12-13), the joint model still performed best in the Extra trees algorithm (AUC 0.73, 95% CI: 0.72-0.74).

[0075] The SCORE2 model identified in previous studies was further applied to the UKB population, and its performance was compared with the aforementioned combined model (Supplementary Tables 14-15). The results showed that the SCORE2 model (AUC range: 0.73-0.74) performed significantly worse than the protein and CMFs combined model in predicting new CVD (P<0.05).

[0076] 4. Explanation of the combined model of 10 CMFs and 5 CVD-related proteins

[0077] The impact of the predictive features of the joint model with the highest AUC in the test set is visualized using a Shap plot. Figure 3 As can be seen, the protein VISIG2 has the highest global importance, followed by SBP and BMI. The IGF2R protein has the broadest response to CVD events, indicating its strong predictive ability for new CVD events in individuals with IBD. Taking the protein VISIG2 as an example, people with higher VISIG2 levels (red) are more likely to develop cardiovascular disease (right), while those with lower VISIG2 levels (blue) tend to remain healthy (left). Similarly, for other characteristics, participants with higher blood pressure (systolic and diastolic), IGF2R, LGALS4 and PIGR, glycated hemoglobin, and cholesterol levels are more likely to develop CVD, while those with lower levels tend to remain healthy.

[0078] 5. Results of replication validation analysis

[0079] The findings of this invention are consistent with the sensitivity analysis. The predictive performance of the five CVD-related protein models and the ten CMFs models remained stable across the four algorithms, and their predictive performance was superior to the other models after excluding individuals with CVD in the first year of follow-up or dividing the population into test / training sets at 75% / 25% (Supplementary Tables 10-14). Notably, for 10-year CVD prediction, the panel for the five CVD-related proteins showed considerable predictive power, with the AUC of the four algorithms ranging from 0.85 to 0.87. Figure 2B (and supplementary tables 8-9).

[0080] 6. The relationship between lifestyle-modified phenotypes and CMFs and protein levels, and the mediating roles of 10 CMFs and 5 proteins.

[0081] To further explore whether the association between CVD, CMFs, and proteins is influenced by modifiable risk factors, this invention investigated the relationship between modifiable lifestyle phenotypes and 10 CMFs and 5 proteins. Obesity, lack of exercise, smoking, poor diet, and insufficient sleep were significantly associated with some CMFs and proteins (Figure 4 and Supplementary Tables 16-17). For example, smoking and alcohol consumption were positively correlated with LGALS4, VSIG2, and PIGR proteins, while obesity was positively correlated with IGF2R and LGALS4. Mediation analysis showed that the TyG index partially mediated the association between IBD and CVD, with a mediating effect of 8.52% (Supplementary Table 18).

[0082] III. Discussion

[0083] In 5248 IBD patients from UKB, this study evaluated the potential of multi-omics data (e.g., CMFs and plasma proteins) based on four machine learning algorithms to predict CVD events. Five protein models showed the best predictive performance, followed by ten CMF models. When proteins and CMFs were combined, this combined model exhibited the highest predictive performance in the Extra trees algorithm, with an AUC (95% CI) of 0.79 (0.79–0.81). Furthermore, the five CVD-related protein models demonstrated satisfactory accuracy and robustness in predicting the 10-year risk of CVD (AUC: 0.85–0.87).

[0084] Previous CVD prediction models, primarily based on diabetic patients or the general population, have limited applicability to the specific population of IBD. Due to the pathophysiological characteristics of IBD patients, including chronic inflammation and gut microbiota dysbiosis, traditional models relying solely on clinical indicators such as age, sex, and blood pressure may not accurately reflect their true CVD risk.

[0085] This study, focusing on patients with impaired blood pressure (IBD), is the first to develop a novel CVD prediction model for this specific population by integrating protein biomarkers and clinical indicators, overcoming the limitations of previous studies in terms of target population and predictive variables. This study identified five key proteins (IGF2R, LGALS4, NUP50, PIGR, and VSIG2) that play important roles in predicting new CVD events in the IBD population. Compared to other single-omics models, these proteins demonstrated superior performance and showed considerable ability in predicting 10-year cardiovascular disease risk.

[0086] Previous studies have shown that IGF2R, LGALS4, PIGR, and VSIG2 proteins play important roles in the pathogenesis of CVD through involvement in pathological mechanisms such as inflammation regulation, metabolic disorders, and immune regulation. Notably, NUP50 protein is a newly discovered CVD-related protein in this study. Previous literature has reported that NUP50 is an important component of the nuclear pore complex, involved in chromatin binding, and may affect cardiomyocyte survival and apoptosis by regulating the transport of signaling molecules into and out of the nucleus. However, current research on the roles of these proteins in CVD in individuals with IBD is limited; this study is the first to demonstrate their important predictive value in CVD.

[0087] The ten CMFs models showed stable performance in the IBD population, but lower than proteomics, indicating that predicting CVD risk in IBD patients solely based on traditional indicators is insufficient. More importantly, both CMFs and proteins demonstrated robust performance across different algorithms, demonstrating methodological robustness in their predictive capabilities and laying the foundation for developing composite models integrating protein biomarkers and clinical indicators.

[0088] The significant advantages of the combined model over single-omics models reflect the synergistic effect of integrating multidimensional biomarkers. Proteomics reveals specific molecular mechanisms of IBD-related CVD, while classic CMFs supplement information on systemic metabolic disorders; the combination of the two enables a comprehensive analysis from molecular mechanisms to clinical phenotypes. Furthermore, both the protein model and the combined model outperform the SCORE2 model, demonstrating the limitations of traditional CVD risk assessment frameworks in specific populations and revealing that incorporating disease-specific biomarkers can significantly improve the sensitivity of disease risk identification.

[0089] These findings not only provide a novel and highly accurate tool for early stratification of CVD risk in IBD patients, but also promote interventional research on the pathogenesis of IBD-related CVD by revealing new protein targets, thus having dual value for developing CVD prevention and control guidelines for the IBD population. The association between modifiable clinical phenotypes and CMFs and proteins further provides methods for prevention strategies.

[0090] This study has several key advantages. First, it establishes a novel and accurate CVD event prediction tool for individuals with IBD, effectively filling a research gap and providing a customized tool for CVD risk monitoring in high-risk IBD populations. Second, it is based on UKB, a large-scale database with long-term follow-up and population-scale data. Third, it combines multi-omics data, including CMFs, proteomics, metabolomics, and genetics, ensuring predictive power, and utilizes various machine learning algorithms to build the model. Finally, it performs 10 replicates of 5-fold cross-validation on a test set geographically different from the training set, along with a series of sensitivity analyses, ensuring the accuracy and stability of the results.

[0091] However, this study has some unavoidable limitations. First, since most of the data used came from participants of European descent, this may affect the model's generalization to other ethnicities. Second, although UKB provides a broad assessment of circulating proteins, it does not include the entire human proteome, which may introduce bias in the selection of secreted proteins for measurement. Third, although the model demonstrated consistent performance in extensive cross-validation and replication analyses, further external validation in a large-scale, long-term prospective cohort is still necessary.

[0092] Based on one of the world's largest long-term prospective cohorts, this study is the first to build an accurate and robust CVD prediction model by integrating proteomics and CMFs of IBD patients, which makes up for the limitations of traditional models and provides a promising tool for personalized diagnosis and intervention.

[0093] While the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the scope of the invention. Any person skilled in the art can make modifications without departing from the scope of the invention; all equivalent modifications made in accordance with the invention should be covered by the scope of the invention.

Claims

1. A CVD prediction model for IBD patients, characterized in that, The CVD prediction model includes: Ten cardiovascular metabolic risk factors, including physical indicators and blood biochemical indicators; physical indicators include systolic blood pressure, diastolic blood pressure, body mass index and waist circumference, and blood biochemical indicators include elevated high-density lipoprotein cholesterol, low-density lipoprotein cholesterol, total cholesterol, elevated glycated hemoglobin, blood glucose level and triglyceride-glucose index. Five CVD-related proteins were identified, including IGF2R, LGALS4, NUP50, PIGR, and VSIG2.

2. The CVD prediction model as described in claim 1, characterized in that, The method for constructing the cardiovascular disease prediction model includes the following steps: A. Using prospective cohort study data from biobanks, we analyzed a large number of adult individuals who had no cardiovascular disease at baseline but had a history of inflammatory bowel disease before baseline, and randomly divided them into a training set (85%) and a test set (15%). B. Identify CVD-related proteins and metabolites using Cox regression analysis; C. Machine learning is performed on the training set using 5-fold cross-validation, and the accuracy of the cardiovascular disease prediction model for CVD is evaluated on the test set.

3. The CVD prediction model as described in claim 1, characterized in that, The cardiovascular disease prediction model further includes one CVD-related metabolite, which is creatinine.

4. An application of the CVD prediction model as described in claim 1, characterized in that, The CVD prediction model was applied to the preparation of a cardiovascular disease early screening kit for IBD patients.

5. An early screening kit for predicting cardiovascular disease in individuals with IBD, characterized in that, The early screening kit includes detection reagents for detecting the concentrations of the following five proteins: IGF2R, LGALS4, NUP50, PIGR, and VSIG2.