CKD early screening marker based on plasma proteomics and application

By selecting IGFBP4, GDF15, EDA2R, HAVCR1, XG and other proteins as early screening markers and combining them with machine learning methods, the problem of insufficient early detection of CKD in existing technologies was solved, highly accurate ultra-early warning and model verification were achieved, and the key proteomic pathways of CKD were revealed.

CN120685912APending Publication Date: 2025-09-23GUANGDONG GENERAL HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510591599.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Currently, diagnostic strategies based on estimated glomerular filtration rate and albuminuria are insufficient for early detection of chronic kidney disease (CKD), limiting the potential for timely therapeutic intervention, and the predictive utility of plasma proteome profiles during long-term follow-up has not been fully explored.

Method used

By utilizing large prospective cohort data and analyzing plasma proteomics, we selected five proteins, including IGFBP4, GDF15, EDA2R, HAVCR1, and XG, as early screening markers. We then combined machine learning methods to establish a model for early prediction of CKD.

Benefits of technology

It achieved ultra-early warning within 16 years before onset, improved the accuracy of CKD prediction, revealed key proteomic pathways, and provided a reference for clinical application and prevention strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120685912A_ABST
    Figure CN120685912A_ABST
Patent Text Reader

Abstract

The invention provides a CKD early screening marker based on plasma proteomics. The early screening marker comprises the following five kinds of protein: IGFBP4, GDF15, EDA2R, HAVCR1 and XG. The early screening marker provided by the invention can provide ultra-early warning within 16 years before disease attack, is high in prediction accuracy, and is superior to the conventional model in short-term and long-term prediction aspects; meanwhile, a key proteome pathway is disclosed, and reference can be provided for clinical application and prevention strategies. The invention also provides application of the early screening marker based on the CKD.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a CKD early screening marker based on plasma proteomics and its application. Background Art

[0002] Chronic kidney disease (CKD) is a major global health concern, affecting approximately 10% of the world's adult population, leading to increased morbidity, mortality, and healthcare costs. CKD is characterized by a gradual decline in kidney function, which often ultimately leads to end-stage renal disease (ESRD) and the need for dialysis or kidney transplantation. Early detection and intervention are critical to slowing disease progression and improving patient outcomes. However, current diagnostic strategies, which are primarily based on estimated glomerular filtration rate (eGFR) and albuminuria, are often insufficient to detect early CKD, limiting the potential for timely therapeutic intervention.

[0003] Plasma proteomics has emerged as a promising tool for the early prediction of CKD. Proteins are dynamic molecules that reflect genetic and environmental influences, and their circulating levels can provide insights into pathophysiological processes that precede the clinical manifestation of CKD. Recent advances in high-throughput proteomic technologies have enabled the identification of protein biomarkers associated with the risk and progression of chronic kidney disease, providing a new approach for early disease prediction. Some studies have found that specific proteins associated with inflammation, oxidative stress, and metabolic dysfunction are early indicators of CKD, highlighting the potential of proteomic analysis for risk stratification and early diagnosis.

[0004] Despite these advances, the utility of plasma proteomic profiles in predicting CKD during long-term follow-up remains underexplored. Most previous studies have been limited by small sample sizes, short follow-up durations, or cross-sectional designs, which reduce their generalizability and predictive value. Longitudinal studies with extended follow-up are essential to validate the predictive power of plasma proteomic biomarkers and determine their clinical utility in detecting early CKD.

[0005] The present invention aims to investigate the predictive value of plasma proteome profiles for chronic kidney disease 16 years before clinical diagnosis; to identify protein biomarkers associated with future CKD risk and evaluate their performance in early disease prediction by utilizing data from a large prospective cohort with comprehensive proteomic and clinical follow-up information; the results may improve the early diagnosis of chronic renal failure, develop targeted prevention strategies, and provide assistance in understanding the potential development mechanisms of chronic renal failure. Summary of the Invention

[0006] In order to address the deficiencies in the prior art, the main purpose of the present invention is to provide a CKD early screening marker based on plasma proteomics and its application.

[0007] In order to achieve the above main objectives, on the one hand, the present invention provides CKD early screening markers based on plasma proteomics, which include the following five proteins: IGFBP4, GDF15, EDA2R, HAVCR1, and XG.

[0008] According to another specific embodiment of the present invention, the method for selecting early screening markers includes the following steps:

[0009] A. Utilize prospective cohort study data from biobanks (e.g., UK Biobank) to analyze a large number (e.g., tens of thousands) of adults without CKD at baseline and measure several (e.g., thousands) of plasma proteins;

[0010] B. Proteins associated with CKD were determined by Cox regression analysis;

[0011] C, Machine learning using 10-fold cross-validation to quantify the importance of the proteins identified in step B.

[0012] On the other hand, the present invention provides an application of the above-mentioned plasma proteomics-based CKD early screening marker, which is used in the preparation of a CKD early screening kit.

[0013] Early identification of chronic kidney disease (CKD) is crucial for preventing disease progression, but effective predictive tools are currently lacking. The inventors analyzed proteomic data from 37,792 patients without baseline CKD in southern England. Using the multivariate Cox-Boruta algorithm, proteins were selected and modeled, and SHAP values ​​were calculated. The top-ranked proteins were combined with clinical data and polygenic risk scores to construct a model.

[0014] These models were internally validated using 1,000 bootstrap resets in the South UK Cohort and externally validated in the Northern UK Cohort (N = 14,956 participants) and the China Kadoorie Biobank (CKB) (N = 3,977 participants). Cumulative CKD incidence rates were compared across quintiles of baseline protein concentrations, and temporal trends in protein levels were assessed. Of the 2,737 proteins analyzed, five were reliable predictors of CKD incidence within 5 years: IGFBP4 (AUC = 0.906), GDF15 (AUC = 0.875), EDA2R (AUC = 0.893), HAVCR1 (AUC = 0.805), and XG (AUC = 0.736). Participants with higher baseline levels of these proteins had a significantly increased risk of developing chronic kidney disease (HR = 1.97-5.70). In the Northern England cohort, a combination of five proteins showed excellent predictive accuracy across different timeframes: 5 years (AUC = 0.921), 10 years (AUC = 0.884), more than 10 years (AUC = 0.790), and all time (AUC = 0.858). Incorporation of demographic and clinical predictors further improved the predictive power of 5 years (AUC = 0.942), 10 years (AUC = 0.927), more than 10 years (AUC = 0.842), and all time (AUC = 0.907).

[0015] The 5-protein panel also performed well in the southern UK cohort (AUC = 0.860). The inventors' study identified five key proteins and established and validated an important prediction model for CKD more than 16 years before onset, providing a new potential strategy for ultra-early prediction and intervention.

[0016] The present invention has the following beneficial effects:

[0017] The early screening markers of the present invention can provide ultra-early warning within 16 years before onset of the disease. They have high predictive accuracy and are superior to previous models in both short-term and long-term predictions. At the same time, they reveal key proteomic pathways, which can provide a reference for clinical applications and prevention strategies.

[0018] The present invention will be described in further detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is an overview diagram of the study process in Example 1. The UK Biobank and the China Kadoorie Biobank excluded individuals with a baseline diagnosis of CKD or self-reported CKD, as well as individuals who did not undergo plasma proteomic analysis. The remaining participants were grouped by the year of initial CKD diagnosis after baseline, and the type of diagnostic stage was determined.

[0020] Figure 2A This is a volcano plot of the association between plasma proteins and CKD incidence; the X-axis is beta value, and the Y-axis is -log 10 (P value). It shows the relationship between 2,737 proteins and the incidence of CKD after Cox multivariate regression. After applying the Bonferroni correction, proteins above the horizontal line were significantly associated with the incidence of CKD (n = 1,169) (P < 0.05).

[0021] Figure 2B This is a diagram of the plasma protein core metabolic network, where each dot represents a protein; the lines connecting the nodes represent different interactions: purple indicates co-expression, indicating that genes or proteins are expressed together; red indicates physical interaction, indicating direct physical contact between molecules; blue indicates co-localization, indicating that molecules are located in the same cellular region; yellow indicates shared protein domains, highlighting molecules with common structural protein components; green indicates gene interactions, indicating relationships influenced by genetic factors;

[0022] Figure 2C is a Venn diagram showing the results compared with the control group, and the overlapping number of significantly dysregulated proteins in different stages of CKD, stage 1, stage 2, stage 3, stage 4, and stage 5;

[0023] Figure 2D Figure 3 is a heat map showing changes in dysregulated proteins at different stages of CKD. Beta values ​​are shown in color, and asterisks indicate statistical significance after Bonferroni correction (***P < 0.001, **P < 0.01, *P < 0.05, unsigned p ≥ 0.05). P values ​​were calculated using a two-sided test, and statistical significance was defined as P < 0.05 after Bonferroni correction.

[0024] Figure 3A The SHAP visualization of the five proteins ultimately selected is shown. The importance of sequenced proteins is demonstrated by filtering the 15 proteins with the highest SHAP values ​​from the candidate proteins selected by Boruta. The importance of protein sequencing can be seen based on their contribution to predicting future CKD (as judged by information gain). The line chart shows the cumulative AUC values ​​as proteins are added one by one (right axis, the five proteins ultimately selected for prediction are marked in red). The width of the horizontal line range can be interpreted as the degree of contribution to CKD prediction; the wider the range, the greater the contribution. The magnitude of the protein contribution is represented by a gradient color, and the X-axis indicates the likelihood of CKD occurring (right) or not occurring (left).

[0025] Figure 3BReceiver operating characteristic curves showing the performance of various models using plasma proteins alone or in combination with other markers for predicting all CKD and CKD within 5 years, 10 years, and more than 10 years; clinical predictors included age, body mass index, diabetes, and hypertension;

[0026] Figure 3C The receiver operating characteristic curves show the performance of various models for predicting different stages of CKD, either alone or in combination with other indicators; clinical predictors include age, body mass index, diabetes, and hypertension;

[0027] Figure 3D This is the external validation result of the CKB validation cohort; the expression distribution of IGFBP4, GDF15, and EDA2R proteins in the CKD group and the non-CKD group; P values ​​were calculated using a two-sided test without multiple comparisons;

[0028] Figure 3E IGFBP4, GDF15, and EDA2R proteins, alone or in combination with other panels, predicted CKD incidence over 5, 10, and all time periods (quantified by AUC and 95% CI);

[0029] Figure 4A Figure 3. Dynamic changes in plasma IGFBP4, GDF15, EDA2R, HAVCR1, and XG before the onset of CKD. Using age, sex, and race as covariates, the nearest neighbor matching method was used to match CKD patients at a ratio of 1:5. The red curve corresponds to CKD patients, and the purple curve corresponds to the control group. Error bars represent standard errors. The Mann-Kendall trend test was used to assess differences in the temporal trends of plasma protein levels between CKD and non-CKD patients. P values ​​were determined by two-sided tests, without the need for correction for multiple comparisons.

[0030] Figure 4B Figure 5. Unadjusted Kaplan-Meier curves showing clinical progression to CKD over time for individuals grouped by quintiles of baseline plasma IGFBP4, GDF15, EDA2R, HAVCR1, and XG levels. The number of people at risk in a 2.5-year interval is shown below each curve. Cox multivariate regression was used to estimate the risk of incident disease, comparing baseline plasma protein levels in group 5 with those in group 1, and to calculate hazard ratios (HRs) and P values. Shaded areas show standard errors based on survival proportions. P values ​​were calculated using two-sided tests with a Bonferroni correction to determine whether the association was significant (P < 0.05).

[0031] Figure 5Gene Ontology (GO), Kyoto Encyclopedia of Genomes (KEGG), and Reactome pathway enrichment analyses were performed using the DAVID website (https: / / david.ncifcrf.gov / ). Significant proteins in Cox multivariate regression across CKD stages were analyzed after Bonferroni correction, using the Olink protein panel as the background gene set (n=6; n=93; n=988; n=955; n=974). Two-sided P values ​​were used for calculation, and statistical significance was determined when the FDR-corrected P value was less than 0.05. DETAILED DESCRIPTION

[0032] Example 1

[0033] In this study, the inventors selected 52,748 patients with undiagnosed chronic kidney disease from the UK Biobank for a longitudinal study to identify predictors of incident chronic kidney disease. After using the Boruta algorithm to identify predictors from 2,737 plasma proteins and 11 clinical demographic indicators, the calculated SHAP values ​​were used to screen for significant proteins. The inventors combined six machine learning methods—LASSO, Elastic Net, LGBM, XGBOOST, MLP, and CNN—with different panels to develop a CKD prediction model and evaluated its performance on a validation set.

[0034] 1. Methods

[0035] 1. Research subjects

[0036] The UK Biobank (UKB) is a prospective cohort study that provides extensive genetic and phenotypic data on 502,369 UK residents recruited between 2006 and 2010. The full UKB protocol is available online (https: / / www.ukbiobank.ac.uk / media / gnkeyh2q / study-rationale.pdf). The UK Biobank (UKB) dataset includes over half a million individuals (aged 37–73 years) recruited from England, Scotland, and Wales between 2006 and 2010. Details of the study design and methods have been published previously. The inventors restricted the UKB sample to participants with OlinkExplore data at baseline (n=53,018) and excluded those with a baseline diagnosis of CKD, self-reported CKD, and those without plasma proteome data, leaving 52,748 participants for the primary analysis. The median age of participants was 58 years, and 46.0% were male. During 16.6 years of follow-up, 3,090 incident chronic kidney disease cases were identified, including 1,914 within 10 years, 656 within 5 years, and 1,176 over 10 years.

[0037] The CKB is a prospective cohort study of 512,724 adults aged 30–79 years who were recruited from 10 different regions (5 rural and 5 urban) in China between 2004 and 2008. Details of the CKB study design and methods have been described previously. 6 Among the 3,977 participants included in the CKB validation cohort, the mean (standard definition) age was 57.3 (11.6) years, 53.7% were women, and 11.2% had a confirmed diagnosis of diabetes.

[0038] 2. Proteomic Analysis

[0039] Proteomic analysis of UKB and CKB was performed using the Olink Explore 3072 platform, which connects four Olink panels (cardiometabolic, inflammation, neurological, and tumor). Olink data preprocessed from both cohorts were provided to any NPX unit on a log2 scale. In UKB, a random subsample of proteomic participants (n=45,441) was selected by removing random subsamples from batches 0 and 7. Proteomic analysis of randomly selected participants from UKB has been shown to be highly representative of the broader UKB population. UKB Olink data are provided as log2-scaled normalized protein expression (NPX) values, and detailed information on sample selection, processing, and quality control is documented online.

[0040] Details of the analytical performance and validation of OLINK have been reported previously. At the CKB, stored baseline plasma samples from participants were extracted, thawed, and divided into multiple aliquots. One aliquot (100 μl) was used to prepare two sets of 96-well plates (40 μl per well). Both sets of plates were shipped on dry ice. One set was sent to Olink Bioscience Laboratories in Uppsala (first batch, 1,463 unique proteins), and the other set was sent to Olink Laboratories in Boston (second batch, 1,460 unique proteins).

[0041] Proteomic analysis was performed using a multivariate proximity expansion assay, with all 3,977 samples analyzed in each batch. Samples were plated in the order they were retrieved from the long-term storage repository at the Wolfson Laboratory in Oxford, normalized using an internal control (expansion control) and an interplate control, and then converted using a predetermined correction factor. Limits of detection were determined using negative control samples (buffer without antigen). If the median value of all samples on the incubation control deviation plate exceeded a predetermined value (±0.3), the sample was flagged as a quality control warning (but values ​​below the LOD were included in the analysis).

[0042] In the analysis, the inventors excluded three proteins (GLB1, NPHS2, and PCDHB15) with missing values ​​greater than 30% in UK blood samples, resulting in a total of 2,734 proteins available for analysis.

[0043] 3. Chronic Kidney Disease Outcomes

[0044] At baseline, the prevalence of CKD was determined from hospital inpatient records using the International Classification of Diseases, the International Statistical Classification of Diseases, Injuries, and Causes of Death (ICD-9), or the International Statistical Classification of Diseases and Related Health Problems, 10th Revision (ICD-10). CKD was identified using International Classification of Diseases, 10th Revision (ICD-10) codes N18, N18.0–N18.9, I12.0–I12.9, I13.1, and I13.2 and ICD-9 codes 585 and 5859 (as previously described).

[0045] 4. Covariates

[0046] Demographic characteristics and health-related information included age, sex, activity, smoking status, alcohol consumption, household income, healthy diet, ethnicity11, education level12, township poverty index (TDI), body mass index (BMI), total cholesterol (TG), high-density lipoprotein (HDL), low-density lipoprotein (LDL),13 region, glucose14, albumin, hemoglobin, creatinine15, hypertension, and diabetes.

[0047] 5. Model Benchmark

[0048] The inventors compared the performance of six different machine learning models (LASSO, random forest, LGBM, XGBOOST, MLP, CNN) for predicting CKD incidence using plasma proteins. All different models were trained using 10-fold cross-validation on a southern UK cohort (n=37,792). The results were tested on an independent validation set of a northern UK cohort (n=14956) and a CKB validation cohort (n=3977). The hyperparameters of the two neural network models were adjusted and optimized using Optuna for 3 to 10 cross-validations in 100 trials to maximize the average R2 of the model across all folds.

[0049] 6. Missing Data Imputation

[0050] Missing values ​​in all non-proteomics UKB data were imputed using the R package miceforest, which combines random forest imputation with predictor mean matching. Three iterations were performed on a single dataset, with 42 random states and all other parameters set to default. The imputation process included all baseline variables in the UKB as predictors. "NA" values ​​were imputed. Incident health outcomes in the UKB were not imputed.

[0051] Protein expression values ​​for the UKB cohort were imputed using the miceforest package in Python. 16 All proteins, except those missing in more than 30% of participants, were used as predictors for each protein’s imputation. Three iterations of imputation were performed for a single dataset with all other parameters set to default.

[0052] 7. Statistical analysis

[0053] The inventors combined age, sex, activity level, race, education, smoking, alcohol consumption, income, healthy diet, Townsend Deprivation Index (TDI), body mass index (BMI), triglycerides (TG), high-density lipoprotein (HDL), low-density lipoprotein (LDL), and regional variables and then used a multivariate Cox proportional hazards model to analyze the differential expression of 2,737 plasma proteins. The positive proteins identified by the multivariate Cox proportional hazards model were used in protein network interaction analysis to visually display the relationships between proteins. To explore the changes and status of plasma proteins during the progression of CKD, the inventors analyzed the intersection and changes of differentially expressed proteins at each stage.

[0054] Using Boruta, 359 significant proteins were screened and fed into a pre-trained LGBM classifier. These proteins were compared with a control group and ranked according to their contribution to CKD prediction. To comprehensively evaluate the predictive performance of the constructed model, multiple machine learning algorithms were employed. The inventors selected the top five proteins by SHAP value and modeled them using six machine learning methods: LASSO, random forest, LGBM, XGBOOST, MLP, and CNN. Ultimately, the machine learning method with the best predictive performance was selected. The incidence of CKD was predicted based on the following panels: Group 1 included creatinine; Group 2 included other serum markers (glucose, albumin, hemoglobin, high-density lipoprotein, low-density lipoprotein); Group 3 included IGFBP4 (the protein with the highest SHAP value); Group 4 included GDF15 (the protein with the second highest SHAP value); Group 5 included EDA2R (the protein with the third highest SHAP value); Group 6 included IGFBP4 protein, clinical predictors (including age, body mass index, diabetes, hypertension), and creatinine; Group 7 included GDF15 protein, clinical predictors, and creatinine; Group 8 included EDA2R protein, clinical predictors, and creatinine; Group 9 included the five proteins with the highest SHAP values; Group 10 included the five proteins with the highest SHAP values, clinical predictors, and creatinine; Group 11 included PRS data. To evaluate the comprehensive ability of proteins, the inventors further utilized a combination of the above selected proteins as an alternative to single protein modeling. Next, we performed a receiver operating characteristic analysis on the Northern England cohort and the CKB validation cohort to assess the predictive accuracy of the above panel. We report performance metrics obtained by the bootstrap strategy, including median AUC values ​​and 95% confidence intervals.

[0055] The inventors plotted time trajectories to observe the dynamic evolution of plasma proteins over the 16 years prior to CKD diagnosis. All new CKD cases identified during follow-up were considered cases, and the inventors matched controls and cases in a 5:1 ratio according to age, sex, and race. The Mann-Kendall trend test was used to investigate whether plasma levels showed a monotonic trend over time. To assess the predictive role of plasma proteins in the risk of clinical progression of CKD, the inventors constructed Kaplan-Meier curves. Plasma proteins were categorized according to the quartiles of the related protein NPX, and participants were divided into five levels: Q1, Q2, Q3, Q4, and Q518. Cox proportional hazards regression analyses were performed, adjusting for factors such as age, sex, race, education level, TD index, income, body mass index, hypertension, healthy diet, smoking, and alcohol consumption, to determine the relationship between individual proteins and CKD events, and cumulative incidence curves were constructed.

[0056] Functional enrichment analysis was performed using the David platform (KEGG, GO, Reactome). Multivariate Cox analysis was performed on the results of different stages (stages 1 to 5), and proteins with P < 0.05 were selected from the multivariate Cox analysis results of different stages (6, 92, 988, 955, and 974 positive proteins in stages 1 to 5, respectively) for functional enrichment analysis.

[0057] II. Achievements

[0058] 1. Differences in plasma proteomes in patients with chronic renal failure compared with controls

[0059] Cox proportional hazards models were used to analyze the associations among age, sex, activity level, race, education level, smoking status, alcohol consumption, household income, healthy diet, Townsend deprivation index (TDI), body mass index (BMI), total cholesterol (TG), high-density lipoprotein (HDL), low-density lipoprotein (LDL), and region. Of the 2,737 plasma proteins, 1,169 were differentially expressed between CKD patients and controls (P < 0.05 after Bonferroni correction; P < 0.05 after Bonferroni correction). Figure 2A ), these proteins were upregulated in CKD patients, among which the top five proteins with SHAP values ​​by Boruta screening and LGBM modeling included insulin-like growth factor binding protein-4 (IGFBP4) (β=1.55, P=0), growth differentiation factor 15 (GDF15) (β=0.97, P=0), EDA receptor 2 recombinant protein (EDA2R) (β=1.74, P=0), hepatitis A virus cellular receptor 1 (HAVCR1) (β=0.68, P=1.03×10-234), and glycoprotein Xg (XG) (β=1.35, P=0).

[0060] 2. Plasma protein core metabolic network

[0061] To further study the molecular mechanism of CKD, the inventors constructed a molecular interaction network using positive proteins. Figure 2B ) visually displays the relationships between these proteins. Each circular node represents a protein (e.g., YAP1, GFRα1, COL18A1). The connections between nodes are color-coded to represent different types of interactions: purple indicates co-expression, red indicates physical interactions, blue indicates co-localization, yellow indicates shared protein domains, and green indicates genetic interactions. For example, YAP1 is connected to surrounding proteins by lines of different colors, indicating that it may occupy an important position in this network.

[0062] 3. Changes in plasma proteins at different stages of chronic kidney disease

[0063] By linking the changes in plasma proteins with the development of chronic kidney disease, the inventors explored the timing and state of these proteins during the development of chronic kidney disease. The inventors focused on the changes in the above differentially regulated proteins during the CKD development stages (stage 1, stage 2, stage 3, stage 4, stage 5). The top 300 proteins were screened from 1,169 differentially regulated proteins in CKD patients and controls, and a graph of the changes in differentially regulated proteins at each disease stage was drawn ( Figure 2C -D).

[0064] Specifically, six proteins were dysregulated in the first stage, and two of them (EDA receptor 2 recombinant protein (EDA2R) and scavenger receptor class F member 2 (SCARF2)) were also dysregulated in individuals in the second stage ( Figure 2C ), reflecting protein changes that occur early in CKD. Another 91 proteins (such as IGFBP4, KRT19, LAG3, and LAYN) were upregulated in stage 2 compared to stage 1, while the majority of proteins (n=988) were not dysregulated until stage 3 (such as GDF15, XG, PRND, and IL1RL1).

[0065] 4. Importance ranking of proteins in CKD prediction

[0066] In order to better translate the inventors' findings into potential clinical applications, the inventors further screened the minimum number of proteins with the strongest predictive power for CKD. The 359 proteins screened by Boruta were put into the optical gradient boosting machine (LGBM) classifier for modeling and SHAP values ​​were calculated, and then the top five proteins were selected based on the SHAP value ranking. The inventors drew a Shapley additive interpretation (SHAP) diagram to visually evaluate the impact of the selected important proteins by the size of their SHAP values ​​(coded with gradient colors) and trend direction (representing the possibility of suffering from CKD) ( Figure 3A ).

[0067] For example, when the plasma protein IGFBP4 predicted the incidence of CKD, it had the widest range on the horizontal axis, indicating that it had the strongest predictive power and a significant impact on the model output. Compared with people with lower IGFBP4 concentrations (blue), people with higher IGFBP4 concentrations (red) (right figure) were more likely to develop CKD (left figure). Similar explanations were found for the other four proteins.

[0068] 5. Plasma proteins can help accurately predict the incidence of CKD

[0069] The inventors tested the performance of the above selected important proteins in predicting the incidence of CKD. The inventors used the same panel and different level models to compare their performance in predicting the incidence of CKD and finally selected the LGBM with the best overall performance in predicting CKD. The top three plasma proteins with the highest SHAP values ​​had good predictive ability for the incidence of CKD: IGFBP4 (AUC = 0.816), GDF15 (AUC = 0.810) and EDA2R (AUC = 0.824) ( Figure 3B The AUC for creatinine, commonly used to predict CKD incidence, is 0.748. The AUC for several commonly used serum markers (glucose, albumin, hemoglobin, high-density lipoprotein, and low-density lipoprotein) is 0.676, while the AUC for clinical predictive markers is 0.839. To improve predictive power, the inventors combined plasma proteins, clinical markers (age, body mass index, diabetes, hypertension), and creatinine for prediction, resulting in significantly improved accuracy. These included IGFBP4 (AUC = 0.897), GDF15 (AUC = 0.891), and EDA2R (AUC = 0.895).

[0070] The inventors also developed a combined model for the first five proteins to test their predictive power. In predicting the incidence of CKD, the inventors' five proteins, individually and in combination, outperformed previous studies, with the five-protein combination achieving an AUC of 0.858. Further addition of clinical predictors and creatinine to the proteome significantly improved the predictive power of CKD (AUC = 0.907).

[0071] When the same variable model was used to predict the incidence of CKD within 5 years, 10 years, and 10 years or more, IGFBP4 had an AUC of 0.906, 0.857, and 0.730, respectively, demonstrating good performance. The AUCs for the combination of IGFBP4 with clinical predictors and creatinine, or the combination of the five proteins with clinical predictors and creatinine, were superior to those for IGFBP4 alone, demonstrating excellent predictive performance, particularly for predicting the incidence of CKD within 5 years. The AUCs for the two combined panels were 0.936 and 0.942, respectively.

[0072] The inventors also tested the performance of the five protein combinations alone and in combination with clinical predictors and creatinine indicators in predicting CKD incidence in subgroups. The combination of the five proteins, clinical predictors, and creatinine indicators performed well in predicting the 5-year CKD incidence in men (AUC = 0.954). This combination also performed well in predicting the 5-year CKD incidence in people under 60 years old (AUC = 0.951).

[0073] 6. Plasma proteins can help accurately predict the incidence of CKD at different stages

[0074] The inventors also continued to explore the performance of the above selected important proteins in predicting CKD stages. Figure 3C As shown in the data, the top three plasma proteins in terms of SHAP values ​​had good predictive ability in predicting the incidence of CKD stage 4: IGFBP4 (AUC = 0.905), GDF15 (AUC = 0.914), and EDA2R (AUC = 0.903).

[0075] The same five-protein combination model was also applied to predict the onset stage of CKD. The combination had an AUC of 0.936 for predicting stage 4, improving the predictive ability of CKD. The inventors also tested the first three proteins individually and in combination with clinical predictors (age, body mass index, diabetes, hypertension) and creatinine. The results were IGFBP4 (AUC = 0.944), GDF15 (AUC = 0.941), and EDA2R (AUC = 0.945), respectively, and the five-protein combination (AUC = 0.950).

[0076] 7. External validation of models trained in the CKB validation queue

[0077] In the CKB validation cohort, IGFBP4 (β=1.55, P=0), GDF15 (β=0.97, P=0), and EDA2R (β=1.74, P=0) protein levels were upregulated in CKD patients ( Figure 3D The trained models had similar efficacy in predicting the incidence of CKD 5, 10, and 20 years later.

[0078] The SHAP value of plasma protein IGFBP4 ranked highest in predicting CKD incidence, followed by GDF15 ( Figure 3A When the first three proteins were included, the prediction AUC (area under the curve on the right axis) increased dramatically, and as the number of proteins continued to increase, the AUC for predicting CKD incidence gradually stabilized. The first five proteins (IGFBP4, GDF15, EDA2R, HAVCR1, and XG) were identified as important proteins for predicting CKD incidence.

[0079] 8. External validation of models trained in the CKB validation queue

[0080] In the CKB validation cohort, the protein levels of IGFBP4 (β=1.55, P=0), GDF15 (β=0.97, P=0), and EDA2R (β=1.74, P=0) were upregulated in CKD patients ( Figure 3D In the CKB validation cohort, the trained model had similar efficacy in predicting CKD incidence at 5 years, 10 years, and all time ( Figure 3E ), among which the performance in predicting the incidence of CKD within 5 years was as follows: IGFBP4 (AUC=0.778), GDF15 (AUC=0.726), EDA2R (AUC=0.728), the combination of five proteins (AUC=0.781), and the combination of five proteins combined with clinical predictors (AUC=0.750).

[0081] 9. Preclinical trajectory of plasma proteins

[0082] The inventors used a 16-year time scale to plot the time trajectory of plasma proteins from the time of CKD diagnosis, and compared it with the protein changes of patients without CKD during the same period. The curve showed that as early as 16 years before the onset of CKD, plasma proteins IGFBP4, GDF15 and EDA2R had deviated from the normal range ( Figure 4A IGFBP4 levels increased more rapidly over time in patients with CKD compared with those without CKD, whereas the slope for GDF15 was not significantly different.

[0083] 10. The cumulative risk of CKD at different serum levels

[0084] The inventors plotted the progression to CKD over time according to quintiles of baseline levels of five important proteins ( Figure 4B The higher the baseline levels of these five proteins, the greater the risk of clinical progression.

[0085] 11. Bioaccumulation analysis of different proteins at different stages of CKD

[0086] The inventors investigated the enrichment pathways reflected by different CKD regulatory proteins at different stages of clinical CKD progression between stage 1 and stage 5. Cell adhesion, signal transduction, and other pathways were enriched for proteins that were dysregulated at different disease stages ( Figure 5 ).

[0087] Discussion

[0088] In this large prospective cohort study, the inventors demonstrated that plasma proteomic analysis has a strong predictive value for chronic kidney disease (CKD), up to 16 years before clinical diagnosis. The inventors' findings highlight the potential of proteomic biomarkers to improve early CKD detection, risk stratification, and clinical decision-making. Because chronic renal failure is asymptomatic in its early stages and current diagnostic markers (such as estimated glomerular filtration rate (eGFR) and albuminuria) have limited sensitivity, early identification of chronic renal failure remains a major clinical challenge. If chronic renal failure can be predicted before clinical symptoms appear, therapeutic intervention can be performed earlier, delaying the progression of the disease, thereby improving patient prognosis and reducing the burden on the healthcare system.

[0089] After adjusting for demographic and clinical covariates, the inventors identified 1,169 differentially expressed plasma proteins between CKD cases and controls. Among them, IGFBP4, GDF15, EDA2R, HAVCR1, and XG emerged as the most significant predictors based on Boruta screening and SHAP analysis. IGFBP4 was consistently the most significant predictor, highlighting its potential role as a key biomarker for CKD risk. GDF15 and EDA2R also demonstrated strong predictive power, strengthening their relevance in CKD pathophysiology. The identification of these proteins suggests that specific biological processes become dysregulated early in the disease, and that changes in the proteome may reflect the underlying molecular mechanisms driving CKD progression.

[0090] IGFBP4 and GDF15 are known to be involved in cellular stress response, inflammation, and metabolic regulation, all of which contribute to the pathogenesis of CKD. IGFBP4 is a member of the insulin-like growth factor binding protein family that regulates insulin-like growth factor signaling and is associated with metabolic dysfunction and renal fibrosis. GDF15 is a stress-responsive cytokine associated with inflammation and oxidative stress, both of which are core factors in the progression of chronic kidney disease. EDA2R is a receptor involved in the TNF superfamily signaling pathway and may contribute to inflammation and tissue remodeling in the kidney. HAVCR1, also known as kidney injury molecule-1 (KIM-1), is a recognized marker of tubular injury and can reflect early kidney damage. XG is a glycoprotein expressed on red blood cells and is a novel candidate biomarker for CKD that deserves further study. The close association of these proteins with CKD risk supports their potential role in early diagnosis and targeted therapy.

[0091] A prediction model combining these five proteins with clinical predictors and creatinine showed superior accuracy compared to traditional CKD biomarkers. The model had higher discrimination power than creatinine alone and other commonly used plasma indices such as glucose, albumin, hemoglobin, high-density lipoprotein, and low-density lipoprotein. The integration of proteomic data with clinical variables improved predictive performance, suggesting that plasma proteins capture unique biological signals that complement existing clinical markers. This highlights the value of a multi-layered approach in predicting the risk of chronic renal failure, in which proteomic biomarkers improve the sensitivity and specificity of traditional diagnostic models.

[0092] The model's performance remained stable across different CKD stages, with high predictive accuracy even in the earliest stages of the disease. IGFBP4, GDF15, and EDA2R showed particularly strong predictive power for CKD stage 4, highlighting their relevance to late-stage disease progression. Furthermore, the model performed well in predicting CKD progression over various time intervals, with the combination of IGFBP4 and other key proteins showing excellent predictive accuracy for CKD over 5, 10, and 16 years. These findings suggest that proteomic biomarkers can not only predict the development of CKD but also provide valuable prognostic information about the disease trajectory.

[0093] Pathway enrichment analysis provides insights into the biological processes underlying the development and progression of CKD. Dysregulated proteins were enriched in pathways related to cell adhesion and signal transduction, suggesting that structural and functional changes in renal tissue may occur early in the disease. Changes in cell adhesion proteins may reflect alterations in the integrity of the glomerular filtration barrier and renal tubular epithelial cells, leading to albuminuria and renal dysfunction. Dysregulation of signal transduction pathways, particularly those involved in inflammatory and fibrotic responses, highlights the role of chronic low-grade inflammation and tissue remodeling in the progression of chronic kidney disease. These findings are consistent with previous evidence that inflammation, oxidative stress, and metabolic dysfunction are involved in the pathogenesis of chronic renal failure. The identification of these biological signatures strengthens the mechanistic basis for the association between the observed proteomic changes and CKD risk.

[0094] Strengths of this study include the large sample size, long follow-up, and comprehensive proteomic data, which enhance the statistical power and generalizability of our findings. The use of Boruta screening and SHAP analysis allowed for precise identification of key predictors and insight into their relative contributions to CKD risk. The prospective design minimized the risk of reverse causality and allowed for the assessment of the temporal association between proteomic changes and the development of CKD.

[0095] However, several limitations should be acknowledged. First, although the prediction model performed well in the study population, external validation in independent cohorts is still necessary to confirm its generalizability. Second, although the study found a strong association between plasma proteins and the risk of chronic kidney disease, the causal nature of these relationships remains uncertain and requires further investigation through mechanistic and intervention studies. Third, the proteomics platform used in this study may not be able to capture all relevant proteins, and some potential biomarkers may have been missed due to technical limitations. Finally, despite adjustment for various covariates, residual confounding due to unmeasured factors cannot be completely ruled out.

[0096] These findings have important clinical implications for CKD screening, early diagnosis, and personalized risk stratification. The identification of key proteomic biomarkers lays the foundation for the development of targeted diagnostic assays that can be incorporated into routine clinical practice. Early identification of high-risk individuals can allow for timely lifestyle adjustments and pharmacological interventions to delay the progression of chronic renal failure. Future studies should focus on validating these findings in diverse populations, exploring the mechanistic pathways linking proteomic changes to CKD pathophysiology, and evaluating the cost-effectiveness of proteomic-based screening programs.

[0097] In conclusion, this study provides strong evidence that plasma proteome analysis can accurately predict CKD years before clinical diagnosis and improve risk stratification. The identification of key predictive proteins highlights novel biological pathways involved in the development of CKD and suggests potential targets for early intervention. These findings support the incorporation of proteomic biomarkers into clinical practice to enhance early CKD detection and improve patient outcomes.

[0098] Plasma proteome analysis demonstrated strong predictive power for chronic kidney disease (CKD) years before clinical diagnosis. The proteins IGFBP4, GDF15, EDA2R, HAVCR1, and XG were identified as the most important predictors, with IGFBP4 showing the highest predictive value. A model combining these proteins with clinical predictors and creatinine outperformed traditional biomarkers and accurately predicted CKD progression. Pathway analysis highlighted early molecular changes associated with cell adhesion and signaling. These findings support the potential of plasma proteome biomarkers for early CKD detection and risk stratification.

[0099] Although the present invention is disclosed above with reference to preferred embodiments, this is not intended to limit the scope of the present invention. Any person skilled in the art may make slight modifications without departing from the scope of the present invention. In other words, any equivalent modifications made in accordance with the present invention should be included within the scope of the present invention.

Claims

1. CKD early screening markers based on plasma proteomics, characterized by: The early screening markers include the following five proteins: IGFBP4, GDF15, EDA2R, HAVCR1, and XG.

2. The CKD early screening marker according to claim 1, characterized in that The method for selecting early screening markers comprises the following steps: A. Using data from a prospective biobank cohort study, we analyzed a large cohort of adults without CKD at baseline and measured several plasma proteins. B. Proteins associated with CKD were determined by Cox regression analysis; C, Machine learning using 10-fold cross-validation to quantify the importance of the proteins identified in step B.

3. An application of a CKD early screening marker based on plasma proteomics as claimed in claim 1, characterized in that: The early screening marker is used in the preparation of a CKD early screening kit.