Protein combinatorial analysis, prediction models, and applications for predicting the incidence of chronic liver disease based on proteomics.

By constructing a predictive model for chronic liver disease based on a combination of proteins such as VCAM1, and combining machine learning and systems biology, the problem of insufficient prediction by traditional models in Chinese adults has been solved, enabling accurate assessment and early screening of chronic liver disease.

CN121331222BActive Publication Date: 2026-03-10PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing traditional prediction models have poor predictive performance for chronic liver disease in Chinese adults. There is a lack of proteomics-based prediction methods to improve predictive performance, and domestic research has insufficient evidence in this area.

Method used

Using a combination of proteins including VCAM1, CD48, NUDT5, MILR1, TNFSF8, DDAH1, IL17RB, NFASC, FCN2, and RTN4R, and combining machine learning algorithms and systems biology, a predictive model for the incidence of chronic liver disease was constructed. Protein concentrations were quantitatively measured using the Olink platform, and risk assessment was conducted by incorporating factors such as age, sex, urban/rural location, alcohol consumption, and HBV infection status.

Benefits of technology

It significantly improves the predictive performance of chronic liver diseases, enables accurate assessment of the risk of disease in the next 10 years, and provides an innovative solution for early screening and dynamic risk monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121331222B_ABST
    Figure CN121331222B_ABST
Patent Text Reader

Abstract

This invention relates to the field of biomedical technology, and particularly to protein assemblies, predictive models, and applications for predicting the incidence of chronic liver diseases based on proteomics. This invention integrates Olink proteomics technology, machine learning algorithms, and a discovery-internal validation research strategy to successfully screen and validate a set of predictive models for chronic liver diseases containing 10 plasma protein biomarkers (VCAM1, CD48, NUDT5, MILR1, TNFSF8, DDAH1, IL17RB, NFASC, FCN2, and RTN4R). Through the cross-integration of systems biology and artificial intelligence technologies, this invention develops a new approach to full-cycle prediction from early disease screening to dynamic risk monitoring, providing an innovative solution for the precise prevention and control of chronic liver diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical technology, and in particular to protein combinations, prediction models and applications for predicting the incidence of chronic liver diseases based on proteomics. Background Technology

[0002] Chronic liver disease causes a severe disease burden globally. Developing early risk prediction models for chronic liver disease, conducting prospective risk stratification assessments, implementing precise screening of high-risk populations, and combining these with preventative interventions have become one of the most cost-effective public health control strategies for reducing the disease burden.

[0003] Chronic liver disease is essentially a complex chronic disease caused by multiple factors. Proteins, as terminal effector molecules in gene transcription, possess dynamic regulatory characteristics in response to environmental exposures, making them important molecular markers for characterizing the body's homeostatic regulatory efficiency and constructing a risk stratification assessment system for chronic liver disease. Previous published studies in Europe and the United States have reported that plasma proteomics can be used for risk prediction of hepatocellular carcinoma, including four key protein molecules: CHI3L1, GDF5, SELE, and IL1RN. Some studies based on fatty liver patients have also reported that plasma proteomics can be used to identify high-risk fatty liver patients, such as those with steatohepatitis and liver fibrosis. Recent articles report that plasma proteomics can be used to differentiate different subtypes of fatty liver, including metabolic dysfunction-associated steatosis (MASLD) and metabolic and alcoholic steatosis (MetALD).

[0004] However, current proteomic studies related to chronic liver disease are mostly based on general population cohorts (such as the UK Biobank) and disease-specific cohorts (such as the fatty liver patient cohort) in Europe and the United States, while evidence from cohort studies based on the general population in my country remains lacking. Furthermore, previous studies have lacked evaluation of the role of proteomics in enhancing the long-term risk of chronic liver disease beyond traditional predictive models. Existing traditional predictive models show poor predictive performance for chronic liver disease in Chinese adults, with a C-index of only 0.5–0.7. As a low-cost, high-throughput, and highly sensitive and specific emerging omics detection technology, proteomics holds promise for improving predictive performance beyond traditional models, and its early predictive role in the onset of chronic liver disease has significant clinical and public health implications.

[0005] Therefore, designing proteomics-based methods and systems for predicting the incidence of chronic liver diseases, and establishing plasma protein prediction models for chronic liver diseases with good predictive efficacy through reasonable data analysis, in order to achieve precise prevention and control of chronic liver diseases, is an urgent problem to be solved. Summary of the Invention

[0006] In view of this, the technical problem to be solved by the present invention is to provide protein combinations, prediction models and applications for predicting the incidence of chronic liver diseases based on proteomics.

[0007] This invention provides a protein combination for predicting the risk of developing chronic liver disease, comprising: VCAM1, CD48, NUDT5, MILR1, TNFSF8, DDAH1, IL17RB, NFASC, FCN2, and RTN4R.

[0008] In the protein combination described in this invention, the UniProt accession numbers of VCAM1, CD48, NUDT5, MILR1, TNFSF8, DDAH1, IL17RB, NFASC, FCN2, and RTN4R are P19320, P09326, Q9UKK9, Q7Z6M3, P32971, O94760, Q9NRM6, O94856, Q15485, and Q9BZR6, respectively.

[0009] This invention provides the application of the aforementioned protein combination in constructing a predictive model for the risk of chronic liver disease.

[0010] This invention provides a predictive model for chronic liver diseases, the formula of which is:

[0011] Model score = PS × 0.900954 + age factor value + gender factor value + urban / rural factor value + alcohol consumption factor value + HBV infection status factor value - 0.588822;

[0012] The age factor value is: subject's age × 0.011503;

[0013] Of the gender factor values, the gender factor value for males is 0, and the gender factor value for females is 0.067556.

[0014] Among the urban-rural factor values, the urban-rural factor value for rural residents was 0, and the urban-rural factor value for urban residents was -1.4202.

[0015] Of the alcohol consumption factor values, the alcohol consumption factor value for non-weekly drinkers was 0; the alcohol consumption factor value for weekly drinkers was 1.115251.

[0016] Among the HBV infection factor values, the HBV infection factor value of subjects who tested negative for HBV in blood was 0, and the HBV infection factor value of subjects who tested positive for HBV in blood was 1.578727.

[0017] The PS = -NPX1×0.159053+NPX2×0.592269+NPX3×1.007608+NPX4×0.231735+NPX5×0.557392-NPX6×0.134576+NPX7×0.842951+NPX8×0.478478-NPX9×1.143148+NPX10×0.891170-1.780483;

[0018] Wherein, NPX1 to NPX10 are the NPX values ​​of VCAM1, CD48, NUDT5, MILR1, TNFSF8, DDAH1, IL17RB, NFASC, FCN2 and RTN4R in the protein combination described in this invention, respectively.

[0019] The NPX value is a relative quantitative unit that is logarithmically related to protein concentration. In this invention, the expression levels of 10 target proteins are quantitatively determined using the Olink platform's Proximity Extension Analysis (PEA) technology to obtain the NPX value.

[0020] The predictive model described in this invention assesses the risk of developing chronic liver disease based on a cutoff value, and the specific assessment criteria are as follows:

[0021] The model score is higher than the cutoff value of 10% for the risk of developing the disease in the next 10 years, and is therefore judged to be at high risk of developing chronic liver disease.

[0022] The model score was below the cutoff value of 10% for the risk of developing the disease in the next 10 years, and was therefore judged to be at low risk of developing chronic liver disease.

[0023] The cutoff value for the 10% risk of developing the disease in the next 10 years is 3.0037.

[0024] The present invention provides a detection reagent that uses the protein combination described in the present invention as the detection target.

[0025] In the specific technical solution of the present invention, the above-mentioned detection reagent can be a specific antibody against VCAM1, CD48, NUDT5, MILR1, TNFSF8, DDAH1, IL17RB, NFASC, FCN2 and RTN4R.

[0026] This invention provides a detection product comprising the detection reagents described herein, as well as acceptable auxiliaries, carriers, or devices.

[0027] This invention provides the application of the protein combination, the prediction model, the detection reagent, and / or the detection product described herein in the preparation of kits for the prevention, treatment, and / or prediction of chronic liver diseases.

[0028] This invention provides the application of the protein combination, the prediction model, the detection reagent, and / or the detection product described herein in constructing a risk prediction system for chronic liver disease.

[0029] This invention provides a chronic liver disease incidence prediction system, which includes: a data collection unit, a risk scoring unit, and a prediction unit;

[0030] The data collection unit acquires information such as the NPX value of the protein in the protein combination described in this invention, the subject's age, gender, urban / rural status, alcohol consumption status, and HBV infection status.

[0031] The risk scoring unit calculates a model score based on the NPX value, the subject's age, gender, urban / rural status, alcohol consumption status, and HBV infection status using the prediction model described in this invention.

[0032] The prediction unit: determines the risk of developing chronic liver disease based on a cutoff value of 10% for the risk of developing the disease in the next 10 years;

[0033] The cutoff value for the 10% risk of developing the disease in the next 10 years is 3.0037.

[0034] This invention provides the application of the aforementioned protein combination, the aforementioned prediction model, the aforementioned detection reagent, the aforementioned detection product, and / or the aforementioned chronic liver disease incidence prediction system in the preparation of products that predict the future incidence of chronic liver disease.

[0035] This invention integrates Olink proteomics technology, machine learning algorithms, and a discovery-internal validation research strategy to successfully screen and validate a set of predictive models for chronic liver diseases containing 10 plasma protein biomarkers (VCAM1, CD48, NUDT5, MILR1, TNFSF8, DDAH1, IL17RB, NFASC, FCN2, and RTN4R). Through the cross-integration of systems biology and artificial intelligence technologies, this invention develops a novel approach to full-cycle prediction from early disease screening to dynamic risk monitoring, providing an innovative solution for the precise prevention and control of chronic liver diseases. Attached Figure Description

[0036] Figure 1 This invention illustrates the steps of a proteomics-based method for predicting the incidence of chronic liver disease.

[0037] Figure 2 In the illustrated example, the Cox proportional hazards regression model was used to preliminarily screen plasma protein association maps related to chronic liver disease;

[0038] Figure 3The diagram illustrates the relationship between the importance of plasma protein variables (VIMP) and the cumulative C-index in the training set of the Random Survival Forest (RSF) model.

[0039] Figure 4 The 10-year ROC curves of the prediction model in the example are shown in the training and validation sets.

[0040] Figure 5 The example shows the area forest plot under the ROC curve of the prediction model in the training and validation sets.

[0041] Figure 6 The survival curves of the risk population are shown in the example, with a cutoff value of 10% for the incidence of chronic liver disease over 10 years. Detailed Implementation

[0042] This invention provides protein combinations, prediction models, and applications for predicting the incidence of chronic liver diseases based on proteomics. Those skilled in the art can refer to the content of this document and appropriately modify the process parameters to achieve the desired results. It is particularly important to note that all similar substitutions and modifications are obvious to those skilled in the art and are considered to be included in this invention. The methods and applications of this invention have been described through preferred embodiments. Those skilled in the art can clearly modify or appropriately change and combine the methods and applications described herein without departing from the content, spirit, and scope of this invention to implement and apply the technology of this invention.

[0043] The test materials used in this invention are all common commercially available products. The invention is further illustrated below with reference to embodiments:

[0044] Example 1: Construction and Validation of a 10-Year Risk Prediction Model for Chronic Liver Disease Based on Plasma Proteomics

[0045] This embodiment consists of seven parts: (1) selection of research subjects and data collection; (2) risk factor assessment and evaluation of common predictive factors; (3) detection of plasma proteins using the Olink Explore 3072 platform; (4) preliminary screening of traditional predictive factors and plasma proteins; (5) establishment of an RSF model to screen key plasma proteins and construct a PS; (6) predictive performance evaluation and validation; (7) relevant parameters of the protein model for chronic liver disease; and (8) a device for detecting protein levels of predictive factors for chronic liver disease. The process of this embodiment is as follows: Figure 1 As shown ( Figure 1 middle The formula for calculating PS is as follows: PS = -NPX1×0.159053 + NPX2×0.592269 + NPX3×1.007608 + NPX4×0.231735 + NPX5×0.557392 - NPX6×0.134576 + NPX7×0.842951 + NPX8×0.478478 - NPX9×1.143148 + NPX10×0.891170 – 1.780483.

[0046] I. Selection of Research Subjects and Data Collection

[0047] Study Participant Selection: The study population consisted of participants with proteomic data from the China Chronic Disease Prospective Study (CKB). From 2004 to 2008, the CKB project recruited 512,714 adults in five cities and five rural areas across China. All participants completed questionnaires, physical examinations, and blood sample collection. Over 100,000 participants were selected for whole-genome genotyping. From all participants with genotype data, 2,026 were randomly selected for plasma proteomic analysis. Inclusion criteria were as follows: ① genotype data available and no kinship among participants; ② baseline self-reported absence of cardiovascular disease and no use of statins. After further exclusion of participants with missing key variables, 1,970 participants were ultimately included (see [link to relevant documentation]). Figure 1 The baseline characteristics are shown in Table 1.

[0048] Table 1. Baseline characteristics of all study subjects

[0049]

[0050] The CKB project baseline data collection process is as follows: The study strictly follows a standardized protocol, with a systematically trained survey team collecting multidimensional data using electronic questionnaires. The research content includes: ① demographic characteristics and socioeconomic status; ② lifestyle characteristics and behavioral patterns; ③ history of chronic diseases and medication records. Standardized anthropometric measurements are performed simultaneously to obtain anthropometric indicators such as height and weight. In the biosample collection phase, a standardized venous blood collection procedure is used. After immediate sample separation and processing, 10 μL is used for random on-site blood glucose testing, and the remaining samples are cryopreserved (-80℃) to provide a biological resource basis for subsequent biochemical indicator and multi-omics data testing. All procedures adhere to unified quality control standards.

[0051] Disease Outcome Follow-up: From the completion of the baseline survey, the CKB project conducted long-term follow-up on all study participants, comprehensively collecting information on mortality events, major chronic disease events (including chronic liver disease), hospitalization events, and loss to follow-up due to migration. Data collection channels included mortality surveillance systems, population registration systems, routine disease surveillance systems, national health insurance databases, and proactive targeted surveillance by project staff in each project area. All morbidity or cause of death was coded according to the International Classification of Diseases, 10th Revision (ICD-10). As of December 31, 2022, the loss to follow-up rate of the CKB project was less than 0.5%. In this invention, chronic liver disease is defined as the occurrence of any of the following diseases: ① liver cancer; ② portal hypertension-related events, such as esophageal variceal bleeding, ascites, etc.; ③ liver dysfunction-related events, such as liver failure or acute exacerbation of chronic liver failure, significant hepatic encephalopathy, and other serious conditions. The ICD-10 codes corresponding to the above chronic liver diseases include C22, C22.0, C22.7, C22.9, K70.1, K70.2, K70.3, K70.4, K70.9, K71.7, K71.9, K72.0, K72.1, K72.9, K73.0, K73.2, K73.8, K73.9, K74.0, K74.1, K74.2, K74.6, K75.8, K75.9, K76.6, K76.7, K76.8, K76.9, I85.0, I85.9, I98.2, and I98.3. All studies were approved by the Biomedical Ethics Committee of Peking University. All participants signed written informed consent forms prior to data collection.

[0052] II. Risk Factor Assessment and Evaluation of Common Predictive Factors

[0053] This invention comprehensively considers currently recognized risk factors for chronic liver disease, previous literature reports, and accessibility in the CKB project, and ultimately includes 10 traditional risk factors for constructing a predictive model, including gender, age, urban / rural location, education level, smoking status, alcohol consumption status, physical activity level, body mass index (BMI), waist circumference, and hepatitis B virus (HBV) infection status.

[0054] The coding methods for the aforementioned risk factors are shown in Table 2.

[0055] Table 2. Categories and codes of risk factors

[0056]

[0057] Genetic factors are one of the important factors in the development of chronic liver disease. To evaluate the ability of individual genes to improve the prediction of chronic liver disease by traditional factors and to compare it with the subsequent protein prediction ability, this invention uses four single nucleotide polymorphism (SNP) sites that have been validated in the CKB population (Xue H. Gastroenterology. 2025) to construct a weighted gene risk score (GRS), as shown in Formula I.

[0058] Formula I;

[0059] in, It is 4. For the first The weights corresponding to each SNP For the first No. 1 research subject The number of effect alleles of a SNP, calculated in more detail using the same formula as Formula I, is: GRS = Dosage1 × 0.0647 + Dosage2 × 0.2738 + Dosage3 × 0.0629 + Dosage4 × 0.2657, where Dosage1, Dosage2, Dosage3, and Dosage4 are the number of alleles for rs1260326_T, rs58542926_T, rs641738_T, and rs738409_G, respectively; see Table 3 for specific SNP information.

[0060] Table 3. Locus information used in the gene risk score for chronic liver disease

[0061]

[0062] III. Detection of Plasma Proteins using the Olink Explore 3072 Platform

[0063] Baseline levels of 2944 plasma proteins from the study subjects were quantified using the Olink Explore 3072 high-throughput proteomics platform. All samples were stored at -80°C, slowly thawed according to a prescribed procedure before testing, and aliquoted into 96-well plates, with 8 wells reserved per plate for quality control. To minimize batch effects, each plate contained both case samples and sub-cohort samples during aliquoting, arranged chronologically from their return from the Oxford Wolfson laboratory. After aliquoting, samples were transported via dry ice cold chain to the Olink Central Laboratory in Uppsala, Sweden, for testing. Protein concentration data were normalized using intra- and inter-plate controls and converted uniformly using a pre-defined correction factor. Preprocessed plasma protein levels are expressed as dimensionless normalized protein expression (NPX).

[0064] IV. Preliminary screening of traditional predictive factors and plasma proteins

[0065] First, this invention screens the traditional risk factors for the aforementioned chronic liver diseases to determine the predictive factors used to construct a traditional predictive model. A multivariate Cox proportional hazards regression model is used, with the incidence of chronic liver disease as the dependent variable, including gender, age, urban / rural residence, education level, smoking status, alcohol consumption status, physical activity level, BMI, waist circumference, and HBV infection status. Except for gender, age, and urban / rural residence, which are retained as variables in the basic model, the remaining variables are selected based on a significance threshold of P < 0.05 to include variables related to chronic liver disease in the basic model. The results of the multivariate Cox regression analysis are shown in Table 4. Based on these results, the selected variables for the basic model include gender, age, urban / rural residence, alcohol consumption status, and HBV infection status.

[0066] Table 4. Results of multivariate Cox regression analysis of traditional risk factors and the incidence of chronic liver disease.

[0067]

[0068] Subsequently, a preliminary screening of plasma proteins was conducted. A Cox proportional hazards regression model was used, with standardized plasma protein NPX values ​​as independent variables and the incidence of chronic liver disease as dependent variables. After adjusting for traditional factors in the above analysis, 140 proteins associated with chronic liver disease were screened using a false discovery rate (FDR) < 0.05 (see [link to relevant analysis]). Figure 2 ).

[0069] V. Establishing an RSF model for protein screening

[0070] Of the entire population, 1970 participants were randomly split into training and validation sets in a 7:3 ratio (see [link to relevant documentation]). Figure 1 An RSF model was constructed using 1379 subjects in the training set.

[0071] The RSF algorithm is an extension of the Random Forest algorithm for right-censored survival data, achieving strong discriminative power while maintaining low generalization error. The RSF algorithm is implemented through the following steps:

[0072] 1. Resampling: Randomly sampling from the original dataset. Each set of samples retains approximately 63% of the original dataset, while the 37% of samples removed are referred to as out-of-bag (OOB) data.

[0073] 2. Construct a single survival tree: Based on the bootstrap sample set, randomly select... Each predictor variable is used to split the sample from the root node according to the splitting criterion, maximizing the survival difference between child nodes, until the terminal node contains no less than a predetermined number of ( The ending event is >0, and the survival tree grows completely;

[0074] 3. Calculation of Cumulative Risk (CHF) and Incidence Probability: For each terminal node Its CHF can be estimated using the Nelson-Aalen formula:

[0075] , Formula II;

[0076] in, For terminal nodes No. A non-repeating survival time, and Terminal nodes Number of cases and number of people in at-risk populations. CHF is consistent across all sample points at the same terminal node. For each sample point... When the independent variable At that time, its CHF is equal to that of the terminal node. CHF:

[0077] , Formula III;

[0078] Will The average CHF of the entire forest is obtained by averaging the CHF of each surviving tree. :

[0079] , formula IV;

[0080] in, For the number of trees, Then it is the first Trees at sample points CHF estimates.

[0081] The CHF calculation for OOB data is shown in Formula V:

[0082] , formula V;

[0083] in Defined as if individual In OOB data, the value is 1, otherwise the value is 0.

[0084] Due to the terminal node The total CHF It equals the number of individuals in the final event within that node; therefore, for each survival tree, it is the sum of the CHF of all its individuals in terms of observed survival time. The number of participants in the final event of the entire tree is equal to the number of participants in the final event, from which the number of individuals can be calculated. The probability of developing the disease:

[0085] , Formula VI;

[0086] in, To and Individuals at the same terminal node The observation survival time This is the expected value of CHF. Similarly, the individual OOB data can be obtained. Incidence rate .

[0087] 4. Estimate forecast error: Use CHF to represent forecast risk and calculate the C-index for OOB data:

[0088] For each sample pair The observation times were respectively Survival status If a pair is defined as having a survival status of 1 for the pair with shorter observation time, or having a survival status of 1 for the pair with equal observation time, then the total number of valid pairs is as follows: This can be represented as formula VII:

[0089] , formula VII;

[0090] in It is an indicator function; it is denoted as 1 if the condition is met, and 0 otherwise.

[0091] A pair of samples that satisfies the following conditions—consistency between observation duration and predicted risk, equal observation duration and survival status 1, and equal predicted risk, or equal observation duration and survival status 1 with the higher predicted risk—is denoted as a consistent pair. The total number of such pairs is [not specified]. This can be expressed using formula VIII:

[0092] , formula VIII;

[0093] in, , ;

[0094] Pairs that meet the following criteria are unequal in observation time but equal in predicted risk, equal in observation time and both with survival status 1 but different in predicted risk, or equal in observation time and both with survival status 1 having a predicted risk that is not higher, are denoted as knotted pairs. The total number of such pairs is... This can be expressed using formula IX:

[0095] , formula IX;

[0096] The C index can be accessed through... Obtained through calculation.

[0097] The CHF of OOB data can be calculated using the RSF model to obtain the model prediction error. This allows for the evaluation of the model's fit and predictive ability.

[0098] , formula X;

[0099] Plasma protein biomarkers, after initial screening, are further screened using Variable Importance (VIMP). The principle is as follows: during the construction of the RSF model, for the independent variable x to be tested, a node is randomly split into two child nodes, obtaining the CHF value of the entire forest after randomized variable assignment. This independent variable... The VIMP is the difference in prediction error between the original RSF model and the variable randomization assignment model, see Formula XI.

[0100] , formula XI;

[0101] in, This represents the prediction error of the original model. as independent variable Prediction error after randomization.

[0102] This embodiment employs a splitting criterion based on the C-index to determine that each terminal node has at least 10 outcome events, constructing an RSF model with 1000 survival trees to calculate the VIMP of 140 chronic liver disease-related proteins. The top 20% of plasma proteins are added as independent variables in descending order of VIMP, and the C-index is calculated using a Cox proportional hazards model with 5-fold cross-validation. Based on the principles of maximizing the cumulative C-index and simplifying model parameters, 10 key plasma proteins are ultimately selected for inclusion in the Cox proportional hazards-based prediction model. Figure 3 For details, please refer to Table 5.

[0103] Table 5. Proteins screened by the RSF model and their coefficients

[0104]

[0105] VI. Predictive Capability Assessment and Verification

[0106] Based on the above traditional risk factors, GRS, and proteins, three chronic liver disease prediction models were constructed using a Cox proportional hazards regression model in the training set population, and predictions were performed on the validation set:

[0107] ① The base model incorporates five traditional risk factors;

[0108] ② Genetic model (Base+GRS): In addition to the basic model, the GRS for chronic liver disease constructed in "II. Risk Factor Assessment and Common Predictive Factor Evaluation" is added.

[0109] ③ Protein model (Base+PS): In addition to the basic model, a protein score (PS) consisting of 10 proteins is added. The PS weights are obtained from the regression coefficients of the Cox model in Table 5, and the calculation formula is shown in Equation XII.

[0110] PS = -NPX1×0.159053+NPX2×0.592269+NPX3×1.007608+NPX4×0.231735+NPX5×0.557392-NPX6×0.134576+NPX7×0.842951+NPX8×0.478478-NPX9×1.143148+NPX10×0.891170-1.780483, Formula XII.

[0111] In the training and validation sets, the area under the curve (AUC), net weight classification improvement index (NRI), and integrated discriminant improvement index (IDI) of the three prediction models at 10-year time points were calculated to evaluate the actual predictive performance of the models. The AUC results showed that the protein model significantly improved the predictive ability of the base model for the incidence of chronic liver disease in the next 10 years. Figures 4-6 The genetic model did not improve predictive ability compared to the basic model. For NRI, the visible protein model showed significant improvement over the basic model in predicting the incidence of chronic liver disease over the next 10 years on both the training and validation sets, while the genetic model did not show any improvement on either the training or validation sets (Table 6). These results consistently demonstrate that the method and model constructed in this invention can accurately predict the 10-year risk of developing chronic liver disease.

[0112] Table 6. NRI and IDI results of training and validation ensemble chronic liver disease prediction models

[0113]

[0114] VII. Relevant Parameters of Chronic Liver Disease Prediction Model

[0115] The specific formula for calculating the model score for the aforementioned protein model is as follows:

[0116] Protein model score = 0.900954 × PS + 0.011503 × age + 0.067556 (gender = 1) - 0.142015 (urban resident = 1) + 1.115251 (weekly alcohol consumption = 1) + 1.578727 (HBV infection = 1) - 0.588822.

[0117] To facilitate the practical application of this invention, the score calculation formulas and corresponding cutoff values ​​for each model are shown in Table 7. Based on a 10% cutoff value for the 10-year incidence risk, the population is divided into high-risk and low-risk groups, and there are significant differences in the incidence risk of chronic liver disease among the populations. Figure 6 ).

[0118] Table 7. Formulas and cutoff values ​​for each prediction model

[0119]

[0120] VIII. Device for detecting protein levels of predictors of chronic liver disease

[0121] This embodiment provides a protein level detection device for predicting chronic liver disease. Depending on the specific application requirements, this detection device can be prepared into various forms of detection kits. The detection method is not limited to enzyme-linked immunosorbent assay (ELISA) and immunofluorescence detection; mass spectrometry, nucleic acid amplification, and other techniques can also be used. Taking an ELISA kit as an example, its basic components include a solid-phase carrier, preferably an enzyme-labeled plate (such as a polystyrene microplate), a membrane carrier (such as a nitrocellulose membrane or nylon membrane), or microspheres, with a specific antibody against the target protein pre-coated on the surface of the solid-phase carrier. To facilitate the detection process and result interpretation, a positive control can be included on the membrane carrier in some embodiments. The kit should also include a primary antibody that binds to the target protein and an enzyme-labeled secondary antibody, wherein the enzyme label can be horseradish peroxidase (HRP) or alkaline phosphatase (ALP).

[0122] All kits possess the capability for quantitative protein detection, making them particularly suitable for the analysis of multiple proteins in blood samples. Taking the detection of CD48 in serum using an ELISA kit as an example, the sample first binds to CD48-specific antibodies pre-coated on a solid-phase support; subsequently, enzyme-labeled secondary antibody is added, binding to the immune complex and immobilizing on the solid-phase surface. The amount of enzyme bound to the solid-phase surface is directly proportional to the CD48 concentration in the sample. After the addition of a specific substrate, a colored product is generated under catalysis; the intensity of the product's color reflects the concentration of the target protein. In this process, the high catalytic efficiency of the enzyme reaction significantly amplifies the signal, improving the sensitivity and accuracy of detection, thereby enabling qualitative and quantitative detection of the target protein.

[0123] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. Use of a panel of proteins in the construction of a predictive model of the risk of chronic liver disease, characterized in that, The protein combination is: VCAM1, CD48, NUDT5, MILR1, TNFSF8, DDAH1, IL17RB, NFASC, FCN2 and RTN4R; The UniProt accession numbers of the VCAM1, CD48, NUDT5, MILR1, TNFSF8, DDAH1, IL17RB, NFASC, FCN2 and RTN4R are P19320, P09326, Q9UKK9, Q7Z6M3, P32971, O94760, Q9NRM6, O94856, Q15485 and Q9BZR6, respectively.

2. Use according to claim 1, characterized in that, The formula of the prediction model is: The model score = PS x 0.900954 + age factor value + gender factor value + urban-rural factor value + drinking factor value + HBV infection state factor value - 0.588822; The age factor value is: the age of the subject x 0.011503; The gender factor value is 0 for males and 0.067556 for females; The urban-rural factor value is 0 for subjects who are rural residents and -1.4202 for subjects who are urban residents; The drinking factor value is 0 for subjects who do not drink every week and 1.115251 for subjects who drink every week; The HBV infection factor value is 0 for subjects whose blood test is negative for HBV and 1.578727 for subjects whose blood test is positive for HBV; The PS = -NPX1 x 0.159053 + NPX2 x 0.592269 + NPX3 x 1.007608 + NPX4 x 0.231735 + NPX5 x 0.557392 - NPX6 x 0.134576 + NPX7 x 0.842951 + NPX8 x 0.478478 - NPX9 x 1.143148 + NPX10 x 0.891170 - 1.780483; Wherein, NPX1 to NPX10 are the NPX values of VCAM1, CD48, NUDT5, MILR1, TNFSF8, DDAH1, IL17RB, NFASC, FCN2 and RTN4R in the protein combination of claim 1.

3. The use of claim 2, wherein the model score is higher than the cutoff value of the risk of developing the disease in the next 10 years being 10%, and is determined as high risk of developing chronic liver disease; The model score is lower than the cutoff value of the risk of developing the disease in the next 10 years being 10%, and is determined as low risk of developing chronic liver disease; The cutoff value of the risk of developing the disease in the next 10 years being 10% is 3.0037. Specific antibodies to the protein combination in the use of claim 1. The detection reagent as claimed in claim 4, together with an acceptable adjuvant, carrier or device.

4. Test reagent for the risk prediction of chronic liver disease, characterized in that, ​ 5. A test product characterised in that, ​ 6. Use of the detection reagent of claim 4 and / or the detection product of claim 5 in the preparation of a kit for predicting chronic liver disease.

7. Use of the detection reagent of claim 4 and / or the detection product of claim 5 in the construction of a chronic liver disease risk prediction system.

8. A chronic liver disease onset prediction system characterized by, Comprising: a data collection unit, a risk score unit and a prediction unit; the data collection unit: obtaining the NPX value of the protein in the protein combination in the application, the age, gender, urban and rural situation, drinking status and HBV infection information of the subject; the risk score unit: according to the NPX value, the age, gender, urban and rural situation, drinking status and HBV infection information of the subject, using the prediction model in the application of claim 2 to calculate the model score; the prediction unit: according to the 10% cutoff value of the 10-year incidence risk to judge the high and low of the incidence risk of chronic liver disease; the 10% cutoff value of the 10-year incidence risk is 3.0037.

Citation Information

Patent Citations

  • Markers for predicting prognosis of chronic and acute liver failure and application of markers

    CN116875674A

  • Diabetic chronic kidney disease occurrence risk prediction system and storage medium

    CN117711619A