Prediction method and system for NASH advanced hepatic fibrosis
By building a non-invasive diagnostic system based on the XGBoost model, using laboratory examination characteristics to predict advanced liver fibrosis in NASH patients, the problem of lack of non-invasive diagnostic models in the prior art is solved, and the ability of early recognition and intervention is improved.
Patent Information
- Application Number
- CN202510125253.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art lacks a non-invasive diagnostic model, making it difficult to identify advanced liver fibrosis in NASH patients early, resulting in a lag in treatment.
By collecting laboratory examination characteristics of NASH patients, calculating SHAP values and sorting them, building an XGBoost model, selecting laboratory examination characteristics such as triglyceride (TG), albumin (ALB), international standardized ratio (INR) and high-density lipoprotein (HDL) as sample data, predicting advanced liver fibrosis in NASH.
The ability to diagnose advanced liver fibrosis in patients with NASH has been achieved, the recognition rate has been improved, and the basis for early intervention has been provided.
Smart Images

Figure CN119993496A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a survival analysis technology for liver disease treatment, and in particular to a prediction method and system for NASH late-stage liver fibrosis. Background Art
[0002] With the change of modern people's lifestyle, the incidence of NASH (non-alcoholic steatohepatitis, also known as metabolic steatohepatitis) has gradually increased. Its incidence in East Asia is as high as 33%, and the global incidence has also increased from 25.26% in 1990 to 38.2% now. NASH patients will gradually progress to liver fibrosis or even cirrhosis due to the long-term inflammatory state of the liver. The incidence of liver fibrosis of F2 and above can be increased by 10% compared with normal people, and F3 / F4 stage liver fibrosis is considered to be advanced liver fibrosis. NASH patients at this stage are prone to a variety of adverse complications. Studies have shown that the risk of liver cancer in NASH patients with advanced liver fibrosis is 4 times higher than that of NASH patients without advanced liver fibrosis; the presence of advanced liver fibrosis also increases the incidence of cardiovascular and cerebrovascular events and the associated mortality rate. Since NASH patients with advanced liver fibrosis are prone to various adverse events, it is crucial to identify and treat these patients early.
[0003] Currently, the diagnosis of liver fibrosis in NASH patients is mainly performed through liver biopsy, which is an invasive procedure. Patients may experience postoperative adverse reactions such as local bleeding and edema, and it is not conducive to long-term monitoring of the patient's condition and treatment. Due to the limitations of traditional non-invasive diagnostic models, it is necessary to develop new non-invasive diagnostic models to identify advanced liver fibrosis in NASH patients. Summary of the invention
[0004] The main purpose of the present invention is to provide a prediction method and system for NASH advanced liver fibrosis to solve the problem of lack of a non-invasive diagnostic model for identifying advanced liver fibrosis in NASH patients.
[0005] According to an embodiment of the present invention, a method for predicting NASH advanced liver fibrosis is proposed, which includes: collecting laboratory examination features of NASH patients; for each laboratory examination feature, calculating its SHAP value and sorting the laboratory examination features according to the size of the SHAP value; selecting multiple laboratory examination features in sequence as sample data to construct an XGBoost model to predict NASH advanced liver fibrosis; wherein the screened laboratory examination features include: triglyceride TG, albumin ALB, international normalized ratio INR, and high-density lipoprotein HDL.
[0006] Wherein, the laboratory examination characteristics also include: liver stiffness value LSM.
[0007] The step of calculating the SHAP value of each laboratory examination feature includes: using the XGBoost model to calculate the SHAP value of each laboratory examination feature.
[0008] The sample data includes a training set and a validation set.
[0009] Wherein, the advanced liver fibrosis is F3 / F4 stage liver fibrosis.
[0010] According to an embodiment of the present invention, a system for predicting NASH advanced liver fibrosis is also proposed, which includes: an acquisition module, used to collect laboratory test features of NASH patients; a calculation module, used to calculate the SHAP value of each laboratory test feature and sort the laboratory test features according to the size of the SHAP value; a model construction module, used to select multiple laboratory test features in sequence as sample data to construct an XGBoost model to predict NASH advanced liver fibrosis; wherein the screened laboratory test features include: triglyceride TG, albumin ALB, international normalized ratio INR, high-density lipoprotein HDL.
[0011] Wherein, the laboratory examination characteristics also include: liver stiffness value LSM.
[0012] The calculation module is also used to calculate the SHAP value of each laboratory examination feature using the XGBoost model.
[0013] The sample data includes a training set and a validation set.
[0014] Wherein, the advanced liver fibrosis is F3 / F4 stage liver fibrosis.
[0015] According to the technical solution of the present invention, this application constructs a non-invasive model for diagnosing advanced liver fibrosis in NASH patients based on interpretable variable screening by selecting four laboratory examination features such as TG, ALB, INR and HDL, which can improve the recognition rate and perform intervention. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0017] Figure 1 is a flow chart according to an embodiment of the present invention;
[0018] Figure 2A and Figure 2B 3 are respectively a beeswarm diagram and an importance bar chart of the SHAP values of the clinical characteristics according to an embodiment of the present invention;
[0019] Figure 3A and Figure 3B Schematic diagrams of ROC curves of five machine learning models of a training cohort and a validation cohort according to an embodiment of the present invention;
[0020] Figure 4A and Figure 4B Schematic diagrams of ROC curves comparing the XBGoost model, the APRI model, and the FIB-4score model of the training cohort and the validation cohort according to an embodiment of the present invention;
[0021] Figure 5A and Figure 5B Schematic diagrams of decision curve analysis of the XGBoost model, XGBoost+LSM model, APRI model, and FIB-4score model for the training cohort and the validation cohort according to an embodiment of the present invention, respectively;
[0022] FIG6 is a schematic diagram of a cumulative curve of the XGBoost model predicting advanced liver fibrosis in a training cohort and a validation cohort according to an embodiment of the present invention. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0024] Although the present invention specification may include embodiments in many different forms, for some preferred embodiments described in detail in the specification and shown in the drawings, it should be understood that: these contents disclosed in the present invention specification should be regarded as schematic illustrations of the principles of the present invention, and these shown embodiments are not intended to limit the scope of protection of the present invention.
[0025] The technical solutions provided by various embodiments of the present invention are described in detail below in conjunction with the accompanying drawings.
[0026] According to an embodiment of the present invention, a method for predicting NASH late-stage liver fibrosis is provided, referring to Figure 1 , the method comprises the following steps:
[0027] S102, collect laboratory test characteristics of NASH patients.
[0028] S104, for each laboratory test feature, calculate its SHAP value and sort the laboratory test features according to the size of the SHAP value, refer to Figure 2A and Figure 2B .
[0029] Among them, SHAP (Shapley Additive Explanations) is an interpretable machine learning method used to explain the prediction results of the model. It is based on the Shapley value theory, quantifies the contribution of each feature to the model output, and combines properties such as local accuracy and consistency to provide a reliable and comprehensive feature importance evaluation indicator. SHAP can be used for variable screening, and the importance of the variable can be judged by comparing the SHAP values of different variables. In this application, the random forest algorithm or XGBoost algorithm can be used to calculate the SHAP value and weight.
[0030] S106, sequentially selecting multiple laboratory test features as sample data to construct an XGBoost model to predict NASH advanced liver fibrosis; wherein the screened laboratory test features at least include: triglyceride (TG), albumin (ALB), international normalized ratio (INR), high-density lipoprotein (HDL). In addition, the laboratory test features may also include liver stiffness value (LSM).
[0031] Among them, the importance of laboratory examination features can be calculated by XGBoost algorithm or random forest algorithm, and laboratory examination features that have a significant impact on cancer prediction can be screened out, thereby providing better feature selection results with higher early prediction accuracy.
[0032] The present application is described below with reference to examples.
[0033] This application included 1,870 patients diagnosed with NASH by liver biopsy at Beijing Ditan Hospital affiliated to Capital Medical University between January 2010 and January 2020. According to the inclusion and exclusion criteria, 746 patients were finally included and randomly divided into training cohort and validation cohort at a ratio of 0.7.
[0034] Inclusion criteria were: (1) aged over 18 years; (2) NASH confirmed by liver biopsy. Exclusion criteria were: (1) patients with liver diseases caused by other causes such as hepatitis B, hepatitis C, alcohol-related liver disease, autoimmune hepatitis, and genetic liver disease; (2) patients with liver malignancies or other tumors; (3) patients with HIV infection; (4) patients without liver biopsy data.
[0035] Data collection: The basic information, clinical and laboratory data of the patients were collected retrospectively through the electronic medical record system. All data were collected within 48 hours after admission. Basic information included gender, age, personal life history and medical history. Laboratory data included blood routine, liver function, renal function, electrolytes, and coagulation function of the patients. The pathological results of liver biopsy of each patient were collected to grade the degree of liver fibrosis.
[0036] Machine learning models: Five machine learning models, including XGBoost (XGB), logistic regression (LR), random forest (RF), support vector machine (SVM), and naive Bayes (NB), were trained and established in the training cohort to diagnose advanced fibrosis. Each machine learning model was validated in the validation cohort.
[0037] Statistical analysis: Statistical analysis was performed using R software (version 4.2.0). Continuous variables that conformed to normal distribution were analyzed using the t test, continuous variables that did not conform to normal distribution were analyzed using the Mann-Whitney U test, and categorical variables were analyzed using the X2 test (chi-square test) or Fisher's exact test. The variables for diagnosing advanced liver fibrosis were screened using the SHAP-based explanatory variable screening method. The ROC curves of the five machine learning models were plotted and the area under the curve was calculated to compare the clinical diagnostic efficacy. The DeLong test was used to compare the AUROCs. The DCA curve (decision curve) was used to compare the clinical decision efficacy between the models, and the calibration curve was used to determine the diagnostic accuracy of the model. The data were considered statistically significant when P < 0.05.
[0038] There was no statistical difference in the patient age (44.03 vs 44.65, P = 0.632), gender (male: 41.32% vs 45.40%, P = 0.328), F3 and above (F3 / F4) liver fibrosis degree (36.16% vs 37.74%, P = 0.683), laboratory-related indicators and non-invasive liver fibrosis scores between the training cohort and the validation cohort. The details can be seen in Table 1.
[0039] Table 1 Basic clinical characteristics of NASH patients
[0040]
[0041]
[0042] refer to Figure 3AAs shown in Table 2, in the diagnosis of F3 / F4 liver fibrosis, in the training cohort, the area under the ROC curve (AUROC) of the LR model in the constructed machine learning model was 0.805 (0.764-0.846), the sensitivity (SE) was 0.901, the specificity (SP) was 0.593, the positive predictive value (PPV) was 0.782, and the negative predictive value (NPV) was 0.787. The AUROC of the RF model was 0.857 (0.825-0.888), the SE was 0.944, the SP was 0.769, the PPV was 0.869, and the NPV was 0.894. The AUROC of the SVM model was 0.773 (0.738-0.809), the SE was 0.969, the SP was 0.578, the PPV was 0.789, and the NPV was 0.920. The NB model has an AUROC of 0.503 (0.499-0.575), a SE of 0.006, an SP of 1.000, a PPV of 1.000, and an NPV of 0.384. The highest performing machine learning model is XGBoost, with an AUROC of 0.934 (0.914-0.955), a SE of 0.958, an SP of 0.575, a PPV of 0.786, and an NPV of 0.914. The P values of XGBoost compared with other machine learning models are all less than 0.001.
[0043] refer to Figure 3B As shown in Table 2, in the validation cohort, the AUROC of the LR model in the constructed machine learning model was 0.745 (0.669-0.821), SE was 0.959, SP was 0.481, PPV was 0.772, and NPV was 0.864. The AUROC of the RF model was 0.840 (0.788-0.892), SE was 0.959, SP was 0.722, PPV was 0.863, and NPV was 0.905. The AUROC of the SVM model is 0.740 (0.684-0.796), SE is 0.986, SP is 0.494, PPV is 0.784, and NPV is 0.951; the AUROC of the NB model is 0.503 (0.495-0.510), SE is 0.007, SP is 1.000, PPV is 1.000, and NPV is 0.354. The machine learning model with the highest performance is still XGBoost, with an AUROC of 0.917 (0.880-0.953), SE is 0.925, SP is 0.552, PPV is 0.756, and NPV is 0.9884. When XGBoost is compared with other machine learning models, the P value is still less than 0.001.
[0044] In summary, the XBGoost model showed good diagnostic value both in the training cohort and the validation cohort, and its diagnostic efficacy was significantly higher than that of other machine learning diagnostic models.
[0045] Table 2 Diagnostic performance of machine learning models for advanced liver fibrosis
[0046]
[0047] In some embodiments of the present application, the laboratory test features in the constructed XGBoost model may also include LSM data. The XGBoost model without LSM data and the XGBoost model with LSM data are compared with other non-invasive diagnosis models, wherein the other non-invasive diagnosis models include APRI model and FIB-4 model.
[0048] refer to Figure 4A As shown in Table 3, in the training cohort, the AUROC of the XGBoost model (XGB+LSM or XGBoost+LSM) with LSM data added was 0.977 (0.966-0.980), SE was 0.985, SP was 0.758, PPV was 0.837, and NPV was 0.929. The AUROC of the APRI model was 0.803 (0.765-0.841), SE was 0.907, SP was 0.402, PPV was 0.711, and NPV was 0.727. The AUROC of the FIB-4 model was 0.811 (0.774-0.848), SE was 0.898, SP was 0.467, PPV was 0.732, and NPV was 0.738. When compared with the APRI and FIB-4 model scores, the diagnostic value of the XGBoost model was higher (P<0.001), and the diagnostic value of the XGBoost model was further improved after the addition of LSM data (P<0.001).
[0049] refer to Figure 4B Similar results were obtained in the validation cohort as shown in Table 3. The AUROC of the XGBoost model with LSM data was 0.970 (0.950-0.990), SE was 0.954, SP was 0.739, PPV was 0.810, and NPV was 0.908. The AUROC of the APRI model was 0.737 (0.669-0.805), SE was 0.917, SP was 0.329, PPV was 0.715, and NPV was 0.684. The AUROC of the FIB-4 model was 0.752 (0.687-0.816), SE was 0.912, SP was 0.165, PPV was 0.681, and NPV was 0.765.
[0050] Table 3 Diagnostic performance of XGBoost model and other non-invasive diagnostic models for advanced hepatitis
[0051]
[0052] refer to Figure 5A , this application uses DCA curve and calibration curve to evaluate the effectiveness of the constructed machine learning diagnostic model. The results show that in the training set, the decision-making performance of the XGBoost machine learning model is significantly better than APRI and FIB-4; after adding LSM data, the decision-making efficiency of the XGBoost model is further improved. Reference Figure 5B , similar results were also found in the validation cohort, showing that XGBoost and XGBoost+LSM have better decision-making performance in clinical practice. Fig. 6A and Figure 6B ,The calibration curve of the XGBoost machine learning model also showed good consistency with the actual results when predicting F3 / F4 liver fibrosis in the training set and validation set.
[0053] This application develops a machine learning model for diagnosing advanced liver fibrosis in NASH patients by comparing five machine learning models. Among them, the model based on the XGBoost algorithm has the best performance. Based on this, an online webpage can be developed to facilitate the calculation of the risk of advanced liver fibrosis. Different model calculations can be performed for the existence of LSM data, which is more convenient for patient management. Finally, the DCA curve and calibration curve also illustrate the accuracy of the XGBoost model in clinical application.
[0054] In the present application embodiment, the XGBoost model construction is carried out by the xgboost package in R language (version 4.4.1). Through the electronic medical record system of Beijing Ditan Hospital affiliated to Capital Medical University, patients with non-alcoholic fatty hepatitis (NASH) confirmed by liver biopsy from January 2010 to January 2020 were retrospectively collected, and the patients were staged according to the results of liver biopsy pathology according to the Brent standard, and the results of blood routine, electrolytes, liver function, renal function, and coagulation function of these patients were collected. First, the 746 patients included were divided into a modeling group including 522 patients and a validation group including 224 patients in a ratio of 7:3. Based on the Shapley value theory, the SHAPSHapley AdditiveexPlanations value of each variable is calculated and sorted. The variables with large differences in the SHAP values of the first 4 variables and other variables are included in the establishment of the XGBoost model. When building the model, the advanced interface xgb.train() for training the xgboost model is used. First, define the acceptance list params in xgb.train(). In params, define the task as a binary logical variable (binary:logistic) and output the probability value; define the booster type (booster) as gbtree, and define the evaluation index (eval_metric) of the validation set as error; the learning rate (eta), that is, the contribution of each tree in the final solution, is defined as 0.1; the maximum depth of a single number (max_depth) is defined as 1000; the proportion of subsample data to the entire observation (subsample) is defined as 0.5; the number of features randomly extracted when building a tree (colsample_bytree) is defined as 0.5; and the proportion of L2 regularization (lambda) is defined as 3.0. After defining the params list, use xgb.train() to build the model. The model is successfully built. Use the predict() function in the training set to get the probability of each patient using the model to diagnose advanced liver fibrosis, and then verify the model in the validation set. Then calculate the AUROC, DCA curve, and calibration curve of the model in the training set and validation set to evaluate the effectiveness of the model. Export the model as a bin file, and then build an online web diagnostic system. The model is used through an online website. When using it, one only needs to fill in the numerical values of the four variables included in the model. The system can automatically generate a probability of diagnosing advanced liver fibrosis, thereby achieving the purpose of non-invasive diagnosis of advanced liver fibrosis.
[0055] According to an embodiment of the present invention, a system for predicting NASH advanced liver fibrosis is also proposed, which includes: an acquisition module, used to collect laboratory test features of NASH patients; a calculation module, used to calculate the SHAP value of each laboratory test feature and sort the laboratory test features according to the size of the SHAP value; a model construction module, used to select multiple laboratory test features in sequence as sample data to construct an XGBoost model to predict NASH advanced liver fibrosis; wherein the screened laboratory test features include: triglyceride TG, albumin ALB, international normalized ratio INR, high-density lipoprotein HDL.
[0056] Wherein, the laboratory examination characteristics also include: liver stiffness value LSM.
[0057] The calculation module is also used to calculate the SHAP value of each laboratory examination feature using the XGBoost model.
[0058] The sample data includes a training set and a validation set.
[0059] Wherein, the advanced liver fibrosis is F3 / F4 stage liver fibrosis.
[0060] The operation steps of the method of the present invention correspond to the structural features of the system, and can be referenced to each other, and will not be described in detail again.
[0061] Although the present disclosure has been described in detail with reference to the specific embodiments of the present application, it will be understood by those skilled in the art that various changes and modifications may be made therein without departing from the spirit and scope of the embodiments. Therefore, the present application is intended to cover the modifications and variations of the present application, and any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of the claims of the present application and their equivalents.
[0062] In addition, the features disclosed in the above description or claims or drawings, in their specific form or according to the mode for performing the disclosed functions or the method or process for obtaining the disclosed results, can be used to implement the present application in their different forms alone or in any combination of these features, as appropriate. Specifically, one or more features of any embodiment described in the present application can be combined with one or more features of any other embodiment described in the present application.
[0063] Protection may also be sought for any features disclosed in any one or more of the publications cited in conjunction with the present application and / or incorporated by reference.
Claims
1. A method for predicting NASH advanced liver fibrosis, characterized in that: include: Laboratory characteristics of NASH patients were collected; For each laboratory examination feature, its SHAP value was calculated and the laboratory examination features were ranked according to the size of the SHAP value; Multiple laboratory test features were selected in sequence as sample data to construct an XGBoost model to predict NASH advanced liver fibrosis; among them, the screened laboratory test features included: triglyceride TG, albumin ALB, international normalized ratio INR, and high-density lipoprotein HDL.
2. The method according to claim 1, characterized in that The laboratory test characteristics also include: liver stiffness value LSM.
3. The method according to claim 1 or 2, characterized in that: The step of calculating the SHAP value of each laboratory test feature comprises: The XGBoost model was used to calculate the SHAP value of each laboratory examination feature.
4. The method according to claim 1, characterized in that The sample data includes a training set and a validation set.
5. The method according to claim 1, characterized in that The advanced liver fibrosis is F3 / F4 stage liver fibrosis.
6. A system for predicting NASH advanced liver fibrosis, characterized in that: include: The collection module is used to collect laboratory test characteristics of NASH patients; A calculation module, used for calculating the SHAP value of each laboratory examination feature and sorting the laboratory examination features according to the size of the SHAP value; The model building module is used to sequentially select multiple laboratory test features as sample data to build an XGBoost model to predict NASH advanced liver fibrosis; among them, the screened laboratory test features include: triglyceride TG, albumin ALB, international normalized ratio INR, and high-density lipoprotein HDL.
7. The system according to claim 6, characterized in that The laboratory test characteristics also include: liver stiffness value LSM.
8. The system according to claim 6 or 7, characterized in that: The calculation module is also used to calculate the SHAP value of each laboratory examination feature using the XGBoost model.
9. The system according to claim 6, characterized in that The sample data includes a training set and a validation set.
10. The system according to claim 6, characterized in that The advanced liver fibrosis is F3 / F4 stage liver fibrosis.