Method, system, equipment and medium for predicting progress of chronic hepatitis B cirrhosis based on ordered logistic regression

By combining an improved ordered logistic regression model with multilayer perceptron and feature selection methods, and using serum biomarkers to construct an adaptive threshold extension model, the problem of insufficient accuracy of existing liver fibrosis early warning models in hepatitis B patients is solved, and high-precision assessment and management of early liver fibrosis is achieved.

CN121483459APending Publication Date: 2026-02-06NINGBO MEDICAL CENT LIHUILI HOSPITACL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510721623.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing liver fibrosis early warning models have low accuracy in assessing the degree of liver fibrosis in hepatitis B patients, especially in the early stages of liver fibrosis, where they are difficult to monitor effectively. Furthermore, liver biopsy is invasive and has poor reproducibility.

Method used

An improved ordered logistic regression model was adopted, which combined multilayer perceptron and forward stepwise feature selection. An adaptive threshold-extended ordered logistic regression model was constructed using a combination of serum biomarkers (PLT, GGT, FPG, GLB, AFP, IBIL, TBIL and HBeAg) to dynamically adjust the variance and improve prediction accuracy.

Benefits of technology

It enables early and accurate assessment of the degree of liver fibrosis in hepatitis B patients, improves the accuracy and practicality of diagnosis, optimizes patient disease management, and reduces the rate of progression to severe illness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121483459A_ABST
    Figure CN121483459A_ABST
Patent Text Reader

Abstract

The invention discloses a system, equipment and medium for predicting the progress of chronic hepatitis B cirrhosis based on ordered logistic regression, and relates to the technical field of intelligent diagnosis and treatment. According to the method, a traditional ordered logistic regression model is improved, a self-adaptive threshold extension ordered logistic regression model is obtained, and a model for predicting the progress of chronic hepatitis B cirrhosis is constructed based on the quantitative data of the optimal serum marker obtained through screening. According to the present invention, in the two scenes of F < 2 + > and F < 3 + > prediction, the performance exceeds the XGB Classifier, the Random Forest Classifier and the Logistic Regression, and is superior to the classic hepatic fibrosis scoring model APRI model, the FIB-4 model and the GPR model, such that the important clinical application value is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent diagnosis and treatment technology, and in particular to a system, device and medium for predicting the progression of chronic hepatitis B cirrhosis based on ordered logistic regression. Background Technology

[0002] Chronic hepatitis B (CHB) caused by hepatitis B virus (HBV) is a major cause of cirrhosis and liver cancer. Liver fibrosis is a key step in the progression of various chronic liver diseases to cirrhosis, and the degree of liver fibrosis is one of the important factors determining the prognosis of CHB patients and whether they need aggressive treatment.

[0003] Currently, the "gold standard" for diagnosing liver fibrosis is liver biopsy. Clinical practice guidelines use serum alanine aminotransferase (ALT) and serum HBV DNA status to assess the necessity of liver biopsy for antiviral treatment in patients with chronic hepatitis B. However, this method overlooks some chronic hepatitis B patients with normal or slightly elevated ALT levels, who may have underlying liver fibrosis or even cirrhosis. Furthermore, liver biopsy itself has drawbacks, including being invasive, having limited sampling capabilities, low patient acceptance, and poor reproducibility. Therefore, liver fibrosis early warning models have become an effective supplement to liver biopsy.

[0004] A series of liver fibrosis early warning models have been developed both domestically and internationally, such as the APRI model, FIB-4 model, GPR model, and FibroTest model. The APRI model has low sensitivity and specificity in diagnosis, the FIB-4 model has low predictive value in early fibrosis, and the GPR model has a high accuracy in predicting HBeAg-positive patients, but none of them are more advantageous than the APRI model or the FIB-4 model. The indicators of the FibroTest model are not easy to obtain in clinical practice, and its feasibility for promotion in community-level primary care is low.

[0005] Studies have shown that the aforementioned early warning models have low predictive accuracy in assessing the early stage of hepatitis B liver fibrosis, and their diagnostic accuracy is significantly affected by severe liver inflammation. They cannot effectively monitor the early progression or reversal of liver fibrosis in patients in real time.

[0006] To date, there is still a lack of ideal predictive methods for early liver fibrosis in CHB patients. Therefore, there is an urgent clinical need for a novel early warning model for hepatitis B liver fibrosis to compensate for the limitations of liver biopsy, assess the degree of liver fibrosis earlier and more accurately, provide strong support for clinical decision-making, and reduce the rate of severe illness in patients with chronic hepatitis B. Summary of the Invention

[0007] To address at least one of the aforementioned technical problems, this application aims to provide a novel early warning model for hepatitis B liver fibrosis, compensating for the shortcomings of liver biopsy, to assess the degree of liver fibrosis earlier and more accurately, providing strong support for clinical decision-making, and with the goal of reducing the rate of severe illness in patients with chronic hepatitis B. To this end, the technical solution adopted in this application is as follows.

[0008] The first aspect of this application provides a method for constructing a model for predicting the progression of chronic hepatitis B-related cirrhosis, comprising the following steps: S101, obtain quantitative data on the combination of serum biomarkers in the population. The group includes multiple groups of patients with chronic hepatitis B cirrhosis with varying degrees of liver fibrosis, the degree of which is classified according to the METAVIR scoring system; The serum biomarker combination includes PLT, GGT, FPG, GLB, AFP, IBIL, TBIL, and HBeAg; S102, using the quantitative data of the serum biomarker combination of the population obtained in step S101, an ordered logistic regression model is constructed to predict the progression of chronic hepatitis B cirrhosis.

[0009] In this application, an improved ordered logistic regression model was used to screen for eight features (i.e., blood biochemical indicators): Platelet count (PLT): Platelets are closely related to liver fibrosis. Platelets play a key role in inhibiting liver fibrosis by suppressing the activation of hepatic stellate cells. Furthermore, platelets play a dual role in liver fibrosis; higher PLT levels are associated with milder liver fibrosis. A decrease in PLT is often seen in the later stages of liver fibrosis.

[0010] Gamma-glutamyl transferase (GGT): GGT is a sensitive indicator of liver damage and biliary tract disease. Its elevation may reflect cholestasis and hepatocellular damage, and is often positively correlated with the occurrence of liver fibrosis and cirrhosis.

[0011] Fasting plasma glucose (FPG): The liver is a vital organ for maintaining the dynamic balance of blood glucose in the body. When the liver is damaged, glucose metabolism is often affected. Patients with liver fibrosis often have insulin resistance or abnormal glucose metabolism.

[0012] Globulin (GLB): Globulin levels can reflect the body's immune function status. During fibrosis, inflammation and damage to the liver lead to abnormal immune function, affecting the synthesis and secretion of globulins. Globulin levels may be elevated in patients with liver fibrosis and cirrhosis.

[0013] Alpha-fetoprotein (AFP): As liver fibrosis progresses, the liver's ability to convert nutrients into energy decreases, and the body needs a large number of new liver cells to produce energy. A certain amount of alpha-fetoprotein is produced during the regeneration of liver cells, so AFP levels may also be elevated in patients with liver fibrosis and cirrhosis.

[0014] Indirect bilirubin (IBIL): As liver fibrosis worsens, hepatocyte function is impaired, and glucuronyl transferase activity decreases, leading to impaired indirect bilirubin binding and elevated IBIL levels. Simultaneously, fibrous tissue proliferation or pseudolobule formation can compress bile ducts, causing intrahepatic cholestasis and indirectly affecting bilirubin metabolism.

[0015] Total bilirubin (TBIL): Elevated total bilirubin may indicate hepatocellular dysfunction or cholestasis. Liver fibrosis can damage the liver, affecting its ability to process bilirubin and leading to elevated total bilirubin.

[0016] Hepatitis B e antigen (HBeAg): A positive HBeAg result indicates a high replication status of the hepatitis B virus, often accompanied by high hepatitis activity, which may accelerate the progression of liver fibrosis. HBeAg-positive patients are more likely to develop liver fibrosis.

[0017] Liver cirrhosis is caused by long-term liver fibrosis and inflammation, leading to the death of a large number of hepatocytes and the destruction of the liver lobule structure, ultimately resulting in the proliferation of fibrous tissue and the formation of nodular lesions within the liver, thus leading to cirrhosis. In this application, the progression of liver cirrhosis is characterized by the degree of liver fibrosis.

[0018] The METAVIR scoring system is a universal method for determining the grade of liver inflammation and necrosis, as well as the degree of fibrosis. The METAVIR scoring system classifies liver fibrosis into five stages, from F0 to F4. F0 stage refers to the absence of fibrous tissue proliferation in the liver; F1 stage refers to the expansion of fibrosis in the portal areas of the liver, but without the formation of fibrotic septa; F2 stage refers to the expansion of fibrosis in the portal areas of the liver, with the formation of a few fibrous septa; F3 stage refers to the formation of most fibrous septa in the liver, but without sclerotic nodules; F4 stage refers to the development of liver cirrhosis.

[0019] It can be seen that in the F0-F4 stages, the degree of liver fibrosis increases step by step, and the process of cirrhosis also develops accordingly.

[0020] In some embodiments of this application, patients may be grouped according to the degree of fibrosis based on the stages obtained from the METAVIR scoring system. The number of groups may be two to five. In some specific embodiments of this application, the number of groups may be three: (1) patients in stages F0 and F1, (2) patients in stage F2, and (3) patients in stages F3 and F4.

[0021] Logistic regression analysis is used to study the impact of X on Y. There are no requirements on the data type of X; X can be categorical or quantitative data, but Y must be categorical data.

[0022] In some possible embodiments of this application, in the ordered logistic regression model, the logarithm lnσ² of the variance is predicted using a multilayer perceptron. x ): lnσ 2 ( x )=MLPϕ( x ) in, x This refers to serum biomarkers, specifically the MLPϕ( x ) represents the output of the multilayer perceptron. However, a positive variance prediction value can be obtained using the following formula: σ 2 ( x )=exp(MLPϕ( x )) Then use the following formula for prediction:

[0023] in, It is a threshold. It is a prediction function.

[0024] It should be noted that, in the above formula, compared to the traditional ordered logistic regression model, dynamic changes are utilized. Replace fixed This makes the model highly interpretable. Specifically, the coefficient of each independent variable can be directly interpreted as the influence of the log odds of the probability of occurrence of each level of the dependent variable.

[0025] Those skilled in the art will know that a multilayer perceptron is a neural network with at least three layers: (1) an input layer (receiving data), (2) some intermediate layers (one or more hidden layers, computing data), and (3) an output layer (outputting results).

[0026] In this application, an activation function is used to achieve nonlinearity. In some embodiments of this application, the activation function of the multilayer perceptron is ReLU. In some specific embodiments of this invention, the intermediate layer of the multilayer perceptron includes two layers, thereby: MLPϕ( x =W2·ReLU(W1) x +b1)+b2 Where W1 represents the weight of the first layer of the multilayer perceptron, W2 represents the weight of the second layer of the multilayer perceptron, b1 represents the bias value of the first layer of the multilayer perceptron, and b2 represents the bias value of the second layer of the neural network structure of the multilayer perceptron.

[0027] The second aspect of this application provides an ordered logistic regression model for predicting the progression of chronic hepatitis B cirrhosis, which is constructed using any of the methods described in the first aspect of this application.

[0028] A third aspect of this application provides a method for predicting the progression of chronic hepatitis B-related cirrhosis, comprising the following steps: S1, obtain quantitative data of the serum biomarker combination of the patient to be predicted, the serum biomarker combination including PLT, GGT, FPG, GLB, AFP, IBIL, TBIL and HBeAg; S2, the quantitative data of the serum biomarker combination of the patient to be predicted is input into the ordered logistic regression model described in the second aspect of this application to obtain the classification result of the degree of liver fibrosis, thereby obtaining the progression of liver cirrhosis.

[0029] In this application, the classification result of the degree of liver fibrosis is related to the grouping of the population used to construct the ordered logistic regression model. Specifically, if the population is grouped into three groups, the model will predict one of them.

[0030] In this application, by making multiple predictions for patients with chronic hepatitis B, dynamic monitoring of the progression of cirrhosis in these patients can be achieved, thereby enabling timely medical assessment or intervention.

[0031] A fourth aspect of this application provides a system for predicting the progression of chronic hepatitis B-related cirrhosis, the system comprising: The data input module is used to obtain quantitative data on the combination of serum biomarkers in the patient to be predicted; A model storage module is used to store the ordered logistic regression model described in the second aspect of this application; The prediction model is connected to the data input module and the model storage module respectively. It inputs the quantitative data of the serum biomarker combination of the patient to be predicted into the ordered logistic regression model to obtain the classification result of the degree of liver fibrosis, thereby obtaining the progression of liver cirrhosis.

[0032] In some embodiments of this application, the model storage module includes a data storage unit and a model building unit. The data storage unit is used to store quantitative data of the serum biomarker combination of the population, and the model building unit is used to construct an ordered logistic regression model using the quantitative data of the serum biomarker combination of the population.

[0033] The fifth aspect of this application provides an electronic device, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method described in any of the first aspects of this application or the method described in the third aspect of this application.

[0034] The fifth aspect of this application provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the method described in any one of the first aspects of this application or the method described in the third aspect of this application.

[0035] Compared with the prior art, this application has the following advantages: This application improves upon the traditional OLR model, obtaining an adaptive threshold-extended ordered logistic regression model. Specifically, it utilizes a multilayer perceptron to predict variance, replacing the fixed variance in the OLR cumulative probability function with dynamically changing variance. Furthermore, it employs a forward stepwise feature selection method to screen for the optimal feature set suitable for model construction, including PLT, GGT, FPG, GLB, AFP, IBIL, TBIL, and HBeAg. This feature set is used for the first time in the early diagnosis of liver fibrosis, comprehensively assessing liver status from various perspectives such as viral immunology and energy metabolism. It provides clinicians with a highly accurate and practical early warning assessment tool for chronic hepatitis B liver fibrosis, further optimizing patient disease management and improving the overall level of clinical diagnosis and treatment.

[0036] In the model of this application, the coefficients of each independent variable can be directly interpreted as the influence on the log odds of the occurrence probability of each level of the dependent variable, which is simpler and easier to understand than some more complex models.

[0037] The method and system of this application outperformed XGBClassifier, Random Forest Classifier and Logistic Regression in both F2+ and F3+ scenarios, demonstrating significantly superior performance.

[0038] In predicting both F2+ and F3+ scenarios, the method and system of this application outperformed classic liver fibrosis scoring models such as the APRI model, FIB-4 model, and GPR model in AUROC. Furthermore, it demonstrated excellent sensitivity and specificity at different thresholds, indicating that compared to traditional scoring systems that rely on fixed biomarker ratios, the model of this application can more accurately capture patterns in complex data, thereby providing more accurate and reliable predictions of liver fibrosis grades.

[0039] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0040] The above and other objects, features, and advantages of exemplary embodiments of this application will become readily apparent from the following detailed description taken in conjunction with the accompanying drawings. Several embodiments of this application are illustrated in the drawings by way of example and not limitation, in which: Figure 1 The construction and validation of the model used to predict the progression of cirrhosis in CHB patients in Embodiment 1 of this application are shown; Figure 2 The AUROC, Top K feats, mean AUROC, and ±1std are shown in Embodiment 1 of this application when different numbers of features are selected using the forward feature selection method. AUROC is an important indicator for evaluating the performance of a classification model and comprehensively reflects the classification effect of the model under different thresholds. Figure 3 The diagram shows a performance comparison between the HeteroOrdinal model in Embodiment 2 of this application and other machine learning algorithms, evaluating the model performance at F2+, F3+, and AVG (combined mean). XGBoost (Extreme Gradient Boosting), Random Forest, Logistic Regression, and HeteroOrdinal represent different machine learning modeling methods. Figure 4 The diagram shows a comparison of ROC curves between the HeteroOrdinal model in Embodiment 2 of this application and other liver fibrosis diagnostic models. The ROC curve is a graphical tool used to evaluate the performance of a classification model. The horizontal axis is FPR and the vertical axis is TPR. TPR (true positive rate) and FPR (false positive rate) are two commonly used indicators for evaluating the performance of a classification model. Figure 5This document shows a calibration plot of the HeteroOrdinal model prediction results and the actual situation of patients with liver fibrosis in the F2+ stage in Example 2 of this application. The calibration plot includes True Probability, Predicted Probability, Binned Calibration, Perfect Calibration, and Nonparametric (Isotonic Regression). Figure 6 A schematic diagram of the composition structure of an electronic device according to Embodiment 4 of this application is shown. Detailed Implementation

[0041] Unless otherwise stated, implied from the context, or as is customary in the art, all parts and percentages in this application are based on weight, and the testing and characterization methods used are concurrent with the filing date of this application. If any definition of a specific term disclosed in the prior art is inconsistent with any definition provided in this application, the definition provided in this application shall prevail.

[0042] To make the technical problems, technical solutions and beneficial effects solved by this application clearer, the following detailed description is provided in conjunction with embodiments.

[0043] The following examples are used to illustrate preferred embodiments of this application. Those skilled in the art will understand that the techniques disclosed in the examples represent technologies discovered by the inventors that can be used to implement this application, and therefore can be considered preferred embodiments of this application. However, those skilled in the art should understand from this specification that many modifications can be made to the specific embodiments disclosed herein, still yielding the same or similar results, without departing from the spirit or scope of this application.

[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains, and all materials cited herein and referenced by them are incorporated herein by reference.

[0045] Those skilled in the art will recognize, or can learn through routine experimentation, many equivalents of the specific embodiments of the invention described herein. These equivalents will be included in the claims.

[0046] Unless otherwise specified, the experimental methods used in the following examples are conventional methods. Unless otherwise specified, the instruments and equipment used in the following examples are all conventional laboratory instruments and equipment; unless otherwise specified, the experimental materials used in the following examples were all purchased from conventional biochemical reagent stores.

[0047] Example 1: Construction and validation of a model for predicting the progression of cirrhosis in CHB patients Combination Figure 1 This embodiment details the process of constructing and validating a model for predicting the progression of cirrhosis in CHB patients.

[0048] 1. Collect patient data and group them according to the stage of liver fibrosis. The inventors retrospectively collected clinical data of CHB patients aged 18-65 years with ALT≤2ULN from Ningbo Medical Center Li Huili Hospital and the First Affiliated Hospital of Wenzhou Medical University between 2015 and 2024.

[0049] After rigorous screening, patients with liver cancer, malignant liver tumors, those who have undergone liver transplantation, those with uncontrollable metabolic diseases of the heart, kidneys, lungs, blood system, autoimmune system, or gastrointestinal tract, and those with other systemic infections were excluded, and 294 patients were finally included.

[0050] Liver fibrosis is classified into stages F0 to F4 according to the METAVIR scoring system, among which: F0 stage refers to the absence of fibrous tissue proliferation in the liver; F1 stage refers to the expansion of fibrosis in the portal areas of the liver, but without the formation of fibrotic septa; F2 stage refers to the expansion of fibrosis in the portal areas of the liver, with the formation of a few fibrous septa; F3 stage refers to the formation of most fibrous septa in the liver, but without sclerotic nodules; F4 stage refers to the development of liver cirrhosis.

[0051] The inventors categorized the collected patient data into three groups based on the degree of liver fibrosis: F0-1 (mild, n=178), F2 (moderate, n=60), and F3-4 (severe, n=56). Based on this labeling method, the inventors performed F-tests on the collected continuous variables to remove features with insignificant differences between groups, ensuring that the subsequent model uses only statistically significant variables. The selected continuous variables were combined with the original categorical variables to form a candidate feature set for model training, as detailed in Table 1. This method not only improves the scientific rigor and precision of feature selection but also lays a solid foundation for building a more accurate and reliable predictive model.

[0052] Table 1: Candidate feature set of liver fibrosis patients after F-test

[0053] 2. Training machine learning models Given that this embodiment aims to construct a three-classification model for assessing the early stage of liver fibrosis—that is, the classification labels are ordered categorical variables (F0-F1, F2, F3-F4)—a model that can effectively predict ordered categorical variables is needed.

[0054] Unlike traditional machine learning classification models such as Support Vector Machine (SVM), Logistic Regression, and Random Forest, the Ordinal Logistic Regression (OLR) model is particularly suitable for handling data where the dependent variable is ordinal categorical data, and can effectively capture the order relationship between different levels.

[0055] However, the OLR model is not flexible enough when dealing with complex data of liver fibrosis test indicators. For example, it cannot capture the phenomenon that the variance of the error term (residual) changes with the independent variable, i.e. the heteroscedasticity problem of the distribution of test indicator data, which may lead to inaccurate model fitting.

[0056] To address this problem, the inventors introduced a variance prediction network, the output of which is the logarithm of the residual variance to the base e, lnσ². x ): lnσ 2 ( x )=MLPϕ( x ) in, x This refers to serum biomarkers, specifically the MLPϕ( x ) represents the output of the multilayer perceptron.

[0057] Then, it is converted into a positive variance prediction value using an exponential function: σ 2 ( x )=exp(MLPϕ( x )) Therefore, the inventors developed an improved model—the adaptive threshold extended ordered logistic regression model, also known as the heteroscedastic ordered logistic regression model (HeteroOrdinal model).

[0058] The specific implementation process is as follows: (1) Variance prediction network: Input: Features x (Features in Table 1) Output: the logarithm of the variance, lnσ 2 ( x ) Structure: Multilayer Perceptron (MLP), with two intermediate layers. lnσ 2 (x )=MLPϕ( x =W2·ReLU(W1) x +b1)+b2 The variance is guaranteed to be positive using an exponential function: σ 2 ( x )=exp(MLPϕ( x )) (2) Error term modeling OLR: Error term ϵ ~ Logistic(0, σ) 2 ), with a fixed variance.

[0059] Modified: Error term ϵ ~ Logistic(0, σ) 2 ( x The variance changes dynamically.

[0060] (3) Modification of model probability calculation In ordered logistic regression models, a cumulative probability function is typically used to calculate class probabilities, for example:

[0061] in, It is a threshold. It is a prediction function. In the HeteroOrdinal model, Replace with ,Right now:

[0062] The HeteroOrdinal model has good interpretability. Specifically, the coefficient of each independent variable can be directly interpreted as the effect on the log odds of the occurrence probability of each level of the dependent variable. This makes the model results more intuitive and easier to understand. Compared with some more complex models, this method provides clearer insights and a clearer path to understanding.

[0063] The HeteroOrdinal model not only accurately reflects the ordered structure in early-stage liver fibrosis data but also ensures high interpretability and good predictive performance. It provides a powerful yet easy-to-understand tool that helps in-depth exploration and explanation of the mechanisms of liver fibrosis progression.

[0064] 3. Feature Selection The feature sets in Table 1 were trained using the HeteroOrdinal model, and then the optimal features were selected using beam search combined with a forward stepwise feature selection method. The specific steps are as follows: (1) Initialization Define an initial empty feature set.

[0065] Set beam width k to represent the number of best candidate feature subsets retained each time; in this embodiment, it is set to 10.

[0066] (2) First round of selection All individual features are evaluated, and the top k best-performing features (with the largest AUROC) are selected as the initial candidate paths.

[0067] (3) Iterative expansion: In each round, each candidate path is expanded: Add each unselected feature to the current path one by one.

[0068] Evaluate the performance of each new path after expansion.

[0069] Based on the evaluation results, the top k best paths are selected as candidate paths for the next round.

[0070] (4) Termination conditions: Stop when the preset maximum number of features is reached or the model performance no longer improves significantly.

[0071] The best-performing subset of features is selected from the candidate paths in the last round as the final optimal feature set.

[0072] The changes in AUROC values ​​throughout the forward feature selection process are shown in Table 2 and Figure 2 As shown.

[0073] Table 2: Preferred characteristics of patients with liver fibrosis

[0074] Table 2 shows the AUROC mean and standard deviation, representing the AUROC mean and standard deviation for the Top 10 models when selecting the corresponding number of features. Table 2 shows that when the number of features reaches 8, the AUROC mean exceeds 0.83. Further increases in the number of features do not significantly increase the AUROC mean. Therefore, the inventors determined the number of features to be 8, with the feature set exhibiting the highest AUROC including 8 features: PLT, GGT, FPG, GLB, AFP, IBIL, TBIL, and HBeAg.

[0075] Thus, the inventors obtained a HeteroOrdinal model constructed using the above eight features, which can be used to predict the progression of liver fibrosis in chronic hepatitis B. The diagnostic significance of the above eight features is shown in Table 3.

[0076] Table 3: Optimal Feature Set for Patients with Liver Fibrosis

[0077] Example 2 Performance Evaluation of the HeteroOrdinal Model In this embodiment, the inventors conducted an in-depth performance comparison analysis of the HeteroOrdinal model and several common machine learning models in the task of predicting the grade of liver fibrosis. The machine learning models used for comparison include XGB Classifier (Gradient Boosting Decision Tree Classifier), Random Forest Classifier, and Logistic Regression.

[0078] Specifically, the inventors evaluated the performance of the four models at two key distinguishing levels: (1) F2+, which distinguishes between no to mild fibrosis [F0-1] and moderate to severe fibrosis [F2, F3-4]). (2) F3+, which further distinguishes between severe fibrosis or cirrhosis [F3-4].

[0079] To comprehensively measure the classification ability of each model, the inventors compared their AUROC values ​​at both levels and calculated the average AUROC for each model on these two key distinctions (see details). Figure 3 ).

[0080] As shown in the figure, the HeteroOrdinal model achieved an AUROC of 0.8173 in predicting F2+ scenarios, surpassing XGB Classifier's 0.7679, Random Forest Classifier's 0.7597, and Logistic Regression's 0.8069. In predicting F3+ scenarios, our model achieved an AUROC of 0.8672, surpassing XGB Classifier's 0.7754, Random Forest Classifier's 0.7334, and Logistic Regression's 0.8374.

[0081] The comparison results of different models show that the HeteroOrdinal model outperforms the other three models in the AUROC index in all test scenarios, proving its superior performance in predicting the grade of liver fibrosis.

[0082] In addition, to further verify the effectiveness of the HeteroOrdinal model, the inventors also compared it with some classic liver fibrosis scoring models (APRI model, FIB-4 model, and GPR model), including: (1) APRI model, liver fibrosis score based on AST and platelet count ratio

[0083] Wherein, AST represents aspartate aminotransferase, unit: IU / L; ULN represents the upper limit of normal for AST; and PLT represents platelet count, unit: 10-1. 9 / L.

[0084] (2) FIB-4 model, liver fibrosis score based on age, ALT, AST and platelet count

[0085] In this context, Age represents age in years; AST and PLT have the same meaning and units as above; ALT represents alanine aminotransferase in IU / L.

[0086] (3) GPR model, liver fibrosis score based on GGT to platelet ratio

[0087] Where GGT represents gamma-glutamyl transferase, unit: IU / L; ULN represents the upper limit of normal for GGT; and PLT represents platelet count, unit: 10-1. 9 / L.

[0088] By plotting the ROC curves of these traditional models and the HeteroOrdinal model in predicting F2+ and F3+ (see details), we can see the results. Figure 4 It can be clearly observed that the HeteroOrdinal model not only significantly outperforms traditional scoring models in terms of AUROC, but also demonstrates excellent sensitivity and specificity at different thresholds. Specifically, when predicting F2+, the HeteroOrdinal model achieved an AUROC of 0.8162, significantly exceeding the APRI score of 0.7597, the FIB-4 score of 0.6937, and the GPR score of 0.7638; when predicting F3+, the HeteroOrdinal model achieved an AUROC of 0.8975, significantly exceeding the FIB-4 score of 0.7801 and the GPR score of 0.8697.

[0089] These results demonstrate that, compared to traditional scoring systems that rely on fixed biomarker ratios, the HeteroOrdinal model can more accurately capture patterns in complex data, thus providing a more accurate and reliable prediction of liver fibrosis grades. This is of great significance for early identification and intervention in clinical practice, helping to optimize patient treatment plans and improve prognosis.

[0090] Finally, the inventors used a calibration plot (also known as a calibration curve, a type of statistical analysis plot) to evaluate the consistency between the probabilities output by the predictive model and the actual observations. The calibration plot demonstrates the model's calibration performance by dividing the model's predicted probabilities into several intervals and calculating the proportion of actual events occurring within each interval. Ideally, if the model is perfectly calibrated, these points should fall on the diagonal (y=x line), indicating that the model's predicted probabilities accurately reflect the actual probabilities. The Brier Score in the plot measures the error between the model's predicted probabilities and the actual observations. The lower the score, the better the model calibration, meaning the model not only has high accuracy but also provides reliable probability estimates.

[0091] pass Figure 5 The calibration plots show that the predictions of the HeteroOrdinal model are in good agreement with the actual outcomes of patients clinically diagnosed with F2+ stage liver fibrosis via biopsy. Specifically, the actual event rates within each probability interval are close to the probabilities predicted by the model, indicating that the model accurately reflects the patients' true condition. Furthermore, the lower Brier Score further confirms this, showing that the HeteroOrdinal model outperforms other comparative models in terms of calibration performance.

[0092] Therefore, the HeteroOrdinal model can not only effectively distinguish between different stages of liver fibrosis, but also provide accurate probability estimates, meeting the clinical need for non-invasive diagnosis of early-stage liver fibrosis. This is of great significance for optimizing treatment plans and improving patient prognosis. This non-invasive method can reduce unnecessary liver biopsies, lower patient risks and discomfort, while ensuring diagnostic efficiency and accuracy.

[0093] Example 3 Independent Validation of the Model The inventors collected and compiled clinical data from patients with chronic hepatitis B who visited the Affiliated People's Hospital of Ningbo University between 2020 and March 2025 as an external independent validation dataset, with inclusion and exclusion criteria as before. This dataset covers samples of different severity levels, specifically including 61 cases in stage F0-1 (i.e., no fibrosis or mild fibrosis), 37 cases in stage F2 (i.e., moderate fibrosis), and 18 cases in stage F3-4 (i.e., severe to end-stage fibrosis), totaling 116 samples. The validation data is not only representative in quantity but also fully considers the diversity of the disease in terms of sample distribution.

[0094] Based on this, the inventors re-evaluated the model constructed in Example 1, and used the APRI model, FIB-4 model, and GPR model for comparison. The evaluation results are shown in Table 4.

[0095] Table 4: Model Evaluation Results on External Independent Validation Sets

[0096] As shown in Table 4, the HeteroOrdinal model achieved an AUROC of 0.7955 for predicting stages F2 and above (F2+), and an even higher AUROC of 0.8554 for predicting stages F3 and above (F3+). Notably, these performances significantly outperform three classic comparative scoring models: the APRI model, the FIB-4 model, and the GPR model. These classic models represent several different technical approaches widely used in the field. Comparison with these models further highlights the superior performance of the proposed model in identifying disease progression.

[0097] In conclusion, the results of this evaluation strongly demonstrate that the HeteroOrdinal model possesses higher reliability and accuracy.

[0098] Example 4: An electronic device and a readable storage medium Figure 6 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0099] like Figure 6 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0100] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0101] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as methods for constructing models for predicting the progression of chronic hepatitis B cirrhosis or methods for predicting the progression of chronic hepatitis B cirrhosis. For example, in some embodiments, the methods for constructing models for predicting the progression of chronic hepatitis B cirrhosis or methods for predicting the progression of chronic hepatitis B cirrhosis can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by computing unit 801, one or more steps of the method for constructing a model for predicting the progression of chronic hepatitis B cirrhosis or the method for predicting the progression of chronic hepatitis B cirrhosis described above can be performed. Alternatively, in other embodiments, computing unit 801 can be configured by any other suitable means (e.g., by means of firmware) to perform the method for constructing a model for predicting the progression of chronic hepatitis B cirrhosis or the method for predicting the progression of chronic hepatitis B cirrhosis.

[0102] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0103] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0104] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0105] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0106] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0107] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0108] Furthermore, it should be understood that after reading the foregoing teachings of this application, those skilled in the art can make various alterations or modifications to this application, and these equivalent forms also fall within the scope defined by the appended claims.

Claims

1. A method for constructing a model to predict the progression of chronic hepatitis B-related cirrhosis, characterized in that, Includes the following steps: S101, obtain quantitative data on the combination of serum biomarkers in the population. The group includes multiple groups of patients with chronic hepatitis B cirrhosis with varying degrees of liver fibrosis, the degree of which is classified according to the METAVIR scoring system; The serum biomarker combination includes PLT, GGT, FPG, GLB, AFP, IBIL, TBIL, and HBeAg; S102, using the quantitative data of the serum biomarker combination of the population obtained in step S101, an ordered logistic regression model is constructed to predict the progression of chronic hepatitis B cirrhosis.

2. The method for constructing a model for predicting the progression of chronic hepatitis B cirrhosis according to claim 1, characterized in that, The multiple groups are divided into three groups: (1) patients in F0 and F1 stages, (2) patients in F2 stages, and (3) patients in F3 and F4 stages.

3. A method for constructing a model for predicting the progression of chronic hepatitis B cirrhosis according to claim 1 or 2, characterized in that, In the ordered logistic regression model, the logarithm of the variance, lnσ², is predicted using a multilayer perceptron. x ): lnσ 2 ( x )=MLPϕ( x ) in, x This refers to serum biomarkers, specifically the MLPϕ( x ) represents the output of the multilayer perceptron. However, a positive variance prediction value can be obtained using the following formula: s 2 ( x )=exp(MLPϕ( x )) Then use the following formula for prediction: in, It is a threshold. It is a prediction function.

4. The method for constructing a model for predicting the progression of chronic hepatitis B cirrhosis according to claim 3, characterized in that, The activation function of the multilayer perceptron is ReLU.

5. An ordered logistic regression model for predicting the progression of chronic hepatitis B-related cirrhosis, characterized in that, It is constructed using the method described in any one of claims 1 to 4.

6. A method for predicting the progression of chronic hepatitis B-related cirrhosis, characterized in that, Includes the following steps: S1, obtain quantitative data of the serum biomarker combination of the patient to be predicted, the serum biomarker combination including PLT, GGT, FPG, GLB, AFP, IBIL, TBIL and HBeAg; S2, the quantitative data of the serum biomarker combination of the patient to be predicted is input into the ordered logistic regression model of claim 5 to obtain the classification result of the degree of liver fibrosis, thereby obtaining the progression of liver cirrhosis.

7. A system for predicting the progression of chronic hepatitis B-related cirrhosis, characterized in that, The system includes: The data input module is used to obtain quantitative data on the combination of serum biomarkers in the patient to be predicted; The model storage module is used to store the ordered logistic regression model as described in claim 5; The prediction model is connected to the data input module and the model storage module respectively. It inputs the quantitative data of the serum biomarker combination of the patient to be predicted into the ordered logistic regression model to obtain the classification result of the degree of liver fibrosis, thereby obtaining the progression of liver cirrhosis.

8. The system for predicting the progression of chronic hepatitis B cirrhosis according to claim 7, characterized in that, The model storage module includes a data storage unit and a model building unit. The data storage unit is used to store quantitative data of the serum biomarker combination of the population, and the model building unit is used to construct an ordered logistic regression model using the quantitative data of the serum biomarker combination of the population.

9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-4 or the method of claim 6.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-4 or the method according to claim 6.