A method, system and device for constructing an esophagogastric varices prediction model

By constructing a machine learning-based XGB model, and utilizing the basic personal information and test data of cirrhosis patients, the problem of insufficient accuracy in non-invasive assessment of esophageal and gastric varices risk in existing technologies has been solved, achieving efficient and accurate non-invasive prediction and supporting clinical decision-making.

CN119581033BActive Publication Date: 2025-11-25PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510131158.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-11-25
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

Current technology lacks a unified, non-invasive, and accurate method to predict the risk of esophageal and gastric varices in patients with cirrhosis, and gastroscopy is invasive and costly.

Method used

We employed machine learning-based methods, particularly the XGB model, to construct a predictive model using basic personal information and test data from cirrhosis patients, such as age, spleen thickness, platelets, albumin, and cholinesterase. The model was then optimized through training, validation, and testing of sample sets to improve predictive accuracy.

Benefits of technology

This provides a non-invasive and accurate method to assess the risk of esophageal and gastric varices in patients with cirrhosis, improving the accuracy and reliability of predictive models and helping clinicians make better clinical decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119581033B_ABST
    Figure CN119581033B_ABST
Patent Text Reader

Abstract

The application discloses a kind of for esophagus stomach varicosity prediction model construction method, system and equipment.The present application establishes a kind of prediction model of subject esophagus stomach varicosity, it can predict the occurrence of esophagus stomach varicosity of cirrhotic patient, with higher prediction ability.In addition, the present application proves the feasibility of the prediction model in clinical evaluation of esophagus stomach varicosity of cirrhotic patient by clinical decision curve.The model constructed by the present application has higher accuracy, and can be used as a non-invasive tool for evaluating esophagus stomach varicosity of cirrhotic patient.The prediction model of the present application is expected to play an important role in evaluating the degree of esophagus stomach varicosity of cirrhotic patient, and can help clinicians make clinical decisions.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence disease prediction, in particular to a method and system for constructing an esophagogastric varices prediction model. BACKGROUND

[0002] Liver cirrhosis is a common chronic liver disease caused by viral infection, alcoholism, non-alcoholic fatty liver, biliary disease, autoimmune disease or genetic factors, etc. Due to the long-term effect of different factors, extensive necrosis of liver cells occurs, on this basis, diffuse proliferation of liver fibrous tissue occurs, forming nodules, false lobules, and further destroying the normal structure and blood supply of the liver. Liver cirrhosis is the 11th most common cause of death worldwide, with about 1.16 million people dying from liver cirrhosis each year, of which 1 million people die from complications of liver cirrhosis.

[0003] Portal hypertension is a major consequence of liver cirrhosis and is the cause of most complications, including ascites, esophagogastric varices and their rupture and bleeding, and encephalopathy. Esophagogastric varices (GOV) are caused by increased portal vein pressure in liver cirrhosis, leading to anastomosis of the gastric coronary vein with the esophageal vein and azygos vein system, causing retrograde blood flow and dilation of the esophageal and gastric veins. Esophagogastric variceal bleeding (EVB) is one of the most common fatal complications of liver cirrhosis GOV. Nearly 50% of patients with liver cirrhosis can develop GOV, and the degree of liver damage is directly proportional to the degree of GOV. About 40% of Child-Pugh grade A and 85% of grade C liver cirrhosis patients have GOV, and the degree of varices will increase at a rate of 10%-15% per year. The mortality rate of the first bleeding caused by GOV is 30%-50%, and early detection and treatment of esophagogastric varices in clinical practice can reduce the risk of esophageal variceal bleeding by 5%-15%. Therefore, early identification and early detection are of great significance to improve the quality of life of patients with liver cirrhosis.

[0004] Currently, gastroscopy is the gold standard for evaluating the presence of GOV and bleeding in patients with cirrhosis. Guidelines recommend that patients with cirrhosis without GOV undergo gastroscopy every 2 years, and patients with mild varices undergo gastroscopy every year. However, gastroscopy is an invasive procedure that can induce bleeding risk, cause patient discomfort, and result in poor compliance and high costs. Therefore, the challenge is to seek non-invasive, accurate and feasible methods to predict the risk of GOV. With the development of technology, detection methods have also gradually developed. Spleen length, width and thickness can be measured by ultrasound. A meta-analysis showed that the platelet count to spleen diameter ratio (PSR) can be used as a non-invasive diagnostic for GOV, with high diagnostic efficiency. The AUC values for diagnosing varices and high-risk varices were 0.872 and 0.813, respectively, with sensitivities of 84.0% and 78.0%, and specificities of 78.0% and 67.0%, respectively. CT is also a widely used non-invasive examination method. Yu-Jen Tseng et al. showed through a meta-analysis that, based on 11 studies, the AUC value for diagnosing GOV by CT was 0.860, with sensitivities and specificities of 89.6% and 72.3%, respectively. Non-invasive predictors such as liver stiffness and spleen stiffness measured by transient elastography have high predictive value. Pons, Monica MD, et al. found that liver stiffness < 15 kPa and platelet count ≥ 150 × 10 9 / L can exclude clinically significant portal hypertension from most etiologies. The Baveno VII consensus states that patients with cACLD have liver stiffness < 20 kPa and platelet count > 150 × 10 9 / L can be exempted from GOV screening by internal diameter, or spleen stiffness ≤ 40 kPa can identify the risk of high-risk GOV. Although there are many indicators to explore GOV prediction, there is no unified non-invasive diagnostic standard.

[0005] Machine learning (ML) is based on existing data to select algorithms and build models based on algorithms and data to predict the occurrence of events. More and more research is applying ML technology to facilitate the prediction of GOV, other complications of cirrhosis and other disease progression. For example, Tien S Dong et al. established an EVendo scoring model using random forests to screen high-risk GOV. The EVendo scoring model mainly includes international normalized ratio, aspartate aminotransferase, platelets, urea nitrogen, hemoglobin and the presence or absence of ascites. The AUC values for mild GOV screening in the training and validation cohorts were 0.84 and 0.82, respectively, and the AUC values for high-risk GOV screening in the two cohorts were 0.74 and 0.75, respectively. ML-based models have become valuable tools for implementation in clinical practice.

[0006] The information in the background section is merely intended to illustrate the general background of the invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art. Summary of the Invention

[0007] This invention aims to establish a machine learning-based diagnostic model for assessing GOV in patients with cirrhosis, providing a simple, convenient, and non-invasive predictive method for GOV patients. Specifically, this invention includes the following:

[0008] A first aspect of the present invention provides a method for constructing a predictive model based on machine learning, comprising the following steps:

[0009] (1) Provide a training sample set containing multiple samples, each of which includes the subject's basic personal information and detection information data and labels, wherein the subjects include patients with cirrhosis, the basic personal information and detection data are obtained by introducing a penalty coefficient into the regression model, and the labels are esophageal and gastric varices.

[0010] (2) Input the personal basic information and detection data into the XGB machine learning model for training to obtain a predictive model for esophageal and gastric varices in patients with cirrhosis.

[0011] In some implementations, according to the method for constructing a prediction model based on machine learning according to the present invention, the number of samples in the training sample set is 50 or more.

[0012] In some implementations, the method for building a predictive model based on machine learning according to the present invention further includes providing a validation sample set and using the validation sample set for validation, thereby adjusting or optimizing the predictive model based on the validation results.

[0013] In some embodiments, the method for building a predictive model based on machine learning according to the present invention further includes providing a test sample set and testing the predictive model using the test sample set.

[0014] In some implementations, according to the method for constructing a predictive model based on machine learning according to the present invention, the personal basic information and detection data include parameter data selected from age, spleen thickness, platelets, albumin, prealbumin and cholinesterase.

[0015] A second aspect of the present invention provides a system for constructing predictive models based on machine learning, comprising:

[0016] The data acquisition unit is configured to acquire data from a training sample set and optional validation and test sample sets. The training sample set includes a training sample set of multiple samples, and each sample contains the subject's basic personal information, detection data, and labels. The subjects include patients with cirrhosis. The basic personal information and detection data are obtained by introducing a penalty coefficient into a regression model for screening. The labels are esophageal and gastric varices.

[0017] The building unit includes an XGB machine learning model and is configured to input the personal basic information and detection data from the training sample set into the XGB machine learning model for training. Optionally, it is further configured to use the validation sample set to validate the trained prediction model, thereby adjusting or optimizing the prediction model based on the validation results. Optionally, it further includes using a test sample set to test, thereby evaluating the prediction model.

[0018] A third aspect of the present invention provides a method for predicting esophageal and gastric varices in a subject, wherein the method includes the step of obtaining the subject's basic personal information and test data and inputting them into a prediction model to obtain a predicted probability, wherein the subject is a patient with cirrhosis, and the basic personal information and test data include parameters such as age, spleen thickness, platelet count, albumin, prealbumin, and cholinesterase, and the prediction model is obtained according to the method described in the present invention.

[0019] A fourth aspect of the present invention provides a system for predicting esophageal and gastric varices in a subject, comprising:

[0020] The data acquisition unit is configured to acquire the basic personal information and test data of the subjects, including patients with cirrhosis. The basic personal information and test data include parameters such as age, spleen thickness, platelet count, albumin, prealbumin, and cholinesterase.

[0021] A prediction unit is configured to input the data acquired by the data acquisition unit into a prediction model and output a prediction result, wherein the prediction model is obtained according to the method described in this invention.

[0022] A fifth aspect of the present invention provides an apparatus comprising: a memory, a processor, and a program stored in the memory and capable of running on the processor, the program being configured to implement the steps of the method or prediction method for constructing a prediction model according to the present invention.

[0023] A sixth aspect of the present invention provides a computer-readable storage medium, wherein a program is stored on the computer-readable storage medium, and when the program is executed by a processor, it implements the steps of the method for constructing a prediction model or the prediction method described in the present invention.

[0024] This invention establishes a novel predictive model for esophageal and gastric varices in subjects, which can predict the occurrence of GOV in patients with cirrhosis and has high predictive ability. The feasibility of this predictive model in clinical assessment of GOV was confirmed using clinical decision curves (DCA). The model constructed in this invention has high accuracy and can serve as a non-invasive tool for assessing GOV in patients with cirrhosis. Therefore, the predictive model of this invention is expected to play an important role in assessing the degree of esophageal and gastric varices in patients with cirrhosis, and can assist clinicians in making clinical decisions. Attached Figure Description

[0025] Figure 1 Data processing workflow for patients with cirrhosis.

[0026] Figure 2 Distribution of causes of cirrhosis.

[0027] Figure 3 This is a path diagram of the regression coefficients.

[0028] Figure 4 This is the Lasso regression cross-validation curve.

[0029] Figure 5 This section compares variables selected by Lasso regression across different GOV (Government of Volunteer) classifications. A. Comparison of Age across different GOV classifications; B. Comparison of ST (Spleen Thickness) across different GOV classifications; C. Comparison of PLT (Platelet-Rich Plasma) across different GOV classifications; D. Comparison of ALB (Albumin) across different GOV classifications; E. Comparison of PA (Prealbumin) across different GOV classifications; F. Comparison of CHE (Chlorolinesterase) across different GOV classifications. Age: Age; ST: Spleen Thickness; PLT: Platelet-Rich Plasma; ALB: Albumin; PA: Prealbumin; CHE: Cholinesterase.

[0030] Figure 6 ROC curves for different machine learning models on the test set. A: Training set; B: Test set; C: External validation. RF: Random Forest; SVM: Support Vector Machine; XGB: XGBoost; GNB: Gaussian Naive Bayes; KNN: K-Nearest Neighbors.

[0031] Figure 7ROC curves were compared between XGB and LSM, PLT, and a combination of both. A. Test set; B. External validation. XGBtest: XGB model in the test set; XGBval: XGB model in the external validation set; LSM: Liver stiffness measurement >20 kPa; PLT: Platelet count >150 × 10^9; LP: Liver stiffness measurement >20 kPa and platelet count >150 × 10^9.

[0032] Figure 8 This is the clinical decision curve for the XGB model. A: Test set; B: External validation. Detailed Implementation

[0033] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as a limitation of the present invention, but rather as a more detailed description of certain aspects, features, and embodiments of the present invention.

[0034] It should be understood that the terminology used in this invention is merely for describing particular embodiments and is not intended to limit the invention. Furthermore, with respect to numerical ranges in this invention, it should be understood that the upper and lower limits of the range and each intermediate value between them are specifically disclosed. Any stated value or intermediate value within a stated range, as well as each smaller range between any other stated value or intermediate value within said range, are also included in this invention. The upper and lower limits of these smaller ranges may be independently included or excluded from the range.

[0035] Unless otherwise stated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. While only preferred methods and materials have been described herein, any methods and materials similar or equivalent to those described herein may be used in the implementation or testing of this invention. All references to this specification are incorporated by way of citation to disclose and describe methods and / or materials associated with those references. In the event of any conflict with any incorporated reference, the content of this specification shall prevail.

[0036] Method of constructing a predictive model based on machine learning

[0037] One aspect of the present invention provides a method for constructing a predictive model based on machine learning, comprising the following steps:

[0038] (1) Provide a training sample set containing multiple samples, each of which includes the subject's basic personal information, detection data, and labels, wherein the subjects include patients with cirrhosis, the basic personal information and detection data are obtained by introducing a penalty coefficient into the regression model, and the labels are esophageal and gastric varices.

[0039] (2) Input the personal basic information and detection data into the XGB machine learning model for training to obtain a predictive model for esophageal and gastric varices in patients with cirrhosis.

[0040] In step (1) of this invention, a training sample set containing multiple samples is first provided. The sample type is not particularly limited and can be a fluid sample (blood sample) or a tissue sample from the subject. The subject includes patients with cirrhosis. The personal basic information and test data are obtained by introducing a penalty coefficient into the regression model. The label is esophageal and gastric varices. In a preferred embodiment, the personal basic information and test data include or consist of the following variables: age, spleen thickness (ST), platelet count (PLT), albumin (ALB), prealbumin (PA), and cholinesterase (CHE). It should be noted that the above variables are not arbitrarily selected. Through extensive research, the inventors used 30 influencing factors with intergroup differences in GOV as independent variables. By introducing a penalty coefficient λ, overfitting can be avoided. Furthermore, ten-fold cross-validation was used to select the optimal penalty coefficient λ. In one specific embodiment, λ is 0.004. In another specific embodiment, λ is 0.049.

[0041] In step (1) of the present invention, the number of samples in the training sample set is not particularly limited, and can be as large as possible to ensure the prediction or diagnostic performance of the final prediction model. For example, the number of samples can be more than 50, more than 100, more than 500, or more than 1000.

[0042] In this invention, the samples are cirrhosis patients from the Chinese population. Therefore, in some embodiments, the construction method of this invention is a method for constructing a predictive model of esophageal and gastric varices based on machine learning, applicable to the Chinese subject population.

[0043] In step (2) of this invention, the personal basic information and detection data are input into the XGB machine learning model for training to obtain a predictive model for esophageal and gastric varices in patients with cirrhosis. The machine learning model in the construction method of this invention is not randomly selected. The final selected XGB machine learning model is suitable for the six variables selected in step (1) and has significantly improved diagnostic or predictive efficacy compared to other commonly used machine learning models (including but not limited to Random Forest (RF), Support Vector Machine (SVM), Gradient Boosting Classification (GBC), Gaussian Naive Bayes (GNB), K-Nearest Neighbor (KNN), etc.). Predictive efficacy indicators include the AUC, sensitivity, and specificity of the constructed model.

[0044] The method for constructing a prediction model based on machine learning according to the present invention further includes providing a validation sample set and using the validation sample set for validation, thereby adjusting or optimizing the prediction model based on the validation results.

[0045] The method for constructing a prediction model based on machine learning according to the present invention further includes providing a test sample set and using the test sample set to test, thereby evaluating the prediction model.

[0046] System of constructing a predictive model based on machine learning

[0047] In one aspect, the present invention provides a predictive model for esophageal and gastric varices constructed by the above method.

[0048] Another aspect of the present invention provides a system for constructing predictive models based on machine learning (also referred to as an "apparatus for constructing predictive models based on machine learning"), comprising:

[0049] The data acquisition unit is configured to acquire data from a training sample set and optional validation and test sample sets. The training sample set includes a training sample set of multiple samples, and each sample contains the subject's basic personal information, detection data, and labels. The subjects include patients with cirrhosis. The basic personal information and detection data are obtained by introducing a penalty coefficient into a regression model. The labels are esophageal and gastric varices.

[0050] The building unit includes an XGB machine learning model and is configured to input the personal basic information and detection data from the training sample set into the XGB machine learning model for training. Optionally, it is further configured to use the validation sample set to validate the trained prediction model, thereby adjusting or optimizing the prediction model based on the validation results. Optionally, it further includes using a test sample set to test, thereby evaluating the prediction model.

[0051] In the machine learning-based predictive model system of the present invention, the data acquisition unit can be any suitable device or instrument capable of collecting parameter data of the subject selected from age, spleen thickness, platelets, albumin, prealbumin, and cholinesterase. For example, it could be an electronic information storage unit containing basic patient information, medical records, follow-up data, etc., or a reagent kit or instrument capable of measuring or quantifying spleen thickness, platelets, albumin, prealbumin, and cholinesterase. An exemplary data acquisition unit includes a multi-mode input port. The multi-mode input port can receive information data from various sources, including the subject's age, spleen thickness, platelets, albumin, prealbumin, and cholinesterase levels, including data from online and offline sources. Another exemplary data acquisition unit includes a port for connection to a serological marker measuring instrument (e.g., a chemiluminescence immunoassay analyzer).

[0052] In the system for building a prediction model based on machine learning according to the present invention, the building unit includes an XGB machine learning model and is configured to input the personal basic information and detection data in the training sample set into the XGB machine learning model for training. Optionally, it is further configured to use the validation sample set to validate the trained prediction model, thereby adjusting or optimizing the prediction model according to the validation results. Optionally, it further includes using a test sample set to test, thereby evaluating the prediction model.

[0053] In the system for constructing predictive models based on machine learning according to the present invention, the form of the construction unit is not limited, as long as it includes the predictive model of the present invention. For example, the construction unit is a processor. Preferably, the construction unit is communicatively connected to the data acquisition unit, thereby enabling it to retrieve data from the data acquisition unit and input it into the construction model for computation.

[0054] System for predicting esophagogastric varices in a subject

[0055] The present invention further provides a system (or device) for predicting esophageal and gastric varices in a subject, comprising:

[0056] The data acquisition unit is configured to acquire the basic personal information and test data of the subjects, including patients with cirrhosis. The basic personal information and test data include parameters such as age, spleen thickness, platelets, albumin, prealbumin, and cholinesterase.

[0057] A prediction unit is configured to input the data acquired by the data acquisition unit into a prediction model and output a prediction result, wherein the prediction model is obtained according to the method described in this invention.

[0058] In the system for predicting esophageal and gastric varices in a subject according to the present invention, the data acquisition unit can be any suitable device or instrument capable of acquiring parameter data of the subject selected from age, spleen thickness, platelets, albumin, prealbumin, and cholinesterase. This could be an electronic information storage unit containing basic patient information, medical records, follow-up data, etc., or a reagent kit or instrument capable of measuring or quantifying spleen thickness, platelets, albumin, prealbumin, and cholinesterase. An exemplary data acquisition unit includes a multi-mode input port. The multi-mode input port can receive information data from various sources, including the subject's age, spleen thickness, platelets, albumin, prealbumin, and cholinesterase levels, including data from online and offline sources. Another exemplary data acquisition unit includes a port for connection to a serological marker measuring instrument (e.g., a chemiluminescence immunoassay analyzer).

[0059] In the system for predicting esophageal and gastric varices of the present invention, the prediction unit includes an XGB prediction model obtained by the construction method of the present invention, which is configured to input the data acquired by the data acquisition unit into the prediction model and output the prediction result.

[0060] In the system for predicting esophageal and gastric varices of the present invention, the form of the prediction unit is not limited, as long as it includes the prediction model of the present invention. Exemplarily, the prediction unit is a processor. Preferably, the prediction unit is communicatively connected to the data acquisition unit, thereby enabling it to retrieve data from the data acquisition unit and input it into the prediction model for calculation.

[0061] Apparatus and computer-readable storage medium

[0062] The present invention further provides an apparatus comprising: a memory, a processor, and a program stored in the memory and capable of running on the processor, the program being configured to implement a method for constructing a prediction model or steps of a prediction method.

[0063] The present invention also provides a computer storage medium (readable storage medium) or cloud for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this invention, a "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples of computer-readable media include: an electrical connection (electronic device) having one or more wires, a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0064] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0065] Those skilled in the art will understand that all or part of the steps of the method implementing the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0066] Furthermore, the functional units in the various embodiments of this invention can be integrated into a single processing module, or each unit can exist physically separately, or two or more units can be integrated into a single module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The aforementioned storage medium can be a read-only memory, a hard disk, or an optical disk, etc.

[0067] Example

[0068] 1. Research Subjects and Methods

[0069] 1.1 Research Subjects

[0070] This embodiment includes 578 patients with cirrhosis from the Fifth Medical Center of the PLA General Hospital between 2006 and 2015, who meet the criteria of the "Guidelines for the Diagnosis and Treatment of Cirrhosis (2019)" and have a diagnosis of GOV. Patients from the Fifth Medical Center between May and October 2024 were also collected for external validation.

[0071] Exclusion criteria: 1. Cases without endoscopic examination: CT, MRI, blank results (endoscopy is the gold standard for GOV) (53 cases); 2. Cases where varices disappeared after treatment such as banding, Tips, or sclerotherapy (6 cases); 3. Duplicate cases: case numbers were duplicated (5 cases); 4. Cases with unclear definitions: esophageal varices showing a tendency to dilate and unable to be graded, or gastric varices without degree of dilation and unable to be graded (46 cases). Finally, 468 patients with cirrhosis remained, including those with hepatitis B, hepatitis C, alcoholic, primary biliary, drug-induced, and autoimmune cirrhosis. The data processing flowchart and the distribution of cirrhosis etiologies are shown below. Figure 1 As shown.

[0072] 1.2 Endoscopic examination

[0073] GOV grading criteria: ① No varicose veins; ② Mild: GOV is slightly tortuous and has no red signs; ③ Moderate: GOV is slightly tortuous and has red signs; or it is serpentine and tortuous; ④ Severe: GOV is serpentine and tortuous and has red signs; or it is beaded, nodular, or tumor-like (with or without red signs).

[0074] GOV grouping: Cirrhosis patients were divided into GOV group, including mild, moderate and severe GOV (360 cases) and no GOV group (108 cases). The external validation group included 71 GOV patients and 40 no GOV patients.

[0075] 1.3 Machine Learning Model Selection

[0076] This embodiment first selects six models, including Random Forest (RF), Support Vector Machine (SVM), Gradient Boosting Classification (GBC), eXtreme Gradient Boosting (XGB), Gaussian Naive Bayes (GNB), and K-Nearest Neighbor (KNN), to predict cirrhosis (GOV). RF is an ensemble learning algorithm that improves prediction accuracy by constructing multiple decision trees and combining them. First, a bootstrap sampling method is used to randomly sample from the training set to form the training set of the trees. After obtaining the Random Forest, a new sample is input, and each tree makes its own judgment and votes to determine the category to which the sample belongs. If a certain category is predicted by most trees, then that category is designated as the final predicted category for the sample. The core idea of ​​SVM is to find the maximum distance between the lines or hyperplanes containing the two nearest points of variables of different categories, thereby separating different data points and solving problems such as classification and regression. GBC is an ensemble learning method that combines multiple weak predictors (usually decision trees) to build a strong predictor. Each subsequent predictor attempts to correct the errors of the previous predictor, thus progressively improving the model's accuracy and generalization ability. XGB is a machine learning framework based on gradient boosting decision tree algorithms, aiming to create a strong learner by integrating a large number of weak learners (such as decision trees). XGB's objective function includes a loss function and a regularization term (L2 regularization). The loss function measures the accuracy of the model's predictions, while the regularization term controls the model's complexity and prevents overfitting. XGB uses a Taylor series to expand the loss function to second order for more accurate estimation of the objective function and more efficient optimization. GNB is a machine learning model based on probability and Gaussian distribution. It assumes that features are independent and that they conform to a Gaussian distribution, meaning that the probability density function of a continuous random variable follows a normal distribution with mean μ and variance δ². GNB calculates the posterior probability by multiplying the prior probability and conditional probability of each feature, thus obtaining the probability of each class, and finally selects the class with the highest probability as the predicted class. The basic principle of KNN is: when predicting a new variable, the category of the new variable is determined by the category of the K nearest neighbors. In regression analysis, the KNN algorithm typically takes the average of the K nearest samples as the output.

[0077] 1.4 Data Statistical Analysis

[0078] Patient data were presented as either continuous or categorical variables. The Kolmogorov-Smirnov test was used to assess whether the data followed a normal distribution. For normally distributed continuous variables, data were described as mean ± standard deviation and compared using a t-test. If continuous variables did not conform to a normal distribution, the Mann-Whitney U test was used, and results were expressed as median (interquartile range). Categorical data were presented as numbers and frequencies and compared using a chi-square test or Fisher's exact test. All statistical tests were two-tailed, and a p-value < 0.05 was considered statistically significant. Analysis was performed using SPSS version 25.0 (IBM SPSS Statistics 25).

[0079] Thirty variables were selected as the total dataset, with the presence or absence of GOV (Goal, Volume, and Value) as the dependent variable. Independent variables included: general characteristics (gender, age, BMI); liver function indicators (alanine aminotransferase, aspartate aminotransferase, alkaline phosphatase, gamma-glutamyl transferase, total bile acids, cholinesterase, albumin, globulin, prealbumin, direct bilirubin, total bilirubin, Golgi protein, alpha-fetoprotein, alpha-fetoprotein isoforms); coagulation indicators (prothrombin time, prothrombin time activity, international normalized ratio of prothrombin time, fibrinogen); complete blood count indicators (red blood cells, hemoglobin, white blood cells, neutrophils, lymphocytes, monocytes, platelets); and imaging indicators (portal vein diameter, spleen thickness). LASSO regression was first used to screen variables. 20% of the selected data was randomly selected as the test set. The remaining data was fed into a machine learning model, and the trained model was evaluated using 5-fold cross-validation and adjusted to appropriate parameters. Finally, the test set data was fed into the adjusted model for testing. K-fold cross-validation splits the data K times, selecting a portion as the validation set in each split and using the remaining portion as the training set, thus obtaining K evaluation results for the validation sets. Each portion can serve as a validation, reducing the error caused by the randomness of the dataset and improving the generalization ability of the model.

[0080] ROC curves were used to compare the models, and sensitivity, specificity, positive predictive values ​​(PPV), and negative predictive values ​​(NPV) were used to evaluate the model performance. Decision Curve Analysis (DCA) was used to compare the benefits of different predictive models at different treatment decision thresholds.

[0081] 2. Results

[0082] 2.1 Comparison of patients' basic information

[0083] Of the 468 patients with cirrhosis, 360 had GOV (Government-Important Viral Disease) (including 118 with mild, 118 with moderate, and 124 with severe GOV), of whom 257 were male (79.10%) and 103 were female (72.00%). The remaining 108 patients did not have GOV; of these, 68 were male (20.90%) and 40 were female (28.00%). The average age of patients with GOV was 52 years, while the average age of patients without GOV was 48 years; this difference was statistically significant (P < 0.05). Univariate analysis showed that there were no statistically significant differences between the two groups in terms of gender, BMI, AFPL3, MONO, and ALT (P > 0.05). However, there were statistically significant differences between the two groups in the following indicators: age, AST, ALP, GGT, TBA, CHE, ALB, GLB, PA, DBiL, TBiL, GP73, AFP, PT, PR, INR, FIB, RBC, HGB, WBC, NEUT, LY, PLT, PVD, and ST (P < 0.05), as shown in Table 1.

[0084] Table 1 Comparison of basic patient information

[0085]

[0086] Note: *The difference is statistically significant at the α=0.05 level.

[0087] ALT: Alanine aminotransferase; AST: Aspartate aminotransferase; ALP: Alkaline phosphatase; GGT: Gamma-glutamyl transferase; TBA: Total bile acids; CHE: Cholinesterase; ALB: Albumin; GLB: Globulin; PA: Prealbumin; DBiL: Direct bilirubin; TBiL: Total bilirubin; PT: Prothrombin time; PR: Prothrombin time activity; INR: International Normalized Ratio of Prothrombin Time; FIB: Fibrinogen; RBC: Red blood cells; HGB: Hemoglobin; WBC: White blood cells; NEUT: Neutrophils; LY: Lymphocytes; MONO: Monocytes; PLT: Platelets; GP73: Golgi protein; AFP: Alpha-fetoprotein; AFP-L3: Alpha-fetoprotein isoform; PVD: Portal vein diameter; ST: Spleen thickness.

[0088] 2.2 LASSO Regression Variable Selection

[0089] With the presence or absence of the government (GOV) as the dependent variable, a total of 30 independent variables were included. LASSO regression was performed on all independent variables. Figure 3It can be seen that as the penalty coefficient λ increases, the coefficients of the included independent variables are compressed, and some independent variable coefficients are compressed to 0, thus avoiding overfitting. Ten-fold cross-validation is used to select the optimal penalty coefficient λ. lambda.min refers to the λ value when the bias is minimized (0.004), and lambda.1se is the λ value (0.049) obtained within a variance range of lambda.min, representing the simplest model. Figure 4 In this embodiment, the lambda.1se value was selected, and the final selected variables were age, spleen thickness (ST), platelet count (PLT), albumin (ALB), prealbumin (PA), and cholinesterase (CHE).

[0090] 2.3 Comparison of the selected variables with and without GOV and in four different levels of GOV

[0091] In this study, patients with GOV (52 years old) had a higher age (52.00 mm) and shorter streak (ST) than those without GOV (48 years old and 40.00 mm). Patients with GOV had lower platelet counts (60.00 × 10^9 / L), alanine aminotransferase (31.00 g / L), acetylcholine (PA) (91.50 mg / L), and chemoluminescence immunoassay (CHE) (3157.00 U / L) than those without GOV (101.00 × 10^9 / L, 40.00 g / L, 167.00 mg / L, and 6125.00 U / L, respectively) (Table 1). The age of patients without GOV (48 years old) was lower than that of patients with mild (52 years old), moderate (51 years old), and severe (52 years old) GOV (P < 0.05). Figure 5 (See Figure A). Pairwise comparisons of ST lengths across the four groups showed statistically significant differences (P < 0.05). The ST length was smallest in patients without GOV (40.00 mm), and increased progressively from the GOV-free group to the severe GOV group (40.00 mm, 48.00 mm, 53.00 mm, and 54.50 mm, respectively). Figure 5 (Figure B). The platelet (PLT) level in patients without GOV (101.00 × 10^9 / L) was higher than that in patients with mild (67.50 × 10^9 / L), moderate (59.00 × 10^9 / L), and severe (54.00 × 10^9 / L) GOV. The PLT level in patients with mild GOV was higher than that in patients with severe GOV (P < 0.05). Figure 5(Figure C). The levels of ALB (40.00 g / L), PA (167.00 mg / L), and CHE (6125.00 U / L) in patients without GOV were significantly higher than those in the other three groups (mild GOV: 31.00 g / L, 99.00 mg / L, 3306.00 U / L; moderate GOV: 32.00 g / L, 86.50 mg / L, 2999.00 U / L; severe GOV: 31.00 g / L, 92.00 mg / L, 3194.50 U / L) (P < 0.05). Figure 5 (D, E, F diagrams).

[0092] 2.4 Comparison of Machine Learning Results

[0093] The results of the six selected variables were fed into six machine learning models. The AUC values ​​of the six models in the training set ranged from 0.8791 to 0.9776. Figure 6 (Figure A), sensitivity range: 58.70%-96.09%, specificity range: 75.71%-100.00%, positive predictive value range: 92.50%-100.00%, negative predictive value range: 42.40%-87.70% (Table 2). In the test set, the AUC values ​​of the six models ranged from 0.7414 to 0.9423 (…). Figure 6 (Figure B), sensitivity range: 74.29%-98.57%, specificity range: 45.83%-100.00%, positive predictive value range: 84.10%-100.00%, negative predictive value range: 56.10%-91.70% (Table 3). In external validation, the AUC values ​​of the six models ranged from 0.5000 to 0.8384, with the XGB model showing better performance ( Figure 6 (See Figure C). Sensitivity range: 35.21%-100.00%, specificity range: 0.00%-92.50%, positive predictive value range: 75.80%-91.40%, negative predictive value range: 36.00%-75.00% (Table 4). In summary, the XGB model has good predictive efficacy.

[0094] Table 2 Performance evaluation of machine learning models on the training set

[0095]

[0096] Table 3 Performance evaluation of different machine learning models on the test set

[0097]

[0098] Table 4 Performance evaluation of different machine learning models in external validation

[0099]

[0100] Note: RF: Random Forest; SVM: Support Vector Machine; XGB: XGBoost; GNB: Gaussian Naive Bayes; KNN: K-Nearest Neighbors; PPV: Positive Prediction Value; NPV: Negative Prediction Value.

[0101] 2.5 Comparison of the XGB model with liver stiffness <20 kPa (LSM) and platelet count >150×10^9 (PLT) and the combination of both (LP).

[0102] In the test set, the XGB model's performance (AUC=0.9423) was higher than that of LSM (AUC=0.7609), PLT (AUC=0.6005), and LP (AUC=0.7860). Figure 7 (Figure A) In external validation, the XGB model also performed well (AUC=0.8384). Figure 7 (See Figure B). The AUC for predicting GOV using LSM was 0.7609 (95% CI: 0.6681–0.8536, sensitivity: 75.25%, specificity: 76.92%, positive predictive value: 92.70%, negative predictive value: 44.40%). The AUC for predicting GOV using PLT was 0.6005 (95% CI: 0.5163–0.6848, sensitivity: 97.03%, specificity: 23.08%, positive predictive value: 83.10%, negative predictive value: 66.70%). The AUC for predicting GOV using both methods was 0.7860 (95% CI: 0.6873–0.8847, sensitivity: 75.25%, specificity: 76.92%, positive predictive value: 92.70%, negative predictive value: 44.40%) (Table 5).

[0103] Table 5 Comparison of XGB model with LSM, PLT and LP

[0104]

[0105] XGBtest: XGB model in the test set; XGBval: XGB model in the external validation set; LSM: Liver stiffness measurement >20 kPa; PLT: Platelet count >150 × 10^9; LP: Liver stiffness measurement >20 kPa and platelet count >150 × 10^9; PPV: Positive predictive value; NPV: Negative predictive value.

[0106] 2.6 Evaluation of the XGB model's performance in predicting GOV using clinical decision curve analysis

[0107] In DCA, the reference line "All" represents the net benefit gained from intervening in all individuals at different risk thresholds; as the risk threshold increases, the net benefit decreases. The reference line "None" represents no benefit from no intervention in any individual. The intersection of the two reference lines represents the actual prevalence of GOV in this study's data. In the test set, the actual prevalence was 76.92%. When the risk threshold was higher than 0.2, the net benefit of the intervention was consistently higher than the reference line. Figure 8 (Figure A) In external validation, the actual prevalence rate was 63.96%, and the net benefit of intervention was higher than the reference line when the risk threshold was higher than 0.4. Figure 8 (Figure B) shows that the threshold range is relatively large, and all of them have good clinical benefits.

[0108] 3. Discussion

[0109] GOV (Gestational Varicose Veins) is the most common complication in patients with cirrhosis. Early and accurate diagnosis of GOV can prevent the disease from progressing further. However, the current gold standard for GOV assessment is invasive, leading to poor patient compliance. Therefore, a non-invasive, convenient, reliable, accurate, and safe method for early detection is needed. With the advancement of machine learning, clinicians can now build feasible models from large amounts of data to improve disease prediction performance. An increasing number of studies have established machine learning (ML) models to predict the risk of GOV and EVB (Extravasary Varicose Veins). For example, Azadeh Bayani et al. found that ML had higher predictive power for various grades of GOV than regression methods, and the RF model performed better in predicting each grade (AUC 0.990), outperforming the SVM model (AUC 0.860) and the ANN model (AUC 0.800). Chul-Min Lee et al. used a deep learning model to analyze the ratio of spleen volume to platelets, which helped detect high-risk esophageal varices; the model showed high sensitivity (69.4%) and specificity (78.5%). Zhang Qun et al. developed a neural network model to construct a predictive model for rebleeding in EVB patients. The predictive model had an AUC value of 0.782, which was significantly better than the Cox regression model (AUC value 0.672), Childpugh score (AUC value 0.557), and MELD score (AUC value 0.562).

[0110] Compared to traditional methods, ML models demonstrate superior performance in disease prediction. In this embodiment, an ML-based model was established to predict the occurrence of GOV in patients with cirrhosis. The training set underwent cross-validation to adjust parameters. Multiple adjustments during this process resulted in the leakage of more and more information, leading to models with good performance but weak generalization ability. Therefore, a test set was needed to evaluate the model's performance and improve its generalization ability. On the test set, the XGB model exhibited good predictive ability (AUC=0.9423, 95%CI: 0.8980-0.9865, sensitivity: 81.43%, specificity: 100.00%, positive predictive value: 100.00%, negative predictive value: 65.00%).

[0111] Baveno VII consensus states that cACLD patients have an LSM < 20 kPa and PLT > 150 × 10⁻⁶. 9 / L may exempt patients from GOV screening, but according to the "Expert Consensus on Diagnosis of Portal Hypertension in Cirrhosis by Ultrasound Elastography in China (2023 Edition)," this standard can exempt cACLD patients from gastroscopy screening but is not applicable to cirrhosis patients of all etiologies. Further research is needed to explore predictive models suitable for cirrhosis patients. In the data of this embodiment, the etiologies of cirrhosis patients included hepatitis B, hepatitis C, alcoholic, primary biliary, drug-induced, and autoimmune cirrhosis. Liver stiffness was measured in 127 patients, among whom LSM < 20 kPa and PLT > 150 × 10⁻⁶. 9 Only 6 patients (4.72%) with / L had no GOV, and the combined efficacy of both methods for predicting GOV in cirrhotic patients was 0.7860 (AUC). The XGB model had an AUC of 0.9423 and the highest specificity (100.00%). Furthermore, in external validation data, the XGB model also showed good statistical power (AUC = 0.8384, 95% CI: 0.7627-0.9141, sensitivity: 81.69%, specificity: 80.00%, positive predictive value: 87.90%, negative predictive value: 71.10%), and the DCA curves of the XGB model in both the test set and external validation were above the curve of the reference strategy, indicating high clinical benefit. This helps clinicians to detect GOV early, quickly, and conveniently, facilitating further examination and treatment decisions.

[0112] This embodiment screened six features: age, ST, PLT, ALB, PA, and CHE. In patients with cirrhosis, increased portal vein pressure obstructs splenic blood return, leading to blood stasis and splenomegaly. The spleen plays a crucial role in predicting GOV (Government-Occupied Vascular Vulnerability). Prediction is achieved by measuring spleen stiffness, diameter, volume, and other indicators, with spleen thickness also playing a significant role. In this embodiment, spleen thickness (ST) is displayed using an oblique section between the costal ribs. ST varies across different GOV grades; the more severe the GOV, the larger the ST. Increased portal vein pressure in cirrhotic patients easily leads to worsening GOV and also increases ST. In this embodiment, the PLT in the no-GOV group was higher than in the other three groups. Furthermore, the PLT in the mild GOV group was higher than in the severe GOV group. The spleen stores a large amount of PLT; splenomegaly leads to hyperfunction, resulting in decreased PLT. Other variables screened in the study include ALB, PA, and CHE, possibly due to their influence on hepatic synthesis and catabolism. This embodiment also found that the age of patients with and without GOV differed, with the average age at the onset of GOV being 52 years, which was higher than that of patients without GOV.

[0113] In summary, this invention establishes a novel XGB model that can predict the occurrence of GOV in cirrhotic patients, exhibiting high predictive ability. This invention uses DCA to demonstrate the feasibility of using the XGB model in clinical assessment of GOV. The XGB model constructed in this invention has high accuracy and can serve as a non-invasive tool for predicting GOV in patients with cirrhosis.

[0114] Although the invention has been described with reference to exemplary embodiments, it should be understood that the invention is not limited to the disclosed exemplary embodiments. Various adjustments or changes may be made to the exemplary embodiments described in this specification without departing from the scope or spirit of the invention. The scope of the claims should be interpreted in the broadest possible sense to cover all modifications and equivalent structures and functions.

Claims

1. A non-invasive system for predicting esophageal and gastric varices in a subject, characterized in that, include: The data acquisition unit is configured to acquire the basic personal information and test data of the subjects, including patients with cirrhosis. The basic personal information and test data include parameters such as age, spleen thickness, platelets, albumin, prealbumin, and cholinesterase. The prediction unit is configured to input the data acquired by the data acquisition unit into the XGB prediction model and output the prediction result. The prediction model is obtained by the following method: (1) Provide a training sample set containing multiple samples, each of which includes the subject's basic personal information, detection data, and labels. The subject is a patient with cirrhosis. The basic personal information and detection data use factors with intergroup differences in esophageal and gastric varices as independent variables. The parameters of age, spleen thickness, platelet count, albumin, prealbumin, and cholinesterase are obtained by introducing a penalty coefficient lambda.1se into the regression model. The label is esophageal and gastric varices, and the λ value of lambda.1se is 0.

049. (2) Input the data into the XGB machine learning model for training to obtain a predictive model for esophageal and gastric varices in patients with cirrhosis.

2. The system for non-invasively predicting esophageal and gastric varices in a subject according to claim 1, characterized in that, The number of samples in the training sample set is 50 or more.

3. The system for non-invasive prediction of esophageal and gastric varices in a subject according to claim 1, characterized in that, It further includes providing a validation sample set and using the validation sample set for validation, thereby adjusting or optimizing the prediction model based on the validation results.

4. The system for non-invasively predicting esophageal and gastric varices in a subject according to claim 1, characterized in that, Further, it includes providing a test sample set and using the test sample set to test, thereby evaluating the prediction model.

5. A device, characterized in that, The device includes: a memory, a processor, and a program stored in the memory and capable of running on the processor, the program being configured to perform: The study obtained basic personal information and test data from subjects, including patients with cirrhosis. The basic personal information and test data included parameters such as age, spleen thickness, platelet count, albumin, prealbumin, and cholinesterase. The data acquired by the data acquisition unit is input into the XGB prediction model, and the prediction result is output. The prediction model is obtained by the following method: (1) Provide a training sample set containing multiple samples, each of which includes the subject's basic personal information, detection data, and labels. The subjects include patients with cirrhosis. The basic personal information and detection data use factors with intergroup differences in esophageal and gastric varices as independent variables. The parameters of age, spleen thickness, platelet count, albumin, prealbumin, and cholinesterase are obtained by introducing a penalty coefficient lambda.1se into the regression model. The label is esophageal and gastric varices, and the λ value of lambda.1se is 0.

049. (2) Input the data into the XGB machine learning model for training to obtain a predictive model for esophageal and gastric varices in patients with cirrhosis.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that is implemented when executed by a processor: The study obtained basic personal information and test data from subjects, including patients with cirrhosis. The basic personal information and test data included parameters such as age, spleen thickness, platelet count, albumin, prealbumin, and cholinesterase. The data acquired by the data acquisition unit is input into the XGB prediction model, and the prediction result is output. The prediction model is obtained by the following method: (1) Provide a training sample set containing multiple samples, each of which includes the subject's basic personal information, detection data, and labels. The subjects include patients with cirrhosis. The basic personal information and detection data use factors with intergroup differences in esophageal and gastric varices as independent variables. The parameters of age, spleen thickness, platelet count, albumin, prealbumin, and cholinesterase are obtained by introducing a penalty coefficient lambda.1se into the regression model. The label is esophageal and gastric varices, and the λ value of lambda.1se is 0.

049. (2) Input the data into the XGB machine learning model for training to obtain a predictive model for esophageal and gastric varices in patients with cirrhosis.

Citation Information

Patent Citations

  • Portal vein thrombosis prediction model construction method and prediction system

    CN117672536A

  • Predicting disease progression in portal hypertension using machine learning

    WO2023227942A1