A method and device for predicting muscle loss after liver transplantation for liver cancer

Through the combination of the imbalanced random forest model and the decision tree model, the probability of muscle loss after liver transplantation is predicted, and the patient's prognosis stratification and survival rate prediction are carried out, which solves the problem of difficult prediction of muscle loss after surgery and improves the treatment effect and satisfaction.

CN119581008BActive Publication Date: 2025-05-30MCCONDI (SHAOXING) TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510142065.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

Muscle loss after liver transplantation plays a key role in patients' rehabilitation and long-term survival results, but the existing technology is difficult to accurately predict dynamic changes in postoperative muscle loss, resulting in a lack of effective interventions by clinicians.

Method used

An imbalanced random forest model containing several decision trees is used to predict multiple clinical key indicators of patients to obtain the predicted probability of postoperative muscle loss, and combined with patient prognosis stratification and survival rate prediction results, helping clinicians develop personalized treatment plans.

Benefits of technology

It achieves accurate prediction of muscle loss level after liver transplantation, provides a more scientific basis for decision-making, helps clinicians formulate personalized treatment plans, and improves the treatment effect and satisfaction of patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119581008B_ABST
    Figure CN119581008B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for predicting muscle loss after liver transplantation for liver cancer, belonging to the field of biomedical technology. The method includes: constructing an imbalanced random forest model containing a number of decision trees, independently predicting a number of binary classification prediction results of high risk or low risk by the number of decision trees, and taking the proportion of the decision trees with a high-risk prediction result among all decision trees as the prediction probability of muscle loss after liver transplantation for liver cancer finally output by the model; performing risk classification according to the prediction probability, and determining the postoperative muscle loss level of the patient as high risk, medium risk or low risk; constructing a decision tree model to stratify the patient's prognosis, and dividing the patient into a high-risk group, a medium-risk group and a low-risk group; combining the liver transplantation criteria and the patient's prognosis stratification results for survival rate prediction. The present invention can achieve accurate prediction of the muscle loss level after liver transplantation, and helps clinicians to formulate personalized treatment plans and scientific management after transplantation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of biomedical technology, and particularly relates to a method and device for predicting muscle loss after liver transplantation for liver cancer. Background Art

[0002] Hepatocellular Carcinoma (HCC), as an important cause of tumor-related deaths globally, cannot be ignored. Among treatment methods, liver transplantation is a radical treatment for some HCC patients. During the rehabilitation process after liver transplantation, muscle loss is a key factor that has not been fully emphasized but affects the overall recovery and long-term survival outcomes of patients.

[0003] Muscle loss, also known as sarcopenia in medicine, is a phenomenon of muscle mass reduction and muscle function decline related to multiple factors such as aging, disease progression, and surgical treatment. In liver transplant patients, this phenomenon is particularly significant. It may not only lead to a decrease in patients' physical strength and limited mobility but also further exacerbate the risk of other complications such as infection and thrombosis, thus seriously affecting the postoperative rehabilitation process and long-term quality of life of patients.

[0004] Although previous studies have emphasized the prognostic significance of sarcopenia in patients undergoing liver transplantation, research in this area is still insufficient. Especially for those patients who did not have sarcopenia before surgery, the dynamic process of their muscle loss has not been fully explored. This gap in understanding hinders the development of effective prediction models, resulting in clinicians often lacking reliable tools to accurately judge which patients may experience significant muscle loss after surgery, and thus it is difficult to take effective intervention measures in a timely manner to prevent and reduce this risk. This not only affects the treatment effect and rehabilitation process of patients but also may increase the waste of medical resources to a certain extent.

[0005] In recent years, with the rapid development and wide application of machine learning technology, its application prospects in the field of healthcare have also become increasingly broad, opening up new ways for predicting complex clinical outcomes. In the prediction and intervention of muscle loss after liver transplantation, machine learning has also demonstrated great potential and value. By collecting and analyzing multi-dimensional clinical data including patients' basic information, preoperative and postoperative examination results, treatment records, etc., machine learning models can construct more refined and accurate prediction models, identify high-risk patient groups who may experience significant muscle loss after surgery, and thus provide more scientific decision-making basis for clinicians.

[0006] In summary, the severity of hepatocellular carcinoma and the impact of muscle loss after liver transplantation on patient recovery and long-term outcomes cannot be ignored. The introduction of machine learning technology provides new ideas and means to solve this problem. In the future, with the continuous in-depth research and improvement of related technologies, it is necessary to further develop a model that can predict muscle loss after liver transplantation in patients with hepatocellular carcinoma, play a more important clinical role in the prediction and intervention of muscle loss after liver transplantation, and stimulate further exploration in this field. Summary of the Invention

[0007] In view of the above, the object of the present invention is to provide a method and device for predicting muscle loss after liver transplantation for hepatocellular carcinoma, which can accurately predict the level of muscle loss after liver transplantation, and help clinicians formulate personalized treatment plans and scientific management after transplantation in combination with the patient prognosis stratification and survival rate prediction results.

[0008] To achieve the above object of the invention, the technical solutions provided by the present invention are as follows:

[0009] A method for predicting muscle loss after liver transplantation for hepatocellular carcinoma provided by an embodiment of the present invention includes the following steps:

[0010] Construct an unbalanced random forest model containing a number of decision trees, and independently predict a number of high-risk or low-risk binary classification prediction results based on a number of preprocessed key clinical indicators of the patient input by the number of decision trees. The proportion of decision trees with a high-risk prediction result among all decision trees is used as the prediction probability of muscle loss after liver transplantation for hepatocellular carcinoma finally output by the unbalanced random forest model;

[0011] If the prediction probability is greater than or equal to the first risk threshold, the muscle loss level after surgery is determined to be high risk. If the prediction probability is less than the first risk threshold and greater than or equal to the second risk threshold, the muscle loss level after surgery is determined to be medium risk. If the prediction probability is less than the second risk threshold, the muscle loss level after surgery is determined to be low risk, where the first risk threshold is greater than the second risk threshold;

[0012] Construct a decision tree model to stratify the patient prognosis. Among them, patients with SMI less than or equal to the stratification threshold are directly classified into the high-risk group, patients with SMI greater than the stratification threshold and a high-risk muscle loss level are classified into the medium-risk group, and patients with SMI greater than the stratification threshold and a medium-risk or low-risk muscle loss level are classified into the low-risk group;

[0013] Survival rate prediction is carried out by combining liver transplantation criteria and the results of patient prognosis stratification. Among them, for patients meeting the liver transplantation criteria, if the patient is in the high-risk group, the 1-year survival rate after surgery is 71.5%, the 3-year survival rate after surgery is 53.2%, and the 5-year survival rate after surgery is 36.3%. If the patient is in the medium-risk group, the 1-year survival rate after surgery is 86.4%, the 3-year survival rate after surgery is 74.6%, and the 5-year survival rate after surgery is 63.9%. If the patient is in the low-risk group, the 1-year survival rate after surgery is 92.3%, the 3-year survival rate after surgery is 85.9%, and the 5-year survival rate after surgery is 75.7%.

[0014] Preferably, multiple key patient clinical indicators are screened from a wide range of patient clinical indicators by using the recursive feature elimination method. The multiple key patient clinical indicators include α-L-fucosidase, cholyglycine dehydrogenase, age, alpha-fetoprotein, international normalized ratio, alanine aminotransferase, eosinophil count, basophil count, alkaline phosphatase, adenosine deaminase, gamma-glutamyl transferase, and body mass index in the patient's preoperative peripheral blood.

[0015] Preferably, the preprocessing performed on the obtained multiple key patient clinical indicators includes missing value processing, outlier processing, feature coding conversion, and data standardization.

[0016] Preferably, in the imbalanced random forest model, the construction process of the decision tree includes:

[0017] Construct the preprocessed multiple key patient clinical indicators collected from different patients into a dataset. These different patients include patients at high risk of postoperative muscle loss and patients at low risk of postoperative muscle loss. Each sample in the dataset represents the preprocessed multiple key patient clinical indicators of a patient, and each key patient clinical indicator serves as a feature;

[0018] For each decision tree, randomly select a part of the categories of features from the features of all categories included in the samples in the dataset as the candidate feature set. Starting from the root node, select an optimal splitting feature and its splitting point from the candidate feature set according to the Gini impurity as the judgment criterion to split the samples in the dataset into two subsets, and generate two new child nodes. Then, recursively repeat the above splitting process for each new child node until the stopping condition is met. The stopping condition is reaching the preset maximum depth or the number of samples in the node is less than the preset value.

[0019] Preferably, the calculation formula for the Gini impurity is:

[0020] ,

[0021] Among them, the high-risk ratio and the low-risk ratio are the proportions of patients at high risk of postoperative muscle loss and the proportions of patients at low risk of postoperative muscle loss in each subset obtained after dividing the dataset for each node. Each decision tree will try various possible splitting features and splitting points, and select the splitting scheme that can minimize the sum of the Gini impurities of each node.

[0022] Preferably, the first risk threshold is set to 0.7, and the second risk threshold is set to 0.3.

[0023] Preferably, the stratification threshold is set to 43.6 cm 2 / m 2 。

[0024] To achieve the above invention purpose, an embodiment of the present invention also provides a device for predicting muscle loss after liver transplantation for liver cancer, which is implemented by using the above method for predicting muscle loss after liver transplantation for liver cancer, and includes: a model prediction module, a risk classification module, a prognosis stratification module, and a survival prediction module;

[0025] The model prediction module is used to construct an unbalanced random forest model containing several decision trees, and independently predict several high-risk or low-risk binary classification prediction results based on several preprocessed clinical key indicators of the input patients through several decision trees. The proportion of decision trees with a high-risk prediction result among all decision trees is used as the prediction probability of muscle loss after liver transplantation for liver cancer finally output by the unbalanced random forest model;

[0026] The risk classification module is used to determine that the postoperative muscle loss level is high risk if the prediction probability is greater than or equal to the first risk threshold, determine that the postoperative muscle loss level is medium risk if the prediction probability is less than the first risk threshold and greater than or equal to the second risk threshold, and determine that the postoperative muscle loss level is low risk if the prediction probability is less than the second risk threshold, where the first risk threshold is greater than the second risk threshold;

[0027] The prognosis stratification module is used to construct a decision tree model to stratify the patient prognosis. Among them, patients with SMI less than or equal to the stratification threshold are directly classified into the high-risk group, patients with SMI greater than the stratification threshold and a high-risk muscle loss level are classified into the medium-risk group, and patients with SMI greater than the stratification threshold and a medium-risk or low-risk muscle loss level are classified into the low-risk group;

[0028] The survival prediction module is used to predict the survival rate by combining the liver transplantation criteria and the results of patient prognosis stratification. Among them, for patients meeting the liver transplantation criteria, if the patient is in the high-risk group, the 1-year survival rate after surgery is 71.5%, the 3-year survival rate after surgery is 53.2%, and the 5-year survival rate after surgery is 36.3%; if the patient is in the medium-risk group, the 1-year survival rate after surgery is 86.4%, the 3-year survival rate after surgery is 74.6%, and the 5-year survival rate after surgery is 63.9%; if the patient is in the low-risk group, the 1-year survival rate after surgery is 92.3%, the 3-year survival rate after surgery is 85.9%, and the 5-year survival rate after surgery is 75.7%.

[0029] Compared with the prior art, the beneficial effects of the present invention at least include:

[0030] (1) By using the recursive feature elimination method, the present invention screens out 12 clinical key indicators from a wide range of patient clinical indicators and incorporates them into the final model input, which not only improves the prediction accuracy and interpretability of the model, but also reduces the computational complexity and promotes the optimization of clinical decision-making. After testing, the AUC of the model can reach 0.8503, effectively avoiding information redundancy and noise interference, and thus significantly improving the prediction accuracy of the imbalanced random forest model.

[0031] (2) The present invention predicts through an imbalanced random forest model to obtain the predicted probability of muscle loss after liver transplantation of patients, determines the risk level according to the predicted probability, further stratifies the patient prognosis according to the risk level, and predicts the survival rate according to the stratification results and liver transplantation criteria. On the basis of being able to accurately predict the level of muscle loss after liver transplantation, combining the patient prognosis stratification and survival rate prediction results can provide more accurate and personalized clinical decision-making support for doctors. This helps doctors more accurately evaluate the patient's condition, formulate more reasonable treatment plans, and thus improve the treatment effect and satisfaction of patients. Brief Description of the Drawings

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0033] Figure 1 It is a schematic flowchart of the method for predicting muscle loss after liver transplantation for liver cancer provided by the embodiment of the present invention.

[0034] Figure 2It is a schematic diagram showing the relationship between the level of muscle loss and the prognosis of liver cancer liver transplantation recipients provided by the embodiments of the present invention; among them, the Kaplan-Meier survival curve shown in A shows the possibility of recurrence-free survival of the pre-operative sarcopenia group and non-sarcopenia group based on sarcopenia in the patient cohort; the Kaplan-Meier survival curve shown in B shows the possibility of recurrence-free survival of the pre-operative mild muscle loss group and severe muscle loss group based on muscle loss in the patient cohort; the Kaplan-Meier survival curve shown in C shows the possibility of recurrence-free survival of the pre-operative mild muscle loss group and severe muscle loss group based on muscle loss (after propensity matching) in the patient cohort; D represents the relative risk ratio between the recurrence results of severe muscle loss and mild muscle loss and the death results of severe muscle loss and mild muscle loss.

[0035] Figure 3 It is a schematic diagram showing the establishment and performance test of the unbalanced random forest model provided by the embodiments of the present invention; among them, A is an overview of 46 features included in model development. The inner circle represents the feature type division and its proportion of 46 indicators, and the outer circle represents the data nature and its proportion corresponding to each feature type; B is a curve showing the change of the average AUC with the number of features during the recursive feature elimination process; C is the ROC curve of the model for the pre-operative group without sarcopenia; D is a model waterfall plot showing the contribution of each variable to the model prediction. For example, when AFU = 16.9, the contribution to the model prediction is +0.155 (positive contribution), and when GPDA = 44, the contribution to the model prediction is -0.0955 (negative contribution). f(x)=0.838 represents the predicted value of the model after fusing 12 indicators, and E[f(x)] = 0.729 represents the expectation of f(x); E is a schematic diagram of the decision tree model; the Kaplan-Meier survival curve shown in F shows the recurrence-free survival rates of the high-risk group, medium-risk group, and low-risk group based on the decision tree model.

[0036] Figure 4 It is a schematic diagram of the structure of a muscle loss prediction device after liver cancer liver transplantation provided by the embodiments of the present invention. Detailed implementation manners

[0037] To make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific implementation manners described herein are only used to explain the present invention and do not limit the protection scope of the present invention.

[0038] The inventive concept of the present invention is as follows: Aiming at the problem of insufficient development of a model for predicting muscle loss after liver transplantation in patients with hepatocellular carcinoma in the prior art, the embodiments of the present invention provide a method and device for predicting muscle loss after liver transplantation for hepatocellular carcinoma. By inputting 12 clinical key indicators of a patient into an unbalanced random forest model containing a number of decision trees for prediction and obtaining the prediction probability of muscle loss after liver transplantation of the patient, the risk level is determined according to the prediction probability, the prognosis of the patient is further stratified in combination with the risk level, and the survival rate is predicted according to the stratification result and the liver transplantation standard, so as to achieve accurate prediction of the muscle loss level after liver transplantation, and help clinicians formulate personalized treatment plans and scientific management after transplantation in combination with the patient prognosis stratification and the survival rate prediction result.

[0039] Figure 1 It is a schematic flowchart of the method for predicting muscle loss after liver transplantation for hepatocellular carcinoma provided by the embodiment of the present invention. As Figure 1 shown, the embodiment provides a method for predicting muscle loss after liver transplantation for hepatocellular carcinoma, including the following steps:

[0040] S1. Construct an unbalanced random forest model containing a number of decision trees, and through a number of decision trees, independently predict based on the preprocessed multiple patient clinical key indicators input to obtain a number of binary classification prediction results of high risk or low risk, and use the proportion of decision trees with a high-risk prediction result among all decision trees as the prediction probability of muscle loss after liver transplantation for hepatocellular carcinoma finally output by the unbalanced random forest model.

[0041] S1.1. Experimental method.

[0042] (1) Research population and data collection.

[0043] Patients who underwent liver transplantation at the First Affiliated Hospital of Zhejiang University School of Medicine (Hangzhou, China) and Shulan Hospital (Hangzhou, China) from January 2015 to January 2021 and were pathologically confirmed to have hepatocellular carcinoma (HCC) after surgery were included in the study. Through subject selection, a total of 248 patients who received complete preoperative and postoperative CT scans were finally included. Peripheral blood indicators collected within 2 weeks before surgery, including complete blood count and liver function tests, were also included in the study. After liver transplantation, the recipients were followed up until December 31, 2022. During the follow-up period, the recipients were managed according to the standard protocol. The alpha-fetoprotein (AFP) level was measured every 1-2 months. Ultrasonography, computed tomography (CT), or magnetic resonance imaging (MRI) examinations were performed every 3-6 months after liver transplantation. The recurrence-free survival (RFS) was defined as the number of months from the surgery date to the date of the first confirmed recurrence or death. The metastatic sites of HCC were determined according to the radiological evidence in accordance with the standard HCC guidelines. Clinical and anthropometric parameters were collected from the Chinese Liver Transplant Registry (CLTR) in strict accordance with the Regulations on Human Organ Transplantation and the requirements of national laws.

[0044] (2) Skeletal muscle assessment.

[0045] Skeletal muscle was assessed at the third lumbar vertebra (L3) on cross-sectional CT images using SliceOmatic software (version 5.0; Tomovision). The Hounsfield unit (HU) threshold for skeletal muscle was -29 to +150 HU. The muscle area was normalized to the square of the patient's height (in meters, m 2 ), yielding the skeletal muscle index (SMI). The measurement of body composition was performed by two trained observers. Before the formal measurement, 50 CT images were randomly selected from the cohort and provided to the observers to evaluate the reproducibility between observers. One week later, the two observers were asked to repeat the same task to evaluate the reproducibility within observers. The images of the above measurement results were verified by an anatomical radiologist. Before performing the dynamic analysis, according to the previously reported method, a precision test (the two observers repeated the measurement of 30 CT images) was performed to avoid random errors caused by measurement. Once the precision error of the measurement was known, the magnitude of muscle change indicating a true biological change could be determined through the least significant change (LSC). The value with the larger error in the results of the two observers was used in the subsequent discussion.

[0046] (3) Model development, feature selection, and model interpretation.

[0047] Data processing adopted a hybrid sampling technique, including oversampling and undersampling, and combined with 5-fold cross-validation to avoid overfitting problems. Grid search benchmark analysis was conducted among 50 machine learning models, and 10-fold cross-validation was performed to obtain a stable model. AUC, accuracy, F1, LogLoss, precision, and recall were used to evaluate model performance.

[0048] The SHAP method was used to obtain the importance of features and interpret the model through the R package "mlr3". Feature selection was based on the strategy of adding features one by one from 1 in order of importance to all features (46 clinically available indicators). The model with reduced features and comparable predictive ability (AUC) was regarded as the final model.

[0049] (4)Statistical analysis.

[0050] Categorical variables were expressed as numbers (percentages), and χ 2 or Fisher's exact test was used for comparison as appropriate. Continuous variables were expressed as the median (interquartile range), and the Kruskal-Wallis test was used for comparison. The comparison of survival curves was examined by the Kaplan-Meier method and the log-rank test.

[0051] The propensity score matching (PSM) analysis was performed using the "MatchIt" R package (version 4.5.5) with a 1:2 matching ratio and the default nearest neighbor matching algorithm. A P value < 0.05 was considered significant for the analysis. All statistical analyses were performed by R language software (version 4.2.1).

[0052] S1.2, Experimental results.

[0053] (1)Patient baseline characteristics.

[0054] A total of 248 patients who received liver transplantation met the inclusion criteria of the study. By examining all possible SMI values in the cohort, patients with an SMI lower than 43.6 were classified as having sarcopenia, and as a result, 69 patients were diagnosed with sarcopenia (27.9%). In the non-sarcopenia cohort, patients with more than 16.8% postoperative muscle loss were classified as having severe muscle loss, including 46 patients (25.7%).

[0055] Kaplan-Meier analysis showed that patients with sarcopenia had a poorer prognosis (as shown in A of Figure 2 ). In addition, the study found that patients with more muscle loss had a poorer prognosis (as shown in B of Figure 2 ), and this trend still persisted after 1:2 PSM (propensity matching) (as shown in Figure 2as shown in C in). Competitive risk analysis showed that the difference in RFS due to muscle loss was attributed to recurrence rather than non-recurrence-related mortality (as shown in D in). Subsequently, the following variables: gender, age, BMI (body mass index), HBV (hepatitis B virus) status, change in SMI, SMI (skeletal muscle index), MELD (model for end-stage liver disease) score, AFP (alpha-fetoprotein) level (>400 μg / L), tumor differentiation, number of tumors, microvascular invasion, and maximum tumor diameter were included in the univariate Cox analysis. Multivariate Cox regression analysis was then performed, which included the following factors: muscle loss, tumor differentiation, number of tumors, microvascular invasion, maximum tumor diameter, and AFP level. The results of the multivariate Cox regression are shown in Table 1 below, where HR represents the hazard ratio, 95%CI represents the 95% confidence interval, and P represents the probability value. Figure 2 as shown in D in). Subsequently, the following variables: gender, age, BMI (body mass index), HBV (hepatitis B virus) status, change in SMI, SMI (skeletal muscle index), MELD (model for end-stage liver disease) score, AFP (alpha-fetoprotein) level (>400 μg / L), tumor differentiation, number of tumors, microvascular invasion, and maximum tumor diameter were included in the univariate Cox analysis. Multivariate Cox regression analysis was then performed, which included the following factors: muscle loss, tumor differentiation, number of tumors, microvascular invasion, maximum tumor diameter, and AFP level. The results of the multivariate Cox regression are shown in Table 1 below, where HR represents the hazard ratio, 95%CI represents the 95% confidence interval, and P represents the probability value.

[0056] Table 1 Results of Multivariate Cox Regression

[0057]

[0058] The research results showed that severe muscle loss was an independent risk factor for recurrence in liver transplant recipients with hepatocellular carcinoma. However, there is still a lack of effective preoperative models in clinical practice to predict the degree of postoperative muscle loss. To address this issue, it is necessary to start developing a model that effectively predicts muscle loss after liver transplantation for hepatocellular carcinoma.

[0059] (2) Model development.

[0060] To develop a machine learning model using clinical indicators to predict postoperative muscle loss in liver transplant patients, first, 46 clinical characteristics of the patients were collected (a summary of the characteristics is as Figure 3as shown in A of , including: AGE (age), height, BMI (body mass index), HBV (hepatitis B virus) infection status, MELD (model for end-stage liver disease) score, AFP (alpha-fetoprotein), tumor differentiation, number of tumors, maximum tumor diameter, total tumor size, microvascular invasion, albumin, albumin-to-globulin ratio, ALT (alanine aminotransferase), AST (aspartate aminotransferase), INR (international normalized ratio), monocyte count, neutrophil count, platelet count, SMI (skeletal muscle index), white blood cell count, lymphocyte count, E (eosinophil count), B (basophil count), hemoglobin concentration, red blood cell count, RDW (red blood cell volume distribution width), MHV (mean corpuscular hemoglobin volume), MCHC (mean corpuscular hemoglobin concentration), HGB (hemoglobin), PCT (plateletcrit), GGT (gamma-glutamyl transferase), ALP (alkaline phosphatase), indirect bilirubin, direct bilirubin, ChE (cholinesterase), TBA (total bile acid), GPDA (glycocholic acid dehydrogenase), AFU (alpha-L-fucosidase), ADA (adenosine deaminase), TG (triglyceride), TC (total cholesterol), HDL (high-density lipoprotein cholesterol), LDL (low-density lipoprotein cholesterol), VLDL (very low-density lipoprotein cholesterol), and fasting blood glucose. The performance of 50 different machine learning models was compared using 5-fold cross-validation. Among them, the imbalanced random forest model adopted in the embodiments of the present invention showed the best performance for this clinical problem. The imbalanced random forest model is specifically designed to handle imbalanced outcome data and is very suitable for the clinical problem in this study.

[0061] To determine the optimal input features of the model, recursive feature elimination (RFE) was used to systematically reduce the number of features from 46 to 12. It was found by comparison that using 12 features could produce performance comparable to that of the model using all features (as shown in B of ). Figure 3 Therefore, these 12 key clinical indicators of patients (AFU, GPDA, AGE, AFP, INR, ALT, E, B, ALP, ADA, GGT, and BMI) were incorporated into the final model in the embodiments of the present invention. The AUC of this imbalanced random forest model reached 0.8503 (as shown in C of ). Figure 3 in .

[0062] The SHAP method was used to interpret this imbalanced random forest model to provide insights for future research. Among the final 12 key clinical indicators of patients, sorted by SHAP values, the top 4 important features related to muscle loss are AFU, GPDA, AGE, and AFP (as shown in D of ). Figure 3 in .

[0063] (3) Model training.

[0064] First, after preprocessing the 12 preprocessed clinical key indicators of patients collected from different patients, a dataset is constructed. These different patients include patients at high risk of postoperative muscle loss and patients at low risk of postoperative muscle loss. Each sample in the dataset represents the preprocessed multiple clinical key indicators of a patient, and each clinical key indicator of a patient serves as a feature.

[0065] Data preprocessing includes: Missing value handling: For numerical features (such as indicators like ALT, AFP, etc.), the median of the indicator is used for filling. For categorical features (such as HBV infection status), the most common category is used for filling. Outlier handling: The 3 - standard - deviation (3σ) principle is used to identify outliers. Values outside the range of mean ± 3 standard deviations are processed. Values exceeding the upper limit are set to the upper limit value, and values below the lower limit are set to the lower limit value. Feature encoding conversion: One - hot encoding conversion is performed on categorical features to convert non - numerical features into numerical features. Data standardization: All numerical features are standardized using the standardization formula: (x - μ) / σ, where μ is the mean and σ is the standard deviation, to unify the feature values within the same scale range.

[0066] The model parameters are set as follows: The number of decision trees is set to 100, with an automatically adjusted maximum tree depth. The feature sampling ratio is set to the square root of the total number of features, and the sample sampling strategy uses balanced sampling. The internal structure of the model is actually a complex and precise decision - making system. Imagine a panel of experts consisting of 100 experienced doctors, each with their own unique diagnostic methods. This is similar to the 100 decision trees in the model, and each decision tree is a set of unique decision rules learned from the training data.

[0067] Next, the five - fold cross - validation method is used for training. The same parameter settings are used for each fold of training, and the training performance indicators for each fold are recorded. For each decision tree, a part of the features of various categories included in all samples in each fold of the dataset is randomly selected as the candidate feature set. Starting from the root node, an optimal splitting feature and its splitting point are selected from the candidate feature set according to the Gini impurity as the judgment criterion to split the samples in the dataset into two subsets, and two new child nodes are generated. When patient data enters a decision tree, it undergoes a series of "yes / no" judgments, just like a doctor conducting a medical interview. For example, the first node might ask: "Is the patient's AFP value greater than 400?" If yes, the data flows to the right branch; if no, it flows to the left branch. An optimal splitting feature and splitting point are selected at each branch point (node).

[0068] This optimal splitting feature and splitting point are constructed based on the Gini impurity index during training. The calculation formula for Gini impurity is:

[0069] ,

[0070] Among them, the high-risk ratio and the low-risk ratio are the proportions of patients at high risk of postoperative muscle loss and patients at low risk of postoperative muscle loss in each subset obtained after splitting the dataset for each node. Assuming that 60% of the patient sample data in the input dataset corresponds to high risk and 40% corresponds to low risk, then the Gini impurity = 1 - 0.6² - 0.4² = 0.48. Each decision tree will try various possible splitting features and splitting points and select the splitting scheme that minimizes the sum of the Gini impurities of each node. Such a split will perform better than a random split because it makes the patient risk levels in each child node more "pure". Subsequently, this splitting process will be continuously repeated, recursively repeating the above splitting process for each new child node until a stopping condition is met. The stopping condition is reaching a preset maximum depth (such as 6 levels) or the number of samples in the node being less than a preset value (such as 5). Such a design can prevent the model from overfitting to the training data.

[0071] After the unbalanced random forest model is trained, when making predictions for new patients, first, clinical indicators are collected, and data on 12 key indicators of the patient are collected to ensure the integrity and accuracy of the data, and the data collection time and conditions are recorded. Then, after preprocessing the collected data, it is input into the trained unbalanced random forest model. The 12 feature data will be input into 100 decision trees simultaneously. Each tree will, according to its own constructed decision rules, guide the patient to a leaf node and give a prediction result: high risk or low risk. Finally, the proportion of decision trees with a high-risk prediction among all decision trees is used as the prediction probability of postoperative muscle loss after liver transplantation for hepatocellular carcinoma finally output by the unbalanced random forest model. For example, among 100 decision trees, 65 decision trees predict high risk and 35 decision trees predict low risk, then the final prediction probability of high risk of muscle loss is 0.65.

[0072] This integrated decision-making method has several important advantages: by having 100 decision trees vote together, it reduces the impact of misjudgments that a single tree might produce. Different trees may focus on different combinations of features, which improves the robustness of the model. The final probability output provides confidence information for the prediction, rather than simply a binary classification result. Notably, this imbalanced random forest model also addresses the common class imbalance problem in medical data through an imbalanced sampling strategy. When training each decision tree, it ensures that a sufficient number of minority class samples (high-risk cases) are seen, so that even if high-risk cases are fewer in the actual data, the model can accurately identify them. This complex and sophisticated prediction mechanism ultimately achieves an AUC value of 0.8503, demonstrating its reliability in clinical prediction. More importantly, its prediction results show a significant correlation with the actual prognosis of patients, which is the best proof of its clinical value.

[0073] S2. According to the predicted probability, perform risk classification to determine whether the postoperative muscle loss level of the patient is high risk, medium risk, or low risk.

[0074] In the embodiment, the first risk threshold is set to 0.7, and the second risk threshold is set to 0.3.

[0075] If the predicted probability ≥ 0.7, then determine the postoperative muscle loss level as high risk;

[0076] If 0.3 ≤ predicted probability < 0.7, then determine the postoperative muscle loss level as medium risk;

[0077] If the predicted probability < 0.3, then determine the postoperative muscle loss level as low risk.

[0078] S3. Construct a decision tree model to stratify the patient prognosis and divide the patients into high-risk group, medium-risk group, and low-risk group.

[0079] In the embodiment, in order to use simple preoperative blood tests and preoperative sarcopenia status to stratify the patient prognosis, the predicted probability of the model is combined with the Hangzhou criteria to develop a simple decision tree model (as shown in Figure 3 E). Among them, for patients with SMI ≤ 43.6 cm 2 / m 2 are directly classified into the high-risk group. For patients with SMI > 43.6 cm 2 / m 2 and whose muscle loss level is determined to be high risk are classified into the medium-risk group. For patients with SMI > 43.6 cm 2 / m 2 and whose muscle loss level is determined to be medium risk or low risk are classified into the low-risk group. This decision tree model is both easy to use and clinically effective.

[0080] S4. Combine the liver transplantation criteria and the results of patient prognosis stratification to predict the survival rate.

[0081] In the embodiment, the Hangzhou criteria for liver transplantation were incorporated, which are mainly used to evaluate whether liver cancer patients are suitable for liver transplantation surgery. The Hangzhou criteria are more lenient than the Milan criteria, expanding the scope of indications, but still maintaining good prognostic effects. According to the results of stratifying the patient prognosis by the decision tree model and combining the Hangzhou criteria for liver transplantation, the survival rate of patients meeting the Hangzhou criteria for liver transplantation was predicted (as shown by F in Figure 3 , where the blue curve represents the high-risk group, the red curve represents the medium-risk group, and the green curve represents the low-risk group), and P < 0.0001 indicates that there are significant differences in the survival probabilities among the three groups. If the patient is in the high-risk group, the 1-year survival rate after surgery is 71.5%, the 3-year survival rate is 53.2%, and the 5-year survival rate is 36.3%. If the patient is in the medium-risk group, the 1-year survival rate after surgery is 86.4%, the 3-year survival rate is 74.6%, and the 5-year survival rate is 63.9%. If the patient is in the low-risk group, the 1-year survival rate after surgery is 92.3%, the 3-year survival rate is 85.9%, and the 5-year survival rate is 75.7%. The above survival rates were obtained by analyzing the historical data of liver transplantation patients in a large-scale retrospective study and calculating the survival rates of patients with different prognostic stratifications based on the postoperative follow-up records. These survival rates were calculated using the Kaplan-Meier survival analysis, a statistical method that fully considers various changes in the survival status of patients from surgery to the end of the follow-up, and thus is scientific and reliable, providing an important reference basis for clinical decision-making.

[0082] In summary, the method for predicting muscle loss after liver transplantation for liver cancer provided by the embodiment of the present invention can accurately predict the level of muscle loss after liver transplantation for liver cancer, helping clinicians formulate personalized treatment plans and scientific management after transplantation, thereby providing more personalized and efficient medical services for patients, and having important clinical application value and scientific research significance.

[0083] Based on the same inventive concept, as shown in Figure 4 , the embodiment of the present invention also provides a device 400 for predicting muscle loss after liver transplantation for liver cancer, including: a model prediction module 410, a risk classification module 420, a prognosis stratification module 430, and a survival prediction module 440.

[0084] The model prediction module 410 is used to construct an unbalanced random forest model containing a number of decision trees, and independently predict a number of high-risk or low-risk binary classification prediction results through the number of decision trees based on the input preprocessed multiple key clinical indicators of patients, and use the proportion of decision trees with high-risk prediction results among all decision trees as the prediction probability of muscle loss after liver transplantation for liver cancer finally output by the unbalanced random forest model.

[0085] The risk classification module 420 is used to determine that the postoperative muscle loss level is a high risk if the predicted probability is greater than or equal to the first risk threshold, determine that the postoperative muscle loss level is a medium risk if the predicted probability is less than the first risk threshold and greater than or equal to the second risk threshold, and determine that the postoperative muscle loss level is a low risk if the predicted probability is less than the second risk threshold, wherein the first risk threshold is greater than the second risk threshold.

[0086] The prognosis stratification module 430 is used to construct a decision tree model to stratify the patient's prognosis. Among them, for patients with SMI less than or equal to the stratification threshold, they are directly classified into the high-risk group; for patients with SMI greater than the stratification threshold and the muscle loss level determined to be a high risk, they are classified into the medium-risk group; for patients with SMI greater than the stratification threshold and the muscle loss level determined to be a medium risk or a low risk, they are classified into the low-risk group.

[0087] The survival prediction module 440 is used to predict the survival rate by combining the liver transplantation criteria and the patient's prognosis stratification results. Among them, for patients who meet the liver transplantation criteria, if the patient is in the high-risk group, the 1-year survival rate after surgery is 71.5%, the 3-year survival rate after surgery is 53.2%, and the 5-year survival rate after surgery is 36.3%; if the patient is in the medium-risk group, the 1-year survival rate after surgery is 86.4%, the 3-year survival rate after surgery is 74.6%, and the 5-year survival rate after surgery is 63.9%; if the patient is in the low-risk group, the 1-year survival rate after surgery is 92.3%, the 3-year survival rate after surgery is 85.9%, and the 5-year survival rate after surgery is 75.7%.

[0088] The above specific embodiments have elaborated in detail the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for predicting muscle loss after liver transplantation for liver cancer, characterized in that: The following steps are involved: An unbalanced random forest model including several decision trees is constructed. Several high-risk or low-risk binary prediction results are independently predicted by the decision trees based on multiple pre-processed clinical key indicators of the input patients. The proportion of decision trees with high-risk prediction results among all decision trees is used as the predicted probability of muscle loss after liver transplantation for liver cancer output by the unbalanced random forest model. Among them, multiple clinical key indicators of patients are screened from a wide range of clinical indicators of patients by recursive feature elimination method. Multiple clinical key indicators of patients include α-L-fucosidase, glycocholate dehydrogenase, age, alpha-fetoprotein, international normalized ratio, alanine aminotransferase, eosinophil count, basophil count, alkaline phosphatase, adenosine deaminase, glutamyl transpeptidase and body mass index in the peripheral blood of patients before surgery. If the predicted probability is greater than or equal to the first risk threshold, the postoperative muscle loss level is determined to be high risk; if the predicted probability is less than the first risk threshold and greater than or equal to the second risk threshold, the postoperative muscle loss level is determined to be medium risk; if the predicted probability is less than the second risk threshold, the postoperative muscle loss level is determined to be low risk, wherein the first risk threshold is greater than the second risk threshold; A decision tree model was constructed to stratify the prognosis of patients. Patients with SMI less than or equal to the stratification threshold were directly classified as the high-risk group, patients with SMI greater than the stratification threshold and muscle loss level judged as high risk were classified as the medium-risk group, and patients with SMI greater than the stratification threshold and muscle loss level judged as medium or low risk were classified as the low-risk group. The stratification threshold was set at 43.6 cm. 2 / m 2 ; Survival rate was predicted by combining liver transplantation criteria and patient prognostic stratification results.

2. The method for predicting muscle loss after liver transplantation for liver cancer according to claim 1, characterized in that: The preprocessing of multiple key clinical indicators of patients included missing value processing, outlier processing, feature coding conversion and data standardization.

3. The method for predicting muscle loss after liver transplantation for liver cancer according to claim 1, characterized in that: In the unbalanced random forest model, the decision tree construction process includes: A plurality of key clinical indicators of patients after pretreatment collected from different patients are constructed into a data set, wherein the different patients include patients with high risk of postoperative muscle loss and patients with low risk of postoperative muscle loss. Each sample in the data set represents a plurality of key clinical indicators of patients after pretreatment of a patient, and each key clinical indicator of the patient is used as a feature; For each decision tree, a part of the features of each category contained in all samples in the data set is randomly selected as the candidate feature set. Starting from the root node, an optimal segmentation feature and its segmentation point are selected from the candidate feature set according to the Gini impurity as the judgment criterion to split the samples in the data set into two subsets and generate two new child nodes. After that, the above segmentation process is recursively repeated for each new child node until the stopping condition is met. The stopping condition is that the preset maximum depth is reached or the number of samples in the node is less than the preset value.

4. The method for predicting muscle loss after liver transplantation for liver cancer according to claim 3, characterized in that: The calculation formula of Gini impurity is: , Among them, the high-risk proportion and the low-risk proportion are the proportion of patients with high risk of postoperative muscle loss and the proportion of patients with low risk of postoperative muscle loss in each subset obtained after each node splits the data set. Each decision tree will try various possible split features and split points, and select the split scheme that can minimize the sum of the Gini impurities of each node.

5. The method for predicting muscle loss after liver transplantation for liver cancer according to claim 1, characterized in that: The first risk threshold is set to 0.7, and the second risk threshold is set to 0.

3.

6. A device for predicting muscle loss after liver transplantation for liver cancer, implemented by the method for predicting muscle loss after liver transplantation for liver cancer according to any one of claims 1 to 5, characterized in that: include: Model prediction module, risk classification module, prognosis stratification module and survival prediction module; The model prediction module is used to construct an unbalanced random forest model including a plurality of decision trees, and independently predict a plurality of high-risk or low-risk binary classification prediction results based on a plurality of pre-processed clinical key indicators of the inputted decision trees, and the proportion of the decision trees with the prediction results of high risk to all the decision trees is used as the prediction probability of muscle loss after liver transplantation for liver cancer, which is finally output by the unbalanced random forest model; The risk classification module is used to determine the postoperative muscle loss level as high risk if the predicted probability is greater than or equal to a first risk threshold, determine the postoperative muscle loss level as medium risk if the predicted probability is less than the first risk threshold and greater than or equal to a second risk threshold, and determine the postoperative muscle loss level as low risk if the predicted probability is less than the second risk threshold, wherein the first risk threshold is greater than the second risk threshold; The prognostic stratification module is used to construct a decision tree model to stratify the patient prognosis, wherein patients whose SMI is less than or equal to the stratification threshold are directly classified as a high-risk group, patients whose SMI is greater than the stratification threshold and whose muscle loss level is determined to be high risk are classified as a medium-risk group, and patients whose SMI is greater than the stratification threshold and whose muscle loss level is determined to be medium risk or low risk are classified as a low-risk group; The survival prediction module is used to predict the survival rate by combining liver transplantation criteria and patient prognosis stratification results.

Citation Information

Patent Citations

  • Disease risk prediction method for improving random forest similarity measurement

    CN114091671A