Prediction method, model and system for transformation from acute kidney injury to chronic kidney disease after operation

By constructing a machine learning prediction model combining multi-dimensional medical data, the problem of difficult prediction of the conversion risk of AKI to CKD after liver transplantation is solved, early identification and personalized intervention are achieved, and long-term prognosis of patients is improved.

CN120072291APending Publication Date: 2025-05-30THE THIRD AFFILIATED HOSPITAL OF SUN YAT SEN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510107588.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing technology lacks effective early identification and monitoring methods, making it difficult to accurately predict the risk of transformation of acute renal injury (AKI) into chronic kidney disease (CKD) after liver transplantation, resulting in high-risk patients not being able to receive timely intervention.

Method used

By constructing an efficient and accurate prediction model, combining multi-dimensional medical data, including preoperative, intraoperative and postoperative multi-stage variables, using machine learning technology, especially the support vector machine model, data preprocessing, screening and model training are performed to select the best-performing prediction model.

Benefits of technology

Early identification and accurate prediction of the risk of conversion of AKI patients to CKD after liver transplantation is achieved, helping clinicians to formulate personalized intervention strategies, reduce the incidence of postoperative chronic kidney disease, and improve patients' long-term prognosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072291A_ABST
    Figure CN120072291A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of risk prediction, in particular to a method, a model and a system for predicting transformation from postoperative acute kidney injury to chronic kidney disease. The method for constructing the prediction model for transformation from the acute kidney injury to the chronic kidney disease after the operation comprises the steps that clinical sample data of a patient who has been subjected to an allogeneic liver transplantation operation and is diagnosed to have AKI after the operation are collected, and the clinical sample data cover preoperative, intraoperative and postoperative multi-stage variables of the patient; variables in the clinical sample data are preprocessed and randomly divided into a training set and a verification set according to the proportion; screening variables in the training set to obtain a first training set containing a plurality of final variables; performing model training on the plurality of different machine learning models by using the first training set to obtain a plurality of candidate models; and the prediction performance of the candidate models is compared by using the verification set to obtain a final prediction model, so that a scientific basis is provided for early intervention and personalized management of the patient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of risk prediction, and particularly to a prediction method, model and system for the transformation of postoperative acute kidney injury to chronic kidney disease. Background Art

[0002] In recent years, the application of machine learning technology in the medical field has developed rapidly, especially showing important value in clinical decision support, disease prediction and personalized treatment. Based on machine learning prediction models, through analyzing clinical big data such as electronic medical records, efficient prediction of medical outcomes and early identification of high-risk patients can be achieved.

[0003] However, there is currently a lack of a prediction model for the transformation of acute kidney injury (AKI) to chronic kidney disease (CKD) after liver transplantation. Liver transplantation (LT) is the preferred treatment method for patients with end-stage liver disease. However, it is reported that the incidence of AKI after liver transplantation is as high as 17%-95%, and its severity is closely related to the increase in postoperative mortality, postoperative hospital stay, cardiovascular events and medical costs. Importantly, studies have shown that AKI is not only an important risk factor for the short-term prognosis of postoperative patients, but may also further develop into CKD through complex pathological mechanisms (such as inflammatory response, tubular injury and interstitial fibrosis), seriously affecting the long-term prognosis of patients, resulting in a decline in the quality of life of patients after surgery, an increase in disease and economic burdens, and a further increase in the risk of death. As a clinical syndrome with relatively high morbidity and mortality, even for patients with mild AKI, the risk of developing CKD in the long term is significantly increased after renal function recovery. Early prediction of high-risk populations with the progression of AKI to CKD and early intervention to prevent the chronic transformation of AKI are crucial. However, the progression of AKI to CKD involves multiple complex mechanisms, and there are currently no reliable means for early identification and monitoring.

[0004] Therefore, the development of an accurate prediction model based on machine learning is of great clinical significance for the early identification, risk stratification and intervention of the transformation of AKI patients after liver transplantation to CKD. Summary of the Invention

[0005] The purpose of the present invention is to combine multi-dimensional medical data to construct an efficient and accurate prediction model to solve the key technical problems in the prediction of the transformation of AKI after liver transplantation to CKD.

[0006] The first aspect of the present invention proposes a method for constructing a prediction model for the transformation of postoperative acute kidney injury to chronic kidney disease, including the following steps:

[0007] Collect clinical sample data of patients who have undergone allogeneic liver transplantation surgery and have been diagnosed with postoperative AKI, wherein the clinical sample data covers multi-stage variables of patients before, during and after surgery;

[0008] Preprocess the variables in the clinical sample data and randomly divide them into a training set and a validation set according to a ratio;

[0009] Screen the variables in the training set to obtain a first training set containing several final variables;

[0010] Use the first training set to train several different machine learning models to obtain several candidate models;

[0011] Use the validation set to compare the prediction performances of several candidate models, and select the candidate model with the best performance as the final prediction model.

[0012] A further improvement lies in that the specific method for preprocessing the variables in the clinical sample data includes:

[0013] Perform mapping processing on categorical variables and map them to binary values;

[0014] Use the mean to fill in the missing values of continuous variables and use the mode to fill in the missing values of categorical variables to ensure data integrity;

[0015] Use the standard deviation normalization method to perform data normalization processing on continuous variables.

[0016] A further improvement lies in that the specific method for screening the variables in the training set includes: initially screening the variables, including deleting variables with too many missing values, extremely deviated variables, serial number variables, and specified variables.

[0017] A further improvement lies in that the specific method for screening the variables in the training set also includes: further screening the variables remaining after the initial screening through statistical tests to select variables significantly related to the conversion of AKI to CKD, including:

[0018] Use the t-test for continuous variables;

[0019] Use the chi-square test for categorical variables;

[0020] Obtain the statistical test results of each variable, and select several variables with p-values < the preset value in the statistical test results as candidate variables.

[0021] A further improvement lies in that the specific method for screening the variables in the training set also includes: using the bootstrap method in combination with the LASSO regression method to screen the candidate variables to obtain the final variables.

[0022] A further improvement lies in that the specific method for using the first training set to train several different machine learning models to obtain several candidate models includes:

[0023] Select several different machine learning algorithms to construct several basic models. The several different machine learning algorithms include at least one or more of logistic regression algorithm, multi-layer perceptron classifier, support vector machine, random forest classifier, extreme gradient boosting classification tree, and light gradient boosting machine;

[0024] Use the grid search method combined with K-fold cross-validation to find the best parameter combination for each basic model. The specific method includes: setting the parameters and parameter value lists that need to be grid searched for each basic model, and performing cross-combinations; randomly dividing the first training set into K equal subsets, and taking turns using each equal subset as the test set, and the remaining K-1 equal subsets as the second training set to train and test each basic model under a specific parameter combination. Each basic model generates K evaluation indicators under each parameter combination, and the average value of the K evaluation indicators is taken as the model score of the corresponding basic model under the corresponding parameter combination; judging the best parameter combination of each basic model according to the model score;

[0025] Use the first training set and the best parameter combination to train the corresponding basic models to obtain several candidate models.

[0026] A further improvement is that the support vector machine realizes the classification of data by constructing a hyperplane, and strives to find the optimal hyperplane that can maximize the interval between different classes. The specific implementation path is as follows:

[0027] f(x) = ∑ i∈SV α i y i K(x, x i ) + b;

[0028] If f(x) > 0, then predict as the positive class; if f(x) < 0, then predict as the negative class;

[0029] Among them, x is a new input sample, and each sample includes the following 9 variable indicators: whether it is liver malignant tumor before surgery, the last albumin before surgery, the last estimated glomerular filtration rate before surgery, the last serum creatinine before surgery, the last urea nitrogen before surgery, the last activated partial thromboplastin time before surgery, anesthesia duration, whether the eGFR is lower than 60 ml / min / 1.73m 2 30 days after surgery, and AKD classification;

[0030] α i is the Lagrange multiplier, which is a matrix. Each row corresponds to a class, and each column corresponds to a support vector, and represents the weight of the support vector. A positive value means that this support vector tends to be predicted as the positive class, while a negative value tends to be predicted as the negative class;

[0031] yi y is the true label of the support vector; b is the bias term with a specific value of 0.335; SV represents the set of support vectors, which are the sample points in the training data;

[0032] K(x, x i ) is the kernel function, which is an exponentially decaying function of the distance between two vectors; when the parameter gamma is set to scale, SVM will automatically calculate the γ value according to the variance of all features, and the specific calculation is as follows:

[0033] K(x, x i ) = exp(-γ||x - x i || 2 );

[0034]

[0035] where n features is the number of features, which is 9; X var is the average variance of all features, which is 0.785; finally, γ is calculated to be 0.142;

[0036] is the square of the Euclidean distance between two vectors;

[0037] where x j is the j-th eigenvalue of vector x; x ij is the j-th eigenvalue of vector x i ; n is the number of features.

[0038] A further improvement lies in that the specific method for predicting and comparing the performance of several candidate models using a validation set includes:

[0039] Using the bootstrap resampling method, perform several resamplings with replacement on the validation set to obtain several test data sets, and use the several test data sets to predict and compare the candidate models. The comparison metrics include at least one or more of the area under the ROC curve, accuracy, sensitivity, specificity, and F1 score.

[0040] The second aspect of the present invention proposes a prediction model construction system for the transformation of acute kidney injury after surgery into chronic kidney disease, including:

[0041] A data acquisition module for collecting clinical sample data of patients who have undergone allogeneic liver transplantation surgery and have been diagnosed with postoperative AKI. Among them, the clinical sample data covers multi-stage variables of the patients before, during, and after surgery;

[0042] A data preprocessing module for preprocessing the variables in the clinical sample data and randomly dividing them into a training set and a validation set according to a ratio;

[0043] A data screening module, configured to screen variables in a training set to obtain a first training set containing a number of final variables;

[0044] A first processing module, configured to use the first training set to train a number of different machine learning models to obtain a number of candidate models;

[0045] A second processing module, configured to compare the prediction performances of a number of candidate models using a validation set, and select the candidate model with the best performance as the final prediction model.

[0046] In a third aspect of the present invention, a prediction system for the transformation of acute kidney injury to chronic kidney disease after surgery is proposed. The system includes an input device, a processor, and a computer-readable medium. The computer-readable medium stores a plurality of instructions. The input device is configured to obtain measurement values of relevant detection indexes of a liver transplant patient to be measured; the processor is connected to the input device, and the processor is configured to process the data obtained by the input device and output a prediction value of the risk of transformation of acute kidney injury to chronic kidney disease; the instructions instruct the input device and the processor to execute a method for predicting the transformation of acute kidney injury to chronic kidney disease after liver transplantation surgery. The method includes the following steps: 1) Obtain the measurement values of 9 indexes including whether the patient is a liver malignancy before liver transplantation, the last albumin before surgery, the last estimated glomerular filtration rate before surgery, the last serum creatinine before surgery, the last blood urea nitrogen before surgery, the last activated partial thromboplastin time before surgery, the anesthesia duration, and whether the eGFR is lower than 60 ml / min / 1.73m 2 and the AKD classification within 30 days after surgery; 2) Standardize the measurement values of the 9 indexes in step 1), and load a trained support vector machine model. Input the result parameters of the standardized 9 indexes into the trained support vector machine model for calculation to obtain a prediction value of the risk of transformation of acute kidney injury to chronic kidney disease.

[0047] The beneficial effects of the present invention are as follows:

[0048] By comprehensively integrating multi-stage variables before, during, and after surgery, the present invention captures the key features of the deterioration of renal function after liver transplantation; through the LASSO regression and multi-step screening methods for variables, the problem of variable multicollinearity is effectively avoided; through the performance comparison of multiple machine learning algorithms, the optimal support vector machine model is selected, providing a scientific basis for the early intervention and personalized management of patients. Clinicians can more accurately identify high-risk patients and formulate optimized intervention strategies to reduce the incidence of postoperative chronic kidney disease, thereby improving the long-term prognosis of patients. Description of the Drawings

[0049] Figure 1Flow chart of a method for constructing a prediction model for the conversion of acute kidney injury to chronic kidney disease in the embodiments of the present invention;

[0050] Figure 2 Importance ranking diagram of 9 final variables obtained in the embodiments of the present invention;

[0051] Figure 3 AUC curve comparison diagram of 6 candidate models in the embodiments of the present invention;

[0052] Figure 4 AUC curve comparison diagram of the final prediction model (SVM model) of the present invention and the traditional AKI-to-CKD model of the prior art. Detailed implementation manners

[0053] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0054] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0055] The following is an explanation of the algorithm models, performance indicators, and other terms involved in the present invention:

[0056] Logistic Regression (LR): Logistic regression is a generalized linear model used for binary or multi-class classification problems. It predicts the class label by estimating the probability of an event occurring. Logistic regression assumes a linear relationship between the features and the log-odds.

[0057] Multilayer Perceptron Classifier (MLP): A multilayer perceptron is a feedforward artificial neural network composed of multiple layers of nodes, where each node is connected to all nodes in the next layer. An MLP can have multiple hidden layers, enabling it to capture complex non-linear relationships in the data. The MLP is trained using the backpropagation algorithm to minimize the difference between the predicted output and the actual output.

[0058] Support Vector Machine (SVM): A support vector machine is a supervised learning model applicable to classification and regression analysis. Its basic idea is to find a hyperplane that can best separate the data points of two or more classes. The SVM attempts to maximize the margin between the data points closest to the hyperplane (referred to as support vectors) and the hyperplane, thereby improving the generalization ability of the classifier. For non-linearly separable data, the SVM can be mapped to a high-dimensional space through the kernel trick for linear separation.

[0059] Random Forest Classifier (RF): A random forest is an ensemble learning method that improves the accuracy of classification or regression by constructing multiple decision trees and taking the average of their results.

[0060] Extreme Gradient Boosting Tree (XGBoost): Extreme gradient boosting is an optimized algorithm based on the gradient boosting framework. It corrects the errors of the existing model by sequentially adding new models, and each new model focuses on the data that was misclassified previously.

[0061] Light Gradient Boosting Machine (LightGBM): Light Gradient Boosting Machine is an efficient gradient boosting framework developed by Microsoft. It adopts a leaf-node-based splitting strategy and histogram algorithm, which enable LightGBM to have faster speed and lower memory consumption when dealing with large datasets.

[0062] AUC (Area Under the ROC Curve): A comprehensive metric for measuring the discrimination ability of a model, ranging from 0.5 to 1.0. A higher value indicates stronger discrimination ability of the model.

[0063] Accuracy: The proportion of correct predictions, but its performance on imbalanced datasets may not be ideal.

[0064] Sensitivity: Also known as Recall, it refers to the proportion of true positives correctly identified.

[0065] Specificity: It refers to the proportion of true negatives that are correctly identified.

[0066] F1-score: It is the harmonic mean of Precision and Recall, and is particularly suitable for classification problems with imbalanced classes.

[0067] Grid Search is an exhaustive search method that traverses all possible combinations in the specified parameter space to find the optimal parameter settings.

[0068] K-fold Cross-Validation is to divide the training set into K mutually exclusive subsets (or "folds"), and then take turns using each subset as the validation set and the remaining K-1 subsets as the training set for K times of training and validation.

[0069] The early identification and intervention of the risk of acute kidney injury (AKI) transforming into chronic kidney disease (CKD) after liver transplantation remains a medical problem that urgently needs to be solved. Existing prediction tools do not fully consider the special pathophysiological characteristics of AKI patients after liver transplantation, such as the use of postoperative immunosuppressants, postoperative hemodynamic changes, and perioperative drug exposure, resulting in low applicability of the model in specific populations. Existing models are difficult to accurately predict the risk of AKI progressing to CKD after liver transplantation, which may lead to the inability to timely intervene in high-risk patients. Currently, most methods are based on single time points or limited indicators for prediction, ignoring the dynamic change laws of multi-stage indicators before, during, and after surgery in patients, which may affect the accurate screening and personalized management of high-risk populations. The occurrence of CKD after liver transplantation is latent, and once the patient enters the middle and late stages of CKD, the pathological changes are irreversible, posing a serious threat to the patient's quality of life and long-term prognosis. Therefore, existing methods lack effective risk assessment tools to guide timely intervention in the early stage of CKD occurrence.

[0070] Embodiments of the present invention aim to construct an efficient and accurate prediction model by combining multi-dimensional medical data to solve the key technical problems in predicting the transformation of AKI to CKD after liver transplantation. Through in-depth mining of a large-scale perioperative database, the present invention can provide a tool for clinically dynamically evaluating the risks of patients, helping doctors achieve early identification of high-risk patients, formulating targeted prevention and intervention strategies, and improving the long-term prognosis of postoperative patients.

[0071] The technical solutions of the embodiments of the present invention are introduced in detail below:

[0072] Please refer to the appendix Figure 1 - Appendix Figure 4, in the first aspect of the embodiments of the present invention, a prediction method, model and system for the transformation of postoperative acute kidney injury into chronic kidney disease are proposed. Generally speaking, the prediction of the transformation of postoperative acute kidney injury into chronic kidney disease is realized by constructing a prediction model for the transformation of postoperative acute kidney injury into chronic kidney disease. Among them, in this embodiment, as Figure 1 shown, a method for constructing a prediction model for the transformation of postoperative acute kidney injury into chronic kidney disease includes the following steps:

[0073] Step S1: Collect the clinical sample data of patients who have undergone allogeneic liver transplantation and have been diagnosed with postoperative AKI. Among them, the clinical sample data covers multi-stage variables of patients before, during and after surgery, and all variables are obtained through the hospital's electronic medical record system.

[0074] Specifically, the clinical sample data of the embodiments of the present invention is derived from the perioperative specialty database of the Third Affiliated Hospital of Sun Yat-sen University. The database covers the clinical and laboratory data of patients before, during and after surgery, including but not limited to demographic information (gender, age, BMI), preoperative medical history (such as diabetes, hypertension), intraoperative parameters (surgery time, intraoperative input and output), postoperative biochemical indicators (serum creatinine, estimated glomerular filtration rate), and postoperative medication conditions (immunosuppressants, contrast agents).

[0075] The inclusion criteria for clinical sample data are: (1) Patients who underwent liver transplantation at the Third Affiliated Hospital of Sun Yat-sen University from January 2015 to February 2023; (2) Patients diagnosed with AKI after liver transplantation according to the above diagnostic criteria; (3) Patients aged 18 or above. In addition, patients with the following conditions are excluded: (1) Patients who underwent combined liver and kidney transplantation; (2) Patients who received living donor transplantation; (3) Patients with pre-existing chronic kidney disease; (4) Patients who died within 90 days after surgery; (5) Patients lacking sufficient serum creatinine (SCr) and glomerular filtration rate (eGFR) records after surgery.

[0076] Step S2: Preprocess the variables in the clinical sample data and randomly divide them into a training set and a validation set according to a ratio.

[0077] Specifically, preprocessing can improve data quality and reduce the impact of noise on the model. In this embodiment, to obtain reliable evaluation and avoid overfitting, the collected clinical sample data is randomly divided into a training set (80%) and a validation set (20%). The training set contains 354 patients, and the validation set contains 89 patients, which are used for subsequent model training and performance evaluation. The distribution of AKI converted to CKD events in the training set and the test set is kept consistent to ensure the reliability of model evaluation. By using an independent validation set, the generalization ability of the model, that is, the performance of the model on unseen data, can be objectively evaluated, thus avoiding overfitting and ensuring good external validity of the model. Of course, those skilled in the art can also adjust the ratio according to actual needs, such as the training set (70%) and the validation set (30%).

[0078] Step S3: Screen the variables in the training set to obtain a first training set containing several final variables.

[0079] It can be understood that the purpose of screening is to identify the features that contribute most to the prediction target and remove redundant or irrelevant variables.

[0080] Step S4: Use the first training set to train several different machine learning models to obtain several candidate models.

[0081] It can be understood that different types of models may be good at capturing different types of data patterns. By comparing multiple models, the most suitable model for solving a specific problem can be selected, thereby improving the prediction performance.

[0082] Step S5: Compare the prediction performances of several candidate models using the validation set, and select the candidate model with the best performance as the final prediction model.

[0083] It can be understood that through prediction comparison, it is ensured that the selected final prediction model not only performs well on the training data, but also can maintain good prediction effects on new and unknown data, enhancing the practicability and credibility of the model.

[0084] In a preferred solution of this embodiment, the specific method for preprocessing the variables in the clinical sample data in step S2 includes:

[0085] Step S21: Perform mapping processing on categorical variables and map them to binary numerical values.

[0086] For example, map "yes", "is", "male", etc. to 1, and map "no", "no", "female", etc. to 0. Through this mapping, non-numerical data can be correctly interpreted by machine learning algorithms and thus participate in feature selection and model training.

[0087] Step S22: Use the mean to fill in the missing values of continuous variables and the mode to fill in the missing values of categorical variables to ensure data integrity.

[0088] It can be understood that for continuous variables, using the mean to fill in can maintain the distribution characteristics of the original data; for categorical variables, using the mode (the value with the highest frequency of occurrence) to fill in, because it reflects the most common category and reduces the impact on the data distribution. By filling in, it is ensured that all variables have complete values, avoiding model training failures or biases caused by missing values, retaining as many sample sizes as possible, and improving the robustness and generalization ability of the model.

[0089] Step S23: Use the standard deviation normalization method to perform data normalization processing on continuous variables.

[0090] It should be understood that standard deviation normalization (also known as Z-score normalization) is a common data scaling technique that adjusts the mean of each feature to 0 and the standard deviation to 1. The calculation formula for standard deviation normalization is x' = (x - mean) / std, where x' is the normalized value, x is the original value, mean is the mean, and std is the standard deviation. Through data normalization processing, the impact of features with different scales on the model can be made consistent, making the model training more stable.

[0091] In a preferred solution of this embodiment, the specific method for screening variables in the training set in step S3 includes: performing a preliminary screening of variables, including deleting variables with excessive missing values, extremely deviated variables, serial number variables, and specified variables.

[0092] In this embodiment, after preliminary screening, 68 variables remain, including albumin (ALB), estimated glomerular filtration rate (eGFR), serum creatinine (SCr), blood urea nitrogen (BUN), activated partial thromboplastin time (APTT), anesthesia time, eGFR 30 days after surgery, AKD classification, etc.

[0093] It should be understood that the preliminary screening of variables can be processed by combining the summary of previous technical literature and the actual situation of the Third Affiliated Hospital of Sun Yat-sen University. Variables with excessive missing values refer to variables with a missing value ratio exceeding a certain threshold (such as 30%). Extremely deviated variables refer to variables with a severely skewed data distribution that cannot be corrected by conventional methods. Serial number variables refer to variables used only as identifiers, such as medical record numbers. Specified variables refer to variables that are clearly irrelevant to the research objective according to the opinions of medical experts.

[0094] Through preliminary screening, the number of irrelevant or low-quality variables can be significantly reduced, focusing on those key factors that are most likely to affect the conversion of AKI to CKD.

[0095] In a preferred solution of this embodiment, the specific method for screening variables in the training set in step S3 further includes: further screening the variables remaining after the preliminary screening through statistical tests to select variables significantly related to the conversion of AKI to CKD, including:

[0096] Use a t-test for continuous variables.

[0097] Use a chi-square test for categorical variables.

[0098] Obtain the statistical test results of each variable, and select several variables with p-values < the preset value in the statistical test results as candidate variables. In this embodiment, the preset value is 0.05. The p-value represents the probability that the observed data conforms to the null hypothesis (i.e., the variable has no association with the target event). In this embodiment, there are 20 variables with p-values less than 0.05, that is, there are 20 candidate variables.

[0099] It should be understood that the t-test is a statistical test used to compare whether the means of two independent samples are equal. In the embodiments of the present invention, it can be used to compare the mean differences of continuous variables between two groups of patients with AKI not converted to CKD (control group) and AKI converted to CKD (experimental group). The t-test can help identify those continuous variables that have significant differences between the two groups, and these variables may be closely related to the conversion of AKI to CKD. By controlling the p-value threshold, variables that seem relevant due to random fluctuations can be excluded, improving the reliability and accuracy of the model.

[0100] The chi-square test is a non-parametric statistical test used to evaluate whether there is a significant association between two categorical variables. For categorical variables, the difference between the expected frequency and the observed frequency can be calculated by constructing a contingency table. The chi-square test helps to identify those categorical variables that play an important role in the process of AKI conversion to CKD. Similarly, by setting the p-value threshold, truly meaningful categorical variables can be selected, avoiding overfitting the model to irrelevant or accidentally related features.

[0101] Through the screening method combining the t-test and the chi-square test, variables that are significantly related to the conversion of AKI to CKD can be further selected from the variables after the preliminary screening. This method not only considers the statistical significance of the variables but also ensures the rationality and interpretability of the finally selected variables in practical applications. In addition, the strict p-value threshold setting helps prevent model overfitting and improves the generalization ability and prediction performance of the model.

[0102] In a preferred solution of this embodiment, the specific method for screening variables in the training set in step S3 further includes: using the bootstrap method in combination with the LASSO regression method to screen the candidate variables to obtain the final variables.

[0103] In this embodiment, after screening the candidate variables by combining the bootstrap method and the LASSO regression method, 9 final variable indicators are obtained. The 9 final variable indicators are respectively:

[0104] Whether it was liver malignant tumor before surgery (0 for no, 1 for yes).

[0105] The last albumin (ALB, g / L) before surgery.

[0106] The last estimated glomerular filtration rate (eGFR, mL / min / 1.73m 2 ) before surgery.

[0107] The last serum creatinine (SCr, μmol / L) before surgery.

[0108] The last blood urea nitrogen (BUN, mmol / L) before surgery.

[0109] The last activated partial thromboplastin time (APTT, s) before surgery.

[0110] The duration of anesthesia (min).

[0111] Whether the eGFR was lower than 60 ml / min / 1.73m 30 days after surgery 2 (0 for no, 1 for yes).

[0112] AKD classification (0 for no AKD, 1 for mild AKD, 2 for severe AKD).

[0113] It should be understood that the bootstrap method is a resampling technique that creates multiple new "bootstrap samples" by randomly sampling with replacement from the original dataset. By analyzing multiple bootstrap samples, the stability and variability of estimators can be evaluated, thus providing more robust results. LASSO (Least Absolute Shrinkage and Selection Operator) regression is a regularized linear regression method that shrinks coefficients by introducing an L1-norm penalty term. LASSO regression can directly shrink the coefficients of unimportant variables to zero, thus automatically eliminating these variables.

[0114] For example, the original dataset can be resampled 1000 times with replacement, randomly extracting the same number of samples from the dataset each time to construct 1000 different test subsets. Apply LASSO regression to each test subset, and finally select the variables with non-zero coefficients in the majority of the test subsets. Simply put, it is to generate multiple bootstrap samples through the bootstrap method and apply the LASSO model to each bootstrap sample. Then, it can be checked which variables are selected in most of the bootstrap samples. If a variable is selected in most of the bootstrap samples, then it is very likely to be an important predictor.

[0115] By combining the bootstrap method with LASSO regression in the present invention, the feature selection results of LASSO regression become more stable, reducing the fluctuations caused by a single dataset; combining the bootstrap method and LASSO regression can ensure that the finally selected variables not only perform well in a single dataset, but also have good predictive ability under multiple different data distributions. The variables with the most predictive ability and stability can be further selected from the candidate variables, improving the accuracy and generalization ability of the model and enhancing the interpretability and practicality of the model.

[0116] In a preferred solution of this embodiment, in step S4, the specific method of using the first training set to train several different machine learning models to obtain several candidate models includes:

[0117] Step S41: Select several different machine learning algorithms to construct several basic models.

[0118] Step S42: Use the grid search method combined with K-fold cross-validation to find the best parameter combination for each basic model.

[0119] Step S43: Use the first training set and the best parameter combination to train the corresponding basic models to obtain several candidate models.

[0120] It can be understood that by selecting multiple types of machine learning algorithms to construct basic models, different algorithms have their own advantages and disadvantages and are suitable for different types of data and problems. The most suitable model for the current dataset and prediction task can be found through comparison. By using K-fold cross-validation to find the best parameter combination for each basic model, the limited data can be utilized more fully and a more robust estimate of the model performance can be provided. By combining grid search with K-fold cross-validation, the parameter space can be systematically explored while maintaining the data utilization rate, so as to find the best parameter combination for each basic model. After determining the best parameter combination for each basic model, the entire first training set and these best parameters are used to retrain the model. This step ensures that the model is optimized on all available training data, rather than just based on the subsets in the cross-validation process.

[0121] Specifically, in step S42, the specific method of using the grid search method combined with K-fold cross-validation to find the best parameter combination for each base model includes:

[0122] Step S421: Set the parameters and parameter value lists that each base model needs to perform grid search, and perform cross-combinations.

[0123] Step S422: Randomly divide the first training set into K equal subsets, and take turns using each equal subset as the test set, and the remaining K-1 equal subsets as the second training set to train and test each base model under a specific parameter combination. Each base model generates K evaluation indicators under each parameter combination, and the average value of the K evaluation indicators is taken as the model score of the corresponding base model under the corresponding parameter combination.

[0124] Simply put, for each round (k = 1, 2,..., K), take turns selecting a subset as the test set, merge the remaining K-1 subsets into the second training set, use the second training set to train the base model, and evaluate it with the test set, and record the evaluation indicators of this round (such as accuracy, F1 score, AUC, etc.). In this embodiment, the value of K is preferably 5.

[0125] Step S423: Judge the best parameter combination of each base model according to the model score.

[0126] It can be understood that the average value of the K evaluation indicators is taken as the model score of the corresponding base model under the corresponding parameter combination. The model score reflects the overall level of the model performance under this parameter combination. At the same time, the K-fold cross-validation reduces the accidental error caused by a single division and improves the reliability of the scoring.

[0127] In this embodiment, the several different machine learning algorithms in step S41 include at least one or more of logistic regression algorithm, multi-layer perceptron classifier, support vector machine, random forest classifier, extreme gradient boosting classification tree, and light gradient boosting machine.

[0128] Preferably, in this embodiment, the above 6 machine learning algorithms are included at the same time, and a total of 6 candidate models are generated in step S43.

[0129] Specifically, in step S5, the specific method of using the validation set to predict and compare the performance of several candidate models includes:

[0130] Using the bootstrap resampling method, the validation set is resampled 1000 times with replacement to obtain 1000 test data sets. The candidate models are predicted and compared using the 1000 test data sets. The comparison metrics include at least one or more of the area under the ROC curve, accuracy, sensitivity, specificity, and F1 score. The evaluation metrics for the 1000 tests are expressed in the form of "median (2.5% quantile, 97.5% quantile)".

[0131] It can be understood that by testing the 6 candidate models with the validation set, the performance of each candidate model can be obtained, and the best-performing prediction model can be selected as the final prediction model.

[0132] In this embodiment, after prediction and comparison, the support vector machine (SVM) model performs excellently in all evaluation metrics. Its AUC is 0.838, sensitivity is 0.85, specificity is 0.73, and accuracy is 0.764, showing the best performance among all candidate models and demonstrating strong classification ability and stability.

[0133] Specific implementation mechanism of the support vector machine (SVM): The support vector machine is a supervised learning algorithm mainly used for classification and regression tasks. It performs well in processing data in high-dimensional spaces and can effectively handle the situation where the number of features is greater than the number of samples. SVM realizes data classification by constructing a hyperplane and strives to find the optimal hyperplane that can maximize the margin between different classes.

[0134] The specific implementation path is as follows:

[0135] f(x) = ∑ i∈SV α i y i K(x, x i ) + b;

[0136] If f(x) > 0, it is predicted as the positive class; if f(x) < 0, it is predicted as the negative class.

[0137] Among them, x is the new input sample, and each sample includes the following 9 variable indicators: whether it is liver malignant tumor before surgery (0 for no, 1 for yes), the last albumin (ALB, g / L) before surgery, the last estimated glomerular filtration rate (eGFR, mL / min / 1.73m 2 ) before surgery, the last serum creatinine (SCr, μmol / L) before surgery, the last blood urea nitrogen (BUN, mmol / L) before surgery, the last activated partial thromboplastin time (APTT, s) before surgery, the duration of anesthesia (min), whether the eGFR is lower than 60 ml / min / 1.73m 30 days after surgery 2(No is 0, Yes is 1), AKD classification (0 means no AKD, 1 means mild AKD, 2 means severe AKD).

[0138] α i is the Lagrange multiplier, which is a matrix where each row corresponds to a class (for this binary classification problem, it has only one row). Each column corresponds to a support vector and represents the weight of that support vector. A positive value means this support vector tends to be predicted as the positive class, while a negative value tends to be predicted as the negative class.

[0139] y i is the true label (0 / 1) of the support vector; b is the bias term, with a specific value of 0.335; SV represents the set of support vectors, which are the sample points in the training data.

[0140] K(x,x i ) is the kernel function, which is an exponentially decaying function of the distance between two vectors. When the parameter gamma is set to scale, the SVM will automatically calculate the γ value based on the variance of all features. Specifically, it will use the variance of all features in the training set to determine an appropriate γ value to ensure that the kernel function has appropriate sensitivity to features of different scales.

[0141] The specific calculation is as follows:

[0142] K(x,x i ) = exp(-γ||x - x i || 2 );

[0143]

[0144] where, n features is the number of features, which is 9; X var is the average variance of all features, which is 0.785; finally, γ is calculated to be 0.142.

[0145] is the square of the Euclidean distance between two vectors;

[0146] where, x j is the j-th eigenvalue of vector x; x ij is the j-th eigenvalue of vector x i ; n is the number of features.

[0147] To further verify the clinical applicability of the present invention, the performance of the SVM model of the present invention was directly compared with the traditional AKI-to-CKD model of James in the prior art. When applied to the dataset of the present invention, the AUC of the traditional model was 0.745 (95% CI: 0.59–0.879), and the sensitivity was 0.062 (95% CI: 0.0–0.214), indicating its deficiency in the ability to identify specific high-risk patients. In contrast, the SVM model of the present invention not only showed an advantage in AUC (0.838 vs 0.745), but also had a significant improvement in sensitivity (0.85 vs 0.062), indicating that the SVM model of the present invention can more accurately identify high-risk patients with the transformation of AKI to CKD after liver transplantation and significantly reduce the missed diagnosis rate.

[0148] In addition, to enhance the clinical applicability of the model, the embodiments of the present invention use SHAP values (Shapley Additive Explanations) for feature importance analysis, intuitively showing the contribution degree of key variables to the prediction results, and further clarifying the contribution degree of key variables to help clinicians quickly understand and apply the model results. The results show that AKD classification, the last preoperative eGFR, the last preoperative BUN, the postoperative 30-day eGFR less than 60 ml / min / 1.73 m 2 and the last preoperative SCr are the main five influencing factors for the model prediction results. Through the result display, not only the transparency of the model is enhanced, but also clear risk prompts are provided for clinicians to support personalized treatment decisions.

[0149] Using the prediction model of the present invention, clinicians can identify high-risk patients in the early postoperative stage, thus providing a scientific basis for personalized intervention. For example, the preoperative drug dosage can be optimized according to the prediction results, the postoperative treatment plan can be adjusted, and the follow-up frequency can be increased to reduce the incidence of CKD and improve the prognosis of patients. At the same time, the examination and treatment frequency of low-risk patients can be reasonably reduced to further optimize resource allocation and reduce medical costs.

[0150] The second aspect of the embodiments of the present invention proposes a prediction model construction system for the transformation of acute kidney injury to chronic kidney disease after surgery, which is used to execute any one of the prediction model construction methods for the transformation of acute kidney injury to chronic kidney disease after surgery in the first aspect of the embodiments, including:

[0151] A data acquisition module, configured to collect clinical sample data of patients who have undergone allogeneic liver transplantation surgery and have been diagnosed with postoperative AKI, wherein the clinical sample data covers multi-stage variables of patients before, during, and after surgery.

[0152] A data preprocessing module for preprocessing variables in clinical sample data and randomly dividing them into a training set and a validation set according to a ratio.

[0153] A data screening module for screening variables in the training set to obtain a first training set containing a number of final variables.

[0154] A first processing module for training a number of different machine learning models using the first training set to obtain a number of candidate models.

[0155] A second processing module for comparing the prediction performances of a number of candidate models using the validation set and selecting the candidate model with the best performance as the final prediction model.

[0156] In a third aspect of the embodiments of the present invention, a prediction system for the transformation of postoperative acute kidney injury into chronic kidney disease is proposed. The system includes an input device, a processor, and a computer-readable medium. The computer-readable medium stores a plurality of instructions. The input device is used to obtain measurement values of relevant detection indicators of the liver transplant patient to be measured. The processor is connected to the input device and is used to process the data obtained by the input device and output a prediction value of the risk of transformation of acute kidney injury into chronic kidney disease. The instructions direct the input device and the processor to execute a method for predicting the transformation of acute kidney injury into chronic kidney disease after liver transplantation. The method includes the following steps: 1) Obtain the measurement values of 9 indicators, including whether the patient had liver malignancy before liver transplantation, the last albumin before surgery, the last estimated glomerular filtration rate before surgery, the last serum creatinine before surgery, the last blood urea nitrogen before surgery, the last activated partial thromboplastin time before surgery, the duration of anesthesia, and whether the eGFR was lower than 60 ml / min / 1.73m 2 and the AKD classification 30 days after surgery; 2) Standardize the measurement values of the 9 indicators in step 1), load the trained support vector machine model, and input the result parameters of the standardized 9 indicators into the trained support vector machine model for calculation to obtain a prediction value of the risk of transformation of acute kidney injury into chronic kidney disease. In addition, the training method and implementation mechanism of the support vector machine model adopt those described in the first aspect of the embodiments of the present invention.

[0157] The above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for constructing a prediction model for the transformation of postoperative acute kidney injury to chronic kidney disease, characterized in that: The following steps are involved: Collect clinical sample data of patients who have undergone allogeneic liver transplantation and are diagnosed with postoperative AKI, wherein the clinical sample data covers multiple stage variables of the patients before, during and after the operation; The variables in the clinical sample data were preprocessed and randomly divided into training and validation sets according to the proportion; Screen the variables in the training set to obtain a first training set containing several final variables; Using the first training set to perform model training on several different machine learning models to obtain several candidate models; The prediction performance of several candidate models is compared using the validation set, and the prediction model with the best performance is selected as the final prediction model.

2. The method for constructing a prediction model for the transformation of postoperative acute kidney injury to chronic kidney disease according to claim 1, characterized in that: Specific methods for preprocessing variables in clinical sample data include: Map the categorical variables into binary values; Use the mean to fill in the missing values ​​of continuous variables and the mode to fill in the missing values ​​of categorical variables to ensure data integrity; The standard deviation method was used to standardize the data for continuous variables.

3. The method for constructing a prediction model for the transformation of postoperative acute kidney injury to chronic kidney disease according to claim 1, characterized in that: The specific method for screening variables in the training set includes: preliminary screening of variables, including deleting variables with too many missing items, extremely biased variables, ordinal variables, and specified variables.

4. The method for constructing a prediction model for the transformation of postoperative acute kidney injury to chronic kidney disease according to claim 3, characterized in that: The specific method for screening the variables in the training set also includes: further screening the variables significantly associated with the conversion of AKI to CKD through statistical tests on the variables remaining after the initial screening, including: The t test was used for continuous variables; Chi-square test was used for categorical variables; Obtain the statistical test results of each variable, and select several variables with p-values ​​< preset values ​​in the statistical test results as candidate variables.

5. The method for constructing a prediction model for the transformation of postoperative acute kidney injury to chronic kidney disease according to claim 4, characterized in that: The specific method for screening the variables in the training set also includes: screening the candidate variables by combining the bootstrap method with the LASSO regression method to obtain the final variables.

6. The method for constructing a prediction model for the transformation of postoperative acute kidney injury to chronic kidney disease according to claim 1, characterized in that: The specific method of using the first training set to perform model training on several different machine learning models to obtain several candidate models includes: Selecting several different machine learning algorithms to build several basic models, wherein the several different machine learning algorithms include at least one or more of a logistic regression algorithm, a multilayer perceptron classifier, a support vector machine, a random forest classifier, an extreme gradient boosting classification tree, and a lightweight gradient boosting machine; A grid search method combined with K-fold cross validation is used to find the best parameter combination for each basic model. The specific method includes: setting the parameters and parameter value list required for grid search for each basic model, and performing cross-combination; randomly dividing the first training set into K equal subsets, using each equal subset as a test set in turn, and using the remaining K-1 equal subsets as the second training set to train and test each basic model under a specific parameter combination, each basic model generates K evaluation indicators under each parameter combination, and the average of the K evaluation indicators is taken as the model score of the corresponding basic model under the corresponding parameter combination; judging the best parameter combination for each basic model according to the model score; The first training set and the best parameter combination are used to train the corresponding basic model to obtain several candidate models.

7. The method for constructing a prediction model for the transformation of postoperative acute kidney injury to chronic kidney disease according to claim 6, characterized in that: The support vector machine classifies data by constructing a hyperplane, striving to find the optimal hyperplane that can maximize the interval between different categories. The specific implementation path is as follows: f(x)=∑ i∈S Vα i yes i K(x,x i )+b; If f(x)>0, the prediction is positive; if f(x)<0, the prediction is negative; Among them, x is a new input sample, each sample includes the following 9 variable indicators: whether it is liver malignancy before surgery, the last albumin before surgery, the last estimated glomerular filtration rate before surgery, the last serum creatinine before surgery, the last urea nitrogen before surgery, the last activated partial thromboplastin time before surgery, the duration of anesthesia, whether the eGFR is less than 60ml / min / 1.73m 30 days after surgery 2 , AKD classification; α i It is the Lagrange multiplier, which is a matrix in which each row corresponds to a category and each column corresponds to a support vector, and represents the weight of the support vector. A positive value means that the support vector tends to be predicted as a positive class, while a negative value tends to be predicted as a negative class. y i is the true label of the support vector; b is the bias term, with a specific value of 0.335; SV represents the support vector set, which is the sample point in the training data; K(x,x i ) is the kernel function, which is an exponential decay function of the distance between two vectors; when the parameter gamma is set to scale, SVM will automatically calculate the gamma value based on the variance of all features, and the specific calculation is as follows: K(x,x i )=exp(-γ||x-x i || 2 ); Among them, n features is the number of features, which is 9; X var is the average variance of all features, which is 0.785; the final calculated γ is 0.142; is the square of the Euclidean distance between two vectors; Among them, x j is the jth eigenvalue of vector x; x ij is the vector x i is the j-th eigenvalue of ; n is the number of features.

8. The method for constructing a prediction model for the transformation of postoperative acute kidney injury to chronic kidney disease according to claim 1, characterized in that: Specific methods for using a validation set to compare the performance of several candidate models include: Using the bootstrap resampling method, the validation set is resampled several times with replacement to obtain several test data sets. The candidate models are compared using the several test data sets. The comparison indicators include at least one or more of the area under the ROC curve, accuracy, sensitivity, specificity and F1 score.

9. A prediction model construction system for the transformation of postoperative acute kidney injury to chronic kidney disease, characterized in that: include: A data acquisition module, used to collect clinical sample data of patients who have undergone allogeneic liver transplantation and are diagnosed with postoperative AKI, wherein the clinical sample data covers multi-stage variables of the patients before, during and after the operation; A data preprocessing module is used to preprocess the variables in the clinical sample data and randomly divide them into training sets and validation sets according to proportion; A data screening module is used to screen the variables in the training set to obtain a first training set containing several final variables; A first processing module is used to perform model training on a plurality of different machine learning models using a first training set to obtain a plurality of candidate models; The second processing module is used to compare the prediction performance of several candidate models using the validation set, and select the prediction model with the best performance as the final prediction model.

10. A prediction system for the transformation of postoperative acute kidney injury to chronic kidney disease, characterized in that: The system includes an input device, a processor and a computer-readable medium, wherein the computer-readable medium stores a plurality of instructions, wherein the input device is used to obtain the measured values ​​of the relevant detection indicators of the tested liver transplant patient; the processor is connected to the input device, and the processor is used to process the data obtained by the input device and output the predicted value of the risk of transformation from acute kidney injury to chronic kidney disease; the instructions instruct the input device and the processor to execute a method for predicting the transformation from acute kidney injury to chronic kidney disease after liver transplantation; the method includes the following steps: 1) obtaining whether the liver transplant patient has liver malignancy before surgery, the last albumin before surgery, the last estimated glomerular filtration rate before surgery, the last serum creatinine before surgery, the last urea nitrogen before surgery, the last activated partial thromboplastin time before surgery, the duration of anesthesia, and whether the eGFR is lower than 60 ml / min / 1.73 m3 30 days after surgery 2 , AKD classification these 9 indicators' measurement values; 2) standardizing the measurement values ​​of the 9 indicators in step 1), loading the trained support vector machine model, inputting the result parameters of the standardized 9 indicators into the trained support vector machine model for calculation, and obtaining the predicted value of the risk of transformation from acute kidney injury to chronic kidney disease.

Citation Information

Cited By

  • Liver function prediction method and device, electronic equipment and medium

    CN120473187A

  • Prediction method and system for postoperative early-stage bad results of craniopharyngeal tubuloma patient

    CN120895266A