Machine learning-based pressure damage interpretable prediction method and system
By processing clinical data based on machine learning, constructing and optimizing stress damage prediction models, the problems of data quality and sample imbalance in the prior art are solved, and the accuracy and interpretability of predictions are improved.
Patent Information
- Application Number
- CN202510063867.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art has problems with data quality, sample imbalance, and insufficient feature selection and data preprocessing in the prediction of pressure damage, which affects the reliability and accuracy of the prediction results.
Using a machine learning-based method, by acquiring and processing clinical data, feature coding, missing value processing, outlier value screening, sample equalization processing and feature selection are carried out, predictive models are constructed, and model parameters are optimized through cross-validation and grid search, feature contribution degree is calculated and feature importance summary charts are generated.
Improve the accuracy and reliability of the predictive model, enhance the generalization ability of the model, provide interpretability for the prediction of stress injury, help medical professionals better understand the model decision-making process, and guide clinical practice.
Smart Images

Figure CN119964805A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pressure injury prediction, and specifically relates to an explainable prediction method and system for pressure injuries based on machine learning. Background Art
[0002] In surgical inpatients, the occurrence of pressure ulcers (PI) may cause serious complications, interfere with wound healing, and even induce infection. Early identification and prediction of PI are crucial to prevent these complications. Currently, the prediction of PI mainly relies on assessment tools such as Waterlow Scale and Braden Scale, which rely heavily on the clinical experience and judgment of doctors. However, in the era of big data, this assessment method that relies on personal experience has limited the clinical value of medical data to a certain extent.
[0003] There are still areas for improvement:
[0004] Data quality issues: The data used in the study contain a large number of missing values and outliers, which usually stems from lax data registration and incomplete data collection and storage. This affects the model's ability to learn accurate patterns to a certain extent, which in turn affects the reliability of the prediction results. Sample imbalance problem: The incidence of pressure injuries is low relative to the overall data, and there is an imbalance between positive and negative samples, which may cause the model to tend to predict categories with higher occurrence frequencies and ignore categories with lower occurrence frequencies. Feature selection and data preprocessing: During the feature selection process, all selected features are relatively complete features, which may lead to the model's predictive ability being too idealized. In addition, due to the relative incompleteness of the data, some features were excluded, which had a positive impact on the subsequent construction of a more stable model. Summary of the invention
[0005] In order to solve the above problems existing in the prior art, the present invention provides an interpretable prediction method and system for pressure injury based on machine learning, comprising:
[0006] Acquire clinical data of hospitalized patients and perform feature encoding on the data to construct a data set, perform sample balancing on the data set and then perform feature selection to obtain feature variables;
[0007] Constructing a prediction model based on a machine learning algorithm, dividing the data set into a training set and a test set, adjusting the optimal parameter combination of the prediction model on the training set by a grid search method, and characterizing the prediction model with the optimal parameter combination; verifying the prediction model by a cross-validation method to obtain a performance evaluation index;
[0008] The contribution of each feature to the model prediction is calculated and a feature importance summary graph is generated, and the feature variables that have the greatest impact on pressure sores are identified based on the feature importance summary graph.
[0009] Specifically, the feature encoding method is: obtaining clinical data and performing missing value processing and outlier screening; for missing value processing, for numerical data, mean filling or median filling is used; for categorical data, mode filling is used; outlier screening uses a box plot method to identify outliers, and selects a threshold based on the distribution of the data to remove outliers; non-numeric categorical data in clinical data are coded with labels of 0 and 1, and one-hot coding is used for numerical data.
[0010] Specifically, the sample balancing processing method is as follows: minority class samples and majority class samples are defined by the number of samples of each category in the data set, the sample imbalance rate is calculated according to the minority class samples and the majority class samples, the sampling multiplier is set according to the sample imbalance rate, and the sampling multiplier is used to represent the number of synthetic samples generated; the distance from each sample in the minority class samples to other samples is calculated by Euclidean distance and 10 nearest neighbor samples are extracted to obtain the neighbor sample set corresponding to each sample; the samples in the minority class samples are linearly interpolated with samples randomly selected from the corresponding neighbor sample set to generate synthetic samples, and the synthetic samples are used to replace the corresponding samples in the minority class samples involved in the calculation, until the number of synthetic samples is reached and the sample balancing processing process is ended.
[0011] Specifically, the feature selection method is to measure the importance of the average gain obtained when the feature is split during the decision tree construction process through the XGBoost model and the number of samples covered when the feature is split, sort the features according to the importance measurement and select the important features.
[0012] Specifically, the grid search method is: exhaustively search the hyperparameters of the prediction model, evaluate the model performance under different parameter combinations through cross-validation, and thus determine the optimal parameter combination.
[0013] Specifically, the cross-validation method is k-fold cross-validation, which randomly divides the data set into k subsets, takes one of the subsets as the test set and the rest as the training set in turn, repeats k times, and finally takes the average of the k evaluation results as the performance evaluation indicator of the model.
[0014] Specifically, the performance evaluation indicators include accuracy, recall, F1 score and AUC value; the accuracy is used to characterize the proportion of correct predictions of the model; the recall is used to characterize the ability of the model to capture positive samples; the F1 score is used to comprehensively evaluate the accuracy and recall under conditions of sample imbalance; the AUC value is used to record the area under the characteristic curve drawn with the false positive rate as the horizontal axis and the true positive rate as the vertical axis.
[0015] Specifically, the contribution calculation method is: add or delete features of different combinations, create an interpreter object, call the interpreter for the incoming prediction model and test set, and calculate the SHAP value of the test set sample. The calculation formula is:
[0016]
[0017] Among them, φ i To represent the contribution of feature i, N represents the set of all features, S represents any feature subset excluding feature i, |S| represents the number of features in the feature subset, |N| represents the number of features in the total features, f(S) represents the contribution of only the feature subset, and f(S∪i) represents the contribution of the feature subset after adding feature i.
[0018] Specifically, the feature importance summary graph generation process includes collecting the prediction results and corresponding feature values of the model on the test set; using the collected data, generating a feature importance ranking by calculating the contribution of each feature to the model prediction results, and displaying the feature name labels and contribution value labels in an intuitive graphical form.
[0019] A machine learning-based interpretable prediction system for pressure injuries, comprising a data preparation module, a model prediction module, and a model interpretation module;
[0020] The data preparation module is used to obtain clinical data of hospitalized patients and perform feature encoding on the data to construct a data set, process missing values and filter out abnormal values on the data set, and perform feature selection to obtain feature variables;
[0021] The model prediction module is used to construct a prediction model based on a machine learning algorithm, divide the data set into a training set and a test set, adjust the optimal parameter combination of the prediction model on the training set by a grid search method, and characterize the prediction model with the optimal parameter combination; verify the prediction model by a cross-validation method to obtain a performance evaluation index;
[0022] The model interpretation module calculates the contribution of each feature to the model prediction and generates a feature importance summary graph, and identifies the feature variables that have the greatest impact on pressure sores based on the feature importance summary graph.
[0023] The beneficial effects of the present invention are as follows: by constructing a comprehensive data preparation module, the system can efficiently process and analyze a large amount of clinical data, ensure data quality, and thus improve the accuracy and reliability of the prediction model. Secondly, the model prediction module adopts advanced machine learning algorithms and grid search technology, which can automatically find the optimal combination of model parameters, which not only improves the prediction performance of the model, but also reduces the time cost of manual intervention and optimization. In addition, the application of the cross-validation method further ensures the generalization ability of the model, making the performance of the model on unknown data more stable. The introduction of the model interpretation module calculates the feature contribution and generates a feature importance summary diagram, and the system can intuitively show which factors have a significant impact on the prediction of pressure injuries. This not only helps medical professionals better understand the decision-making process of the model, but also guides clinical practice and provides a scientific basis for the prevention and treatment of pressure injuries. Ultimately, this interpretability enhances the transparency and credibility of medical decisions and helps improve the quality of patient care. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings.
[0025] Figure 1 A schematic diagram of a process flow of an explainable prediction method for pressure injury based on machine learning of the present invention;
[0026] Figure 2 This is a schematic diagram of the structure of an explainable prediction system for pressure injuries based on machine learning of the present invention. DETAILED DESCRIPTION
[0027] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the specific implementation methods, structures, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.
[0028] See also Figure 1-2 , a method and system for interpretable prediction of pressure injuries based on machine learning, comprising:
[0029] Acquire clinical data of hospitalized patients and perform feature encoding on the data to construct a data set, perform sample balancing on the data set and then perform feature selection to obtain feature variables;
[0030] Constructing a prediction model based on a machine learning algorithm, dividing the data set into a training set and a test set, adjusting the optimal parameter combination of the prediction model on the training set by a grid search method, and characterizing the prediction model with the optimal parameter combination; verifying the prediction model by a cross-validation method to obtain a performance evaluation index;
[0031] The contribution of each feature to the model prediction is calculated and a feature importance summary graph is generated, and the feature variables that have the greatest impact on pressure sores are identified based on the feature importance summary graph.
[0032] In this embodiment, the model is implemented using Python 3.7.4 version, using Jupyter Notebook as the development environment. Required libraries include Scikit-learn, XGBoost, SHAP, etc. Environment configuration is completed by installing the corresponding library through pip. The compilation method of the model involves steps such as data preprocessing, feature selection, model training and evaluation. Data preprocessing includes missing value processing and outlier removal, feature selection uses XGBoost model, and model training uses grid search and 10-fold cross validation. The present invention selects seven machine learning algorithms, including decision tree (DT), K nearest neighbor (KNN), logistic regression (LR), random forest (RF), support vector machine (SVM), extreme gradient boosting (XGBoost) and extra tree (ET). Model training uses a grid search method to find the optimal parameter combination. For example, the parameters of the XGBoost model include learning rate (0.01), sample ratio (0.8), number of base learners (100) and maximum depth (5). Use `XGBClassifier` to perform feature selection and model training from the XGBoost library. Use the `train_test_split` function from the Scikit-learn library to split the dataset into training and test sets. Use `cross_val_score` to perform cross-validation of the model. Use the `Explainer` class from the SHAP library to calculate SHAP values, such as `shap.TreeExplainer`. Use the `shap_values_numpy` function to draw a feature importance summary plot to show the contribution of each feature in the overall model prediction. Use the `shap.decision_plot` function to highlight only misclassified samples and analyze them in conjunction with the graphical representation of the experimental results.
[0033] Specifically, the feature encoding method is: obtaining clinical data and performing missing value processing and outlier screening; for missing value processing, for numerical data, mean filling or median filling is used; for categorical data, mode filling is used; outlier screening uses a box plot method to identify outliers, and selects a threshold based on the distribution of the data to remove outliers; non-numeric categorical data in clinical data are coded with labels of 0 and 1, and one-hot coding is used for numerical data.
[0034] In this embodiment, the gender characteristics in the clinical data are coded with labels of 0 and 1, the age, albumin, postoperative hospital stay, operation duration, blood sugar, body temperature, and respiratory rate in the clinical data are coded with numerical values, and the diagnostic information in the clinical data, such as hypertension, hyperlipidemia, diabetes, blood transfusion during surgery, surgical position, intraoperative dressing, oxygen saturation, malignant tumor, BMI, anesthesia method, anesthesia grade, pulse, diastolic blood pressure, systolic blood pressure, smoking, drinking, humidity, self-care ability grade, and intraoperative hypotension, are coded with unique hot encoding.
[0035] Specifically, the sample balancing processing method is as follows: minority class samples and majority class samples are defined by the number of samples of each category in the data set, the sample imbalance rate is calculated according to the minority class samples and the majority class samples, the sampling multiplier is set according to the sample imbalance rate, and the sampling multiplier is used to represent the number of synthetic samples generated; the distance from each sample in the minority class samples to other samples is calculated by Euclidean distance and 10 nearest neighbor samples are extracted to obtain the neighbor sample set corresponding to each sample; the samples in the minority class samples are linearly interpolated with samples randomly selected from the corresponding neighbor sample set to generate synthetic samples, and the synthetic samples are used to replace the corresponding samples in the minority class samples involved in the calculation, until the number of synthetic samples is reached and the sample balancing processing process is ended.
[0036] Specifically, the feature selection method is to measure the importance of the average gain obtained when the feature is split during the decision tree construction process through the XGBoost model and the number of samples covered when the feature is split, sort the features according to the importance measurement and select the important features.
[0037] Specifically, the grid search method is: exhaustively search the hyperparameters of the prediction model, evaluate the model performance under different parameter combinations through cross-validation, and thus determine the optimal parameter combination.
[0038] Specifically, the cross-validation method is k-fold cross-validation, which randomly divides the data set into k subsets, takes one of the subsets as the test set and the rest as the training set in turn, repeats k times, and finally takes the average of the k evaluation results as the performance evaluation indicator of the model.
[0039] Specifically, the performance evaluation indicators include accuracy, recall, F1 score and AUC value; the accuracy is used to characterize the proportion of correct predictions of the model; the recall is used to characterize the ability of the model to capture positive samples; the F1 score is used to comprehensively evaluate the accuracy and recall under conditions of sample imbalance; the AUC value is used to record the area under the characteristic curve drawn with the false positive rate as the horizontal axis and the true positive rate as the vertical axis.
[0040] Specifically, the contribution calculation method is: add or delete features of different combinations, create an interpreter object, call the interpreter for the incoming prediction model and test set, and calculate the SHAP value of the test set sample. The calculation formula is:
[0041]
[0042] Among them, φ i To represent the contribution of feature i, N represents the set of all features, S represents any feature subset excluding feature i, |S| represents the number of features in the feature subset, |N| represents the number of features in the total features, f(S) represents the contribution of only the feature subset, and f(S∪i) represents the contribution of the feature subset after adding feature i.
[0043] Specifically, the feature importance summary graph generation process includes collecting the prediction results and corresponding feature values of the model on the test set; using the collected data, generating a feature importance ranking by calculating the contribution of each feature to the model prediction results, and displaying the feature name labels and contribution value labels in an intuitive graphical form.
[0044] A machine learning-based interpretable prediction system for pressure injuries, comprising a data preparation module, a model prediction module, and a model interpretation module;
[0045] The data preparation module is used to obtain clinical data of hospitalized patients and perform feature encoding on the data to construct a data set, process missing values and filter out abnormal values on the data set, and perform feature selection to obtain feature variables;
[0046] The model prediction module is used to construct a prediction model based on a machine learning algorithm, divide the data set into a training set and a test set, adjust the optimal parameter combination of the prediction model on the training set by a grid search method, and characterize the prediction model with the optimal parameter combination; verify the prediction model by a cross-validation method to obtain a performance evaluation index;
[0047] The model interpretation module calculates the contribution of each feature to the model prediction and generates a feature importance summary graph, and identifies the feature variables that have the greatest impact on pressure sores based on the feature importance summary graph.
[0048] The computer storage medium of the embodiment of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device.
[0049] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0050] The program code included on the computer readable medium can be transmitted with any appropriate medium, including but not limited to wireless, electric wire, optical cable, RF, etc., or any suitable combination of the above. The computer program code for performing the operation of the present invention can be written in one or more programming languages or their combinations, and the programming language includes object-oriented programming languages-such as Java, Smalltalk, C++, and also includes conventional procedural programming languages-such as "C" language or similar programming languages. The program code can be executed completely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on the remote computer, or completely on the remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, using an Internet service provider to connect through the Internet).
[0051] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technical personnel in this field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A method for interpretable prediction of pressure injuries based on machine learning, characterized in that: include: Acquire clinical data of hospitalized patients and perform feature encoding on the data to construct a data set, perform sample balancing on the data set and then perform feature selection to obtain feature variables; Constructing a prediction model based on a machine learning algorithm, dividing the data set into a training set and a test set, adjusting the optimal parameter combination of the prediction model on the training set by a grid search method, and characterizing the prediction model with the optimal parameter combination; verifying the prediction model by a cross-validation method to obtain a performance evaluation index; The contribution of each feature to the model prediction is calculated and a feature importance summary graph is generated, and the feature variables that have the greatest impact on pressure sores are identified based on the feature importance summary graph.
2. The method according to claim 1, characterized in that The feature encoding method is: Clinical data were obtained and missing values and outlier screening were performed. For numerical data, missing values were filled with mean or median values. For categorical data, mode was used for filling. Outlier screening used the box plot method to identify outliers, and selected thresholds based on the distribution of the data to eliminate outliers. Non-numeric categorical data in clinical data were coded with labels of 0 and 1, and one-hot coding was used for numerical data.
3. The method according to claim 1, characterized in that The sample balancing method is as follows: minority class samples and majority class samples are defined by the number of samples of each category in a data set, a sample imbalance rate is calculated according to the minority class samples and the majority class samples, a sampling multiplier is set according to the sample imbalance rate, and the sampling multiplier is used to represent the number of synthetic samples generated; the distance between each sample in the minority class samples and other samples is calculated by Euclidean distance and 10 nearest neighbor samples are extracted to obtain a neighbor sample set corresponding to each sample; linear interpolation is performed between the samples in the minority class samples and the samples randomly selected from the corresponding neighbor sample set to generate synthetic samples, and the synthetic samples are used to replace the corresponding samples in the minority class samples that participate in the calculation, until the number of synthetic samples is reached and the sample balancing process is terminated.
4. The method according to claim 1, characterized in that: The feature selection method is to measure the importance of the average gain obtained when the feature is split during the decision tree construction process through the XGBoost model and the number of samples covered when the feature is split, sort the features according to the importance measurement and select the important features.
5. The method according to claim 1, characterized in that The grid search method is to perform an exhaustive search on the hyperparameters of the prediction model, evaluate the model performance under different parameter combinations through cross-validation, and thus determine the optimal parameter combination.
6. The method according to claim 1, characterized in that The cross-validation method is k-fold cross-validation, which randomly divides the data set into k subsets, takes one of the subsets as the test set and the rest as the training set in turn, repeats k times, and finally takes the average of the k evaluation results as the performance evaluation indicator of the model.
7. The method according to claim 1, characterized in that The performance evaluation indicators include accuracy, recall, F1 score and AUC value; the accuracy is used to characterize the proportion of correct predictions of the model; the recall is used to characterize the ability of the model to capture positive samples; the F1 score is used to comprehensively evaluate the accuracy and recall under the condition of sample imbalance; the AUC value is used to record the area under the characteristic curve drawn with the false positive rate as the horizontal axis and the true positive rate as the vertical axis.
8. The method according to claim 1, characterized in that The contribution calculation method is: add or delete features of different combinations, create an interpreter object, call the interpreter for the incoming prediction model and test set, and calculate the SHAP value of the test set sample. The calculation formula is: Among them, φ i To represent the contribution of feature i, N represents the set of all features, S represents any feature subset excluding feature i, |S| represents the number of features in the feature subset, |N| represents the number of features in the total features, f(S) represents the contribution of only the feature subset, and f(S∪i) represents the contribution of the feature subset after adding feature i.
9. The method according to claim 1, characterized in that: The feature importance summary graph generation process includes collecting the prediction results of the model on the test set and the corresponding feature values; using the collected data, generating a feature importance ranking by calculating the contribution of each feature to the model prediction results, and displaying the feature name labels and contribution value labels in the form of intuitive charts.
10. A machine learning-based pressure injury explainable prediction system for executing the method according to claims 1-9, characterized in that: Includes data preparation module, model prediction module, and model interpretation module; The data preparation module is used to obtain clinical data of hospitalized patients and perform feature encoding on the data to construct a data set, process missing values and filter out abnormal values on the data set, and perform feature selection to obtain feature variables; The model prediction module is used to construct a prediction model based on a machine learning algorithm, divide the data set into a training set and a test set, adjust the optimal parameter combination of the prediction model on the training set by a grid search method, and characterize the prediction model with the optimal parameter combination; verify the prediction model by a cross-validation method to obtain a performance evaluation index; The model interpretation module calculates the contribution of each feature to the model prediction and generates a feature importance summary graph, and identifies the feature variables that have the greatest impact on pressure sores based on the feature importance summary graph.