Patient perioperative period adverse event risk prediction system and method based on electronic medical record data
Through in-depth mining and machine learning of electronic medical record data, a perioperative adverse event risk prediction model was established, which solved the problem of untimely and inaccurate perioperative risk identification in the existing technology, achieved accurate prediction of surgical complications and unplanned reoperation risks, and improved the quality of medical services.
Patent Information
- Application Number
- CN202510255093.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
AI Technical Summary
In the medical process related to surgery, patients' perioperative adverse event risk identification is not timely and inaccurate, resulting in frequent occurrence of surgical complications, unplanned reoperation, etc., and patients need to bear additional pain and financial burden.
Through in-depth mining and machine learning of electronic medical record data related to specific surgical risks, a patient's perioperative adverse event risk prediction model is established, and a comprehensive analysis is carried out in combination with multi-source data to provide comprehensive risk information warning.
Accurate prediction of the risks of perioperative adverse events such as unplanned resurgence and surgical complications of hospitalized patients has been achieved, providing doctors with scientific and reliable risk assessment basis, reducing medical disputes, and improving the quality of medical services.
Smart Images

Figure CN120183693A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical information technology, and particularly relates to a risk prediction system and method for perioperative adverse events of patients based on electronic medical record data. Background Art
[0002] With the advancement of medical informatization, electronic medical records have been widely used in hospitals, but there are still many problems in the process of use. Especially in the medical process related to surgery, the risk identification of perioperative adverse events of patients is not timely and accurate, and situations such as surgical complications and unplanned reoperations often occur. Patients need to bear additional pain and economic burdens, and it also has an adverse impact on the hospital. In the prior art, although some hospitals have begun to try to use data analysis for risk prediction and management, most technical solutions rely on a single data source and fail to conduct comprehensive analysis.
[0003] Therefore, it is of great practical significance to develop a medical system that can effectively predict the risks of perioperative adverse events such as unplanned reoperations and surgical complications of inpatients, and give early warning prompts. Summary of the Invention
[0004] The technical problem to be solved by the present invention is: to overcome the problems existing in the above-mentioned prior art, a risk prediction system for perioperative adverse events of patients based on electronic medical record data, through in-depth mining and machine learning of electronic medical record data related to specific surgical risks, establishes a risk prediction model for perioperative adverse events of patients, effectively predicts the risks of perioperative adverse events such as unplanned reoperations and surgical complications of inpatients, provides comprehensive risk information early warning for doctors, assists them in fully communicating with patients, improves the quality of medical services, reduces the risk of doctor-patient disputes, and safeguards the safety and rights of patients.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is:
[0006] A risk prediction system for perioperative adverse events of patients based on electronic medical record data, comprising:
[0007] A data layer, used to dock with various hospital information systems, collect, clean and store the electronic medical record data of inpatients, specifically including a data acquisition module, a data cleaning module and a data storage module which are electrically connected.
[0008] A feature layer, used to extract features related to risk prediction from the original medical record data stored in the data layer, and perform annotation and coding, specifically including a feature extraction module, a feature annotation module and a feature storage module which are electrically connected.
[0009] The model layer is used to construct, train, and optimize the risk prediction model for perioperative adverse events of patients. Specifically, it includes a model selection module, a model training module, a model evaluation module, and a model storage module that are electrically connected. The model layer extracts the feature information in the feature storage module to train the risk prediction model for perioperative adverse events of patients and conduct risk prediction.
[0010] The application layer is used to predict the risk of perioperative adverse events for new patients based on the risk prediction model for perioperative adverse events of patients and issue a warning prompt for the prediction results with a prediction probability above 0. Specifically, it includes a risk prediction module, a risk assessment module, and a warning prompt module that are electrically connected. After the medical record data of a new patient is entered into the hospital's various information systems, the data layer first collects and preprocesses the medical record data of the new patient, and then the feature layer extracts and labels the various features in the medical record data of the new patient. The risk prediction module conducts risk prediction on the medical record data of the new patient through the trained risk prediction model for perioperative adverse events of patients. The risk assessment module interprets and ranks the risk prediction results of perioperative adverse events, and the warning prompt module issues a warning prompt for the risk prediction results of perioperative adverse events.
[0011] The feedback layer is used to regularly compare the prediction evaluation issued by the risk prediction model for perioperative adverse events of patients with the actual situation data of the patients to conduct model training and optimization.
[0012] In another improved technical solution, the data collection module is docked with the hospital's electronic medical record system, surgical anesthesia management system, laboratory information system, and picture archiving and communication system to collect the medical record data of inpatients. The data cleaning module cleans, denoises, and standardizes the collected various types of electronic medical record data. The data storage module uses a relational database to store structured data and uses a NoSQL database to store unstructured data.
[0013] In another improved technical solution, the feature extraction module is used to extract the admission diagnosis, disease stage, past medical history, patient signs, test results, imaging results, and surgical information in the electronic medical record data of patients. The feature annotation module is used to perform one-hot encoding on categorical variables and standardize numerical data. The feature storage module is used to store the extracted features in the feature library for the model layer to conduct model training and for the application layer to conduct risk prediction of perioperative adverse events of patients.
[0014] In another improved technical solution, the model selection module adopts a hybrid model architecture that combines a neural network model based on deep learning and ensemble learning.
[0015] The model training module divides the data into a training set, a validation set, and a test set, trains a deep learning model using the backpropagation algorithm, and performs ensemble learning using random forests and gradient boosting decision trees.
[0016] The model evaluation module evaluates the model performance using the validation set and evaluates the model stability using cross-validation.
[0017] The model storage module stores the trained model in the model library for the risk prediction module to call.
[0018] In another improved technical solution, the process of training the patient perioperative adverse event risk prediction model in the model layer includes:
[0019] Data preprocessing: Clean, denoise, and standardize the collected medical record data.
[0020] Feature importance analysis: Feature importance scoring based on random forests and feature importance analysis based on gradient boosting decision trees.
[0021] Weight adjustment during model training: During the model training process, the neural network automatically adjusts the weight of each feature through the backpropagation algorithm.
[0022] Model training: Divide the preprocessed medical record data into a training set, a validation set, and a test set according to the ratio of 70:15:15. Use the training set to train the constructed hybrid model, adjust the weight parameters and biases of the neural network model through the backpropagation algorithm, and at the same time use random forests and gradient boosting decision trees in the ensemble learning algorithm to supervise and optimize the training process.
[0023] Model evaluation and optimization: Use the validation set to verify and optimize the model during the training process. By comparing the performance metrics such as accuracy, recall, F1 value, receiver operating characteristic curve, and the area under the curve of different model versions on the validation set, select the model version with the best performance for further optimization and adjustment, and use cross-validation technology to evaluate the stability and reliability of the model.
[0024] In another improved technical solution, after receiving the medical record data of a new patient, the risk prediction module calls the patient perioperative adverse event risk prediction model to predict the perioperative adverse event risk of the new patient, and outputs the probability values of surgical complications, unplanned reoperation, prolonged hospital stay, risk of returning to the ICU, and blood transfusion.
[0025] The risk assessment module is used to interpret the perioperative adverse event risk prediction results and rank the occurrence probabilities, and provide the types of possible complications and their probability values.
[0026] The early warning prompt module pops up an early warning window in the electronic medical record system according to the prediction results of the risk assessment module. The early warning content includes surgical complications, unplanned reoperation, prolonged hospital stay, risk of returning to the ICU, and blood transfusion requirements, and provides detailed risk explanations, treatment suggestions, and clinical guideline links.
[0027] In another improved technical solution, the risk prediction probability calculation formula for surgical complications by the risk prediction module is:
[0028]
[0029] Where P(complication) is the probability of complication occurrence, W is the weight vector, x is the input feature vector, and b is the bias term.
[0030] The risk prediction probability calculation formula for unplanned reoperation is:
[0031]
[0032] Where P(reoperation) is the probability of reoperation occurrence, W’ is the reoperation weight vector, x is the input feature vector, and b’ is the reoperation bias term, which are obtained through independent training respectively.
[0033] The present invention also provides a method for predicting the risk of perioperative adverse events of patients based on electronic medical record data. The risk prediction is carried out by using the above-mentioned risk prediction system for perioperative adverse events of patients based on electronic medical record data, including the steps:
[0034] S1. Obtain the electronic medical record data of the patient to be predicted, extract the feature data related to risk prediction from the original data, and perform annotation and coding. The feature data related to risk prediction specifically includes the patient's admission diagnosis information, disease stage, past medical history, general physical signs, personal information, examination results, imaging results, surgical information, and hospitalization information;
[0035] S2. Based on the perioperative adverse event risk prediction model of the patient and the feature data of the patient to be predicted, output the perioperative adverse event risk prediction result of the patient, that is, calculate and output the probability values of surgical complications, unplanned reoperation, prolonged hospital stay, risk of returning to the ICU, and blood transfusion;
[0036] The construction process of the perioperative adverse event risk prediction model of the patient includes:
[0037] S21. Collect the electronic medical record information of inpatients in each information system of the hospital at the data layer to construct a data set, and the data set contains the patient's medical feature data and the patient's medical record information;
[0038] S22. Use the medical record information of the patients in the dataset as training labels, process the patients' medical feature data in the dataset according to the following steps, and perform model training to obtain a risk prediction model for perioperative adverse events. The model training process is as follows:
[0039] Data preprocessing: Clean, denoise, and standardize the collected medical record data. For numerical data, use the Z-score standardization method to make its mean 0 and variance 1. For categorical variables, use one-hot encoding for conversion.
[0040] Formula: For numerical data x, the standardized data is:
[0041] where μ is the mean and σ is the standard deviation.
[0042] Feature importance analysis: Before model training, perform importance analysis on features, including feature importance scoring based on random forest and feature importance analysis based on gradient boosting decision tree, and preliminarily determine important features.
[0043] Weight adjustment in model training: During model training, the neural network automatically adjusts the weights of each feature through the backpropagation algorithm. The adjustment of weights is based on the gradient descent of the loss function, and continuously optimizes the weights to make the prediction results more accurate.
[0044] Model training: Divide the processed medical record data into a training set, a validation set, and a test set according to the ratio of 70:15:15. Use the training set to train the constructed hybrid model, adjust the weight parameters and biases of the neural network model through the backpropagation algorithm, and at the same time use the random forest and gradient boosting decision tree in the ensemble learning algorithm to supervise and optimize the training process.
[0045] Model evaluation and optimization: Use the validation set to verify and tune the model during the training process. By comparing the performance indicators of different model versions on the validation set, such as accuracy, recall rate, F1 value, receiver operating characteristic curve, and area under the curve, select the model version with the best performance for further optimization and adjustment, and use cross-validation technology to evaluate the stability and reliability of the model.
[0046] S3. Pop up a warning message in the electronic medical record system for patients with a predicted probability of perioperative adverse events above 0.
[0047] In another improved technical solution, the model calculation formula in step S22 is:
[0048] The multi-layer perceptron (MLP) model is a typical feedforward neural network model, and the calculation formula of its output layer is:
[0049] y = σ(W·x + b)
[0050] where y is the output result, W is the weight matrix, x is the input feature vector, b is the bias term, and σ is the activation function.
[0051] For the Random Forest (RF) model, its output is the voting result of multiple decision trees:
[0052] y = mode(T1(x), T2(x),..., T n (x))
[0053] where T i (x) is the prediction result of the i-th decision tree, and mode is the mode function.
[0054] For the decision-making process of Gradient Boosting Decision Tree (GBDT), different from the random forest, the gradient boosting decision tree generates the prediction result through an iterative additive model, fitting the residual of the previous round of prediction in each round, and the final output is the weighted sum of all decision trees:
[0055]
[0056] where M is the total number of decision trees, h m (x) is the prediction result of the m-th decision tree, and x m is the weight coefficient.
[0057] In another improved technical solution, it further includes step S4: updating and optimizing the risk prediction model for perioperative adverse events of patients. The feedback layer incorporates new discharge medical record data into the existing data set, and then repeats the steps of data preprocessing, feature extraction, model training, and evaluation to retrain and optimize the risk prediction model.
[0058] The beneficial technical effects achieved by the technical solution of the present invention are:
[0059] The present invention is a risk prediction system for perioperative adverse events of patients based on electronic medical record data. By comprehensively analyzing multi-source data such as the electronic medical record data of inpatients, the surgical anesthesia management system, the laboratory information system, and the picture archiving and communication system, it can predict in advance the potential risks of perioperative adverse events such as surgical complications and unplanned reoperations, so as to provide support for doctors and patients, assist in doctor-patient communication, reduce doctor-patient disputes, and improve the quality of medical services.
[0060] Accurate risk prediction: Through in-depth mining of a large amount of surgical-related medical record data and machine learning, the prediction system of the present invention can accurately identify the risks of potential perioperative adverse events, effectively predict the occurrence probabilities of risk events such as surgical complications, unplanned reoperations, and blood transfusions, provide a scientific and reliable basis for risk assessment for doctors, help doctors formulate countermeasures in advance, reduce medical risks, and improve medical quality.
[0061] Improve doctor-patient communication: The system provides doctors with detailed risk information and personalized doctor-patient communication suggestions, enabling doctors to communicate fully and effectively with patients before surgery, allowing patients to fully understand the surgical risks and possible situations, enhancing patients' trust in doctors, improving patients' treatment compliance, and thus effectively reducing doctor-patient disputes and misunderstandings caused by information asymmetry.
[0062] Enhance medical efficiency: Through an automated risk assessment and early warning mechanism, doctors can timely discover the risks of potential perioperative adverse events, avoid medical delays caused by negligence or untimely information, reasonably arrange medical resources, optimize the diagnosis and treatment process, and improve the overall medical efficiency and service level of the hospital.
[0063] Continuous learning and optimization: The system has the ability of self-learning and updating. As new patient medical record data accumulates, it can continuously optimize and update the risk prediction model, adapt to the changes and developments in clinical practice, always maintain a high prediction accuracy and practicality, and provide strong support for the long-term improvement of the hospital's medical quality. Brief Description of the Drawings
[0064] The drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0065] Figure 1 It is a schematic diagram of the overall architecture of the patient perioperative adverse event risk prediction system based on electronic medical record data of the present invention.
[0066] Figure 2 It is a schematic diagram of the internal architecture of the patient perioperative adverse event risk prediction system based on electronic medical record data of the present invention.
[0067] Figure 3 It is a schematic diagram of the main steps of the patient perioperative adverse event risk prediction method based on electronic medical record data of the present invention. Detailed Description of the Invention
[0068] The present invention will be described in detail below with reference to the accompanying drawings. The description in this part is merely exemplary and explanatory, and should not have any restrictive effect on the protection scope of the present invention. In addition, those skilled in the art can make corresponding combinations of the features in the embodiments and different embodiments according to the description of this document.
[0069] Referring to Figure 1 and Figure 2 As shown, this embodiment is a risk prediction system for perioperative adverse events of patients based on electronic medical record data, specifically including:
[0070] (1) Data layer: Data collection and integration
[0071] Responsible for docking with various hospital information systems, collecting, cleaning, and storing data. The specific modules include:
[0072] 1. Data collection module: Establish seamless docking with multiple data sources such as the hospital's electronic medical record system (EMR), surgical anesthesia management system, laboratory information system (LIS), and picture archiving and communication system (PACS) to ensure that the patient medical record data related to surgical complications, unplanned reoperation, surgical time exceeding 6 hours, readmission to the ICU, hospital stay exceeding 30 days, blood transfusion, etc. can be comprehensively and accurately collected.
[0073] 2. Data cleaning module: Clean, denoise, and standardize various types of collected data, remove duplicate, incorrect, or incomplete data records, and uniformly convert data in different formats into a structured data format for subsequent analysis and processing. For example, unify the date format to "YYYY-MM-DD" and ensure consistent units for numerical data, etc.
[0074] 3. Data storage module: Use relational databases (such as MySQL, Oracle) to store structured data, and use NoSQL databases (such as MongoDB) to store unstructured data (such as imaging reports).
[0075] (2) Feature layer: Feature extraction and annotation
[0076] The feature layer extracts features related to risk prediction from the original data, and performs annotation and encoding. The specific modules include:
[0077] 1. Feature extraction module: Extract admission diagnosis, disease stage, past medical history, patient signs, test results, imaging results, surgical information, etc.
[0078] 2. Feature annotation module: Perform one-hot encoding on categorical variables and z-score standardization on numerical data.
[0079] 3. Feature storage module: Store the extracted features in the feature library for model training and prediction.
[0080] The specific operations are as follows:
[0081] 2.1 Admission diagnosis: Precisely extract the patient's admission diagnosis information from the medical record, including the primary diagnosis, secondary diagnosis, and diagnostic basis. For the primary diagnosis, record in detail the name of the disease, ICD code (International Classification of Diseases code), and the way of diagnosis determination (such as clinical symptoms, laboratory tests, imaging examinations, etc.).
[0082] 2.2 Disease staging: For diseases with clear staging criteria, such as various cancers (using the TNM staging system), cardiovascular diseases (such as NYHA heart function classification), etc., accurately extract the disease staging data from information such as examination reports and pathological results in the medical record and perform digital annotation.
[0083] 2.3 Past medical history: Comprehensively sort out the patient's past medical history, including the duration of chronic diseases (such as hypertension, diabetes, coronary heart disease, chronic obstructive pulmonary disease, etc.), treatment conditions (types of medications, dosages, treatment time, and whether the treatment is regular), and disease control conditions (such as blood pressure control range for hypertensive patients, glycated hemoglobin level for diabetic patients, etc.).
[0084] 2.4 General patient signs: Include education level, income, age, gender, height, weight, blood pressure, etc. The education level is divided into seven levels: primary school and below, junior high school, high school / technical secondary school, junior college, undergraduate, master's degree, and doctoral degree and above, and are coded with numbers "1" to "7" respectively. The income combines the economic development level and income statistics data of the hospital's location area, and divides the patient's family annual income into five intervals: low income, medium - low income, medium income, medium - high income, and high income, and are represented by numbers "1" to "5" respectively.
[0085] 2.5 City of residence, whether born in urban or rural area: Record in detail the patient's residential address information. Classify cities into first - tier, second - tier, third - tier, fourth - tier and below cities, and are represented by numbers "1" to "4" respectively; classify the place of birth into urban and rural areas, and are marked with numbers "0" and "1".
[0086] 2.6 Marital status: Divide the marital status into four categories: unmarried, married, divorced, and widowed, and are marked with numbers "1" to "4" respectively.
[0087] 2.7 Post - admission examinations and test results: Systematically collect various laboratory test data of patients after admission, including blood routine, biochemical indicators, immune indicators, endocrine indicators, etc. Standardize these indicators, and convert the test results into quantitative forms such as the number of abnormal indicators, the score of abnormal degree, and the trend of indicator changes (such as increase, decrease, stability, etc.) according to the normal reference ranges of different indicators.
[0088] 2.8 Imaging examination results: Collect imaging examination reports of patients such as X - ray, CT, MRI, ultrasound, nuclear medicine, etc., and extract key imaging diagnosis results, including detailed information such as the location, size, shape, density, boundary, and relationship with surrounding tissues of the lesion.
[0089] 2.9 Surgical method and name: Accurately record the detailed operation method and standard name of the surgery, and accurately code the surgical method according to the hospital - formulated surgical classification catalog and the authoritative surgical name coding system (such as the ICD - 9 - CM - 3 surgical operation classification coding).
[0090] 2.10 Duration of surgery: Directly obtain the start time and end time of the surgery, accurately calculate the duration of the surgery (unit: minutes), and classify and label it according to a certain time interval.
[0091] 2.11 Professional titles of surgeons: Extract the professional title information of the primary surgeon and assistant surgeons from the surgical record. Classify the professional titles into four levels: resident physician, attending physician, deputy chief physician, and chief physician, and label them with numbers "1" to "4" respectively.
[0092] 2.12 Blood loss during surgery: Accurately record the blood loss value during the surgery (unit: milliliters), and formulate a reasonable blood loss grading standard according to the surgical type and the patient's physical condition (such as weight, underlying diseases, etc.).
[0093] 2.13 Whether blood transfusion: Record whether the patient received blood transfusion during the surgery as a binary variable, using the number "0" to indicate no blood transfusion and "1" to indicate blood transfusion.
[0094] 2.14 Types and durations of antibiotic use: Detailedly record the types of antibiotics used by the patient during the peri - operative period (including information such as the name, dosage form, and specification of the antibiotics) and the duration of use (unit: days).
[0095] 2.15 Length of hospital stay: Accurately record the admission date and discharge date of the patient, and calculate the length of hospital stay.
[0096] (III) Model layer: Machine learning model training and optimization
[0097] The model layer is responsible for constructing, training, and optimizing the risk prediction model. The specific modules include:
[0098] 1. Model Selection Module: A hybrid model framework that combines a neural network model based on deep learning (such as a multi-layer perceptron MLP, Multi-Layer Perceptron) and an ensemble learning algorithm (such as a random forest RF, Random Forest, and gradient boosting decision tree GBDT, Gradient Boosting Decision Tree).
[0099] 2. Model Training Module: Divide the data into a training set, a validation set, and a test set (with proportions of 70:15:15 respectively). Use the backpropagation algorithm to train the multi-layer perceptron MLP model, and use the random forest RF and gradient boosting decision tree GBDT for ensemble learning.
[0100] 3. Model Evaluation Module: Use the validation set to evaluate the model performance (accuracy, recall rate, F1 value, area under the curve AUC), and use cross-validation to evaluate the model stability.
[0101] 4. Model Storage Module: Store the trained model in the model library for the prediction module to call.
[0102] Model Training Process:
[0103] Data Preprocessing: Clean, denoise, and standardize the collected medical record data. For numerical data, use the Z-score standardization method to make its mean 0 and variance 1. For categorical variables, use one-hot encoding (One-Hot Encoding) for conversion.
[0104] Formula: For numerical data x, the standardized data is:
[0105]
[0106] where μ is the mean and σ is the standard deviation.
[0107] Feature Importance Analysis: Before model training, conduct feature importance analysis. Common methods include feature importance scoring based on random forests and feature importance analysis based on gradient boosting decision trees. Through these methods, it is possible to initially determine which features have a greater impact on the model's prediction results.
[0108] Weight Adjustment in Model Training: During model training, a neural network model (such as a multi-layer perceptron MLP) automatically adjusts the weights of each feature through the backpropagation algorithm. The adjustment of weights is based on the gradient descent of the loss function, and the model will continuously optimize the weights to make the prediction results more accurate.
[0109] Specific manifestations of weights: For example, features such as the duration of surgery, the amount of bleeding during surgery, and the age of the patient may have relatively high weights because these features are closely related to the surgical risk; while features such as marital status and educational level have relatively low weights because their direct impact on surgical risk is relatively small.
[0110] Model training: The processed medical record data is divided into a training set, a validation set, and a test set in a ratio of 70:15:15. The constructed hybrid model is trained using the training set, and the weight parameters and biases of the neural network model are adjusted through the backpropagation algorithm. At the same time, the random forest and gradient boosting decision tree in the ensemble learning algorithm are used to supervise and optimize the training process.
[0111] Model evaluation and optimization: The model during the training process is validated and tuned using the validation set. By comparing performance metrics such as accuracy, recall, F1 value, receiver operating characteristic curve (ROC curve), and area under the curve (AUC) of different model versions on the validation set, the model version with the best performance is selected for further optimization and adjustment. Cross-validation techniques (such as 5-fold cross-validation or 10-fold cross-validation) are used to evaluate the stability and generalization ability of the model, and the stability and reliability of the risk prediction model are evaluated. Cross-validation divides the dataset into K equal parts (folds). Each time, K - 1 parts are used as the training set, and the remaining 1 part is used as the validation set. This is repeated K times, and each part of the data is used as the validation set once. Finally, the average value of the K validation results is taken as the model performance metric. For example, 5-fold cross-validation divides the data into 5 parts and trains and validates 5 times; 10-fold cross-validation divides it into 10 parts and loops 10 times. This method can effectively reduce the evaluation bias caused by the randomness of data division and ensure the performance consistency of the model on different data subsets.
[0112] Model calculation formula:
[0113] For the multi-layer perceptron (MLP) model, the calculation formula for its output layer is:
[0114] y = σ(W·x + b)
[0115] where y is the output result, W is the weight matrix, x is the input feature vector, b is the bias term, and σ is the activation function (such as Sigmoid or ReLU).
[0116] For the random forest (RF) model, its output is the voting result of multiple decision trees:
[0117] y = mode(T1(x), T2(x),..., T n (x))
[0118] where T i (x) is the prediction result of the i-th decision tree, and mode is the mode function.
[0119] For the decision-making process of Gradient Boosting Decision Tree (GBDT), different from Random Forest, Gradient Boosting Decision Tree generates prediction results through an iterative additive model, fitting the residuals of the previous round of predictions in each round, and the final output is the weighted sum of all decision trees:
[0120]
[0121] where M is the total number of decision trees, h m (x) is the prediction result of the m-th decision tree, and x m is the weight coefficient.
[0122] (IV) Application Layer: Risk Prediction and Assessment
[0123] The application layer conducts risk prediction based on the trained model and provides early warnings and tips. The specific modules include:
[0124] 1. Risk Prediction Module: Input the medical record data of a new patient, call the model for prediction, and output the probabilities of risks such as surgical complications, unplanned reoperation, prolonged hospital stay, risk of returning to the ICU, blood transfusion, etc.
[0125] 2. Risk Assessment Module: Interpret and rank the prediction results, and provide the types of possible complications and their probabilities.
[0126] 3. Early Warning Module: Trigger early warnings according to the prediction results. The warning content includes surgical complications, unplanned reoperation, prolonged hospital stay, risk of returning to the ICU, blood transfusion requirements, etc.
[0127] 4. Tip Module: Pop up a warning window in the electronic medical record system, providing detailed risk explanations and links to clinical guidelines.
[0128] The specific processing steps of the application layer are as follows:
[0129] After the medical record data of a new patient is entered into the system, the data collection and integration module first collects and preprocesses it, and then the feature extraction and annotation module extracts and annotates the features in the medical record, and inputs the processed data into the trained risk prediction model.
[0130] Risk Prediction Formula: For the risk prediction of surgical complications, the probability output by the model is:
[0131]
[0132] where P(Complication) is the probability of the complication occurring, W is the weight vector, x is the input feature vector, and b is the bias term.
[0133] For the risk prediction of unplanned reoperation, the probability output by the model is:
[0134]
[0135] Among them, P(reoperation) is the probability of reoperation, W' is the reoperation weight vector, x is the input feature vector, and b' is the reoperation bias term.
[0136] Risk assessment results: The risk assessment results are presented in a quantitative form. For example, the probability of surgical complications is expressed as a percentage, and at the same time, the specific types of possible complications are predicted and ranked, and the top several possible complications are listed in descending order of occurrence probability; for the risk of unplanned reoperation, it is also expressed as a probability value, and the reasons that may lead to reoperation (such as surgical site infection, postoperative bleeding, anastomotic leakage, postoperative wound dehiscence, etc.) and their occurrence probabilities are analyzed.
[0137] Once the risk prediction and assessment module determines that a patient is a high-risk patient (that is, the probability of perioperative adverse events occurring in this patient is greater than 0), the early warning prompt module is immediately activated, and early warning information is sent through means such as a prompt window and an alarm. The early warning content includes but is not limited to:
[0138] Surgical complication early warning: including but not limited to postoperative infection, bleeding, thrombosis, organ failure, etc.
[0139] Unplanned reoperation early warning: including surgical site infection, postoperative bleeding, anastomotic leakage, postoperative wound dehiscence, etc.
[0140] Prolonged hospital stay early warning: predicting that the patient's hospital stay may exceed 30 days.
[0141] Return to ICU early warning: predicting that the patient may need to return to the ICU after surgery.
[0142] Blood transfusion early warning: predicting that the patient may need blood transfusion during or after surgery.
[0143] Anesthesia risk early warning: predicting high-risk situations that may occur during the patient's anesthesia, such as respiratory depression, arrhythmia, etc.
[0144] Early warning method: A prominent prompt window pops up on the interface of the hospital's electronic medical record system, displaying the patient's basic information, surgical arrangement information, and risk assessment results, including key information such as possible complications and the probability of unplanned reoperation, and at the same time providing detailed risk explanations and links to relevant clinical guidelines.
[0145] (V) Feedback layer: System update and optimization
[0146] The feedback layer updates and optimizes the model regularly to adapt to the new data distribution, specifically including the system update and optimization module. As new medical record data in the hospital accumulates continuously, the system update and optimization module regularly compares the risk prediction model with the actual patient outcome data for update and optimization. This module first incorporates the new medical record data into the existing dataset, and then repeats steps such as data preprocessing, feature extraction, model training, and evaluation to retrain and optimize the model, so as to adapt to the changes in the new data distribution and clinical situations, and continuously improve the prediction accuracy and generalization ability of the model.
[0147] The patient perioperative adverse event risk prediction system based on electronic medical record data of the present invention forms a computer program and is stored in a processor. The memory calls the computer program to execute the steps of the risk prediction method. The hardware environment configuration required for the operation of the embodiments of the present invention: A high-performance server is selected as the operation platform of the system, equipped with sufficient memory (such as 64GB or more), a multi-core CPU (such as an Intel Xeon Gold series processor), and a large-capacity storage space (such as a solid-state drive of 1TB or more) to ensure that the system can operate efficiently and stably. At the same time, a secure and reliable network environment is built, and technologies such as firewalls and encrypted transmission protocols are used to ensure the security of data transmission between the system and various hospital information systems. Software environment setup: Install an operating system (such as Windows Server 2019 or Linux Ubuntu 20.04) and a database management system (such as MySQL 8.0 or Oracle 19c) on the server for storing and managing medical record data. Install relevant development tools and software libraries, including Python 3.8 and above, TensorFlow 2.5 and above for building deep learning models, Scikit-learn for implementing traditional machine learning algorithms, NLTK for natural language processing, Pandas and NumPy for data processing and analysis, etc., as well as a web development framework (such as Django 3.2 or Spring Boot 2.5) for developing and deploying each module of the present invention and building a user-friendly interaction interface.
[0148] The model training and optimization of the present invention mainly includes:
[0149] (1) Feature engineering: Extract features related to risk prediction from the collected medical record data, including patients' basic information, medical history, diagnosis information, surgical information, test results, etc., as described in detail above. Encode and normalize these features. For example, for categorical variables (such as gender, marital status, professional title, etc.), use one-hot encoding for conversion; for numerical variables (such as age, height, weight, laboratory test indicators, etc.), perform standardization so that the mean is 0 and the variance is 1 to meet the input requirements of the machine learning model and improve the training efficiency and performance of the model.
[0150] (2) Model selection and training: Adopt a hybrid model architecture that combines a deep learning-based neural network model (such as a multi-layer perceptron MLP) and an ensemble learning algorithm (such as random forest RF, gradient boosting decision tree GBDT). First, construct an MLP model with multiple hidden layers, set appropriate activation functions (such as the ReLU function) and initialization methods (such as He initialization), determine that the dimension of the input layer matches the number of features, and the dimension of the output layer is determined according to the number of risk prediction categories (for example, whether a surgical complication occurs is a binary classification problem, and the dimension of the output layer is 1; for the prediction of multiple complication types, the dimension of the output layer is the number of complication types). At the same time, set the relevant parameters of the random forest and gradient boosting decision tree, such as the number of decision trees, the depth of the tree, the splitting criterion, etc. Divide the processed medical record data into a training set, a validation set, and a test set in a ratio of 70:15:15. Use the training set to train the constructed hybrid model, adjust the weight parameters and biases of the neural network model through the backpropagation algorithm, and at the same time use the random forest and gradient boosting decision tree in the ensemble learning algorithm to supervise and optimize the training process, and continuously adjust the hyperparameters of the model (such as the number of layers, the number of nodes, the learning rate, the number of iterations, etc. of the neural network, and the number of decision trees, the depth of the tree, the splitting criterion, etc. in the ensemble learning algorithm) to improve the performance of the model. During the training process, use the early stopping method to prevent overfitting, that is, stop training when the performance of the model on the validation set no longer improves.
[0151] (3) Model evaluation and optimization: Use the validation set to validate and tune the model during the training process. By comparing performance metrics such as accuracy, recall, F1-score, and the area under the receiver operating characteristic curve (ROC curve) of different model versions on the validation set, select the model version with the best performance for further optimization and adjustment. At the same time, adopt cross-validation techniques (such as 5-fold cross-validation or 10-fold cross-validation) to evaluate the stability and reliability of the model, ensuring that the model can exhibit good performance on different data subsets. After the model training is completed, use the test set to evaluate the performance of the final model to verify the accuracy and reliability of the model in practical applications. If the performance of the model on the test set does not meet the expected requirements, further optimize and improve the model, such as increasing the amount of training data, adjusting the feature selection strategy, trying other machine learning algorithms or model architectures, etc., until the performance of the model on the test set meets the requirements of practical applications.
[0152] The system integration and testing of the present invention mainly includes integrating each module (data collection and integration module, feature extraction and annotation module, machine learning model training and optimization module, risk prediction and assessment module, early warning prompt module, system update and optimization module) to ensure normal data transmission and interaction between modules. Conduct comprehensive functional testing on the integrated system, including medical record data collection and storage function testing, feature extraction and annotation function testing, risk prediction function testing, early warning reminder function testing, and user interaction function testing, etc., to ensure that all functions of the system can operate normally without obvious loopholes and errors.
[0153] Conduct performance testing on the system, simulate the scenario of multi-user concurrent access to the system, test performance metrics such as the response time, throughput, and resource utilization rate of the system, and optimize the performance of the system according to the test results, such as optimizing database query statements, adjusting server parameters, adopting caching technology, etc., to improve the performance and stability of the system and ensure that the system can meet the actual business needs of the hospital.
[0154] See the appendix Figure 3 For the schematic illustration, another embodiment of the present invention provides a method for predicting the risk of perioperative adverse events of patients based on electronic medical record data. Use the risk prediction system for perioperative adverse events of patients based on electronic medical record data in the above embodiment to conduct risk prediction, including the steps:
[0155] S1. Obtain the electronic medical record data of the patient to be predicted, extract the feature data related to risk prediction from the original data, and perform annotation and coding. The feature data related to risk prediction specifically includes the patient's admission diagnosis information, disease stage, past medical history, general physical signs, personal information, examination results, imaging results, surgical information, and hospitalization information;
[0156] S2. Based on the risk prediction model for perioperative adverse events of patients and the characteristic data of the patients to be predicted, output the risk prediction results of perioperative adverse events of patients, that is, calculate and output the probability values of surgical complications, unplanned reoperation, prolonged hospital stay, risk of returning to the ICU, and blood transfusion;
[0157] The construction process of the risk prediction model for perioperative adverse events of patients includes:
[0158] S21. Collect the electronic medical record information of inpatients in each information system of the hospital at the data layer to construct a data set, and the data set contains the medical characteristic data of patients and the medical record information of patients;
[0159] S22. Use the medical record information of patients in the data set as training labels, process the medical characteristic data of patients in the data set according to the following steps and perform model training to obtain the risk prediction model for perioperative adverse events of patients. The model training process is as follows:
[0160] Data preprocessing: Clean, denoise, and standardize the collected medical record data. For numerical data, use the Z-score standardization method to make its mean 0 and variance 1. For categorical variables, use one-hot encoding for conversion.
[0161] Formula: For numerical data x, the standardized data is:
[0162] where μ is the mean and σ is the standard deviation.
[0163] Feature importance analysis: Before model training, perform importance analysis on features, including feature importance scoring based on random forest and feature importance analysis based on gradient boosting decision tree, and preliminarily determine important features.
[0164] Weight adjustment in model training: During the model training process, the neural network automatically adjusts the weights of each feature through the backpropagation algorithm. The adjustment of the weights is based on the gradient descent of the loss function, and continuously optimizes the weights to make the prediction results more accurate.
[0165] Model training: Divide the processed medical record data into a training set, a validation set, and a test set according to the ratio of 70:15:15. Use the training set to train the constructed hybrid model, adjust the weight parameters and biases of the neural network model through the backpropagation algorithm, and at the same time use random forest and gradient boosting decision tree in the ensemble learning algorithm to supervise and optimize the training process.
[0166] Model evaluation and optimization: Use the validation set to validate and tune the model during the training process. By comparing the performance metrics such as accuracy, recall, F1-score, receiver operating characteristic curve, and area under the curve of different model versions on the validation set, select the model version with the best performance for further optimization and adjustment. Adopt cross-validation techniques to evaluate the stability and reliability of the model;
[0167] S3. Pop up a warning prompt message in the electronic medical record system for patients with a predicted probability of perioperative adverse event risk above 0;
[0168] S4. Update and optimize the patient perioperative adverse event risk prediction model. The feedback layer incorporates new discharged medical record data into the existing dataset, and then repeats the steps of data preprocessing, feature extraction, model training, and evaluation to retrain and optimize the risk prediction model.
[0169] As can be seen from the above embodiments, the patient perioperative adverse event risk prediction system based on electronic medical record data of the present invention can effectively achieve its invention purpose in practical applications, has high practical value and promotion prospects, and can bring significant social and economic benefits to the medical industry.
[0170] The present invention combines deep learning and ensemble learning algorithms to construct an efficient and accurate risk patient prediction system, which can effectively identify the potential risks of surgical patients, provide detailed warning information, help doctors formulate countermeasures in advance, reduce surgical risks, improve doctor-patient communication, and enhance medical efficiency. The system has the ability of self-learning and updating, can continuously optimize the model with the accumulation of new medical record data, adapt to the changes and developments of clinical practice, and always maintain high prediction accuracy and practicality.
[0171] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A patient perioperative adverse event risk prediction system based on electronic medical record data, characterized in that: include: The data layer is used to connect with various information systems of the hospital to collect, clean and store the electronic medical record data of inpatients. It specifically includes an electrically connected data collection module, a data cleaning module and a data storage module. The feature layer is used to extract features related to risk prediction from the original medical record data stored in the data layer, and to mark and encode them, and specifically includes an electrically connected feature extraction module, a feature marking module, and a feature storage module. The model layer is used to construct, train and optimize the patient perioperative adverse event risk prediction model, specifically including an electrically connected model selection module, a model training module, a model evaluation module and a model storage module. The model layer extracts the feature information in the feature storage module to perform patient perioperative adverse event risk prediction model training and risk prediction. The application layer is used to predict the risk of perioperative adverse events for new patients according to the patient perioperative adverse event risk prediction model, and issue a warning prompt for the prediction result with a prediction probability of 0 or above. It specifically includes an electrically connected risk prediction module, a risk assessment module, and a warning prompt module. When the medical record data of the new patient is entered into various information systems of the hospital, the data layer first collects and preprocesses the medical record data of the new patient, and then the feature layer extracts and annotates various features in the medical record data of the new patient. The risk prediction module predicts the risk of the new patient's medical record data through the trained patient perioperative adverse event risk prediction model. The risk assessment module interprets and sorts the perioperative adverse event risk prediction results. The warning prompt module issues a warning prompt for the perioperative adverse event risk prediction results. The feedback layer is used to regularly compare the predicted evaluation issued by the patient's perioperative adverse event risk prediction model with the patient's actual situation data to perform model training and optimization.
2. The patient perioperative adverse event risk prediction system based on electronic medical record data according to claim 1, characterized in that: The data acquisition module is connected to the hospital's electronic medical record system, surgical anesthesia management system, laboratory information system, image archiving and communication system to collect medical record data of inpatients. The data cleaning module cleans, denoises and standardizes the collected various electronic medical record data. The data storage module uses a relational database to store structured data and uses a NoSQL database to store unstructured data.
3. The patient perioperative adverse event risk prediction system based on electronic medical record data according to claim 2, characterized in that: The feature extraction module is used to extract the admission diagnosis, disease stage, past medical history, patient physical signs, test results, imaging results and surgical information from the patient's electronic medical record data; the feature annotation module is used to perform one-hot encoding on categorical variables and standardize numerical data; the feature storage module is used to store the extracted features in a feature library for use by the model layer for model training and the application layer for predicting the risk of perioperative adverse events in patients.
4. The patient perioperative adverse event risk prediction system based on electronic medical record data according to claim 3 is characterized in that: The model selection module adopts a hybrid model architecture that combines a neural network model based on deep learning and ensemble learning. The model training module divides the data into training set, validation set and test set, uses the back propagation algorithm to train the deep learning model, and uses random forest and gradient boosting decision tree for ensemble learning. The model evaluation module uses a validation set to evaluate model performance and uses cross-validation to evaluate model stability. The model storage module stores the trained model in the model library for the risk prediction module to call.
5. The patient perioperative adverse event risk prediction system based on electronic medical record data according to claim 4 is characterized in that: The process of training the patient perioperative adverse event risk prediction model at the model layer includes: Data preprocessing: Clean, denoise and standardize the collected electronic medical record data. Feature importance analysis: Feature importance scoring based on random forests, feature importance analysis based on gradient boosting decision trees, Weight adjustment during model training: During the model training process, the neural network model automatically adjusts the weight of each feature through the back propagation algorithm. Model training: The pre-processed electronic medical record data is divided into training set, validation set and test set in a ratio of 70:15:
15. The constructed hybrid model is trained using the training set. The weight parameters and bias of the neural network model are adjusted through the back propagation algorithm. At the same time, the random forest and gradient boosting decision tree in the ensemble learning algorithm are used to supervise and optimize the training process. Model evaluation and optimization: Use the validation set to verify and tune the model during the training process. By comparing the performance indicators of accuracy, recall rate, F1 value, receiver operating characteristic curve, and area under the validation set of different model versions, select the model version with the best performance for further optimization and adjustment. Use cross-validation technology to evaluate the stability and reliability of the model.
6. The patient perioperative adverse event risk prediction system based on electronic medical record data according to claim 5, characterized in that: After receiving the medical records of a new patient, the risk prediction module calls the patient perioperative adverse event risk prediction model to predict the risk of perioperative adverse events for the new patient, and outputs the probability values of surgical complications, unplanned reoperation, prolonged hospital stay, risk of returning to the ICU, and blood transfusion. The risk assessment module is used to interpret the risk prediction results of perioperative adverse events and rank the probability of occurrence, and provide possible complication types and their probability values. The early warning prompt module pops up an early warning window in the electronic medical record system according to the prediction results of the risk assessment module. The early warning content includes surgical complications, unplanned reoperation, prolonged hospitalization, risk of returning to the ICU, and blood transfusion requirements, and provides detailed risk explanations, treatment suggestions and clinical guideline links.
7. The patient perioperative adverse event risk prediction system based on electronic medical record data according to claim 6, characterized in that: The risk prediction module calculates the risk prediction probability of surgical complications using the following formula: Among them, P(complication) is the probability of complication occurrence, W is the weight vector, x is the input feature vector, and b is the bias term. The risk prediction probability calculation formula for unplanned reoperation is: Among them, P (reoperation) is the probability of reoperation, W' is the reoperation weight vector, x is the input feature vector, and b' is the reoperation bias term, which are obtained through independent training.
8. A method for predicting the risk of adverse events during perioperative period of patients based on electronic medical record data, characterized in that: The patient perioperative adverse event risk prediction system based on electronic medical record data according to any one of claims 1 to 7 is used to perform risk prediction, comprising the following steps: S1. Obtain the electronic medical record data of the patient to be predicted, extract the characteristic data related to risk prediction, and mark and encode them. The characteristic data related to risk prediction specifically includes the patient's admission diagnosis information, disease stage, past medical history, general signs, personal information, examination results, imaging results, surgery information and hospitalization information; S2. Based on the patient perioperative adverse event risk prediction model and the characteristic data of the patient to be predicted, output the patient perioperative adverse event risk prediction result, that is, calculate and output the probability values of surgical complications, unplanned reoperation, prolonged hospitalization, risk of returning to the ICU, and blood transfusion; The process of constructing the patient perioperative adverse event risk prediction model includes: S21, the data layer collects the electronic medical record information of inpatients from various information systems of the hospital to construct a data set, wherein the data set includes the medical characteristic data of the patients and the medical record information of the patients; S22. The medical records of the patients in the data set are used as training labels. The medical characteristic data of the patients in the data set are processed according to the following steps and the model training is performed to obtain a risk prediction model for adverse events during perioperative period of the patients. The model training process is as follows: Data preprocessing: The collected medical record data were cleaned, denoised and standardized. The Z-score standardization method was used for numerical data to make its mean 0 and variance 1. The one-hot encoding was used for categorical variables. Formula: For numerical data x, the standardized data is: Where μ is the mean and σ is the standard deviation. Feature importance analysis: Before model training, perform feature importance analysis, including feature importance scoring based on random forests and feature importance analysis based on gradient boosting decision trees, to preliminarily determine important features. Weight adjustment in model training: During the model training process, the neural network automatically adjusts the weight of each feature through the back propagation algorithm. The weight adjustment is based on the gradient descent of the loss function, and the weight is continuously optimized to make the prediction result more accurate. Model training: The processed medical record data is divided into training set, validation set and test set in a ratio of 70:15:
15. The constructed hybrid model is trained using the training set. The weight parameters and bias of the neural network model are adjusted through the back propagation algorithm. At the same time, the random forest and gradient boosting decision tree in the ensemble learning algorithm are used to supervise and optimize the training process. Model evaluation and optimization: Use the validation set to verify and tune the model during training. By comparing the performance indicators of accuracy, recall, F1 value, receiver operating characteristic curve, and area under the validation set of different model versions, select the model version with the best performance for further optimization and adjustment. Use cross-validation technology to evaluate the stability and reliability of the model. S3. For patients whose predicted probability of perioperative adverse event risk is above 0, a warning message will pop up in the electronic medical record system.
9. The method for predicting the risk of perioperative adverse events of patients based on electronic medical record data according to claim 8, characterized in that: The model calculation formula in step S22 is: For the multi-layer perceptron model in the neural network model, the calculation formula of its output layer is: y=σ(W·x+b) Among them, y is the output result, W is the weight matrix, x is the input feature vector, b is the bias term, σ is the activation function, For the random forest model, the output is the voting results of multiple decision trees: y=mode(T1(x),T2(x),...,Tn(x)) Among them, T i (x) is the prediction result of the i-th decision tree, and mode is the mode function. For the decision process of gradient boosting decision tree, unlike random forest, gradient boosting decision tree generates prediction results through iterative addition model, each round fits the residual of the previous round of prediction, and the final output is the weighted sum of all decision trees: Among them, M is the total number of decision trees, h m (x) is the prediction result of the mth decision tree, x m is the weight coefficient.
10. The method for predicting the risk of perioperative adverse events of patients based on electronic medical record data according to claim 9, characterized in that: The method also includes step S4 of updating and optimizing the patient perioperative adverse event risk prediction model. The feedback layer incorporates the new discharge medical record data into the existing data set, and then repeats the steps of data preprocessing, feature extraction, model training and evaluation to retrain and optimize the patient perioperative adverse event risk prediction model.
Citation Information
Cited By
State prediction method and device for traumatic patient, electronic equipment and storage medium
CN121171613A
Surgical site infection early warning method and device, terminal and storage medium
CN121350850A