A medical data management method based on general surgery big data
Through the improved Shapley formula and random forest model, combined with multi-time data to calculate the improved contribution value of pathological indicators, the problem of inaccurate key indicators in the existing technology is solved, and more accurate prediction and more optimized treatment plans are achieved.
Patent Information
- Application Number
- CN202510195663.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-21
AI Technical Summary
After using a random forest model to predict disease risk, the prior art has errors when interpreting the prediction results through Shapley values, which leads to inaccurate key indicators, which in turn affects the optimization effect of the treatment plan.
The improved Shapley formula combined with the random forest model is used to calculate the improved contribution value of pathological indicators through multi-time data, and key indicators with high contribution value are screened out, preliminary plans are generated, and the optimization scheme is iterated through optimization algorithms.
The accuracy and interpretability of the predicted results are significantly improved, and the core factors affecting the predicted results are accurately positioned. The generated treatment plan is clearly targeted and feasible, ensuring that the final plan achieves the best results in practical applications.
Smart Images

Figure CN119673432B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data management, and particularly to a medical data management method based on general surgery big data. Background Art
[0002] With the rapid development of information technology, the medical industry has gradually realized the transformation from traditional paper records to electronic and digital forms. The wide application of information systems such as hospital information system (HIS), electronic medical record (EMR), laboratory information system (LIS), and picture archiving and communication system (PACS) enables the efficient storage and management of a large amount of medical data. However, these systems usually operate independently, resulting in serious data island phenomena and making it difficult to achieve cross-system data integration and analysis. Big data technology provides a new solution to solve the above problems. By integrating and storing the scattered data in a unified data warehouse, valuable information can be mined using advanced data analysis tools and technologies (such as machine learning, natural language processing, and deep learning). This not only helps with clinical decision support and the formulation of personalized treatment plans but also optimizes the allocation of medical resources. General surgery, as an important department in the hospital, involves a large number of surgeries and medical treatments, accumulating rich case materials and surgical data. These data include the patient's medical history, diagnosis results, surgical procedures, and postoperative recovery conditions, etc. Effectively managing and utilizing these data is crucial for improving the quality and efficiency of medical services. Especially in the optimization of treatment plans, how to use the existing data to improve and optimize treatment plans becomes the key.
[0003] In the prior art, the risk probability of affecting a disease is predicted through a random forest model and key factors are obtained, and intervention measures are generated based on the key factors. However, after the random forest model obtains the prediction result, the Shapley value is also used to interpret the prediction result, that is, the contribution to the prediction result is calculated for the model input data. However, there are certain errors in the prediction process of the model, so that the data with a high risk probability may have a relatively low contribution value, resulting in inaccurate determination of key indicators. The plan generated based on inaccurate key indicators cannot minimize the disease risk to the greatest extent, thus unable to ensure that the intervention effect of the plan on the disease risk reaches the ideal state. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a medical data management method based on general surgery big data, including:
[0005] Step 1, obtaining patient data and performing preprocessing; the patient data includes several pathological indicators;
[0006] Step 2: Select several training samples from the preprocessed patient data and label each training sample; the labels include the disease type and the risk probability of the pathological indicators for the disease type;
[0007] Step 3: Construct a random forest model, and use all the training samples with calibrated labels as the training set to train the random forest model to obtain a trained random forest model;
[0008] Obtain the real-time patient data and input it into the trained random forest model to get the prediction results; the prediction results include the risk probability of each pathological indicator for the current disease type; calculate the improved Shapley formula based on the Shapley formula, calculate the improvement contribution value of each pathological indicator in the real-time patient data to the prediction results through the improved Shapley formula, calculate the comprehensive score according to the improvement contribution value of each pathological indicator in the real-time patient data to the prediction results and the risk probability of each pathological indicator in the real-time patient data for the current disease type, and arrange them from high to low, and take the pathological indicators corresponding to the top 2 / 3 comprehensive scores as the key indicators;
[0009] Step 4: Generate a preliminary plan based on the prediction results: disease type, real-time patient data, key indicators, the risk probability of the key indicators for the current disease type, and intervention measures; and execute the preliminary plan to obtain the implemented key indicators;
[0010] Step 5: Calculate the intervention eigenvalue of the key indicators based on the pathological indicators before implementing the preliminary plan and each key indicator after implementing the preliminary plan; define an optimization objective, construct an optimization function based on the intervention eigenvalue, and use an optimization algorithm to solve the optimal solution of the optimization function, and the optimal solution is the optimization plan.
[0011] Furthermore, calculating the improved Shapley formula based on the Shapley formula specifically includes the following steps:
[0012] Step 31: Extract each pathological indicator from the real-time patient data and generate a feature subset that does not include the current pathological indicator based on each pathological indicator; obtain the prediction values of each pathological indicator and each feature subset at multiple times through the trained random forest model;
[0013] Step 32: Calculate the initial contribution values of the current pathological indicator and the feature subset that does not include the current pathological indicator through the Shapley formula based on the prediction values of each pathological indicator and the feature subset that does not include the current pathological indicator at multiple times;
[0014] Step 33: Determine a combined weight for the current pathological indicator and the feature subset that does not include the current pathological indicator;
[0015] Step 34, calculate the mutual information between the current pathological index and the feature subset excluding the current pathological index at the current moment;
[0016] Step 35, obtain the improved Shapley formula based on the initial contribution value, combined weight, and mutual information to calculate the improved contribution value.
[0017] Further, determining the combined weight includes the following steps:
[0018] Step 331, calculate the variance based on the patient data in Step 1 and the current pathological index at the current moment;
[0019] Step 332, determine the initial weights of the current pathological index and each pathological index in the feature subset excluding the current pathological index;
[0020] Step 333, calculate the correlation strength values between the current pathological index and each pathological index in the feature subset excluding the current pathological index at the current moment;
[0021] Step 334, calculate the combined weight based on the initial weight, correlation strength value, and variance.
[0022] Further, the calculation formula for the comprehensive score is:
[0023] ;
[0024] In the formula, represents the comprehensive score of the i-th pathological index, represents the initial contribution value of the i-th pathological index, represents the risk probability of the i-th pathological index for the current disease type, represents the mutual information between the i-th pathological index and the feature subset S excluding the current pathological index, represents the maximum mutual information between the i-th pathological index and the feature subset S excluding the current pathological index, represents the variance of the i-th pathological index at the current moment t, represents the maximum variance of the i-th pathological index at the current moment t, represents the combined weight of the i-th pathological index and the feature subset excluding the current pathological index at the current moment t, and F represents the set of feature subsets and / or pathological indices.
[0025] Further, the calculation formula for the combined weight is:
[0026] ;
[0027] In the formula, represents the correlation intensity value between the i-th pathological index at the current moment t and the j-th pathological index in the feature subset excluding the current pathological index. represents the initial weight of the i-th pathological index and the j-th pathological index in the feature subset excluding the current pathological index. represents the adjustment factor of the correlation intensity value. represents the adjustment factor of the variance.
[0028] Furthermore, based on the pathological indexes before implementing the preliminary plan and the key indexes after implementing the preliminary plan, the intervention eigenvalue of the key index is calculated, which specifically includes the following steps:
[0029] Step 51: Obtain the eigenvalue of each pathological index before implementing the preliminary plan.
[0030] Step 52: Obtain the eigenvalue of the key index after implementing the preliminary plan. Based on the eigenvalue of the key index after implementing the preliminary plan and the eigenvalue of the key index before implementing the preliminary plan, calculate the feature change amount of the key index.
[0031] Step 53: Based on the feature change amount and the eigenvalue of each pathological index before implementing the preliminary plan, obtain the intervention eigenvalue of the key index.
[0032] Furthermore, the optimization function is:
[0033] ;
[0034] In the formula, E represents the optimization function, min represents minimization processing, represents the weight of the k-th key index, represents the weight of the k-th key index and the L-th key index, represents the eigenvalue of the k-th key index, represents the feature change amount of the k-th key index, represents the eigenvalue of the L-th key index, represents the feature change amount of the L-th key index, n represents the total number of key indexes, and b is a constant greater than 0.
[0035] Furthermore, the optimization goal is: to minimize the health risk of patients.
[0036] Furthermore, the improved Shapley formula is:
[0037] ;
[0038] In the formula, represents the improved contribution value of the i-th pathological index at the current moment t, F represents the set of feature subsets and / or pathological indexes, The combined weight of the i-th pathological index representing the current moment t and the feature subset excluding the current pathological index Represents the initial contribution value of the i-th pathological index Represents the mutual information between the i-th pathological index and the feature subset S excluding the current pathological index Represents the adjustment coefficient
[0039] The embodiments of the present invention have the following technical effects:
[0040] Through the innovative method of combining the random forest model and the improved Shapley formula, the present invention significantly improves the accuracy and interpretability of the prediction results. First, accurate prediction results are generated using the random forest model. Second, based on the improved Shapley formula, the specific contribution value of each pathological index to the prediction result is calculated. This process not only quantifies the importance of each index but also overcomes the problem of low calculation efficiency of the traditional Shapley value, greatly improving the analysis speed and practicality. By screening out the key indexes with high improved contribution values, the present invention can accurately locate the core factors affecting the prediction results and provide a scientific basis for subsequent decision-making. Further, the preliminary scheme generated based on the key indexes has clear pertinence and feasibility. Through the dynamic monitoring and feedback of the key indexes after implementation, an optimization function is constructed to iteratively optimize the preliminary scheme to ensure that the final scheme achieves the best effect in practical applications. In summary, the present invention shows significant advantages in improving prediction accuracy, enhancing the ability to optimize the scheme, and reducing implementation risks, and has important theoretical value and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0042] Figure 1 It is a flowchart of a medical data management method based on general surgery big data provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0044] Figure 1 It is a flowchart of a medical data management method based on general surgery big data provided by an embodiment of the present invention. Refer to Figure 1 , specifically including:
[0045] Step 1, obtain patient data and perform preprocessing; the patient data includes a number of pathological indicators.
[0046] Obtain patient data from an electronic health information system, laboratory testing equipment, wearable medical devices, medical imaging data, etc. The preprocessing includes:
[0047] Data cleaning: delete or fill in missing values: for missing data, it can be filled in using mean filling, interpolation method or model-based prediction method.
[0048] Remove outliers: use statistical methods to identify and eliminate data points that significantly deviate from the normal range.
[0049] Data standardization / normalization: convert data with different dimensions to the same scale for subsequent modeling and analysis.
[0050] Feature selection and dimensionality reduction: use methods such as principal component analysis (PCA) or Lasso regression to screen out the features that have the most influence on the prediction target.
[0051] The pathological indicators at least include: blood-related indicators: blood routine: white blood cell count (WBC), red blood cell count (RBC), hemoglobin (Hb), platelet count (PLT).
[0052] Biochemical indicators: glucose (Glucose), total cholesterol (TC), triglyceride (TG), high-density lipoprotein (HDL), low-density lipoprotein (LDL).
[0053] Liver function indicators: alanine aminotransferase (ALT), aspartate aminotransferase (AST), alkaline phosphatase (ALP), total bilirubin (TBIL).
[0054] Renal function indicators: creatinine (Creatinine), blood urea nitrogen (BUN), uric acid (Uric Acid).
[0055] Cardiovascular indicators: heart rate (HR), blood pressure (SBP, DBP), blood lipid level (TC, TG), C-reactive protein (CRP).
[0056] Immunological indicators: C-reactive protein (CRP), immunoglobulins (IgA, IgG, IgM), complements (C3, C4).
[0057] Tumor markers: carcinoembryonic antigen (CEA), alpha-fetoprotein (AFP), carbohydrate antigen (CA125, CA19-9).
[0058] Other special indicators: Inflammatory factors: interleukin-6 (IL-6), tumor necrosis factor-α (TNF-α).
[0059] Genetic information: gene mutation sites, single nucleotide polymorphisms (SNPs), etc.
[0060] Step 2: Screen a number of training samples from the preprocessed patient data and label each training sample; the label includes the disease type and the risk probability of the pathological indicators for the disease type.
[0061] Divide the patient data into different groups according to the disease type. For example, for the diabetes prediction task, patients can be divided into two groups: "diabetes patients" and "non-diabetes patients". Example: Extract the patient data diagnosed with diabetes in the past year from the hospital database as positive samples, and randomly select the same number of healthy population data as negative samples. Clearly label the disease category to which each sample belongs, and label the risk contribution probability of each pathological indicator to a specific disease type through clinical expert evaluation or statistical analysis methods.
[0062] Step 3: Construct a random forest model, use all the training samples with calibrated labels as the training set to train the random forest model, and obtain the trained random forest model.
[0063] Obtain the real-time patient data and input it into the trained random forest model to get the prediction results; the prediction results include the risk probability of each pathological indicator for the current disease type. Calculate the improved Shapley formula based on the Shapley formula, and calculate the improved contribution value of each pathological indicator in the real-time patient data to the prediction results through the improved Shapley formula:
[0064] Step 31: Extract each pathological indicator from the real-time patient data and generate a feature subset that does not include the current pathological indicator based on each pathological indicator; obtain the prediction values of each pathological indicator and each feature subset at multiple times through the trained random forest model.
[0065] Step 32: Based on the prediction values of each pathological indicator and the feature subset that does not include the current pathological indicator at multiple times, calculate the initial contribution values of the current pathological indicator and the feature subset that does not include the current pathological indicator through the Shapley formula.
[0066] In this embodiment, predictions are made by obtaining real-time data of a patient at multiple moments. This not only enables capturing the changing trends of pathological indicators over time but also provides more comprehensive and dynamic data support for improving the calculation of the Shapley value. Compared with data at a single moment, multi-moment data can reflect the volatility and stability of indicators over different time periods, thereby generating more accurate initial contribution values. These initial contribution values, as the basis for improving the Shapley value calculation, can significantly enhance the accuracy and reliability of subsequent analyses. The initial contribution values generated based on multi-moment data can better reflect the long-term impact of each pathological indicator on the disease prediction result, avoiding biases caused by the contingency or noise of single-timepoint data. In addition, this dynamic analysis method helps identify the time-dependent characteristics of key indicators. For example, certain indicators may only have a significant contribution to disease risk at specific stages. Therefore, multi-moment data is used to optimize the calculation of the initial contribution value and improve the calculation efficiency of the improved Shapley value.
[0067] The Shapley formula is as follows:
[0068] ;
[0069] Among them, represents the initial contribution value of the i-th pathological indicator, represents the number of pathological indicators in the feature subset, F represents the set of feature subsets and / or pathological indicators, represents the number of pathological indicators included in the set of feature subsets and / or pathological indicators, represents the predicted value of all pathological indicators, represents the predicted value of the feature subset that does not include the current pathological indicator i.
[0070] Step 33: Determine a combined weight for the current pathological indicator and the feature subset that does not include the current pathological indicator:
[0071] Step 331: Calculate the variance based on the patient data in Step 1 and the current pathological indicator at the current moment.
[0072] The variance is used to quantify the degree of change of each pathological indicator. All data of each pathological indicator are collected from the patient data in Step 1 and the real-time data of the patient at the current moment, and the variance is calculated.
[0073] Step 332: Determine the initial weights of the current pathological indicator and each pathological indicator within the feature subset that does not include the current pathological indicator.
[0074] The initial weights are obtained by normalizing the variances according to the above steps. The variance contribution rate of each index is obtained through normalization. The larger the variance, the wider the range of changes of the index, and the more important it may be for disease prediction. Therefore, in this embodiment, the initial weights are assigned based on the variance contribution rate.
[0075] Step 333: Calculate the correlation strength values between the current pathological index at the current moment and each pathological index in the feature subset that does not include the current pathological index.
[0076] Calculate the correlation strength values through the Pearson correlation coefficient.
[0077] Step 334: Calculate the combined weights based on the initial weights, the correlation strength values, and the variances. ;
[0078] ;
[0079] In the formula, represents the correlation strength value between the i-th pathological index at the current moment t and the j-th pathological index in the feature subset that does not include the current pathological index. represents the initial weight of the i-th pathological index and the j-th pathological index in the feature subset that does not include the current pathological index. represents the adjustment factor of the correlation strength value. represents the adjustment factor of the variance.
[0080] Among them, , this part represents the exponential decay of the correlation strength value between the current pathological index i and other pathological indices j in the feature subset. In machine learning, overfitting is a common problem, that is, the model fits the training data too well, so that it performs poorly when facing new data. Highly correlated indicators may cause the model to rely too much on certain specific data patterns or noises, rather than the real underlying relationships. By reducing the weights of highly correlated indicators, the model can be prompted to consider more those relatively independent indicators that can provide additional information, thereby enhancing the generalization ability of the model and making its performance better on unseen data. Therefore, in this embodiment, for the pathological indices with relatively strong correlation strength, the weights are reduced through exponential decay.
[0081] , this part represents the influence of the variance of the current pathological index i at the current moment on the weight. An index with a large variance indicates that its values are more dispersed in the dataset, rather than concentrated around a specific value. This dispersion may reflect the contribution of the index to the diversity and complexity of the target variable. Therefore, from the perspective of information content, a larger variance usually means that the index contains more information about the prediction target, so a higher weight should be given.
[0082] Step 34, calculate the mutual information between the current pathological index and the feature subset excluding the current pathological index at the current moment.
[0083] Step 35, obtain the improved Shapley formula based on the initial contribution value, combined weight, and mutual information to calculate the improved contribution value. The specific formula is as follows:
[0084] ;
[0085] In the formula, represents the improved contribution value of the i-th pathological index at the current moment t, F represents the set of feature subsets and / or pathological indices, represents the combined weight of the i-th pathological index at the current moment t and the feature subset excluding the current pathological index, represents the initial contribution value of the i-th pathological index, represents the mutual information between the i-th pathological index and the feature subset S excluding the current pathological index, represents the adjustment coefficient.
[0086] The combined weight reflects the correlation and variance between the current pathological index and other indices in the feature subset, ensuring the dynamic adjustment of the importance of each index at different time points; the introduction of time series data enables the model to capture the trend of pathological indices changing over time, enhancing the timeliness and accuracy of the prediction results; the mutual information quantifies the dependence relationship between the current pathological index and the feature subset, further optimizing the model's ability to handle complex interaction effects. This comprehensive consideration improves the prediction accuracy of the model.
[0087] Calculate the comprehensive score based on the improved contribution value of each pathological index in the patient's real-time data to the prediction result and the risk probability of each pathological index in the patient's real-time data for the current disease type, and arrange them from high to low. The pathological indices corresponding to the top 2 / 3 of the comprehensive scores are used as key indices.
[0088] The calculation formula for the comprehensive score is as follows:
[0089] ;
[0090] In the formula, represents the comprehensive score of the i-th pathological index, represents the initial contribution value of the i-th pathological index, represents the risk probability of the i-th pathological index for the current disease type, represents the mutual information between the i-th pathological index and the feature subset S excluding the current pathological index, represents the maximum mutual information between the i-th pathological index and the feature subset S excluding the current pathological index, represents the variance of the i-th pathological index at the current time t, represents the maximum variance of the i-th pathological index at the current time t, represents the combined weight of the i-th pathological index and the feature subset excluding the current pathological index at the current time t, where F represents the set of feature subsets and / or pathological indices.
[0091] where, , combining the initial contribution value and the risk probability ensures the balance of the basic importance and disease relevance. , adjusted by the exponential function of mutual information, reflects the dependence relationship between the current pathological index and other indices, and ensures the comparability between different indices through normalization processing. , introducing variance enhances the weight of high-variance indices through standardization processing, ensuring the manifestation of the importance of dynamic changes.
[0092] Step 4, based on the prediction results, generate a preliminary plan: disease type, patient real-time data, key indices, the risk probability of key indices for the current disease type, and intervention measures; and execute the preliminary plan to obtain the executed key indices.
[0093] Exemplarily, the corresponding intervention measures generated for diabetes are:
[0094] Intervention measures for fasting blood glucose
[0095] Diet control: It is recommended that patients reduce the intake of high-sugar and high-carbohydrate foods and increase the proportion of dietary fiber and high-quality protein.
[0096] Exercise plan: Develop a personalized exercise plan, such as performing 30 minutes of moderate-intensity aerobic exercise every day.
[0097] Drug treatment: Use hypoglycemic drugs according to the doctor's advice and regularly monitor blood glucose levels.
[0098] Health education: Popularize diabetes management knowledge to patients and improve their self-management ability.
[0099] Intervention measures for glycated hemoglobin
[0100] Long-term blood glucose control: Combine diet, exercise, and drug treatment to control fasting blood glucose within the target range.
[0101] Regular monitoring: Measure glycated hemoglobin levels every 3 months to evaluate the blood glucose control effect.
[0102] Psychological support: Provide psychological counseling or support group services to help patients cope with the stress brought by chronic diseases.
[0103] Intervention measures for insulin level
[0104] Insulin therapy: For patients with insufficient insulin secretion, exogenous insulin can be supplemented by injection.
[0105] Weight management: Reduce weight through a reasonable diet and exercise to improve insulin sensitivity.
[0106] Lifestyle adjustment: Avoid staying up late and overworking, and maintain a regular daily routine.
[0107] Intervention measures for C-reactive protein
[0108] Anti-inflammatory diet: Recommend foods rich in antioxidants to reduce the level of inflammation in the body.
[0109] Smoking cessation and alcohol restriction: Encourage patients to quit smoking and limit alcohol intake to reduce the additional burden on the body.
[0110] Drug assistance: Use anti-inflammatory drugs to reduce the level of C-reactive protein when necessary.
[0111] Apply the intervention measures in the preliminary plan to the patient and guide them to execute according to the plan. After the intervention measures have been implemented for a period of time, re-collect the patient's pathological index data as the key execution indicators. Exemplarily: Fasting blood glucose: decreased from 126 mg / dL to 110 mg / dL; Glycated hemoglobin: decreased from 7.5% to 6.8%; Insulin level: increased from 10 µIU / mL to 12 µIU / mL; C-reactive protein: decreased from 5 mg / L to 3 mg / L.
[0112] Step 5, based on the pathological indicators before implementing the preliminary plan and each key indicator after implementing the preliminary plan, calculate the intervention eigenvalue of the key indicator; define an optimization goal, construct an optimization function based on the intervention eigenvalue, and use an optimization algorithm to solve the optimal solution of the optimization function. The optimal solution is the optimization plan. The optimization goal is: to minimize the patient's health risk.
[0113] Among them, calculating the intervention eigenvalue of the key indicator is through each key indicator after implementing the preliminary plan and the pathological indicators corresponding to the key indicators before implementing the preliminary plan, rather than all pathological indicators.
[0114] Furthermore:
[0115] Step 51, obtain the eigenvalue of each pathological indicator before implementing the preliminary plan;
[0116] The pathological indicators in this step refer to the pathological indicators corresponding to the key indicators.
[0117] Step 52: Obtain the eigenvalue of the key indicator after implementing the preliminary solution. Based on the eigenvalue of the key indicator after implementing the preliminary solution and the eigenvalue of the key indicator before implementing the preliminary solution, calculate the characteristic change amount of the key indicator.
[0118] Step 53: Based on the characteristic change amount and the eigenvalue of each pathological indicator before implementing the preliminary solution, obtain the intervention eigenvalue of the key indicator.
[0119] The optimization function is:
[0120] ;
[0121] In the formula, E represents the optimization function, min represents minimization processing, represents the weight of the k-th key indicator, represents the weight between the k-th key indicator and the L-th key indicator, represents the eigenvalue of the k-th key indicator, represents the characteristic change amount of the k-th key indicator, represents the eigenvalue of the L-th key indicator, represents the characteristic change amount of the L-th key indicator, n represents the total number of key indicators, and b is a constant greater than 0.
[0122] , this part reflects the individual contribution of each key indicator, , this part reflects the interaction between all key indicators, and minimizes the objective function to seek the optimal solution to obtain the optimization plan. The implementation of the optimization plan can minimize the patient's disease risk to the greatest extent.
[0123] As described above, only the preferred specific implementation manners of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A medical data management method based on general surgery big data, characterized in that: include: Step 1: Obtain patient data and perform preprocessing; Patient data include several pathological indicators; Step 2: select a number of training samples from the preprocessed patient data and label each training sample; the label includes the disease type and the risk probability of the pathological index for the disease type; Step 3, construct a random forest model, use all training samples with calibration labels as training sets to train the random forest model, and obtain a trained random forest model; Obtaining real-time patient data and inputting it into a trained random forest model to obtain prediction results; the prediction results include the risk probability of each pathological indicator for the current disease type; The improved Shapley formula is calculated based on the Shapley formula, and the improved contribution value of each pathological indicator in the patient's real-time data to the prediction result is calculated by the improved Shapley formula: Step 31, extracting each pathological index from the patient's real-time data, and generating a feature subset that does not include the current pathological index based on each pathological index; obtaining the predicted values of each pathological index and each feature subset at multiple moments through the trained random forest model; Step 32, based on the predicted values of each pathological index and the feature subset excluding the current pathological index at multiple moments, the initial contribution value of the current pathological index and the feature subset excluding the current pathological index is calculated by Shapley's formula; Step 33, determining a combination weight for the current pathological index and the feature subset that does not include the current pathological index; Step 34, calculating the mutual information between the current pathological index at the current moment and the feature subset that does not include the current pathological index; Step 35, based on the initial contribution value, the combination weight and the mutual information, an improved Shapley formula is obtained to calculate the improved contribution value: ; In the formula, represents the improvement contribution value of the i-th pathological index at the current time t, F represents the set of feature subsets and / or pathological indicators, represents the combined weight of the i-th pathological index at the current time t and the feature subset that does not contain the current pathological index, represents the initial contribution value of the i-th pathological index, represents the mutual information between the i-th pathological index at the current time t and the j-th pathological index in the feature subset S that does not contain the current pathological index, represents the adjustment factor; According to the improvement contribution value of each pathological indicator in the patient's real-time data to the prediction result and the risk probability of each pathological indicator in the patient's real-time data for the current disease type, a comprehensive score is calculated and arranged from high to low, and the pathological indicators corresponding to the top 2 / 3 of the comprehensive scores are used as key indicators; Step 4: Generate a preliminary plan based on the prediction results: disease type, patient real-time data, key indicators, risk probability of key indicators for the current disease type, and intervention measures; And implement the preliminary plan and obtain key execution indicators; Step 5, based on the pathological indicators before the implementation of the preliminary plan and each key indicator after the implementation of the preliminary plan, the intervention characteristic value of the key indicator is calculated; Define an optimization objective, build an optimization function based on the intervention eigenvalue, and use the optimization algorithm to solve the optimal solution of the optimization function. The optimal solution is the optimization plan.
2. A medical data management method based on general surgery big data according to claim 1, characterized in that: Determining the combination weights includes the following steps: Step 331, calculating the variance according to the patient data in step 1 and the current pathological index at the current moment; Step 332, determining the initial weight of the current pathological indicator and each pathological indicator in the feature subset that does not include the current pathological indicator; Step 333, calculating the correlation strength value between the current pathological index at the current moment and each pathological index in the feature subset that does not include the current pathological index; Step 334, calculating the combined weight according to the initial weight, the correlation strength value and the variance.
3. The medical data management method based on general surgery big data according to claim 2 is characterized in that: The formula for calculating the comprehensive score is: ; In the formula, represents the comprehensive score of the i-th pathological index, represents the initial contribution value of the i-th pathological index, represents the risk probability of the i-th pathological indicator for the current disease type, represents the mutual information between the i-th pathological index at the current time t and the j-th pathological index in the feature subset S that does not contain the current pathological index, represents the maximum mutual information between the i-th pathological index at the current time t and the j-th pathological index in the feature subset S that does not contain the current pathological index, represents the variance of the i-th pathological index at the current time t, represents the maximum variance of the i-th pathological index at the current time t, represents the combined weight of the i-th pathological index at the current time t and the feature subset that does not include the current pathological index, and F represents the set of feature subsets and / or pathological indexes.
4. The medical data management method based on general surgery big data according to claim 3 is characterized in that: The formula for calculating the combined weight is: ; In the formula, represents the correlation strength value between the i-th pathological index at the current time t and the j-th pathological index in the feature subset that does not contain the current pathological index, represents the initial weight of the i-th pathological index and the j-th pathological index in the feature subset that does not contain the current pathological index, represents the adjustment factor for the correlation strength value, Represents the adjustment factor for the variance.
5. The medical data management method based on general surgery big data according to claim 1 is characterized in that: Based on the pathological indicators before the implementation of the preliminary plan and the key indicators after the implementation of the preliminary plan, the intervention characteristic values of the key indicators are calculated, which specifically includes the following steps: Step 51, obtaining the characteristic value of each pathological indicator before executing the preliminary plan; Step 52, obtaining characteristic values of the key indicators after the preliminary plan is executed, and calculating characteristic changes of the key indicators based on the characteristic values of the key indicators after the preliminary plan is executed and the characteristic values of the key indicators before the preliminary plan is executed; Step 53, based on the characteristic change amount and the characteristic value of each pathological indicator before executing the preliminary plan, the intervention characteristic value of the key indicator is obtained.
6. The medical data management method based on general surgery big data according to claim 1 is characterized in that: The optimization function is: ; In the formula, E represents the optimization function, min represents the minimization process, represents the weight of the kth key indicator, Represents the weight of the kth key indicator and the Lth key indicator, represents the eigenvalue of the kth key indicator, Represents the characteristic change of the kth key indicator, represents the eigenvalue of the Lth key indicator, represents the characteristic change of the Lth key indicator, n represents the total number of key indicators, and b is a constant greater than 0.
7. The medical data management method based on general surgery big data according to claim 1 is characterized in that: The optimization goal is to minimize the health risks to patients.
Citation Information
Patent Citations
Electronic data analysis method and system for digestive system department
CN117747113A
Chronic disease screening and follow-up visit data collection and management method and system
CN119008010A