Operation duration prediction method and system based on department distribution adjustment, electronic equipment and readable storage medium
By applying PCA and SHAP techniques to screen features in surgery duration prediction and designing a modular prediction model, the problems of high complexity and poor adaptability of prediction models in existing methods are solved, and surgery duration prediction with higher accuracy and reliability is achieved.
Patent Information
- Application Number
- CN202510763738.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-12
AI Technical Summary
Existing surgical duration prediction methods have many shortcomings, including reliance on the subjectivity of empirical judgment, failure of machine learning models to be personalized for departmental differences, and insufficiently optimized processing of input data features, resulting in high complexity, poor adaptability, and unstable performance of the prediction models.
Principal component analysis (PCA) and SHAP techniques were used to screen out features that significantly affect operation duration. Combined with the distribution characteristics of operation duration in each department, a modular prediction model was designed, and the optimal prediction model was selected through a dynamic calling mechanism to predict operation duration.
It improves the accuracy, adaptability and reliability of surgery duration prediction, reduces model complexity, improves prediction performance and interpretability, and is suitable for the specific needs of different departments.
Smart Images

Figure CN120636845A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method, system, electronic equipment and readable storage medium for predicting operation duration based on department distribution adjustment, and belongs to the technical field of medical data processing. Background Art
[0002] With the rapid development of medical technology and the extension of human life expectancy, the demand for surgery continues to grow. Hundreds of millions of surgeries are performed worldwide each year, becoming a vital part of modern medical services. In my country, the utilization rate of medical resources has increased significantly over the past few decades, and the number of surgeries has also maintained a steady growth. As one of the core indicators of the surgical process, the duration of surgery has a significant impact on the resource allocation and operational efficiency of the hospital. Accurately predicting the duration of surgery can not only optimize the arrangement of operating rooms and improve the efficiency of the use of medical equipment and human resources, but also reduce the phenomenon of vacant or overloaded operating rooms, providing patients with more efficient and higher-quality medical services.
[0003] However, current surgical duration prediction methods still face many challenges. Traditional surgical duration estimation relies on the experience and judgment of surgeons and anesthesiologists. This method is subjective and limited, and is prone to prediction errors due to differences in personal experience. In addition, existing prediction methods based on machine learning also have shortcomings in practical applications. First, there are significant differences in the types of surgeries, patient groups, and the distribution patterns of surgical durations in different departments. For example, the surgical durations of some specialized departments are concentrated in a specific range, while the distribution of surgical durations in general surgery is wider and more random. Existing prediction models have failed to be personalized for departmental differences, and a unified prediction model often cannot take into account the needs of all departments, resulting in uneven performance of prediction models.
[0004] Secondly, most existing methods use mean squared error (MSE) or mean absolute error (MAE) as metrics for predictive model performance. These metrics have limitations when measuring predictive effectiveness across a range of surgical durations. Especially for long surgeries, large errors can mask prediction issues for shorter surgeries, preventing a fair and comprehensive assessment of predictive model performance. These limitations reduce the practical guiding value of predictive models.
[0005] Furthermore, existing methods also leave room for improvement in how they handle input data features. Medical data is highly multimodal, including numerical features (such as patient age and height), categorical features (such as surgery type and anesthesia method), and text features (such as surgery name and diagnosis). Existing methods directly integrate these features into prediction models, lacking a thorough assessment and screening of their contributions. This not only increases the complexity of the prediction model but can also reduce prediction accuracy and efficiency due to the presence of redundant or irrelevant features.
[0006] Based on the above problems, there is an urgent need for an innovative surgical duration prediction method that can be optimized according to the characteristics of surgeries in different departments, while introducing more reasonable and fair prediction model evaluation indicators and combining feature screening technology to improve the adaptability, interpretability and performance of the prediction model, providing a more scientific and efficient solution for medical resource management.
[0007] In real-world scenarios, accurate prediction of surgical duration is crucial for hospital resource allocation and efficient operation. However, existing surgical duration prediction methods still suffer from multiple deficiencies. In terms of feature selection, due to the multimodality and complexity of medical data, existing methods fail to effectively select the most important features for prediction. This results in a large amount of redundant information and noise in the prediction model input, increasing the complexity of the prediction model and reducing prediction performance and interpretability. Furthermore, existing methods generally use a single prediction model, failing to fully exploit the specific characteristics of surgical duration distribution across departments and ignoring differences in duration between departments, limiting the adaptability and versatility of the prediction model. Regarding evaluation metrics, traditional methods often use mean squared error (MSE) or mean absolute error (MAE). However, these absolute error metrics cannot eliminate the impact of surgical duration range on prediction evaluation and can easily mask model deficiencies, especially in short-duration surgeries. Therefore, a surgical duration prediction solution that combines feature selection, department-specific optimization, and relative error evaluation is urgently needed to improve the accuracy, adaptability, and reliability of the prediction model. Summary of the Invention
[0008] In response to the problems existing in the above-mentioned prior art, the present invention provides a method, system, electronic device, and readable storage medium for predicting operation duration based on department distribution adjustment. The present invention improves the accuracy, adaptability, and reliability of operation duration prediction.
[0009] The technical solution of the present invention is: a method for predicting the duration of surgery based on department distribution adjustment, the method comprising:
[0010] Step 1: Use principal component analysis (PCA) and SHAP techniques to identify features that significantly affect the duration of surgery.
[0011] Step 2. Based on feature screening, a modular prediction model is designed for the distribution characteristics of operation duration in different departments. Through a dynamic calling mechanism, the optimal prediction model is selected according to the department to which the surgery belongs to predict the operation duration.
[0012] In surgery duration prediction, input data includes text, categorical, and numerical features. However, due to differences in physician input habits and surgery types, many features may be missing or irrelevant. For example, in complex surgeries, physicians may pay more attention to the patient's age and medical history, while they may ignore this information in simpler procedures. If all features are used in the prediction, it will not only introduce noise but also increase the computational complexity of the model, reducing prediction performance. Therefore, the core goal of feature screening is to extract key features, eliminate redundant information, and improve the efficiency and accuracy of the prediction model.
[0013] Furthermore, the Step 1 includes:
[0014] Step 1.1. Principal component analysis (PCA) is used to reduce the dimensionality of numerical and categorical features. This is done to retain features that are important for predicting surgical duration, while removing features with low correlation and high collinearity to obtain a reduced-dimensional feature representation. The covariance matrix calculation process in PCA is as follows:
[0015]
[0016] in is the covariance matrix, x i is the feature vector of the i-th sample, is the feature mean vector, n is the total number of samples;
[0017] By performing eigenvalue decomposition on the covariance matrix, the first k principal components with large eigenvalues are selected as the features after dimensionality reduction; the features after dimensionality reduction are expressed as:
[0018]
[0019] Where W is the matrix containing the first k eigenvectors, is the feature representation after dimensionality reduction, x is the data point in the feature vector, which is used to transform the feature representation during the dimensionality reduction process. Through the matrix W containing the principal components and the transformation formula, the data is mapped to the low-dimensional space to obtain the feature representation z after dimensionality reduction;
[0020] Step 1.2: To further quantify the contribution of features to prediction, SHAP technology is used to analyze the importance of features. The core principle of SHAP is based on cooperative game theory. By calculating the marginal contribution of each feature to the output of the prediction model, the importance of each feature is defined. For a feature vector x i , whose SHAP value φ i The calculation formula is:
[0021]
[0022] Among them, S represents the feature set, which is a subset of all feature sets N. Represents a partial feature combination selected from N. In the SHAP value calculation, S is used to represent only some of the features used by the current model. Its purpose is to analyze the contribution of these features to the model prediction results; N represents the entire feature set, which includes the complete input features required for model prediction. For example, if the input features of the model include patient age, surgery type, anesthesia method, etc., then N includes all these features; the core of the SHAP value calculation is to evaluate the incremental contribution of a specific feature {i} (a feature that does not belong to S) to the model prediction after being added to S; f(S) is the prediction output of the prediction model when it only contains the feature set S, which means that the model only considers the features in S and ignores the impact of other features on the results; f(S∪{i}) represents the model output after adding feature {i} to the feature set S, and the meaning of f(S∪{i})-f(S) is: it represents the marginal contribution of feature {i}, that is, the incremental change brought about by adding feature {i} to the model output on the basis of feature set S.
[0023] By combining PCA dimensionality reduction with SHAP analysis, we screened out the top m features that contribute most to surgical duration prediction. Combined with actual physician survey results, we focused on retaining the patient characteristics that physicians are most concerned about during surgery. This strategy reduces redundant information and noise while making the model more aligned with clinical needs.
[0024] Furthermore, the Step 2 includes:
[0025] Designing and selecting the optimal prediction model based on the surgical characteristics of different departments aims to significantly improve the accuracy and adaptability of surgery duration prediction. Different departments have significant differences in surgery types, patient populations, and the distribution of surgery durations. For example, the duration distribution of gynecological surgeries is relatively concentrated, indicating a high degree of standardization of surgical processes. However, due to the diverse types of surgeries, the duration distribution of general surgery is more random and complex. These differences make it difficult for a single prediction model to fully adapt to the needs of all departments. Therefore, optimizing the prediction model by department is the key to solving this problem.
[0026] Step 2.1. Through statistical analysis of the historical data of surgical duration in each department, the distribution characteristics of surgical duration in each department are identified, and the results of the department characteristic analysis are obtained. The distribution characteristics of surgical duration in each department include mean, standard deviation, skewness, and kurtosis. The department characteristic analysis clarifies the characteristics of different departments and provides a theoretical basis for model design. For example, if the skewness of a department is close to 0 and the kurtosis is high, it indicates that its surgical duration distribution is close to a normal distribution and is suitable for a simple linear model. However, if the skewness and kurtosis values are large, a more complex nonlinear model is required to capture the long-tail characteristics in the distribution.
[0027] Through statistical analysis of the historical data of operation time of each department, the distribution characteristics of the operation time of each department (including mean, standard deviation, skewness and kurtosis) are obtained and department characteristics analysis is performed; the statistical analysis of the historical data of operation time of each department mainly focuses on the distribution characteristics of the operation time of each department, and does not involve the feature screening results in Step 1; in Step 2.1, only the operation time data of each department are used to calculate its mean, standard deviation, skewness and kurtosis, so as to understand the basic distribution of operation time in each department; for example, the mean reflects the average level of operation time, the standard deviation measures the volatility of operation time, and the skewness and kurtosis further describe the symmetry and concentration of the data distribution; these statistical characteristics help understand the distribution pattern of operation time in different departments and provide a basis for subsequent model design;
[0028] The feature screening in Step 1 (important features screened using PCA and SHAP techniques) was not directly used in Step 2.1. In Step 2.1, the statistical analysis of departmental surgical duration was performed independently to provide a quantitative basis for departmental characteristics for subsequent model design.
[0029] Step 2.2: Design a modular prediction model based on the departmental characteristics analysis results: The features that significantly influence surgery duration identified through PCA and SHAP in Step 1 will be used as model inputs in this step to optimize the prediction model design for different departments. Based on the distribution characteristics of surgery duration in each department, the most suitable model architecture will be selected. These models will be trained and optimized using the features identified in Step 1 to ensure that the models can accurately capture the characteristics of surgery duration in each department.
[0030] In Step 2.2, based on the department-specific analysis results from Step 2.1, the process of designing a modular prediction model involves using the department's surgical duration distribution characteristics (e.g., mean, standard deviation, skewness, and kurtosis) to guide model selection and design. Specifically, these analysis results help us understand the distribution pattern of surgical duration in each department and thus select the most appropriate model type for different departments.
[0031] For example, for departments with a relatively concentrated distribution of surgical duration, a simple regression model (such as linear regression or a neural network with a small number of hidden layers) can be used as the model. For example, if the standard deviation of surgical duration in a department is small and the data distribution is close to a normal distribution, we can choose linear regression or a shallow neural network to fit the data.
[0032] For departments with a more dispersed and volatile distribution of surgical duration, more complex neural network models (such as multi-layer perceptrons (MLPs) or deep neural networks) are used for prediction. These models can better capture the complex nonlinear relationships and interactive features in surgical duration. For example, models such as the GLU activation function and multi-layer linear transformation can be used to model these complex relationships.
[0033] Through these designs, Step 2.2 ensures that each department uses its most appropriate model, which can improve prediction accuracy based on department characteristics and operation duration distribution;
[0034] The modular prediction models all contain embedding layers, multi-layer linear networks, and activation functions, which can flexibly adapt to the input characteristics of different departments;
[0035] The embedding layer is used to map classification features (such as surgeon ID, department ID, etc.) into a fixed-dimensional vector space to achieve feature densification;
[0036] The multi-layer linear network is used to perform feature fusion on the standardized numerical features and classification features;
[0037] The activation function is used to enhance the nonlinear expression ability of the prediction model. Different activation functions are selected according to the distribution characteristics of different departments. The selected activation functions include GELU, Sigmoid or Tanh.
[0038] Furthermore, in Step 2.2, the modular prediction model further includes:
[0039] (1) To improve the adaptability to extreme values, the modular prediction model introduces output adjustment logic to improve the adaptability to extreme values. The output adjustment logic includes range correction of the predicted value. The range correction of the predicted value is performed as follows:
[0040] y pred =y base +σ(y adjust )·α (4)
[0041] Among them, y base is the basic prediction value, σ is the activation function (such as Sigmoid), y adjust is the adjustment value, which is based on the initial prediction result y base The correction is related to the distribution characteristics of operation time in different departments (such as mean, standard deviation, etc.); the distribution of operation time in each department may be different, so the basic prediction value needs to be adjusted; for example, the operation time in some departments may be generally longer (larger mean), while that in other departments may be shorter (smaller mean); by adjusting y baseDepartment-based adjustments can more accurately reflect the predictions of specific departments. This adjustment is not just a simple shift of the basic prediction value, but a dynamic adjustment by considering the distribution characteristics of the department (such as standard deviation, skewness, kurtosis, etc.). For example, if the operation time of a department is generally longer, y adjust May be increased accordingly to ensure that the model better predicts the length of surgery in this department; pred is the final prediction result, the activation function σ is used to predict y adjust A nonlinear transformation is performed to compress the prediction results into a reasonable range (for example, to limit the results to a certain interval). α is a dynamic adjustment factor used to capture the special distribution characteristics of long-duration surgeries and is related to the specific characteristics of the department (such as the distribution characteristics of the operation duration). For example, for some departments, α may be larger, which means that their predictions will respond more sensitively to changes in the adjustment value. Through this correction method, y pred It not only reflects the basic prediction value, but also takes into account the adjustments brought about by department characteristics (such as the distribution of surgical duration), making the model more adaptable and accurate when processing data from different departments;
[0042] (2) For departments with less data, the modular prediction model introduces a distribution similarity measurement method to merge the data of departments with similar distributions for joint training; the similarity measurement D sim The calculation formula is:
[0043]
[0044] Among them, P A (t) and P B (t) are the density functions of the operation time distribution of departments A and B, respectively, and t represents the variable of the operation time distribution in the two departments A and B; D sim Larger values indicate higher distribution similarity. This approach, for example, by merging data from orthopedics and pain medicine, can effectively alleviate the problem of insufficient training in small datasets and improve the generalization ability of the model.
[0045] In practice, the system dynamically calls the corresponding prediction model through the identification module of the surgical department, inputs preoperative data into the model, and outputs the predicted surgical duration. This dynamic calling strategy ensures that each department makes predictions based on its optimal model, maximizing model performance. Experimental results show that compared with traditional single-model methods, the department-specific prediction module significantly reduces the mean absolute error and demonstrates higher prediction accuracy in complex departments.
[0046] Furthermore, the dynamic calling mechanism for selecting the optimal prediction model according to the surgical department includes:
[0047] The corresponding prediction model is dynamically called through the identification module of the surgery department, the preoperative data is input into the prediction model and the predicted value of the surgery duration is output.
[0048] Furthermore, the prediction error was normalized to the ratio of the actual value by the mean absolute percentage error, and the ratio was used to reflect the performance of the prediction model in different surgical scenarios;
[0049] The main advantage of the MAPE (mean absolute percentage error) evaluation module is that it provides a fairer evaluation method for different surgical scenarios through the calculation of relative error. Traditional absolute error metrics (such as mean squared error (MSE) or mean absolute error (MAE)) are easily affected by the range of surgical duration, often underestimating the error for long surgeries and overestimating the error for short surgeries. Therefore, using MAPE as the core evaluation metric can effectively avoid this problem and more accurately reflect the actual predictive performance of the model.
[0050] The calculation formula of the mean absolute percentage error MAPE is:
[0051]
[0052] Among them, n is the total number of samples, y i is the actual operation time of the i-th sample, is the operation duration predicted by the prediction model; this formula eliminates the influence of the absolute value of the operation duration on the error evaluation by normalizing the prediction error to the ratio of the actual value, so that the prediction performance of long and short operations can be fairly measured;
[0053] The relative error can accurately reflect the model performance in different surgical scenarios by standardizing the error size. For example, for a complex operation with an actual operation time of 300 minutes, even if the prediction error reaches 30 minutes, the relative error is 10%, indicating that the overall accuracy of the model prediction is still high. For another short operation of only 30 minutes, even if the prediction error is only 3 minutes, the relative error is still 10%, indicating that the prediction of the short operation also achieves similar accuracy. Under this evaluation method, the prediction performance of both long and short operations can be fairly reflected, avoiding the bias of traditional absolute error indicators towards long operations.
[0054] The present invention also provides a surgery duration prediction system based on department distribution adjustment, the system comprising: a module for executing the above-mentioned surgery duration prediction method based on department distribution adjustment.
[0055] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method for predicting surgical duration based on departmental distribution adjustment when executing the program.
[0056] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for predicting the length of surgery based on departmental distribution adjustment is implemented.
[0057] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned method for predicting surgical duration based on department distribution adjustment.
[0058] In the present invention, Step 2 relies on the key features screened out in Step 1 to design and optimize the department-specific prediction model. Specifically, Step 1 uses principal component analysis (PCA) and SHAP techniques to remove redundant and noisy features from multimodal data, retaining the most important features for predicting surgical duration. These screened features provide a clear and optimized data foundation for the model input in Step 2.
[0059] In Step 2, these filtered key features are directly used as input data for the modular prediction model. For example, for a specific department, the model only uses the highly important features identified by PCA and SHAP, excluding features that have little contribution to prediction or may introduce noise. This feature enables the model to more closely capture the key factors affecting surgical duration, improving model training efficiency and prediction accuracy.
[0060] In addition, the feature screening process in Step 1 not only optimizes the input dimension but also provides guidance for model design. The results of the SHAP analysis clarify the relative importance of each feature, thereby helping to select an appropriate model architecture. For example, for certain departments, if patient age and surgical complexity are found to be the main features, the model can focus on feature interaction design for these features while ignoring other irrelevant features. On the other hand, PCA provides principal component features, which can be further used as optimized data for the model input layer to ensure that the department-specific models are consistent in terms of data input dimensions.
[0061] Finally, during the actual model call process, whether in the training or prediction stage, the screened features are directly applied to the models of each department as a unified input standard; this combination of optimized features and dynamic model calls ensures the adaptability and efficiency of predictions in different departments.
[0062] The present invention proposes a method for predicting surgical duration based on departmental optimization, combining feature screening, departmental prediction module design and MAPE evaluation to construct a complete solution. First, the input features are screened using principal component analysis (PCA) and SHAP technology. PCA removes redundant features and noise through dimensionality reduction, retaining the principal components that contribute most to the prediction of surgical duration, significantly reducing the complexity of the prediction model and improving training efficiency. SHAP technology quantifies the contribution of each feature to the prediction results, screens out the features that have the most significant impact on surgical duration, and optimizes feature selection based on the clinical needs of doctors, making the prediction model more efficient and highly interpretable.
[0063] Based on feature screening, the present invention designs a series of modular prediction models for the duration distribution characteristics of surgeries in different departments. By analyzing the mean, standard deviation and distribution characteristics of the operation duration of each department, a variety of candidate prediction models are combined for evaluation and selection to ensure that each department can use the prediction model that best suits its characteristics. In addition, in order to solve the problem of insufficient data in some departments, the present invention analyzes the similarity of the distribution of operation duration, merges the data of departments with similar distributions for training, and improves data utilization and the generalization ability of the prediction model. In the prediction stage, the system automatically selects the corresponding optimal prediction model according to the department to which the operation belongs through a dynamic calling mechanism to ensure the adaptability and prediction accuracy of the prediction model in different department scenarios.
[0064] Finally, this paper uses the mean absolute percentage error (MAPE) as a core performance evaluation metric. By normalizing the prediction error as a proportion of the actual value, MAPE effectively eliminates the influence of the range of surgical durations on the evaluation results, allowing for a fair assessment of the prediction performance of long and short surgeries. During the prediction model optimization process, MAPE provides intuitive feedback for performance analysis, helping to identify the shortcomings of the prediction model in different surgical duration scenarios, further improving the applicability and reliability of the prediction model.
[0065] The beneficial effects of the present invention are:
[0066] 1. This invention effectively reduces the complexity of the prediction model and improves the training efficiency through feature screening strategy, while enhancing the interpretability of the model, making the prediction model more in line with clinical needs;
[0067] 2. This invention analyzes the similarity of surgical duration distribution and merges the data of departments with similar distribution for training. This solves the problem of insufficient data in some departments and improves data utilization and the generalization ability of the prediction model.
[0068] 3. The present invention significantly improves the prediction accuracy and adaptability of the prediction model through department-specific optimization design, ensuring that the prediction performance of operation duration in different departments can reach the best.
[0069] 4. This paper introduces MAPE as the core evaluation indicator, which fairly measures the performance of the prediction model in surgeries of different durations, providing a scientific basis for the optimization and practical application of the prediction model;
[0070] 5. By combining various technologies, the present invention constructs an efficient, accurate and adaptable method for predicting surgical duration in complex clinical scenarios, providing strong support for the scientific allocation and efficient utilization of medical resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 It is the overall flow chart of the scheme in the present invention. DETAILED DESCRIPTION
[0072] Example 1: Figure 1 As shown, a method for predicting operation duration based on department distribution adjustment includes:
[0073] Step 1: Use principal component analysis (PCA) and SHAP techniques to identify features that significantly affect the duration of surgery.
[0074] Furthermore, the Step 1 includes:
[0075] In surgery duration prediction, input data includes text, categorical, and numerical features. However, due to differences in physician input habits and surgery types, many features may be missing or irrelevant. For example, in complex surgeries, physicians may pay more attention to the patient's age and medical history, while they may ignore this information in simpler procedures. If all features are used in the prediction, it will not only introduce noise but also increase the computational complexity of the model, reducing prediction performance. Therefore, the core goal of feature screening is to extract key features, eliminate redundant information, and improve the efficiency and accuracy of the prediction model.
[0076] Step 1.1. Principal component analysis (PCA) is used to reduce the dimensionality of numerical and categorical features. This is done to retain features important for predicting surgical duration, remove features with low correlation and high collinearity, and obtain a reduced-dimensional feature representation. The core principle of PCA is to find the main direction of change in the data through the feature covariance matrix. The calculation process of the covariance matrix in PCA is as follows:
[0077]
[0078] in is the covariance matrix, x i is the feature vector of the i-th sample, is the feature mean vector, n is the total number of samples;
[0079] By performing eigenvalue decomposition on the covariance matrix, the first k principal components with large eigenvalues are selected as the features after dimensionality reduction; the features after dimensionality reduction are expressed as:
[0080]
[0081] Where W is the matrix containing the first k eigenvectors, is the feature representation after dimensionality reduction, x is the data point in the feature vector, which is used to transform the feature representation during the dimensionality reduction process. Through the matrix W containing the principal components and the transformation formula, the data is mapped to the low-dimensional space to obtain the feature representation z after dimensionality reduction;
[0082] In the specific implementation, the present invention selects the principal components that can explain more than 90% of the data variance, and reduces the feature dimensions from hundreds of dimensions to less than 50 dimensions, which significantly improves the training efficiency of the model.
[0083] Step 1.2: To further quantify the contribution of features to prediction, SHAP technology is used to analyze the importance of features. The core principle of SHAP is based on cooperative game theory. By calculating the marginal contribution of each feature to the output of the prediction model, the importance of each feature is defined. For a feature vector x i , whose SHAP value φ i The calculation formula is:
[0084]
[0085] Among them, S represents the feature set, which is a subset of all feature sets N. Represents a partial feature combination selected from N. In the SHAP value calculation, S is used to represent only some of the features used by the current model. Its purpose is to analyze the contribution of these features to the model prediction results. N represents the entire feature set, which includes the complete input features required for model prediction. For example, if the input features of the model include patient age, surgery type, anesthesia method, etc., then N includes all these features. The core of the SHAP value calculation is to evaluate the incremental contribution of a specific feature {i} (a feature that does not belong to S) to the model prediction after being added to S. f(S) is the prediction output of the prediction model when it only contains the feature set S, which means that the model only considers the features in S and ignores the impact of other features on the results. f(S∪{i}) represents the model output after adding feature {i} to the feature set S. The meaning of f(S∪{i})-f(S) is: it represents the marginal contribution of feature {i}, that is, the incremental change brought about by adding feature {i} to the model output on the basis of feature set S.
[0086] By combining PCA dimensionality reduction with SHAP analysis, we identified the top m features that contribute most to surgical duration prediction. Furthermore, combined with actual physician survey results, we prioritized the patient characteristics that physicians most value during surgery. This strategy reduces redundant information and noise while making the model more relevant to clinical needs. Experiments show that after feature selection, model training time was more than halved and prediction accuracy increased by approximately 5%, achieving high efficiency without sacrificing accuracy.
[0087] Step 2. Based on feature screening, a modular prediction model is designed for the distribution characteristics of operation duration in different departments. Through a dynamic calling mechanism, the optimal prediction model is selected according to the department to which the surgery belongs to predict the operation duration.
[0088] Designing and selecting the optimal prediction model based on the surgical characteristics of different departments aims to significantly improve the accuracy and adaptability of surgery duration prediction. Different departments have significant differences in surgery types, patient populations, and the distribution of surgery durations. For example, the duration distribution of gynecological surgeries is relatively concentrated, indicating a high degree of standardization of surgical processes. However, due to the diverse types of surgeries, the duration distribution of general surgery is more random and complex. These differences make it difficult for a single prediction model to fully adapt to the needs of all departments. Therefore, optimizing the prediction model by department is the key to solving this problem.
[0089] Furthermore, the Step 2 includes:
[0090] Step 2.1. Through statistical analysis of the historical data of surgical duration in each department, the distribution characteristics of surgical duration in each department are identified, and the results of the department characteristic analysis are obtained. The distribution characteristics of surgical duration in each department include mean, standard deviation, skewness, and kurtosis. The department characteristic analysis clarifies the characteristics of different departments and provides a theoretical basis for model design. For example, if the skewness of a department is close to 0 and the kurtosis is high, it indicates that its surgical duration distribution is close to a normal distribution and is suitable for a simple linear model. However, if the skewness and kurtosis values are large, a more complex nonlinear model is required to capture the long-tail characteristics in the distribution.
[0091] Through statistical analysis of the historical data of operation time of each department, the distribution characteristics of the operation time of each department (including mean, standard deviation, skewness and kurtosis) are obtained and department characteristics analysis is performed; the statistical analysis of the historical data of operation time of each department mainly focuses on the distribution characteristics of the operation time of each department, and does not involve the feature screening results in Step 1; in Step 2.1, only the operation time data of each department are used to calculate its mean, standard deviation, skewness and kurtosis, so as to understand the basic distribution of operation time in each department; for example, the mean reflects the average level of operation time, the standard deviation measures the volatility of operation time, and the skewness and kurtosis further describe the symmetry and concentration of the data distribution; these statistical characteristics help understand the distribution pattern of operation time in different departments and provide a basis for subsequent model design;
[0092] The feature screening in Step 1 (important features screened using PCA and SHAP techniques) was not directly used in Step 2.1. In Step 2.1, the statistical analysis of departmental surgical duration was performed independently to provide a quantitative basis for departmental characteristics for subsequent model design.
[0093] Step 2.2: Design a modular prediction model based on the departmental characteristics analysis results: The features that significantly influence surgery duration identified through PCA and SHAP in Step 1 will be used as model inputs in this step to optimize the prediction model design for different departments. Based on the distribution characteristics of surgery duration in each department, the most suitable model architecture will be selected. These models will be trained and optimized using the features identified in Step 1 to ensure that the models can accurately capture the characteristics of surgery duration in each department.
[0094] In Step 2.2, based on the department-specific analysis results from Step 2.1, the process of designing a modular prediction model involves using the department's surgical duration distribution characteristics (e.g., mean, standard deviation, skewness, and kurtosis) to guide model selection and design. Specifically, these analysis results help us understand the distribution pattern of surgical duration in each department and thus select the most appropriate model type for different departments.
[0095] For example, for departments with a relatively concentrated distribution of surgical duration, a simple regression model (such as linear regression or a neural network with a small number of hidden layers) can be used as the model. For example, if the standard deviation of surgical duration in a department is small and the data distribution is close to a normal distribution, we can choose linear regression or a shallow neural network to fit the data.
[0096] For departments with a more dispersed and volatile distribution of surgical duration, more complex neural network models (such as multi-layer perceptrons (MLPs) or deep neural networks) are used for prediction. These models can better capture the complex nonlinear relationships and interactive features in surgical duration. For example, models such as the GLU activation function and multi-layer linear transformation can be used to model these complex relationships.
[0097] Through these designs, Step 2.2 ensures that each department uses its most appropriate model, which can improve prediction accuracy based on department characteristics and operation duration distribution;
[0098] The modular prediction models all contain embedding layers, multi-layer linear networks, and activation functions, which can flexibly adapt to the input characteristics of different departments;
[0099] The embedding layer is used to map classification features (such as surgeon ID, department ID, etc.) into a fixed-dimensional vector space to achieve feature densification;
[0100] The multi-layer linear network is used to perform feature fusion on the standardized numerical features and classification features;
[0101] The activation function is used to enhance the nonlinear expression ability of the prediction model. Different activation functions are selected according to the distribution characteristics of different departments. The selected activation functions include GELU, Sigmoid or Tanh.
[0102] Furthermore, in Step 2.2, the modular prediction model further includes:
[0103] (1) To improve the adaptability to extreme values, the modular prediction model introduces output adjustment logic to improve the adaptability to extreme values. The output adjustment logic includes range correction of the predicted value. The range correction of the predicted value is performed as follows:
[0104] y pred =y base +σ(y adjust )·α (4)
[0105] Among them, y base is the basic prediction value, σ is the activation function (such as Sigmoid), y adjust is the adjustment value, which is based on the initial prediction result y base The correction is related to the distribution characteristics of operation time in different departments (such as mean, standard deviation, etc.); the distribution of operation time in each department may be different, so the basic prediction value needs to be adjusted; for example, the operation time in some departments may be generally longer (larger mean), while that in other departments may be shorter (smaller mean); by adjusting y baseDepartment-based adjustments can more accurately reflect the predictions of specific departments. This adjustment is not just a simple shift of the basic prediction value, but a dynamic adjustment by considering the distribution characteristics of the department (such as standard deviation, skewness, kurtosis, etc.). For example, if the operation time of a department is generally longer, y adjust May be increased accordingly to ensure that the model better predicts the length of surgery in this department; pred is the final prediction result, the activation function σ is used to predict y adjust A nonlinear transformation is performed to compress the prediction results into a reasonable range (for example, to limit the results to a certain interval). α is a dynamic adjustment factor used to capture the special distribution characteristics of long-duration operations and is related to the specific characteristics of the department (such as the distribution characteristics of the operation duration). For example, for some departments, α may be larger, which means that their predictions will respond more sensitively to changes in the adjustment value. In this way, y pred It not only reflects the basic prediction value, but also takes into account the adjustments brought about by department characteristics (such as the distribution of surgical duration), making the model more adaptable and accurate when processing data from different departments;
[0106] (2) For departments with less data, the modular prediction model introduces a distribution similarity measurement method to merge the data of departments with similar distributions for joint training; the similarity measurement D sim The calculation formula is:
[0107]
[0108] Among them, P A (t) and P B (t) are the density functions of the operation time distribution of departments A and B, respectively, and t represents the variable of the operation time distribution in the two departments A and B; D sim Larger values indicate higher distribution similarity. This approach, for example, by merging data from orthopedics and pain medicine, can effectively alleviate the problem of insufficient training in small datasets and improve the generalization ability of the model.
[0109] In practice, the system dynamically calls the corresponding prediction model through the identification module of the surgical department, inputs preoperative data into the model, and outputs the predicted surgical duration. This dynamic calling strategy ensures that each department makes predictions based on its optimal model, maximizing model performance. Experimental results show that compared with traditional single-model methods, the department-specific prediction module significantly reduces the mean absolute error and demonstrates higher prediction accuracy in complex departments.
[0110] Furthermore, the dynamic calling mechanism for selecting the optimal prediction model according to the surgical department includes:
[0111] The corresponding prediction model is dynamically called through the identification module of the surgery department, the preoperative data is input into the prediction model and the predicted value of the surgery duration is output.
[0112] Furthermore, the prediction error was normalized to the ratio of the actual value by the mean absolute percentage error, and the ratio was used to reflect the performance of the prediction model in different surgical scenarios;
[0113] The main advantage of the MAPE (mean absolute percentage error) evaluation module is that it provides a fairer evaluation method for different surgical scenarios through the calculation of relative error. Traditional absolute error metrics (such as mean squared error (MSE) or mean absolute error (MAE)) are easily affected by the range of surgical duration, often underestimating the error for long surgeries and overestimating the error for short surgeries. Therefore, using MAPE as the core evaluation metric can effectively avoid this problem and more accurately reflect the actual predictive performance of the model.
[0114] The calculation formula of the mean absolute percentage error MAPE is:
[0115]
[0116] Among them, n is the total number of samples, y i is the actual operation time of the i-th sample, is the operation duration predicted by the prediction model; this formula eliminates the influence of the absolute value of the operation duration on the error evaluation by normalizing the prediction error to the ratio of the actual value, so that the prediction performance of long and short operations can be fairly measured;
[0117] The relative error can accurately reflect the model performance in different surgical scenarios by standardizing the error size. For example, for a complex operation with an actual operation time of 300 minutes, even if the prediction error reaches 30 minutes, the relative error is 10%, indicating that the overall accuracy of the model prediction is still high. For another short operation of only 30 minutes, even if the prediction error is only 3 minutes, the relative error is still 10%, indicating that the prediction of the short operation also achieves similar accuracy. Under this evaluation method, the prediction performance of both long and short operations can be fairly reflected, avoiding the bias of traditional absolute error indicators towards long operations.
[0118] Experimental results show that compared with traditional MSE and MAE evaluation methods, MAPE can more clearly reflect the actual performance of the model; in some departments with a high proportion of short-duration operations, MAPE can accurately capture the distribution of prediction errors and help discover the direction of model optimization; at the same time, the introduction of MAPE makes the model training objectives more in line with actual business needs, providing a scientific basis for improving the performance of the surgery duration prediction model.
[0119] The present invention also provides a surgery duration prediction system based on department distribution adjustment, the system comprising:
[0120] Feature screening module, used to screen out features that significantly affect surgery duration through principal component analysis (PCA) and SHAP techniques;
[0121] The department-specific prediction module is used to design modular prediction models based on feature screening and the distribution characteristics of operation duration in different departments; through a dynamic calling mechanism, the optimal prediction model is selected according to the department to which the surgery belongs to predict the operation duration.
[0122] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method for predicting surgical duration based on departmental distribution adjustment when executing the program.
[0123] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for predicting the length of surgery based on departmental distribution adjustment is implemented.
[0124] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned method for predicting surgical duration based on department distribution adjustment.
[0125] The specific embodiments of the present invention are described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made within the knowledge of ordinary technicians in this field without departing from the scope of the present invention.
Claims
1. A method for predicting surgical duration based on departmental distribution adjustment, characterized by: The method comprises: Step 1: Use principal component analysis (PCA) and SHAP techniques to identify features that significantly affect the duration of surgery. Step 2. Based on feature screening, a modular prediction model is designed for the distribution characteristics of operation duration in different departments. Through a dynamic calling mechanism, the optimal prediction model is selected according to the department to which the surgery belongs to predict the operation duration.
2. The method for predicting surgical duration based on departmental distribution adjustment according to claim 1, characterized in that: Step 1 includes: Step 1.
1. Principal component analysis (PCA) is used to reduce the dimensionality of numerical and categorical features. This is done to retain features that are important for predicting surgical duration, while removing features with low correlation and high collinearity to obtain a reduced-dimensional feature representation. The covariance matrix calculation process in PCA is as follows: in is the covariance matrix, x i is the feature vector of the i-th sample, is the feature mean vector, n is the total number of samples; By performing eigenvalue decomposition on the covariance matrix, the first k principal components with large eigenvalues are selected as the features after dimensionality reduction; the features after dimensionality reduction are expressed as: Where W is the matrix containing the first k eigenvectors, is the feature representation after dimensionality reduction, x is the data point in the feature vector, which is used to transform the feature representation during the dimensionality reduction process; Step 1.2, use SHAP technology to analyze the importance of features; define the importance of each feature by calculating the marginal contribution of each feature to the output of the prediction model; for a feature vector x i , whose SHAP value φ i The calculation formula is: Among them, S represents the feature set, which is a subset of all feature sets N. N represents all feature sets, including the complete input features required for model prediction. f(S) is the prediction output of the prediction model when it only contains the feature set S. f(S∪{i}) represents the model output after adding feature {i} to the feature set S. The meaning of f(S∪{i})-f(S) is: it represents the marginal contribution of feature {i}, that is, the incremental change brought about by adding feature {i} to the feature set S.
3. The method for predicting surgical duration based on departmental distribution adjustment according to claim 1, characterized in that: Step 2 includes: Step 2.
1. Analyze the historical data of operation duration in each department to identify the distribution characteristics of operation duration in each department and obtain the department characteristic analysis results. The distribution characteristics of operation duration in each department include mean, standard deviation, skewness and kurtosis. Step 2.2: Design a modular prediction model based on the departmental characteristics analysis results: The features that significantly influence surgery duration identified through PCA and SHAP in Step 1 will be used as model inputs in this step to optimize the prediction model design for different departments. Based on the distribution characteristics of surgery duration in each department, the most suitable model architecture will be selected. These models will be trained and optimized using the features identified in Step 1 to ensure that the models can accurately capture the characteristics of surgery duration in each department. The modular prediction models all contain embedding layers, multi-layer linear networks, and activation functions; The embedding layer is used to map the classification features to a fixed-dimensional vector space to achieve feature densification; The multi-layer linear network is used to perform feature fusion on the standardized numerical features and classification features; The activation function is used to enhance the nonlinear expression ability of the prediction model. Different activation functions are selected according to the distribution characteristics of different departments. The selected activation functions include GELU, Sigmoid or Tanh.
4. The method for predicting surgical duration based on departmental distribution adjustment according to claim 3, characterized in that: In Step 2.2, the modular prediction model further includes: (1) The modular prediction model introduces output adjustment logic to improve its adaptability to extreme values. The output adjustment logic includes range correction of the predicted values. The range correction of the predicted values is performed as follows: and pred =and base +σ(and adjust )·α (4) Among them, y base is the basic predicted value, y adjust is the adjustment value, y pred is the final prediction result, σ is the activation function, and α is the dynamic adjustment factor used to capture the special distribution characteristics of long-duration operations; (2) The modular prediction model introduces a distribution similarity measurement method to merge the data of departments with similar distributions for joint training; the similarity measurement D sim The calculation formula is: Among them, P A (t) and P B (t) are the operation time distribution density functions of departments A and B respectively, t represents the variable of operation time distribution in the two departments A and B, D sim Larger values indicate higher distribution similarity.
5. The method for predicting surgical duration based on departmental distribution adjustment according to claim 1, characterized in that: The dynamic calling mechanism for selecting the optimal prediction model according to the surgical department includes: The corresponding prediction model is dynamically called through the identification module of the surgery department, the preoperative data is input into the prediction model and the predicted value of the surgery duration is output.
6. The method for predicting surgical duration based on departmental distribution adjustment according to claim 1, characterized in that: Also includes: The prediction error is normalized to the ratio of the actual value by the mean absolute percentage error, and the ratio reflects the performance of the prediction model in different surgical scenarios; The calculation formula of the mean absolute percentage error MAPE is: Among them, n is the total number of samples, y i is the actual operation time of the i-th sample, The operation duration predicted by the prediction model.
7. A surgery duration prediction system based on department distribution adjustment, characterized in that: The system includes: a module for executing a method for predicting surgical duration based on department distribution adjustment as described in any one of claims 1 to 6.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements a method for predicting surgical duration based on department distribution adjustment as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements a method for predicting surgical duration based on department distribution adjustment as described in any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it implements a method for predicting surgical duration based on department distribution adjustment as described in any one of claims 1 to 6.
Citation Information
Cited By
Large model operation duration prediction method based on PCA (Principal Component Analysis) weighted retrieval and Bayesian aggregation
CN121858930A