Fair hospital performance assessment data processing method and system based on big data
By preprocessing and comprehensive scoring of multi-source data, combining partial derivative adjustment and Pareto optimal solution set solution, the problem of traditional assessment methods being unable to handle complex nonlinear relationships and ignoring data noise is solved, and a more accurate and intelligent hospital performance appraisal is achieved.
Patent Information
- Application Number
- CN202510066159.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional hospital performance appraisal methods are limited to a single data dimension or a simple linear model, and cannot fully consider the complex nonlinear relationship between various factors, and ignore the impact of data noise and missing data on the assessment results, resulting in possible deviations in the assessment results.
By collecting multi-source data for preprocessing, calculating comprehensive scores, and adjusting input variables using partial derivatives, building an objective function to solve Pareto's optimal solution set, and finally building a visual interface to display the results and store the data.
It improves the accuracy and real-time performance appraisal of hospitals, enhances the intelligence and adaptability of data processing, solves data quality problems, and provides more scientific and accurate performance evaluation methods.
Smart Images

Figure CN119993420A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing, and in particular to a method and system for processing public hospital performance appraisal data based on big data. Background Art
[0002] With the rapid development of information technology, especially the widespread application of big data, cloud computing and artificial intelligence technology, the medical field has gradually entered a new era driven by data. In the management and operation of public hospitals, performance appraisal, as an important decision-making support tool, has received more and more attention. Traditional hospital performance appraisal mainly relies on manually collected data and simple evaluation indicators, which cannot comprehensively and objectively reflect the hospital's operating status and service quality. With the continuous growth of medical service demand and the increasing tension of medical resources, how to efficiently and accurately evaluate the comprehensive performance of hospitals has become a key issue in the current medical management field. Performance appraisal methods based on big data have gradually emerged. Through comprehensive analysis of multi-dimensional data such as the number of patients, resource consumption, financial status and satisfaction scores, it provides hospitals with more scientific and accurate performance evaluation methods. In particular, the collection and processing technology of multi-source data, such as obtaining real-time data through the hospital information management system (HIS), data cleaning and denoising technology, and analysis methods based on machine learning and deep learning, have become cutting-edge technologies in hospital management.
[0003] The existing hospital performance appraisal methods still have some shortcomings. Traditional methods are often limited to a single data dimension or a simple linear model and cannot fully consider the complex nonlinear relationship between various factors. Most of the existing data processing technologies ignore the noise problem of the data and the impact of missing data on the appraisal results, which may lead to certain deviations in the final appraisal results. Summary of the invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a public hospital performance appraisal data processing method based on big data, which solves the problem that traditional methods are often limited to a single data dimension or a simple linear model and cannot fully consider the complex nonlinear relationship between various factors. Most of the existing data processing technologies ignore the data noise problem and the impact of missing data on the appraisal results, which may lead to certain deviations in the final appraisal results.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a method for processing data for performance appraisal of public hospitals based on big data, which includes collecting multi-source data and preprocessing it, calculating a comprehensive score based on the preprocessed multi-source data; calculating the partial derivatives of the comprehensive score with respect to the input variables and adjusting the input variables, calculating the adjusted input variables, constructing an objective function and solving it to obtain a Pareto optimal solution set; constructing a visual interface to display the Pareto optimal solution set, and storing, collecting and analyzing the generated multi-source data.
[0008] As a preferred solution of the data processing method for public hospital performance evaluation based on big data described in the present invention, wherein: the collecting of multi-source data and pre-processing thereof refers to collecting multi-source data from a hospital management system through an API and pre-processing thereof;
[0009] The multi-source data includes patient volume, resource consumption, financial and satisfaction score data;
[0010] The preprocessing includes standardizing the multi-source data and performing denoising using wavelet transform;
[0011] Use the Gaussian kernel function to calculate the local density of data points in multi-source data, and calculate the bucket width w based on the quantile i , randomly sample from a standard normal distribution and generate a random vector a from the interval [0,w i ] to perform uniform sampling and generate a dynamic offset b;
[0012] According to the barrel width w i , random vector a and dynamic offset b, construct a hash function to calculate data point p i The hash value h(p i );
[0013] Randomly generate k independent hash functions, calculate the hash value of the data point in turn, and generate a hash signature H(p i );
[0014] Using hash bucket classification will have the same hash signature H(p i ) are grouped into the same bucket, defined as the candidate point set C(p i );
[0015] The number of nearest neighbors is set using the sparsity adjustment method, according to the candidate point set C(p i ), use the Euclidean distance method to calculate the Euclidean distance between the data point and the candidate point, sort them from large to small according to the Euclidean distance, filter out the nearest Euclidean distance according to the number of nearest neighbors, use the Gaussian kernel function to perform nonlinear transformation on the nearest Euclidean distance, calculate the Gaussian weight and G(q i ) where q i is a candidate point;
[0016] Use the reverse nearest neighbor calculation to count the number of times the candidate point is selected as the nearest neighbor, and get the reverse nearest neighbor count Nk(q i );
[0017] According to the nearest Euclidean distance, the contribution value CR(q i );
[0018] According to the calculated Gaussian weights and G(q i ), inverse neighbor count Nk(q i ) and contribution value CR(q i ), calculate the data point density D(q i ), the formula is:
[0019]
[0020] Use statistical distribution methods to set anomaly thresholds, compare data point density with the anomaly threshold, filter out data point density greater than the anomaly threshold and delete them;
[0021] Use k-nearest neighbor imputation to fill missing values.
[0022] As a preferred solution of the data processing method for public hospital performance appraisal based on big data described in the present invention, wherein: the calculation of the comprehensive score based on the pre-processed multi-source data refers to collecting historical satisfaction score data and sorting them by time to generate a time series, using the sliding average method to smooth the time series, and screening out the maximum smoothed time series, which is recorded as the saturation value B;
[0023] According to the preprocessed multi-source data, the number of patients, resource consumption, and financial data are defined as input variables X(t). Where X 1 (t), X 2 (t) and X 3 (t) are the number of patients, resource consumption and financial data, respectively. is the transpose, t is the time, and the satisfaction score is defined as the output variable Y(t).
[0024] Based on the input variable X(t) and the output variable Y(t), a nonlinear dynamic equation is constructed using the Logistic nonlinear growth model;
[0025] Use numerical integration methods to solve nonlinear dynamic equations and obtain numerical integration results;
[0026] Define the data in the input variable X(t) and the output variable Y(t) as nodes, extract the variable sequence of each time step from the numerical integration results, calculate the Pearson correlation using the time-sensitive correlation method, define the absolute value of the Pearson correlation as the element W[j,i] of the interaction matrix, where j is the index of all variables, construct the dynamic interaction matrix W, traverse each row of the dynamic interaction matrix W, extract the column index of the non-zero element, and generate each node s according to the non-zero column index. i The neighbor set N(s i );
[0027] According to the element W[j,i] of the interaction matrix, all elements are calculated and summed to obtain the node s j The weighted degree of deg(s j );
[0028] Calculate the nodes s using the iterative formula j PageRank value PR(s i );
[0029] Slave nodes j PageRank value PR(s i ) to extract the PageRank value PR(Y 1 );
[0030] According to the PageRank value PR(Y 1 ), the comprehensive score R of public hospital performance is calculated using the time integral method, the formula is:
[0031]
[0032] Where T is the time frame for performance appraisal of public hospitals.
[0033] As a preferred solution of the data processing method for public hospital performance evaluation based on big data described in the present invention, wherein: the partial derivative of the comprehensive score to the input variable is calculated and the input variable is adjusted, and the adjusted input variable is calculated by using the finite difference method to calculate the partial derivative of the comprehensive score R to the input variable X(t)
[0034] According to the partial derivative Use the sensitivity analysis method to calculate the sensitivity coefficient A of the comprehensive score to the input variable X(t) Xi ;
[0035] Use the percentile method to set the sensitivity threshold, compare the sensitivity coefficient with the sensitivity threshold, and if the sensitivity coefficient is greater than or equal to the sensitivity threshold, it is judged as a high-efficiency hospital and continues to be monitored;
[0036] If the sensitivity coefficient is less than the sensitivity threshold, it is judged as an inefficient hospital, and the proportion method is used to calculate the adjustment proportion ΔC of the input variable X(t) Xi , use the resource adjustment formula to calculate the adjusted input variable X ′ i ;
[0037] The stopping threshold is set using the asymptotic stopping threshold setting method. When the adjusted input variable is greater than the stopping threshold, the adjustment is stopped. The adjusted input variable is calculated into the comprehensive score formula to calculate the comprehensive score of the adjusted input variable.
[0038] As a preferred solution of the data processing method for public hospital performance appraisal based on big data described in the present invention, wherein: constructing the objective function and solving it to obtain the Pareto optimal solution set refers to constructing the objective function using a linear programming method based on the comprehensive score of the adjusted input variables;
[0039] Randomly generate N particles in the population, each particle represents a solution [X(t), R];
[0040] Perform Pareto sorting on all solutions in the population, identify the dominance relationship, use the non-dominated sorting method to screen, and select the Pareto optimal solution set;
[0041] Calculate the objective function value of the Pareto optimal solution set and calculate the congestion degree of the optimal solution;
[0042] Select the Pareto optimal solution with the highest congestion to enter the next generation;
[0043] The maximum number of iterations is set using the rule of thumb. When the maximum number of iterations is reached, the iteration is stopped and the Pareto optimal solution set is output, including the optimized input variables and the corresponding maximum comprehensive score.
[0044] As a preferred solution of the public hospital performance appraisal data processing method based on big data described in the present invention, wherein: the construction of a visual interface to display the Pareto optimal solution set refers to using the front-end framework React.js to construct a visual interface, including a main chart area and a top information bar;
[0045] The optimized input variables are displayed in the main chart area, and the optimized input variables and the corresponding maximum comprehensive scores are displayed in the top information bar;
[0046] Users who have passed real-name verification are allowed to view the information.
[0047] As a preferred solution of the public hospital performance appraisal data processing method based on big data described in the present invention, the storage of multi-source data generated by collection and analysis refers to storing the collected multi-source data and the Pareto optimal solution set generated by analysis in a central database, and setting security access measures. The central database backs up the stored data in the cloud, and regularly performs integrity checks on the stored data and backup data. After the test is completed, an integrity test record is generated and synchronously stored in the central database.
[0048] In a second aspect, the present invention provides a public hospital performance evaluation data processing system based on big data, comprising:
[0049] A collection and calculation module is used to collect and preprocess multi-source data, and calculate a comprehensive score based on the preprocessed multi-source data;
[0050] The adjustment solution module is used to calculate the partial derivative of the comprehensive score with respect to the input variable and adjust the input variable, calculate the adjusted input variable, construct the objective function and solve it to obtain the Pareto optimal solution set;
[0051] The visualization storage module is used to build a visualization interface to display the Pareto optimal solution set and store the multi-source data generated by collection and analysis.
[0052] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the public hospital performance appraisal data processing method based on big data as described in the first aspect of the present invention is implemented.
[0053] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the public hospital performance appraisal data processing method based on big data as described in the first aspect of the present invention.
[0054] The beneficial effects of the present invention are as follows: the present invention collects multi-source data and performs preprocessing, calculates a comprehensive score based on the preprocessed multi-source data; calculates the partial derivatives of the comprehensive score with respect to the input variables and adjusts the input variables, calculates the adjusted input variables, constructs the objective function and solves it to obtain the Pareto optimal solution set; solves the data quality problems of the prior art, improves the accuracy and real-time performance of public hospital performance appraisal, and enhances the intelligence and adaptability of data processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0056] Figure 1 This is a flow chart of the public hospital performance appraisal data processing method based on big data in Example 1.
[0057] Figure 2 This is a schematic diagram of the public hospital performance appraisal data processing system based on big data in Example 1. DETAILED DESCRIPTION
[0058] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0059] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0060] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0061] Example 1, reference Figure 1 and Figure 2 , which is the first embodiment of the present invention, and provides a method for processing public hospital performance evaluation data based on big data, comprising the following steps:
[0062] S1, collect and preprocess multi-source data, and calculate the comprehensive score based on the preprocessed multi-source data;
[0063] Specifically, collecting multi-source data and pre-processing it means collecting multi-source data from a hospital management system through an API and pre-processing it;
[0064] The multi-source data includes patient volume, resource consumption, financial and satisfaction score data;
[0065] The preprocessing includes standardizing the multi-source data and performing denoising using wavelet transform;
[0066] Use the Gaussian kernel function to calculate the local density of data points in multi-source data, and calculate the bucket width w based on the quantile i , randomly sample from a standard normal distribution and generate a random vector a from the interval [0,w i ] to perform uniform sampling and generate a dynamic offset b;
[0067] According to the barrel width w i , random vector a and dynamic offset b, construct a hash function to calculate data point p i The hash value h(p i ), the formula is:
[0068]
[0069] Randomly generate k independent hash functions, calculate the hash value of the data point in turn, and generate a hash signature H(p i ), the formula is:
[0070] H(p i )=[h 1 (p i ),h 2 (p i ),…,h k (p i )],
[0071] Using hash bucket classification will have the same hash signature H(p i ) are grouped into the same bucket, defined as the candidate point set C(p i ), the formula is:
[0072] C(p i )={q i │H(p i )=H(q i )},
[0073] where q i is a candidate point;
[0074] The number of nearest neighbors is set using the sparsity adjustment method, according to the candidate point set C(p i ), use the Euclidean distance method to calculate the Euclidean distance between the data point and the candidate point, sort them from large to small according to the Euclidean distance, filter out the nearest Euclidean distance according to the number of nearest neighbors, use the Gaussian kernel function to perform nonlinear transformation on the nearest Euclidean distance, calculate the Gaussian weight and G(q i );
[0075] Use the reverse nearest neighbor calculation to count the number of times the candidate point is selected as the nearest neighbor, and get the reverse nearest neighbor count Nk(q i );
[0076] According to the nearest Euclidean distance, the contribution value CR(q i ), the formula is:
[0077]
[0078] Where o is the number of nearest neighbors, i is the index of the data point, and d(q i ,N i ) is the closest Euclidean distance, N i is the nearest neighbor point, ∈ is a small positive number;
[0079] According to the calculated Gaussian weights and G(q i ), inverse neighbor count Nk(q i ) and contribution value CR(q i ), calculate the data point density D(q i ), the formula is:
[0080]
[0081] Use statistical distribution methods to set anomaly thresholds, compare data point density with the anomaly threshold, filter out data point density greater than the anomaly threshold and delete them;
[0082] Gaussian weights and are introduced to describe local density, and inverse neighbor counts are introduced to describe global contribution. Finally, the two are combined through a formula to balance local and global characteristics. Traditional density calculation methods usually only rely on local neighbor points and cannot effectively capture the importance of data points in the global distribution. This formula supplements the global perspective through inverse neighbor counts, improves the comprehensiveness and accuracy of the results, and dynamically adjusts the contribution ratio of each data point to the density through the contribution value, so that the formula can adapt to changes in different density areas, avoiding the underestimation of high-density areas and overestimation of low-density areas in traditional methods. The introduction of dynamic contribution values enhances the adaptive ability of density estimation, making the formula more robust when dealing with outliers and sparse areas. The introduction of logarithmic transformation makes the density estimation result smoother and avoids the influence of extreme values on the result. Compared with the traditional direct summation or product method, the use of logarithmic function significantly improves the stability of the formula. Especially in the scenario of processing data with a large distribution span, the numerator and denominator calculate the Gaussian weight and the inverse neighbor count respectively, and compare and balance them in combination with the contribution value, avoiding the problem of a single feature dominating the density estimation result. Through the design of the density formula and the subsequent statistical distribution method, the threshold of the outlier can be set more accurately. Compared with the traditional density threshold setting method (such as simple mean or variance), this method can adaptively set the outlier screening criteria according to the dynamic characteristics of the formula calculation results, thereby significantly improving the accuracy of outlier detection;
[0083] Use k-nearest neighbor imputation to fill missing values.
[0084] By collecting multi-source data from the hospital management system through API and preprocessing it, important indicators such as the number of patients, resource consumption, financial data and satisfaction scores can be obtained in real time. In the process of data preprocessing, wavelet transform is used for denoising, which can effectively remove high-frequency noise in the data, making the data smoother and more reliable, thereby improving the accuracy and effectiveness of subsequent analysis. The hash value of the data point is generated by an accurate hash function, and hash bucket classification is performed to cluster data points with similar characteristics. Similar data point sets can be quickly identified in large data sets, thereby improving efficiency and reducing computational complexity during data clustering and classification. Through nonlinear adjustment of distance, the similarity of data points can be measured more accurately, which is especially suitable for processing data with nonlinear characteristics. Through density calculation methods, the relative importance of each data point in the overall data can be quantified, thereby improving the precision and reliability of data processing. The k-nearest neighbor interpolation method is used to fill in missing values to ensure data integrity and reduce the computational bottleneck of traditional methods in the case of large data volume and high dimension.
[0085] Furthermore, calculating the comprehensive score based on the preprocessed multi-source data means collecting the historical satisfaction score data and sorting them by time to generate a time series, smoothing the time series using the sliding average method, and selecting the maximum smoothed time series, which is recorded as the saturation value B;
[0086] According to the preprocessed multi-source data, the number of patients, resource consumption, and financial data are defined as input variables X(t). Where X 1 (t), X 2 (t) and X 3 (t) are the number of patients, resource consumption and financial data, respectively. is the transpose, t is the time, and the satisfaction score is defined as the output variable Y(t).
[0087] According to the input variable X(t) and the output variable Y(t), the nonlinear dynamic equation is constructed using the Logistic nonlinear growth model. The formula is:
[0088]
[0089] Where j is the index of all variables;
[0090] Traditional performance appraisal methods (such as simple weighted models or linear regression models) usually assume that each indicator is independent of each other and cannot capture the dynamic relationship between input and output. The global weight of the satisfaction score is quantified by PageRank, so that the scoring result can reflect the influence of nodes in the complex interactive network. The iterative calculation method of PageRank can capture the hidden system characteristics in the dynamic interaction matrix, thereby improving the scientificity and accuracy of the scoring. Traditional methods usually perform performance appraisal based on static data (such as the score or resource input at a certain point in time), ignoring the cumulative effect in the time dimension. The time integration method can fully reflect the cumulative effect of indicators within the time range, avoiding the limitations of static analysis. The numerator and denominator respectively perform time integration on the satisfaction score and resource consumption, making the performance scoring result more fair and credible. The traditional scoring method directly linearly weights the input and output, ignoring the efficiency relationship between the two. The efficiency of the satisfaction score and resource consumption is directly quantified in the form of a ratio, which can more intuitively evaluate the resource utilization efficiency of the hospital. Compared with the existing technology, this formula is more scientific and reasonable, and can provide reliable data support for hospital management and decision-making.
[0091] Use numerical integration methods to solve nonlinear dynamic equations and obtain numerical integration results;
[0092] Define the data in the input variable X(t) and the output variable Y(t) as nodes, extract the variable sequence of each time step from the numerical integration result, calculate the Pearson correlation using the time-sensitive correlation method, define the absolute value of the Pearson correlation as the element W[j,i] of the interaction matrix, construct the dynamic interaction matrix W, traverse each row of the dynamic interaction matrix W, extract the column index of the non-zero element, and generate each node s according to the non-zero column index. i The neighbor set N(s i );
[0093] According to the element W[j,i] of the interaction matrix, all elements are calculated and summed to obtain the node s j The weighted degree of deg(s j );
[0094] Calculate the nodes s using the iterative formula j PageRank value PR(s i ), the formula is:
[0095]
[0096] Where d is the damping factor;
[0097] Slave nodes j PageRank value PR(s i) to extract the PageRank value PR(Y 1 );
[0098] According to the PageRank value PR(Y 1 ), the comprehensive score R of public hospital performance is calculated using the time integral method, the formula is:
[0099]
[0100] Where T is the time frame for performance appraisal of public hospitals.
[0101] By sorting the historical satisfaction rating data and using the sliding average method to smooth the time series, short-term fluctuations and outliers can be effectively eliminated, and the stability and reliability of the data can be enhanced. By screening out the maximum smoothed time series (saturation value B), the long-term trend of hospital performance can be better captured, avoiding the interference of a single fluctuation on the final result. The Logistic nonlinear growth model can flexibly reflect the nonlinear dynamic changes of hospital performance. By constructing this model, the change pattern of hospital performance over time can be more accurately described, rather than simply relying on a linear model. This makes the present invention more in line with actual conditions when predicting and evaluating hospital performance, avoiding overly simplified analysis methods. By numerically integrating the Logistic nonlinear dynamic equation, the variable value of each time step can be obtained, providing accurate data support for subsequent analysis. The Pearson correlation is calculated by the sensitive correlation method, and a dynamic interaction matrix is constructed, which can reveal the complex interactive relationship between input variables (such as the number of patients, resource consumption, and financial data) and output variables (satisfaction scores). The weighted calculation of nodes (variables) by the PageRank algorithm can evaluate the relative contribution of each input variable to hospital performance (satisfaction score). This not only optimizes the accuracy of the performance evaluation model, but also improves the dynamic adaptability of the evaluation process. The importance of each variable can be adjusted according to actual data. The comprehensive score of hospital performance is calculated by the time integration method, which can comprehensively consider the impact of variables at each time step on performance and provide quantitative indicators of long-term performance for hospital managers. It can not only help hospitals understand their own performance status, but also provide in-depth decision-making support for management, and promote continuous improvement of hospitals in resource allocation, service optimization, etc.
[0102] S2, calculate the partial derivative of the comprehensive score with respect to the input variable and adjust the input variable, calculate the adjusted input variable, construct the objective function and solve it to obtain the Pareto optimal solution set;
[0103] Specifically, the partial derivative of the comprehensive score with respect to the input variable is calculated and the input variable is adjusted. Calculating the adjusted input variable means using the finite difference method to calculate the partial derivative of the comprehensive score R with respect to the input variable X(t) The formula is:
[0104]
[0105] According to the partial derivative Use sensitivity analysis method to calculate the sensitivity coefficient of comprehensive score to input variable X(t) The formula is:
[0106]
[0107] Use the percentile method to set the sensitivity threshold, compare the sensitivity coefficient with the sensitivity threshold, and if the sensitivity coefficient is greater than or equal to the sensitivity threshold, it is judged as a high-efficiency hospital and continues to be monitored;
[0108] If the sensitivity coefficient is less than the sensitivity threshold, it is judged as an inefficient hospital, and the proportion method is used to calculate the adjustment proportion of the input variable X(t) Use the resource adjustment formula to calculate the adjusted input variable X ′ i , the formula is:
[0109]
[0110] The stopping threshold is set using the asymptotic stopping threshold setting method. When the adjusted input variable is greater than the stopping threshold, the adjustment is stopped. The adjusted input variable is calculated into the comprehensive score formula to calculate the comprehensive score of the adjusted input variable.
[0111] The finite difference method is used to calculate the partial derivatives of the comprehensive score with respect to the input variables, so that the hospital can identify which variables have a greater impact on performance, thereby providing a solid foundation for subsequent sensitivity analysis. Sensitivity analysis helps hospitals understand the specific impact of each input variable on performance by calculating the sensitivity coefficient of the comprehensive score to the input variables. Sensitivity analysis can not only reveal the key factors in hospital operations, but also provide decision makers with specific directions for optimizing resources and improve the accuracy and efficiency of hospital performance management. By setting the sensitivity threshold by the percentile method, hospitals can be divided into two categories: efficient and inefficient. For inefficient hospitals, the adjustment ratio of the input variables is calculated by the proportion method, thereby providing them with a quantitative improvement plan. The adjusted input variables will help improve the performance of the hospital and thus improve the overall operation level of the hospital. By using the stopping threshold method, the performance adjustment of the hospital can be carried out within a reasonable range, avoiding redundancy and repeated calculations, and improving calculation efficiency and implementation effects. Hospital management can formulate more scientific resource allocation strategies and operation optimization plans based on the adjusted comprehensive score to promote continuous improvement and innovation of the hospital.
[0112] Furthermore, the objective function is constructed and solved to obtain the Pareto optimal solution set, which refers to the objective function constructed using a linear programming method based on the comprehensive scores of the adjusted input variables;
[0113] Randomly generate N particles in the population, each particle represents a solution [X(t), R];
[0114] Perform Pareto sorting on all solutions in the population, identify the dominance relationship, use the non-dominated sorting method to screen, and select the Pareto optimal solution set;
[0115] Calculate the objective function value of the Pareto optimal solution set and calculate the congestion degree of the optimal solution;
[0116] Select the Pareto optimal solution with the highest congestion to enter the next generation;
[0117] The maximum number of iterations is set using the rule of thumb. When the maximum number of iterations is reached, the iteration is stopped and the Pareto optimal solution set is output, including the optimized input variables and the corresponding maximum comprehensive score.
[0118] By constructing and solving the objective function through the linear programming method, the various input variables of the hospital can be optimized under the predetermined constraints, so that hospital management decisions can be based on mathematical optimization, avoiding the limitations of traditional empirical management and improving the scientificity and systematicness of hospital management. Through Pareto sorting, solutions with clear dominance relationships can be screened out, ensuring that the final selected solution set is the solution with the best performance for all objectives, enhancing the comprehensiveness and accuracy of the solution set, and making the optimization results more representative. By selecting non-dominated solutions, it can ensure that the solution set in the optimization process has good globality and avoid the solution space being confined to a specific area. Non-dominated sorting can effectively handle the complex relationships in multi-objective optimization and avoid the local optimal solution problem that may occur in traditional optimization methods. Through the optimized input variables, the hospital can more comprehensively optimize resource allocation and management strategies and improve overall operational efficiency. By limiting the number of iterations, it can avoid the algorithm from continuing to calculate when there is no significant progress, saving computing resources and improving the execution efficiency of the algorithm. According to different solutions, the optimization plan that best meets actual needs is selected, thereby improving the overall operational efficiency and service quality of the hospital.
[0119] S3, build a visual interface to display the Pareto optimal solution set, store the multi-source data generated by collection and analysis;
[0120] Specifically, building a visual interface to display the Pareto optimal solution set refers to using the front-end framework React.js to build a visual interface, including the main chart area and the top information bar;
[0121] The optimized input variables are displayed in the main chart area, and the optimized input variables and the corresponding maximum comprehensive scores are displayed in the top information bar;
[0122] Users who have passed real-name verification are allowed to view the information.
[0123] Through React.js, developers can efficiently handle user input, data updates and view rendering, ensuring that users can view changes in optimization results in real time. The interface is highly responsive and maintainable, especially when displaying complex optimization results. It can ensure the smoothness and stability of the interface. The visual display makes complex data more intuitive. Users can easily understand the relationship between the changes in each input variable and the comprehensive score of the hospital, helping decision makers to make better optimization decisions. The introduction of real-name verification not only ensures the privacy of the data, but also enhances the compliance of the system, avoiding unauthorized users from accessing sensitive information. By authenticating users with real names, the risk of data leakage can be effectively reduced, ensuring that the optimization results can only be viewed and used by relevant authorized personnel, thereby enhancing the security and trust of the system.
[0124] Furthermore, storing the multi-source data generated by collection and analysis means storing the collected multi-source data and the Pareto optimal solution set generated by the analysis in a central database, and setting up security access measures. The central database will back up the stored data in the cloud, and regularly perform integrity checks on the stored data and backup data. After the test is completed, an integrity test record will be generated and stored synchronously in the central database.
[0125] Through this centralized management, hospitals can easily query, analyze and make decisions on data, avoiding the management inconvenience caused by decentralized data storage. The implementation of secure access measures ensures that only authenticated users can access sensitive data. By ensuring the security of data, data leakage, tampering or loss can be avoided, the hospital's data assets can be maintained, and the trust of hospital users in the system can be enhanced. By regularly backing up data to the cloud, hospitals can ensure the safe storage of data and quickly restore data in the event of a disaster, minimizing the impact of data loss on hospital operations. Data integrity testing can promptly detect potential data problems and maintain data accuracy and consistency through repair measures. By centrally storing multi-source data and Pareto optimal solution sets, hospitals can manage data more efficiently and avoid the query inconvenience that may be caused by storing data in multiple places. Cloud backup and integrity testing of data jointly improve the robustness of the system and avoid potential threats to hospital operations caused by data loss or damage.
[0126] This embodiment also provides a public hospital performance evaluation data processing system based on big data, including:
[0127] A collection and calculation module is used to collect and preprocess multi-source data, and calculate a comprehensive score based on the preprocessed multi-source data;
[0128] The adjustment solution module is used to calculate the partial derivative of the comprehensive score with respect to the input variable and adjust the input variable, calculate the adjusted input variable, construct the objective function and solve it to obtain the Pareto optimal solution set;
[0129] The visualization storage module is used to build a visualization interface to display the Pareto optimal solution set and store the multi-source data generated by collection and analysis.
[0130] This embodiment also provides a computer device, which is suitable for the case of a public hospital performance appraisal data processing method based on big data, including: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute computer executable instructions to implement the public hospital performance appraisal data processing method based on big data proposed in the above embodiment.
[0131] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.
[0132] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the method for processing public hospital performance appraisal data based on big data as proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, SRAM for short), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM for short), erasable programmable read-only memory (Erasable Programmable Read Only Memory, EPROM for short), programmable read-only memory (Programmable Red-Only Memory, PROM for short), read-only memory (Read-Only Memory, ROM for short), magnetic storage, flash memory, disk or optical disk.
[0133] In summary, the present invention collects multi-source data and performs preprocessing, calculates a comprehensive score based on the preprocessed multi-source data; calculates the partial derivatives of the comprehensive score with respect to the input variables and adjusts the input variables, calculates the adjusted input variables, constructs the objective function and solves it to obtain the Pareto optimal solution set; solves the data quality problems of the prior art, improves the accuracy and real-time performance of public hospital performance appraisal, and enhances the intelligence and adaptability of data processing.
[0134] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for processing public hospital performance appraisal data based on big data, characterized by: include, Collect and preprocess multi-source data, and calculate the comprehensive score based on the preprocessed multi-source data; Calculate the partial derivative of the comprehensive score with respect to the input variable and adjust the input variable, calculate the adjusted input variable, construct the objective function and solve it to obtain the Pareto optimal solution set; Build a visual interface to display the Pareto optimal solution set and store the multi-source data generated by collection and analysis.
2. The method for processing public hospital performance evaluation data based on big data as claimed in claim 1, characterized in that: The collecting of multi-source data and pre-processing thereof refers to collecting multi-source data from a hospital management system through an API and pre-processing thereof; The multi-source data includes patient volume, resource consumption, financial and satisfaction score data; The preprocessing includes standardizing the multi-source data and performing denoising using wavelet transform; Use the Gaussian kernel function to calculate the local density of data points in multi-source data, and calculate the bucket width w based on the quantile i , randomly sample from a standard normal distribution and generate a random vector a from the interval [0,w i ] to perform uniform sampling and generate a dynamic offset b; According to the barrel width w i , random vector a and dynamic offset b, construct a hash function to calculate data point p i The hash value h(p i ); Randomly generate k independent hash functions, calculate the hash value of the data point in turn, and generate a hash signature H(p i ); Using hash bucket classification will have the same hash signature H(p i ) are grouped into the same bucket, defined as the candidate point set C(p i ); The number of nearest neighbors is set using the sparsity adjustment method, according to the candidate point set C(p i ), use the Euclidean distance method to calculate the Euclidean distance between the data point and the candidate point, sort them from large to small according to the Euclidean distance, filter out the nearest Euclidean distance according to the number of nearest neighbors, use the Gaussian kernel function to perform nonlinear transformation on the nearest Euclidean distance, calculate the Gaussian weight and G(q i ) where q i is a candidate point; Use the reverse nearest neighbor calculation to count the number of times the candidate point is selected as the nearest neighbor, and get the reverse nearest neighbor count Nk(q i ); According to the nearest Euclidean distance, the contribution value CR(q i ); According to the calculated Gaussian weights and G(q i ), inverse neighbor count Nk(q i ) and contribution value CR(q i ), calculate the data point density D(q i ), the formula is: Use statistical distribution methods to set anomaly thresholds, compare data point density with the anomaly threshold, filter out data point density greater than the anomaly threshold and delete them; Use k-nearest neighbor imputation to fill missing values.
3. The method for processing public hospital performance evaluation data based on big data as claimed in claim 2, characterized in that: The calculation of the comprehensive score based on the pre-processed multi-source data refers to collecting historical satisfaction score data and sorting them by time to generate a time series, smoothing the time series using a sliding average method, and screening out the maximum smoothed time series, which is recorded as the saturation value B; According to the preprocessed multi-source data, the number of patients, resource consumption and financial data are defined as input variables X(t), X(t) = [X1(t), X2(t), X3(t)] T , where X1(t), X2(t) and X3(t) are the number of patients, resource consumption and financial data respectively, T is the transpose, t is the time, and the satisfaction score is defined as the output variable Y(t), Y(t) = [Y1(t)] T ; Based on the input variable X(t) and the output variable Y(t), a nonlinear dynamic equation is constructed using the Logistic nonlinear growth model; Use numerical integration methods to solve nonlinear dynamic equations and obtain numerical integration results; Define the data in the input variable X(t) and the output variable Y(t) as nodes, extract the variable sequence of each time step from the numerical integration results, calculate the Pearson correlation using the time-sensitive correlation method, define the absolute value of the Pearson correlation as the element W[j,i] of the interaction matrix, where j is the index of all variables, construct the dynamic interaction matrix W, traverse each row of the dynamic interaction matrix W, extract the column index of the non-zero element, and generate each node s according to the non-zero column index. i The neighbor set N(s i ); According to the element W[j,i] of the interaction matrix, all elements are calculated and summed to obtain the node s j The weighted degree of deg(s j ); Calculate the nodes s using the iterative formula j PageRank value PR(s i ); Slave nodes j PageRank value PR(s i ) extracts the PageRank value PR(Y1) of the satisfaction score; According to the PageRank value PR(Y1) of the satisfaction score, the comprehensive score R of public hospital performance is calculated using the time integration method. The formula is: Where T is the time frame for performance appraisal of public hospitals.
4. The method for processing data for public hospital performance evaluation based on big data as claimed in claim 3, characterized in that: The method of calculating the partial derivative of the comprehensive score with respect to the input variable and adjusting the input variable, and calculating the adjusted input variable means using the finite difference method to calculate the partial derivative of the comprehensive score R with respect to the input variable X(t) According to the partial derivative Use the sensitivity analysis method to calculate the sensitivity coefficient A of the comprehensive score to the input variable X(t) Xi ; Use the percentile method to set the sensitivity threshold, compare the sensitivity coefficient with the sensitivity threshold, and if the sensitivity coefficient is greater than or equal to the sensitivity threshold, it is judged as a high-efficiency hospital and continues to be monitored; If the sensitivity coefficient is less than the sensitivity threshold, it is judged as an inefficient hospital, and the proportion method is used to calculate the adjustment proportion ΔC of the input variable X(t) Xi , use the resource adjustment formula to calculate the adjusted input variable X ′ i ; The stopping threshold is set using the asymptotic stopping threshold setting method. When the adjusted input variable is greater than the stopping threshold, the adjustment is stopped. The adjusted input variable is calculated into the comprehensive score formula to calculate the comprehensive score of the adjusted input variable.
5. The method for processing data of public hospital performance evaluation based on big data according to claim 4, characterized in that: The constructing the objective function and solving it to obtain the Pareto optimal solution set refers to constructing the objective function using a linear programming method based on the comprehensive scores of the adjusted input variables; Randomly generate N particles in the population, each particle represents a solution [X(t), R]; Perform Pareto sorting on all solutions in the population, identify the dominance relationship, use the non-dominated sorting method to screen, and select the Pareto optimal solution set; Calculate the objective function value of the Pareto optimal solution set and calculate the congestion degree of the optimal solution; Select the Pareto optimal solution with the highest congestion to enter the next generation; The maximum number of iterations is set using the rule of thumb. When the maximum number of iterations is reached, the iteration is stopped and the Pareto optimal solution set is output, including the optimized input variables and the corresponding maximum comprehensive score.
6. The method for processing data for public hospital performance evaluation based on big data according to claim 5, characterized in that: The said constructing a visualization interface to display the Pareto optimal solution set refers to using the front-end framework React.js to construct a visualization interface, including a main chart area and a top information bar; The optimized input variables are displayed in the main chart area, and the optimized input variables and the corresponding maximum comprehensive scores are displayed in the top information bar; Users who have passed real-name verification are allowed to view the information.
7. The method for processing data for public hospital performance evaluation based on big data according to claim 6, characterized in that: The storage of multi-source data generated by collection and analysis refers to storing the collected multi-source data and the Pareto optimal solution set generated by the analysis in a central database, and setting security access measures. The central database will back up the stored data in the cloud, and regularly perform integrity checks on the stored data and backup data. After the test is completed, an integrity test record will be generated and stored synchronously in the central database.
8. A public hospital performance appraisal data processing system based on big data, based on the public hospital performance appraisal data processing method based on big data according to any one of claims 1 to 7, characterized in that: include, A collection and calculation module is used to collect and preprocess multi-source data, and calculate a comprehensive score based on the preprocessed multi-source data; The adjustment solution module is used to calculate the partial derivative of the comprehensive score with respect to the input variable and adjust the input variable, calculate the adjusted input variable, construct the objective function and solve it to obtain the Pareto optimal solution set; The visualization storage module is used to build a visualization interface to display the Pareto optimal solution set and store the multi-source data generated by collection and analysis.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the public hospital performance evaluation data processing method based on big data described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the public hospital performance evaluation data processing method based on big data described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Medical performance assessment analysis management method and system
CN122048123A