Patient Condition Assessment Method and System Based on Multi-Source Information Fusion Mechanism

Through the multi-source information fusion mechanism, combined with principal component analysis, orderly clustering analysis and D-S evidence theory, the problem of inability to reflect the impact of seasonal changes and subjective judgment in the existing technology is solved, and a more comprehensive hospital reputation assessment and a more accurate basis for patient selection is achieved.

CN119740934BActive Publication Date: 2025-06-20GENERAL HOSPITAL OF PLA
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510260712.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-20
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

The prior art cannot effectively reflect the seasonal changes in physiological functions or disease spectrum, cannot quantify data on various characteristics, and relies on the subjective judgment of experts or managers to set weights, resulting in insufficient objectivity and accuracy of the evaluation results.

Method used

The patient's condition assessment method based on the multi-source information fusion mechanism is adopted to evaluate the patient's underlying disease through principal component analysis, and the patient's mortality risk is evaluated using orderly cluster analysis. Combined with D-S evidence theory and entropy weight method, a comprehensive consideration factor is formed to guide patients to choose a hospital or physician.

Benefits of technology

The reflection of seasonal changes in physiological functions or disease spectrum is achieved, the data on each characteristic is quantified, the impact of subjective judgment is reduced, the objectivity and accuracy of the evaluation results are improved, and a more comprehensive hospital reputation assessment is provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119740934B_ABST
    Figure CN119740934B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for patient condition assessment based on a multi-source information fusion mechanism; the present invention relates to the technical field of medical service management; a data matrix dm is constructed from the underlying disease index data set D, where the rows represent different historical patients and the columns represent different indices; based on the data matrix d m perform principal component analysis (PCA) to obtain an evaluation factor A; calculate the death rate curve D of historical patients vc , and use the ordered clustering algorithm to cluster patients according to the historical medical record data D' and the death rate curve D vc ; make full use of data analysis techniques, including principal component analysis, ordered clustering analysis, and D-S evidence theory, etc., to organically combine quantitative data from different sources and of different natures to form a comprehensive consideration factor. This data-driven method not only improves the accuracy and efficiency of assessment, but also provides a powerful decision support tool for hospital managers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical service management, specifically to the technical direction of incorporating patients' underlying diseases and mortality rates into the consideration scope of medical reputation, and particularly to a method and system for patient condition assessment based on a multi-source information fusion mechanism. Background Art

[0002] The reputation of a hospital is similar to the corporate credit and is an important intangible asset, which is of great significance for a hospital to attract patients and enhance its competitiveness. The evaluation of reputation is not only a measure of the hospital's external image, but also a comprehensive reflection of various factors such as its internal management, service quality, and technical level.

[0003] Hospital reputation management faces many challenges, such as the asymmetry of medical information, the diversification of patient needs, etc.; for the problem of quantifying hospital reputation, many existing technologies attempt to solve it and further improve the objectivity and accuracy of reputation evaluation. For example:

[0004] (1) CN202110489756.7 discloses a reputation risk management ability evaluation method (publication date: 2024-08-20), which combines business values with weights to obtain the scoring information of each evaluation index, and organizes the scoring information to obtain the evaluation information corresponding to each evaluation dimension. At the same time, CN201310343622.X discloses a service trustworthiness evaluation method based on service quality and reputation (publication date: 2016-07-06), which, based on the credibility problem of service selection, integrates service trust information from multiple sources to implement a reputation evaluation mechanism;

[0005] However, the above existing technologies cannot reflect the seasonal changes of physiological functions or disease spectra, cannot quantify the data of each feature, cannot provide a more comprehensive evaluation and reflect the reputation of the hospital, and combining business values with weights often relies on the subjective judgment of experts or managers to set weights, lacking objectivity.

[0006] (2) CN202110185476.7 discloses a reputation-based V2X node message credibility evaluation method (publication date: 2022-08-26), which comprehensively determines the credibility of the response message based on the time difference between the time information related to the request content and the time information when responding to the request message, the distance between the location information related to the request content and the location information when responding to the request message, and the reputation value of the message response end; at the same time, CN202310883674.X discloses an evaluation method based on big data analysis (publication date: 2023-08-29), which collects information data related to the target group based on big data, extracts features from positive data, negative data, and neutral data, extracts feature keywords and several themes, and mines and then implements a reputation evaluation mechanism;

[0007] However, the above-mentioned existing technologies mainly rely on the time difference between the information related to the request content and the response time, the distance between the location information, and the reputation value of the message response end to determine the credibility of the response message. This evaluation method ignores the performance of the hospital in multiple aspects such as patient treatment, service quality, and medical technology, as well as the death factors caused by the patient's own reasons. The evaluation dimension and accuracy are relatively limited. And although the traditional method collects information data related to the target based on big data, the utilization of the data is not sufficient. Because it mainly focuses on the feature extraction of positive data, negative data, and neutral data.

[0008] (5) CN202210141263.9 discloses a reputation risk monitoring and quantitative evaluation method (publication date: 2022-06-24), which performs feature matching on the to-be-processed bad feature information to obtain the matched to-be-processed bad feature information; performs risk monitoring on the matched to-be-processed bad feature information to obtain corresponding early warning information.

[0009] However, the above-mentioned existing technologies are relatively static and difficult to adjust the evaluation model with the change of the hospital environment. When comprehensively determining the credibility of the response message, there is a lack of a comprehensive evaluation mechanism, which cannot evaluate the combination of different factor weights and quantify the data of each feature, and thus cannot obtain a more comprehensive reputation evaluation result.

[0010] Therefore, the present invention proposes a patient condition evaluation method and system based on a multi-source information fusion mechanism. Summary of the Invention

[0011] In view of this, the present invention hopes to provide a patient condition evaluation method and system based on a multi-source information fusion mechanism to solve or alleviate the technical problems existing in the prior art, that is:

[0012] (1) How to reflect the seasonal changes of physiological functions or disease spectra and quantify the data of each feature;

[0013] (2) How to avoid the subjectivity problem of setting weights by relying on the subjective judgment of experts or managers;

[0014] (3) How to comprehensively consider different quantified data based on a data-driven method, better adapt to the change of the hospital environment, and quantify the data of each feature, so as to obtain a more comprehensive reputation evaluation result.

[0015] The technical solution of the present invention is realized as follows:

[0016] In the first aspect, a patient condition evaluation method based on a multi-source information fusion mechanism:

[0017] (I) Overview:

[0018] The present invention aims to construct a comprehensive hospital reputation evaluation system by comprehensively evaluating various factors such as the severity of patients' underlying diseases, the risk of mortality, and the reputation of the hospital. The principal component analysis (PCA) is used to evaluate patients' underlying diseases, and the ordered clustering analysis is used to analyze the risk of patients' mortality. At the same time, a number of characteristic data of the hospital are collected and quantified, and the D-S evidence theory is combined with the treatment factors and the comprehensive reputation eigenvalue to form a comprehensive consideration factor, so as to guide patients to reasonably select hospitals or physicians. It aims to provide a more accurate and comprehensive basis for patients to choose medical services through scientific data analysis and comprehensive evaluation methods.

[0019] (2) Technical solution:

[0020] To achieve the above technical objectives, when patients call the medical reputation of a hospital / physician, after collecting the dataset D of underlying disease indicators of historical patients from the hospital information system, including basic information (such as age, gender, and BMI index), past medical history, and physical examination data, the following steps are executed.

[0021] 2.1 Step S1, data collection:

[0022] Construct the dataset D of underlying disease indicators into a data matrix d with rows representing different historical patients and columns representing different indicators m .

[0023] 2.1.1 Step S100, data cleaning:

[0024] Fill in the missing values in the dataset D of underlying disease indicators, delete or correct the outliers. Then, use the Z-score standardization method to standardize all indicators in the dataset D of underlying disease indicators to ensure that all variables have the same mean and variance.

[0025] (1) Filling in missing values: ;

[0026] where x missing is the missing value, X i is the non-missing value, and n is the number of non-missing values;

[0027] (2) Outlier processing (Z-score method): ;

[0028] where, x i is the i-th original value in the dataset D of underlying disease indicators, μ a is the mean of the current indicator for all non-outlier patients, and σ a is the standard deviation of the current indicator for all non-outlier patients. Data points with |Z| > 3 are regarded as outliers;

[0029] (3) Z-score standardization: ;

[0030] Among them, x i ′ is the i-th element after standardization in the underlying disease index dataset D, and x i is its original value, and μ d is the mean value of the current index for all patients with non-missing values; σ d is the standard deviation of the current index for all patients with non-missing values.

[0031] 2.1.2 Step S101, construct the data matrix d m :

[0032] ;

[0033] Among them, x ij ′ represents the j-th index (standardized value) of the i-th historical patient, n is the number of patients, and m is the number of indices.

[0034] 2.2 Step S2, patient underlying disease assessment:

[0035] Based on the data matrix d m perform principal component analysis (PCA) to obtain the evaluation factor A.

[0036] 2.2.1 Step S200, calculate the covariance matrix:

[0037] Convert the data matrix d m into the covariance matrix C that describes the strength and direction of the linear relationship between each index, and then perform eigenvalue decomposition to obtain the eigenvalues E (expressed in matrix form) and the corresponding eigenvectors F (expressed in diagonal matrix form):

[0038] (1) Convert to the covariance matrix C: ;

[0039] Among them, represents the mean vector of the data matrix d m , represents the i-th row in the data matrix d m , that is, the index data of the i-th patient (a 1*m vector, where m is the number of index j), and n is the number of patients; T represents the transpose operation of the matrix.

[0040] (2) Eigenvalue decomposition: C = F * E * F T ;

[0041] Among them, F is the eigenvector matrix, E is the diagonal matrix containing the eigenvalues, and T represents the transpose operation of the matrix.

[0042] 2.2.2 Step S201, select the principal components:

[0043] Select the principal component P that makes the cumulative contribution rate reach the preset threshold according to the magnitude of the eigenvalue E; then project the original data into the new feature space F formed by the selected principal component P n to calculate the new coordinates P of each data point c (i.e., the principal component scores); normalize the new coordinates P c and perform weighted summation according to the contribution rate to obtain the evaluation factor A:

[0044] (1) Assume that selecting k principal components can make the cumulative contribution rate R C reach the preset threshold;

[0045] (2) Calculate the cumulative contribution rate R C : ;

[0046] where E jj represents the j-th diagonal element of the diagonal matrix E; m is the number of indices j;

[0047] (3) Project the original data into the new feature space F formed by the selected principal components n : P c = d m * F k ;

[0048] where F k is the matrix formed by the eigenvectors corresponding to the selected k principal components; P c are the principal component scores;

[0049] (4) Normalize the principal component scores P c : ;

[0050] where P cn represents the value of the j-th principal component score after normalization;

[0051] (5) Perform weighted summation according to the contribution rate to obtain the evaluation factor A: ;

[0052] where represents the sum of the variance contribution rates of the first k principal components, which is used to calculate the weight of each principal component in the weighted summation.

[0053] 2.3 Step S3, patient mortality assessment:

[0054] Calculate the death rate curve D of historical patients vc , and use the ordered clustering algorithm to cluster patients according to the historical medical record data D' and the death rate curve D vc to calculate the evaluation factor B.

[0055] 2.3.1 Step S300, collect historical medical record data D':

[0056] Collect disease types De and death time d t Then perform encoding, unify the time format, and calculate the length of hospital stay to form the historical medical record data D'.

[0057] 2.3.2 Step S301, determine the time window and death time distribution:

[0058] Set the time window T for the patients in the historical medical record data D' w (such as within one year after admission), count the death month M d ; For each historical patient, analyze the distribution of different disease types De and their corresponding death times d t and conduct time series analysis, use the exponential model for curve fitting to obtain the death rate curve D vc :

[0059] (1) Exponential model: D vc (t)=a·exp(b·t)+c;

[0060] Among them, a, b, and c are parameters to be fitted; t is the time variable (hours, days, weeks, or months, etc.);

[0061] (2) Use the least squares method to minimize the sum of the squared differences between the observed values and the model predicted values:

[0062] That is ;

[0063] Among them, N is the number of observed values. Use relevant functions or libraries in statistical software or programming languages (such as Python, R, etc.) to perform the fitting and parameter estimation of the exponential model, solve the above least squares problem, and obtain the values of parameters a, b, and c.

[0064] 2.3.3 Step S302, ordered clustering analysis:

[0065] Use the ordered clustering algorithm (Fisher optimal segmentation method) to cluster the patients according to the historical medical record data D' and the death rate curve D vc to obtain different risk groups. Quantify the evaluation indicators and calculate the evaluation factor B;

[0066] The goal of the Fisher optimal segmentation method is to divide n patients into K ordered groups (i.e., risk groups) so that the similarity within the ordered groups is maximized and the difference between the ordered groups is maximized:

[0067] (1) Objective function: Let G1, G2,..., GK A partition that divides n patients into K groups, and the objective function L(n, K) is defined as:

[0068] ;

[0069] where n i is the number of patients in the i-th group G i , is the mean vector of the characteristics of the patients in group G i , is the mean vector of the characteristics of all patients.

[0070] (2) Solving the optimal partition: Use the dynamic programming algorithm to solve the optimal partition. Define L(j, i) as the objective function value of the optimal partition that divides the first j patients into i groups. The recurrence relation is obtained as:

[0071] ;

[0072] where is the mean vector of the characteristics from patient s + 1 to patient j; represents the difference between the mean vector of the characteristics from patient s + 1 to patient j and the mean vector of the characteristics of all patients. represents the transpose of the above difference vector (in the real number case, it is the vector itself; in the complex number or more general vector space case, the conjugate operation should be performed). s represents a boundary of the patient number used to divide different patient groups. L(s, i - 1) represents the objective function value of the optimal partition that divides the first s patients into i - 1 groups.

[0073] (3) Initialization: Since there is no internal difference when there is only one group, so when i = 1, L(j, 1) = 0;

[0074] (4) Calculation: Starting from i = 2, gradually increase the number of groups until the predetermined number of groups K is reached;

[0075] (5) Obtain the optimal partition: Through backtracking the dynamic programming table, obtain the optimal partition that divides n patients into K groups;

[0076] (6) Quantification: ;

[0077] where is the proportion of the number of patients in group G i to the total number of patients, and R i is the risk score of group G i ; The evaluation factor B (after normalization) reflects the risk level of the overall patients.

[0078] 2.4 Step S4, Comprehensive Reputation Feature Collection and Quantification Phase:

[0079] Collect data on social reputation, treatment level, service attitude, and patient satisfaction to form a comprehensive reputation feature value S.

[0080] 2.4.1 Step S400, Collect data on social reputation, treatment level, service attitude, and patient satisfaction:

[0081] Quantify the characteristic sequence R1 of reputation through industry rankings;

[0082] Quantify the characteristic sequence R2 of treatment level through the cure rate;

[0083] Quantify the characteristic sequence R3 of service attitude through the results of the complaint rate survey;

[0084] Directly quantify the characteristic sequence R4 of patient satisfaction through the results of the satisfaction survey.

[0085] 2.4.2 Step S401, Determine weights and calculate the comprehensive reputation feature value:

[0086] Determine the corresponding weights W1, W2, W3, and W4 for the characteristic sequence R1, characteristic sequence R2, characteristic sequence R3, and characteristic sequence R4 according to statistical methods. Then calculate the comprehensive reputation feature value S: S = W1⋅R1 + W2⋅R2 + W3⋅R3 + W4⋅R4.

[0087] 2.5 Step S5, D-S Evidence Theory Algorithm Combining Phase:

[0088] Use the D-S evidence theory algorithm to combine evaluation factor A and evaluation factor B to form a treatment factor S with patient underlying disease assessment information and patient mortality assessment information aving .

[0089] 2.5.1 Step S500, Define the frame of discernment and assign basic probability assignments (BPAs):

[0090] Define the frame of discernment I = {LR, MR, HR}, where LR represents low risk, MR represents medium risk, and HR represents high risk; then assign basic probability assignments (BPAs) to evaluation factor A and evaluation factor B, denoted as m A and m B ; where:

[0091] (1) m A (LR) represents the degree of support of evaluation factor A for the low-risk proposition, m A (MR) represents the degree of support of evaluation factor A for the medium-risk proposition, m A (HR) represents the degree of support of evaluation factor A for the high-risk proposition;

[0092] (2)m B (LR) represents the degree of support of evaluation factor B for low - risk propositions, m B (MR) represents the degree of support of evaluation factor B for medium - risk propositions, m B (HR) represents the degree of support of evaluation factor B for high - risk propositions;

[0093] 2.5.2 Step S501, Combine Basic Probability Assignments (BPAs):

[0094] (1)Use Dempster's combination rule to combine the basic probability m A and the basic probability m B , and calculate the normalization constant D:

[0095] ;

[0096] where X and Y are both subsets of the frame of discernment I; the symbol represents the empty set;

[0097] (2)The combined basic probability BPA value m A,B is calculated as:

[0098] ;

[0099] where Z is also a subset of the frame of discernment I.

[0100] 2.5.3 Step S502, Evaluation and Decision - Making:

[0101] Calculate the belief function Bel and the plausibility function Pl to form the treatment factor S aving . The method is as follows:

[0102] (1)Calculate the belief function Bel and the plausibility function Pl:

[0103] ;

[0104] ;

[0105] (2)According to the values of the belief function Bel and the plausibility function Pl, form the treatment factor S aving . The rule is that if the belief degree of the low - risk LR proposition is greater than the plausibility degree of the high - risk proposition HR, then it can be considered that the treatment effect is better, so the value of the treatment factor S aving will also be larger:

[0106] S aving = Bel(LR) - Pl(HR);

[0107] Among them, Bel(LR) represents the degree of belief in the low-risk proposition, and Pl(HR) represents the likelihood of the high-risk proposition. The treatment factor S aving The larger the value of, the better the treatment effect, and the lower the overall risk after combining the basic disease assessment information and mortality assessment information of the patient.

[0108] 2.6 Step S6, Hospital / Physician Selection Phase:

[0109] Combine the treatment factor S aving and the comprehensive reputation eigenvalue S into a comprehensive consideration factor C that can help patients select a hospital / physician CF .

[0110] 2.6.1 Step S600, Determine Weights:

[0111] Use the entropy weight method to determine the weight W5 of the treatment factor S aving and the weight W6 of the comprehensive reputation eigenvalue S:

[0112] (1) Construct an evaluation matrix EM: Suppose there are n hospitals / physicians to be evaluated, and each hospital / physician has two evaluation indicators: the treatment factor S aving and the comprehensive reputation eigenvalue S. Construct an n*2 evaluation matrix EM, where any element x ij in the evaluation matrix EM represents the value of the i-th hospital / physician on the j-th evaluation indicator.

[0113] (2) Data standardization: Since both the treatment factor S aving and the comprehensive reputation eigenvalue S are positive indicators (the larger the value, the better), use the same standardization operation: ;

[0114] Among them, x ij ′ is the value of the i-th hospital / physician on the j-th evaluation indicator after standardization.

[0115] (3) For each evaluation indicator j, calculate the entropy value e j : ;

[0116] Among them, p ij represents the proportion of the standardized value of the i-th hospital / physician on the j-th evaluation indicator ( ); k is a constant, equal to , where n is the number of hospitals / physicians to be evaluated.

[0117] (4) For each evaluation indicator j, calculate its difference coefficient g j : g j =1 - e j ;

[0118] (5) For each evaluation index j, calculate the corresponding weight w j : ;

[0119] 2.6.2 Step S601, calculate the comprehensive consideration factor:

[0120] Combine the treatment factor S aving and the comprehensive reputation eigenvalue S, and calculate the comprehensive consideration factor C through weighted summation CF : CCF = W5 * Saving + W6 * S;

[0121] Patients can select a hospital / physician according to the value of the comprehensive consideration factor. The higher the value, the better the hospital / physician performs in terms of medical technology, service quality, reputation, etc.

[0122] (III) Mechanism for solving technical problems:

[0123] 3.1 How to reflect the seasonal changes of physiological functions or disease spectra and quantify the data of each feature:

[0124] The present invention collects the detailed medical records of patients and sets a time window, analyzes the distribution of the death months of patients, and observes whether there are concentrated death months or specific death rate curves. That is, time series analysis and curve fitting techniques are used to describe the trend of the death rate changing with time, so as to reflect the seasonal changes of physiological functions or disease spectra. At the same time, for the quantification of the data of each feature, the present invention uses a variety of specific indicators to quantify features such as reputation, ensuring the objectivity and comparability of the data.

[0125] 3.2 How to avoid the subjectivity problem of setting weights depending on the subjective judgment of experts or managers:

[0126] When setting weights, the present invention adopts an objective data-driven method (entropy weight method), rather than relying on the subjective judgment of experts or managers. The entropy weight method determines the weights by analyzing the dispersion degree of the data of each feature itself. The greater the dispersion degree, the greater the amount of information provided by the feature, and thus the greater the weight in the comprehensive evaluation. It effectively avoids the deviation caused by subjective weight setting and makes the distribution of weights more scientific and reasonable.

[0127] 3.3 How to comprehensively combine and consider different quantified data based on a data-driven method, better adapt to the changes in the hospital environment, and quantify the data of each feature, so as to obtain a more comprehensive reputation evaluation result:

[0128] The present invention adopts a variety of data-driven methods. Among them, principal component analysis is used to extract important comprehensive indicators of patients' underlying diseases, while ordered clustering analysis is used to evaluate the risk of patient mortality. At the same time, D-S evidence theory is used to combine treatment factors and comprehensive reputation eigenvalues. It can not only effectively quantify the data of each feature, but also organically combine the quantified data from different sources and of different natures to form a comprehensive consideration factor. By continuously collecting and analyzing the latest medical data, the present invention can dynamically reflect the changes in the hospital environment and patients' needs, thereby obtaining a more comprehensive and accurate reputation assessment result, providing a more accurate and scientific basis for patients to select medical services.

[0129] Second aspect, a patient condition assessment system based on a multi-source information fusion mechanism:

[0130] This system is used to implement the patient condition assessment method based on the multi-source information fusion mechanism described above, including:

[0131] (1) A data collection module responsible for automatically collecting the dataset of patients' underlying disease indicators from the hospital information system: It has functions of interface development, data scraping, data cleaning and preprocessing.

[0132] (2) A data analysis module that applies principal component analysis to extract comprehensive indicators of patients' underlying diseases and applies ordered clustering analysis to evaluate the mortality risk of patients: It has functions of PCA algorithm implementation, ordered clustering algorithm implementation (Fisher optimal segmentation method), time series analysis and curve fitting.

[0133] (3) A feature quantification and weight assignment module responsible for implementing the entropy weight method: It has functions of formulating quantification standards, calculating weights by the entropy weight method and weight assignment strategies.

[0134] (4) A fusion module responsible for implementing the D-S evidence theory algorithm: It uses the Dempster combination rule to fuse the treatment factors (evaluation factors A and B) based on patients' underlying diseases and mortality risk with the comprehensive reputation eigenvalues of the hospital to form a comprehensive consideration factor. It includes functions of D-S evidence theory algorithm implementation, basic probability assignment (BPA) calculation, belief function and plausibility function calculation.

[0135] (5) A result interpretation module for evaluating the reputation of a hospital or a physician: It interprets the fused consideration factor and provides suggestions for patients to choose medical treatment.

[0136] Compared with the prior art, the beneficial effects of the present invention are:

[0137] 1. Improve the comprehensiveness and objectivity of the evaluation: By comprehensively considering the severity of the patient's underlying disease, mortality risk, hospital reputation, treatment level and other factors, the present invention can provide a more comprehensive and objective hospital reputation evaluation result. It helps patients to more accurately understand the overall strength and service quality of the hospital, and provide strong support for medical selection.

[0138] Second, reflecting seasonal changes in physiological functions and disease spectrum: By analyzing the distribution of the months of death of patients and fitting the death rate curve, the present invention can reveal the seasonal changes in physiological functions or disease spectrum. This is of great significance for disease prevention, allocation of medical resources and improvement of medical services, and helps hospitals better cope with seasonal health challenges.

[0139] 3. Reduce subjective influence: In terms of weight setting, the present invention adopts objective data-driven methods such as entropy weight method to avoid the bias caused by relying on the subjective judgment of experts or managers. This objective and scientific weight allocation method helps to improve the fairness and credibility of the evaluation results.

[0140] 4. Enhanced data-driven decision support: This invention makes full use of data analysis techniques, including principal component analysis, ordered cluster analysis and DS evidence theory, to organically combine quantitative data from different sources and properties to form a comprehensive consideration factor. This data-driven approach not only improves the accuracy and efficiency of the evaluation, but also provides a powerful decision support tool for hospital managers.

[0141] 5. Adapt to changes in hospital environment: Since the present invention is based on continuous data collection and analysis, it can dynamically reflect changes in hospital environment and patient needs. This makes the evaluation results closer to the actual situation, helping hospitals to adjust service strategies and optimize resource allocation in a timely manner to better meet patients' needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0142] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0143] Figure 1 It is a schematic diagram of the method flow of the present invention;

[0144] Figure 2 Schematic diagram of the execution method of steps S2 to S5 of the present invention;

[0145] Figure 3 It is a schematic diagram of the system composition of the present invention;

[0146] Figure 4 Schematic diagram for comparing the principal component analysis effects of step S2 in the first embodiment of the present invention;

[0147] Figure 5 Visualization diagram of the clustering effect of step S3 in the first embodiment of the present invention. Detailed implementation manners

[0148] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention with reference to the accompanying drawings. Many specific details are set forth in the following description to facilitate a thorough understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below;

[0149] It should be noted that the various embodiments in this specification are described in a progressive manner, and the key point of each embodiment is to describe the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0150] Embodiment 1: As Figure 1 shown, this embodiment discloses a method for evaluating a patient's condition based on a multi-source information fusion mechanism. When a patient calls the medical reputation of a hospital / physician, the hospital information system first calls out the basic disease index dataset D of historical patients, including basic information (age and BMI index), past medical history, physical examination data, laboratory test data (such as blood routine or liver and kidney function), and health records (eating habits, exercise habits, and sleep conditions).

[0151] Specifically, the basic disease index dataset D is an n-row and m-column matrix, where n is the number of patients and m is the number of corresponding basic disease indexes. Let D ij represent the value of the j-th basic disease index of the i-th patient:

[0152] ;

[0153] For each basic disease index D ij in the basic disease index dataset D, the following normalization is performed: (1)

[0154] where D ij ′ is the value after standardization, min(D j ) and max(D jThey are respectively the minimum and maximum values of the j-th underlying disease index among all patients.

[0155] (2) For some special underlying disease indices, such as the BMI index or blood pressure, they both have standard normal ranges. In this case, these indices are selected to be standardized within their normal ranges instead of using the minimum and maximum values of the entire dataset. For example, if the normal range of the BMI index is from 18.5 to 24.9, then: ;

[0156] Or if it is necessary to take into account the index values outside the normal range, a piecewise function is used for processing;

[0157] (3) For laboratory test data with different units and magnitudes, the Z-score standardization method is used in step S100 to eliminate the influence of units and magnitudes.

[0158] Subsequently, the following steps S1 to S6 of this method are executed.

[0159] In this embodiment, regarding step S1, data collection: The underlying disease index dataset D is constructed into a data matrix d where rows represent different historical patients and columns represent different indices m . Its specific scheme includes the following steps S100 to S101.

[0160] Specifically, in step S100, data cleaning: The purpose of data cleaning is to ensure the integrity and accuracy of the data for subsequent analysis and modeling. This step includes three links: missing value filling, outlier handling, and Z-score standardization. Among them, the Z-score standardization method standardizes all indices to ensure that all variables have the same mean and variance:

[0161] (1) Missing value filling: ;

[0162] Where x missing is the missing value, X i are the non-missing values, and n is the number of non-missing values; The missing values are filled by calculating the mean of the non-missing values, which can maintain the overall distribution characteristics of the data.

[0163] (2) Outlier handling (Z-score method): ;

[0164] Where, x i is the i-th original value in the underlying disease index dataset D, μ a is the mean of the current index among all non-outlier patients, σ ais the standard deviation of the current indicator among all patients without outliers. Data points with ∣Z∣>3 are considered outliers; for outliers, deletion, correction, or other methods can be selected based on a preset dictionary to ensure data accuracy and reliability.

[0165] (3) Z-score standardization: First, calculate the mean and standard deviation of each indicator. Then, for each data point, we subtract the mean from its original value and divide by the standard deviation to obtain its Z-score value. This Z-score value reflects the degree of deviation of the data point from the mean, in units of the standard deviation:

[0166] ;

[0167] where, x i ′ is the i-th element after standardization in the dataset D of comorbidity indicators, x i is its original value, μ d is the mean of the current indicator among all patients without missing values; σ d is the standard deviation of the current indicator among all patients without missing values. Through standardization, we can eliminate the dimensional differences between different variables and make them comparable.

[0168] It can be understood that since Z-score standardization is based on the statistical characteristics of the data, it is applicable to various types and distributions of data. For laboratory test data, although different test items may have different units and magnitudes, through Z-score standardization, we can convert them into standard scale data with comparability, thus facilitating subsequent analysis and modeling.

[0169] Specifically, in step S101, construct the data matrix d m :

[0170] ;

[0171] where, x ij ′ represents the j-th indicator (standardized value) of the i-th historical patient, n is the number of patients, and m is the number of indicators. The data matrix d m will serve as the basis for subsequent multi-source information fusion and hospital reputation assessment.

[0172] In this embodiment, regarding step S2, patient comorbidity assessment: Based on the data matrix d m obtained in step S1, perform principal component analysis (PCA) to obtain the evaluation factor A. Its specific scheme includes the following steps S200~S201. At the same time, please refer to Figure 2This step is an important part of helping patients obtain a more objective medical reputation of the hospital / physician, because it quantifies the characteristics of the patients' underlying diseases and provides key inputs for subsequent multi-source information fusion.

[0173] Specifically, in step S200, calculate the covariance matrix: Convert the data matrix d m into a covariance matrix C that describes the strength and direction of the linear relationship between each index, and then perform eigenvalue decomposition to obtain eigenvalues E and corresponding eigenvectors F:

[0174] (1) Convert to covariance matrix C: ;

[0175] Among them, represents the mean vector of the data matrix d m , represents the i-th row in the data matrix d m , that is, the index data of the i-th patient (a 1*m vector, where m is the number of index j); T represents the transpose operation of the matrix. The covariance matrix C describes the strength and direction of the linear relationship between each index in the data matrix d m . By calculating the deviation of each patient's index data from the mean vector and finding the average of its outer product, the covariance matrix can be obtained. This step is to capture the correlation between the indexes and lay a foundation for subsequent eigenvalue decomposition and principal component selection.

[0176] (2) Eigenvalue decomposition: C = F * E * F T ;

[0177] Among them, eigenvalue decomposition decomposes the covariance matrix C into the product of an eigenvector matrix F, a diagonal matrix E containing eigenvalues, and the transpose of F. The eigenvalues and eigenvectors respectively reflect the degree of variation and direction of the data in the corresponding direction. By eigenvalue decomposition, the main variation direction in the data, that is, the principal component, can be identified.

[0178] Specifically, in step S201, select the principal component: Select the principal component P that makes the cumulative contribution rate reach the preset threshold according to the size of the eigenvalue E; then project the original data into the new feature space F n constituted by the selected principal component P, and calculate the new coordinates P c of each data point (that is, the principal component score); normalize the new coordinates P c and perform weighted summation according to the contribution rate to obtain the evaluation factor A:

[0179] (1) Assume that selecting k principal components can make the cumulative contribution rate R C reach the preset threshold; the cumulative contribution rate reflects the proportion of the selected principal components in explaining the total variation of the data, and it can be determined how many principal components need to be selected;

[0180] (2) Calculate the cumulative contribution rate R by computing the ratio of the sum of the first k eigenvalues to the sum of all eigenvalues. C : ;

[0181] where E jj represents the j-th diagonal element of the diagonal matrix E; this step is to check whether the selected principal components reach the preset cumulative contribution rate threshold.

[0182] (3) Project the original data matrix d m onto the new feature space F n composed of the selected k principal components, and the coordinates of each data point in the new space, i.e., the principal component scores P c can be obtained: P c = d m * F k ;

[0183] where F k is the matrix composed of the eigenvectors corresponding to the selected k principal components; P c are the principal component scores; this step is to transform the original data into a representation in the new coordinate system based on the principal components, facilitating subsequent analysis and processing.

[0184] (4) Normalize the principal component scores P c to eliminate the dimensional differences between different principal component scores and make them comparable: ;

[0185] where P cn represents the value of the j-th principal component score after normalization; this step is to ensure that the contributions of each principal component score are fair during subsequent weighted summation.

[0186] (5) Perform weighted summation according to the variance contribution rate of each principal component to obtain the evaluation factor A for comprehensively evaluating the patient's underlying diseases: ;

[0187] where represents the sum of the variance contribution rates of the first k principal components, which is used to calculate the weight of each principal component during weighted summation. This step is to integrate the information of each principal component to form an evaluation index that can comprehensively reflect the patient's underlying disease status.

[0188] It can be understood that through PCA processing, the original high-dimensional data is reduced to several main principal components, thus simplifying the complexity of data analysis while retaining the main information of the data. Since PCA is an objective data analysis method, it is based on the statistical characteristics of the data, avoiding the subjective interference of human factors and improving the accuracy and objectivity of the evaluation. The evaluation factor A, as a comprehensive evaluation index of the patient's underlying disease status, can provide key inputs for subsequent multi-source information fusion and hospital reputation evaluation, thereby improving the accuracy and reliability of the entire evaluation system.

[0189] Furthermore, the Python execution program for the above step S2 is as follows:

[0190] import numpy as np

[0191] from sklearn.decomposition import PCA

[0192] from sklearn.preprocessing import StandardScaler

[0193] # The data matrix dm has been obtained in step S1

[0194] np.random.seed(0)

[0195] dm = np.random.randn(10, 5) # 10 patients, 5 indicators

[0196] # Step S200: Calculate the covariance matrix

[0197] mean_vector = np.mean(dm, axis=0)

[0198] cov_matrix = np.cov(dm, rowvar=False) # rowvar=False indicates that each column in the data matrix represents a variable

[0199] # Step S201: Select the principal components and obtain the evaluation factor A

[0200] # Use the PCA class in scikit-learn to perform PCA

[0201] pca = PCA()

[0202] # Standardize the data

[0203] scaler = StandardScaler()

[0204] dm_standardized = scaler.fit_transform(dm)

[0205] # Fit the PCA model and select the principal components

[0206] pca.fit(dm_standardized)

[0207] # Set the cumulative contribution rate threshold (95%)

[0208] cumulative_variance_ratio_threshold = 0.95

[0209] num_components = np.argmax(pca.explained_variance_ratio_.cumsum() >= cumulative_variance_ratio_threshold) + 1

[0210] # Extract the eigenvectors and principal component scores corresponding to the selected principal components

[0211] principal_components = pca.components_[:num_components]

[0212] principal_scores = pca.transform(dm_standardized)[:, :num_components]

[0213] # Normalize the principal component scores

[0214] principal_scores_normalized = (principal_scores - np.min(principal_scores, axis = 0)) / (np.max(principal_scores, axis = 0) - np.min(principal_scores, axis = 0))

[0215] # Weighted sum according to the contribution rate to obtain the evaluation factor A

[0216] explained_variance_ratios = pca.explained_variance_ratio_[:num_components]

[0217] evaluation_factor_A = np.sum(explained_variance_ratios * principal_scores_normalized, axis=1)

[0218] # Output evaluation factor A

[0219] print(evaluation_factor_A)

[0220] In the above program, the numpy.cov function is used to calculate the covariance matrix of the data matrix d m Note that the rowvar=False parameter indicates that each column in the data matrix represents a variable (i.e., an indicator), and each row represents an observation (i.e., a patient). The PCA class in scikit-learn is used to perform PCA. First, the data is standardized. Then, the PCA model is fitted and the variance contribution rate of each principal component is obtained through the explained_variance_ratio_ attribute. The eigenvectors and principal component scores corresponding to the selected principal components are extracted. The principal component scores are normalized to ensure that their contributions are fair when weighted and summed. The evaluation factor A for each patient is obtained by weighted summation according to the contribution rate.

[0221] Exemplarily, let there be a data matrix d m whose data point distribution is as shown in part A of Figure 4 ; there are multiple dimensions in the original data, and the distribution of data points is difficult to directly observe in two-dimensional or three-dimensional space. PCA projects the data onto new coordinate axes through linear transformation, and these coordinate axes (i.e., principal components) are sorted according to the data variance. In the scatter plot after dimensionality reduction (such as part B of 4), the data points are redistributed along the direction of the principal components. Moreover, if there are obvious clusters or linear relationships in the original data, these relationships are more obvious in the scatter plot after dimensionality reduction.

[0222] In this embodiment, regarding step S3, patient mortality assessment: To more accurately evaluate the medical quality of a hospital, especially for the key indicator of patient mortality, this method introduces the concept of the death rate curve and combines it with the ordered clustering algorithm for analysis. This step aims to understand the patient death pattern in a more refined way by deeply analyzing historical medical records, so as to provide a more objective hospital / physician reputation assessment. This step is mainly responsible for calculating the death rate curve D vc of historical patients, and using the ordered clustering algorithm according to the historical medical record data D' and the death rate curve D vcCluster the patients and calculate the evaluation factor B. The specific solution includes the following steps S300 to S302. Please also refer to Figure 2 .

[0223] Specifically, for step S300, collect historical medical record data D': First, collect the medical records of patients with specific diseases De (such as heart disease, cancer, etc.) from the hospital information system, including information such as the patient's admission date, discharge date (if applicable), and death date (if applicable). Encode this data to ensure a unified data format, especially standardize the time format for subsequent analysis.

[0224] Based on the admission and discharge (or death) dates, calculate the length of stay for each patient, which is the basic data for understanding the disease progression and treatment effect. Integrate the above information to form a historical medical record dataset D' containing key information such as disease type, death time, and length of stay.

[0225] In practice, the factors leading to patient death are complex. Simply considering the mortality rate in the scope of consideration is likely to lead to inaccurate data; specifically, for this reason, perform step S301 to determine the time window and death time distribution: Set the time window T of the patients in the historical medical record data D' w (such as within one year after admission to exclude the interference of deaths caused by long-term chronic diseases on the evaluation of short-term treatment effects), and count the death month M d ;

[0226] At the same time, for each historical patient, analyze the distribution of different disease types De and their corresponding death times d t and perform time series analysis. To achieve this goal, in this step, an exponential model is selected to perform curve fitting on the death time d t under different disease types De to obtain the death rate curve D vc corresponding to this disease type De:

[0227] (1) Exponential model: D vc (t)=a·exp(b·t)+c;

[0228] Among them, a, b, and c are parameters to be fitted; t is the time variable (days, months, etc.); the selection of the exponential model is based on its ability to better capture the trend of growth or decay over time and is suitable for describing the dynamic changes of disease progression and death risk.

[0229] (2) Use the least squares method to minimize the sum of the squared differences between the observed values and the model predicted values:

[0230] That is ;

[0231] Among them, Dvc(ti ) is the predicted value of the exponential model. N is the number of observations, that is, whether the patient died at a specific time point t i (in the fitting process of the death rate curve, the number of deaths or the probability of death corresponding to each time point can be regarded as an observation); y i is the actual death data, including the time of death, cause of death (if related to disease De), etc. The actual death data is used to calculate the death rate curve, that is, by analyzing the death situation at different time points, to reveal the changing trend of the death rate over time.

[0232] It should be particularly noted that the actual death data is the key input required for fitting the death rate curve. It is essentially the same as the observation value, but described from different perspectives. The observation value focuses more on the numerical performance at a specific time point, while the actual death data more comprehensively reflects the patient's death situation.

[0233] It can be understood that the exponential model is fitted and parameter estimates are made through relevant functions or libraries in statistical software or programming languages (such as Python, R, etc.), the above least squares problem is solved, and the values of parameters a, b, and c are obtained. And by introducing the death rate curve and the exponential model, this method can more precisely depict the relationship between patient death and time. Compared with a single mortality rate indicator, it can better reflect the effectiveness and timeliness of medical intervention, thereby improving the accuracy of hospital / physician reputation assessment. Time series analysis helps to discover abnormal fluctuations in the mortality rate within a specific time period, which may indicate problems in medical processes, treatment plans, or resource allocation, providing directions for hospital management improvement.

[0234] Further, the Python execution program for the above steps S300~S301 is as follows:

[0235] import pandas as pd

[0236] import numpy as np

[0237] from scipy.optimize import curve_fit

[0238] import matplotlib.pyplot as plt

[0239] # The CSV file contains historical medical record data

[0240] # Data columns include: patient_id, disease_type, admission_date, discharge_date(or death_date)

[0241] data_file = 'historical_medical_records.csv'

[0242] df = pd.read_csv(data_file)

[0243] # Step S300: Collect historical medical record data D'

[0244] df_disease = df[df['disease_type'] == 'heart_disease']

[0245] # Convert date format, calculate length of hospital stay and time of death

[0246] df_disease['admission_date'] = pd.to_datetime(df_disease['admission_date'])

[0247] df_disease['death_date'] = pd.to_datetime(df_disease['death_date'].fillna(df_disease['discharge_date'])) # If there is no death date, use the discharge date

[0248] df_disease['hospital_stay'] = (df_disease['death_date'] - df_disease['admission_date']).dt.days

[0249] # Only retain data of patients who died within one year after admission (time window Tw)

[0250] df_disease_within_tw = df_disease[df_disease['hospital_stay'] <= 365]

[0251] # Statistic the month of death Md

[0252] df_disease_within_tw['death_month'] = df_disease_within_tw['death_date'].dt.month

[0253] # Step S301: Determine the time window and the distribution of death times, and fit the death rate curve Dvc

[0254] # Aggregate the data and count the number of deaths by month

[0255] monthly_deaths = df_disease_within_tw.groupby('death_month').size()

[0256] # Time variable t (here, months are used as the time unit)

[0257] t = monthly_deaths.index.values

[0258] # Actual death data y (number of deaths per month)

[0259] y = monthly_deaths.values

[0260] # Define the exponential model function

[0261] def exp_model(t, a, b, c):

[0262] return a * np.exp(b * t) + c

[0263] # Use the curve_fit function to perform curve fitting and obtain the parameters a, b, c

[0264] params, covariance = curve_fit(exp_model, t, y)

[0265] a, b, c = params

[0266] # Print the fitted parameters

[0267] print(f"Fitted parameters: a={a}, b={b}, c={c}")

[0268] # Use the fitted model to predict the number of deaths

[0269] y_pred = exp_model(t, a, b, c)

[0270] # Calculate the sum of squared differences (i.e., the residual sum of squares)

[0271] rss = np.sum((y - y_pred) ** 2)

[0272] print(f"Residual Sum of Squares (RSS): {rss}")

[0273] # Visualize the fitting result

[0274] plt.scatter(t, y, label='Observed Deaths')

[0275] plt.plot(t, y_pred, label='Fitted Death Speed Curve', color='red')

[0276] plt.xlabel('Month')

[0277] plt.ylabel('Number of Deaths')

[0278] plt.legend()

[0279] plt.title('Death Speed Curve Fitting')

[0280] plt.show()

[0281] In the above program, the pandas library is used to read the historical medical record data in CSV format. Filter the patient data for a specific disease (such as heart disease). Convert the date format, and calculate the length of hospitalization and the time of death. If the death date is missing, use the discharge date instead (in practical applications, more complex logic may be required to handle this situation). Only retain the patient data of those who died within one year after admission to meet the setting of the time window Tw. Define the exponential model function exp_model to describe the change of death speed over time. Use the scipy.optimize.curve_fit function to perform curve fitting on the actual death data to obtain the optimal values of the model parameters a, b, and c.

[0282] Then perform step S302, ordered clustering analysis: Use the ordered clustering algorithm (Fisher optimal segmentation method) according to the historical medical record data D’ and the death speed curve D vc Cluster the patients to obtain different risk groups. Quantify the evaluation indicators and calculate the evaluation factor B;

[0283] The goal of Fisher's optimal segmentation method is to divide the n patients under each disease type De into K ordered groups (i.e., risk groups), maximizing the similarity within the ordered groups and the difference between the ordered groups. Through this process, we can quantitatively evaluate the risk level of patients and calculate the evaluation factor B, providing an objective basis for the subsequent hospital reputation evaluation.

[0284] (1) Objective function: Let G1, G2, …, G K be a segmentation that divides n patients into K groups. Define the objective function as: ;

[0285] where, n i is the number of patients in the i-th group G i , is the mean vector of the characteristics of the patients in group G i , is the mean vector of the characteristics of all patients. This formula calculates the difference between each group and the overall mean and sums them with weights to reflect the overall difference between groups.

[0286] (2) Solving the optimal segmentation: Use the dynamic programming algorithm to solve the optimal segmentation. Define L(j, i) as the value of the objective function of the optimal segmentation that divides the first j patients into i groups. Use the dynamic programming algorithm to solve the optimal segmentation and gradually calculate the value of the objective function of the optimal segmentation through the recurrence relation. The recurrence relation is:

[0287] ;

[0288] where, is the mean vector of the characteristics from patient s + 1 to patient j; represents the difference between the mean vector of the characteristics from patient s + 1 to patient j and the mean vector of the characteristics of all patients. represents the transpose of the above difference vector (in the real number case, it is the vector itself; in the complex number or more general vector space case, the conjugate operation is performed). s represents a boundary of the patient number used to divide different patient groups. L(s, i - 1) represents the value of the objective function of the optimal segmentation that divides the first s patients into i - 1 groups.

[0289] (3) Initialization: Since there is no difference within a single group, when i = 1, L(j, 1) = 0;

[0290] (4) Calculation: Starting from i = 2, gradually increase the number of groups until the predetermined number of groups K is reached;

[0291] (5) Obtain the optimal segmentation: Through backtracking the dynamic programming table, obtain the optimal segmentation that divides n patients into K groups;

[0292] (6) Quantification: ;

[0293] Among them, is the proportion of the number of patients in group G i in the total number of patients, and R i is the risk score of group G i ; The evaluation factor B (after normalization) reflects the risk level of the overall patients.

[0294] Exemplarily, set the death rate curves D of six disease types De vc as shown in part (A) of Figure 5 , and the visualization effect of the groups segmented by the above clustering method is as shown in area (B) of Figure 5 .

[0295] It can be understood that through ordered clustering analysis, the risk groups of patients can be divided based on objective data such as the historical medical records and death rate curves of patients, avoiding the deviation of subjective evaluation. By calculating the evaluation factor B, the risk level of the overall patients can be accurately quantified, providing reliable data support for the hospital reputation evaluation. Since the dynamic programming algorithm is used, the number of groups K can be flexibly adjusted to adapt to the evaluation needs in different scenarios, improving the flexibility and applicability of the evaluation. Through clustering analysis and quantitative evaluation, an objective basis for the risk level of patients can be provided for the hospital and physicians, assisting them in making more reasonable medical decisions.

[0296] Furthermore, the Python execution program for the above step S302 is as follows:

[0297] import numpy as np

[0298] def fisher_optimal_partitioning(data, K):

[0299] """

[0300] Use the Fisher optimal partitioning method to cluster the data.

[0301] Parameters:

[0302] data (numpy.ndarray): The feature data of the patients, with shape (n_samples, n_features)

[0303] K (int): The number of groups to be divided into

[0304] Returns:

[0305] partitions (list of lists): A list of patient indices for each group

[0306] B (float): Evaluation factor

[0307] """

[0308] n = data.shape[0] # Number of patients

[0309] # Initialize the dynamic programming table

[0310] L = np.zeros((n + 1, K + 1))

[0311] # Initialize the split point record table for backtracking

[0312] split_points = np.zeros((n + 1, K + 1), dtype=int)

[0313] # Calculate the overall mean

[0314] overall_mean = np.mean(data, axis=0)

[0315] # Solve for the optimal split using dynamic programming

[0316] for i in range(2, K + 1): # Start from 2 groups and gradually increase to K groups

[0317] for j in range(i, n + 1): # Traverse the first j patients

[0318] min_L = float('inf')

[0319] best_s = 0

[0320] for s in range(i - 1, j): # Find the optimal split point s

[0321] # Calculate the mean from s+1 to j

[0322] mean_s_j = np.mean(data[s:j], axis=0)

[0323] # Calculate the objective function value

[0324] L_value = L[s, i - 1] + ((j - s) / n) * np.sum((mean_s_j - overall_mean) ** 2)

[0325] # Update the minimum objective function value and the best split point

[0326] if L_value < min_L:

[0327] min_L = L_value

[0328] best_s = s

[0329] L[j, i] = min_L

[0330] split_points[j, i] = best_s

[0331] # Backtrack the dynamic programming table to obtain the optimal split

[0332] partitions = []

[0333] j = n

[0334] for i in range(K, 0, -1):

[0335] s = split_points[j, i]

[0336] partitions.append(list(range(s + 1, j + 1)))

[0337] j = s

[0338] partitions.reverse() # Since we backtrack from K to 1, we need to reverse the list

[0339] # We can subtract 1 from the index list of each group to match Python's 0-based indexing

[0340] partitions = [[idx - 1 for idx in group] for group in partitions]

[0341] # Calculate the evaluation factor B

[0342] B = 0

[0343] for group in partitions:

[0344] group_data = data[group]

[0345] group_mean = np.mean(group_data, axis = 0)

[0346] group_risk_score = np.sum((group_mean - overall_mean) ** 2) # Risk score

[0347] B += (len(group) / n) * group_risk_score

[0348] return partitions, B

[0349] # Perform clustering analysis

[0350] K = 3 # Divide into 3 groups

[0351] partitions, B = fisher_optimal_partitioning(data, K)

[0352] # Output the results

[0353] print("Partitions:", partitions)

[0354] print("Evaluation Factor B:", B)

[0355] In the above program, the dynamic programming table L and the split point record table split_points are initialized. The overall mean overall_mean is calculated. Two nested loops are used to traverse the first j patients and the number of groups i. In the inner loop, the optimal split point s is found, the mean from s + 1 to j and the objective function value are calculated, and the minimum objective function value and the best split point are updated. The dynamic programming table is backtracked from K to 1 to obtain the list of group indices for the optimal partition. Each group is traversed to calculate the mean and risk score of the group (here, in the example, the sum of the squares of the difference between the mean and the overall mean is simply used as the risk score). The evaluation factor B is calculated based on the group size and the risk score.

[0356] In this embodiment, regarding step S4, the comprehensive reputation feature collection and quantification stage: The core objective of this stage is to collect and quantify the reputation features of a hospital or a physician in multiple dimensions, and then form a comprehensive reputation feature value S to comprehensively reflect the medical reputation of the hospital or the physician. This process is subdivided into two sub-steps: S400 and S401.

[0357] Specifically, in step S400, collect data on social reputation, treatment level, service attitude, and patient satisfaction:

[0358] (1) Collect social reputation data (R1): Use authoritative ranking agencies within the industry or publicly available industry reports to collect information on the social reputation of hospitals or physicians. Convert the rankings into a numerical feature sequence R1. The higher the ranking, the higher the value of R1, indicating better social reputation.

[0359] (2) Collect treatment level data (R2): Statistically calculate the cure rate of hospitals or physicians, that is, the proportion of successfully cured cases to the total number of cases. Convert the cure rate into a feature sequence R2. The higher the cure rate, the larger the value of R2, indicating a higher treatment level.

[0360] (3) Collect service attitude data (R3): Measure the service attitude through the survey results of the patient complaint rate. Convert the complaint rate into a feature sequence R3. The lower the complaint rate, the smaller the value of R3 (or use reverse scoring so that a low complaint rate corresponds to a high R3 value), indicating better service attitude.

[0361] (4) Collect patient satisfaction data (R4): Directly adopt the survey results of patient satisfaction. Convert the satisfaction score into a feature sequence R4. The higher the satisfaction, the larger the value of R4.

[0362] Specifically, in step S401, determine the weights and calculate the comprehensive reputation feature value: Determine the corresponding weights W1, W2, W3, and W4 for the feature sequence R1, feature sequence R2, feature sequence R3, and feature sequence R4 according to statistical methods. Then calculate the comprehensive reputation feature value S:

[0363] S = W1 ⋅ R1 + W2 ⋅ R2 + W3 ⋅ R3 + W4 ⋅ R4.

[0364] The higher the value of S, the better the overall reputation of the hospital or physician. By collecting and quantifying reputation features from multiple dimensions, it is possible to more comprehensively reflect the medical reputation of hospitals or physicians, avoiding the one-sidedness of single-index evaluation. Transforming reputation information from different sources and natures into numerical feature sequences, and through weight assignment and comprehensive calculation, the objective quantification of reputation is achieved, facilitating comparison and ranking. The calculation process of the comprehensive reputation feature value S is transparent and traceable, and based on actual data and statistical methods, improving the credibility and persuasiveness of the evaluation results. The comprehensive reputation feature value S can be used as an important reference for patients to choose hospitals or physicians, helping patients make more informed medical decisions.

[0365] In this embodiment, regarding step S5, the D-S evidence theory algorithm combination stage: The D-S evidence theory algorithm is used to combine evaluation factor A and evaluation factor B to form a treatment factor S with patient underlying disease assessment information and patient mortality assessment information. aving Please refer to Figure 2 , this process includes defining the frame of discernment and assigning basic probability BPA (step S500), combining basic probability BPA (step S501), and evaluation and decision-making (step S502).

[0366] Specifically, in step S500, defining the frame of discernment and assigning basic probability BPA: Define the frame of discernment I = {LR, MR, HR}, where LR represents low risk, corresponding to a lower numerical range or probability interval, such as [0, 0.3]; MR represents medium risk, corresponding to an intermediate numerical range or probability interval, such as [0.3, 0.7]; HR represents high risk, corresponding to a higher numerical range or probability interval, such as [0.7, 1];

[0367] Set the scoring criteria for the underlying disease condition, for example: mild (0 - 3 points), moderate (4 - 7 points), severe (8 - 10 points). Then assign basic probability BPA to evaluation factor A and evaluation factor B, denoted as m A and m B respectively; where:

[0368] (1) m A (LR) represents the degree of support of evaluation factor A for the low-risk proposition, m A (MR) represents the degree of support of evaluation factor A for the medium-risk proposition, m A (HR) represents the degree of support of evaluation factor A for the high-risk proposition, that is, according to the patient's underlying disease condition, give the corresponding score S A ;

[0369] To calculate the degree of support for each risk proposition, this embodiment selects to use a bell-shaped function form of the fuzzy membership function (Degree of Membership Function) to describe the membership degree of elements in the fuzzy set, whose shape is similar to the normal distribution curve, and then perform the calculation on the degree of support of evaluation factor A:

[0370] Define three bell-shaped functions to represent the membership degrees of low-risk, medium-risk, and high-risk propositions respectively. Let the following parameters be used to define these functions:

[0371] 1) The center point of the low-risk proposition (LR): c A,LR = 2, width: σ A,LR = 1.5;

[0372] 2) Center point of the medium - risk proposition (MR): c A,MR = 5, width: σ A,MR = 1.5;

[0373] 3) Center point of the high - risk proposition (HR): c A,HR = 8, width: σ A,HR = 1.5;

[0374] ;

[0375] Among them, μ A,i (S A ) represents the membership degree of the evaluation factor A (i.e., the patient's underlying disease condition) for the risk proposition i (i is the low - risk LR, medium - risk MR, or high - risk HR). Its output value is between 0 and 1, indicating the degree of support of S A for the proposition X. C A,i represents the center point of the bell - shaped function corresponding to the risk proposition i. For different risk propositions (low - risk, medium - risk, high - risk), the value of the center point will be different. It represents the most typical value of the evaluation factor A under this risk proposition. σ A,i represents the width of the bell - shaped function corresponding to the risk proposition i. The width determines the shape of the bell - shaped function, that is, the rate at which the membership degree changes with the value of the evaluation factor A. The larger the width, the flatter the function shape and the slower the change of the membership degree; the smaller the width, the steeper the function shape and the faster the change of the membership degree. The denominator part, through the operation of adding 1 to the sum of squares, ensures that the denominator is always positive, and when S A is close to c A,i , the denominator value is small, so the membership degree is large; when S A is far from c A,i , the denominator value is large, so the membership degree is small.

[0376] Therefore, the degrees of support of the evaluation factor A for the low - risk, medium - risk, and high - risk propositions are respectively:

[0377] m A (LR)=μ A,LR (S A );

[0378] m A (MR)=μ A,MR (S A );

[0379] m A (HR)=μ A,HR (S A );

[0380] (2) m B (LR) represents the degree of support of the evaluation factor B for the low - risk proposition, mB (MR) represents the degree of support of evaluation factor B for the medium-risk proposition, m B (HR) represents the degree of support of evaluation factor B for the high-risk proposition; since its value range has been normalized to [0, 1], a similar bell-shaped function can be directly used to define the membership degrees of the low-risk, medium-risk, and high-risk propositions, and the corresponding score S B :

[0381] 1) Center point of the low-risk proposition (LR): c B,LR = 0.2, width: σ B,LR = 0.15;

[0382] 2) Center point of the medium-risk proposition (MR): c B,MR = 0.5, width: σ B,MR = 0.15;

[0383] 3) Center point of the high-risk proposition (HR): c B,HR = 0.8, width: σ B,HR = 0.15;

[0384] ;

[0385] The parameter symbols are the same as those in the membership function of evaluation factor A and will not be elaborated here. Therefore, the degrees of support of evaluation factor B for the low-risk, medium-risk, and high-risk propositions are respectively:

[0386] m B (LR)=μ B,LR (S B );

[0387] m B (MR)=μ B,MR (S B );

[0388] m B (HR)=μ B,HR (S B );

[0389] It should be noted that when calculating the degree of support, it should be ensured that for any given score, the sum of the degrees of support of the low-risk, medium-risk, and high-risk propositions is not 1, because the fuzzy membership function allows the overlap of membership degrees. If the normalization requirement (i.e., the sum of the degrees of support is 1) needs to be met, other types of membership functions or post-processing can be considered. However, this step essentially belongs to a kind of fuzzy logic, and usually, the sum of membership degrees is not forced to be 1.

[0390] Specifically, in step S501, the basic probability assignments (BPAs) are combined. The purpose of this step is to combine the BPA values from different evaluation factors to form a comprehensive BPA value. The Dempster combination rule is a method for combining belief functions that takes into account the overlap (i.e., intersection) between different belief functions; it includes:

[0391] (1) Use the Dempster combination rule to combine m A and m B , and calculate the normalization constant D:

[0392] ;

[0393] where the symbol represents the empty set; X can be any one of {LR}, {MR}, {HR}, {LR, MR}, {LR, HR}, {MR, HR}, or {LR, MR, HR}, and similarly, Y can also be one of these subsets. In the Dempster combination rule, X and Y represent the subsets to which different belief functions (or BPAs) assign non-zero probabilities.

[0394] (2) The combined basic probability assignment BPA value mA,B is calculated as:

[0395] ;

[0396] where Z is also a subset of the frame of discernment I, which represents the subset to which the combined BPA (i.e., m A,B ) assigns non-zero probabilities. In the Dempster combination rule, Z is a certain combination of the intersection of X and Y, and this combination generates a new BPA during the combination process. However, Z is not randomly selected but is determined by the intersection of X and Y. Specifically, we only consider the product of the BPA values of X and Y during the combination process when and only when the intersection of X and Y is equal to Z. Therefore, Z is a specific result of the intersection of X and Y, which reflects the degree of support of the combined belief function for different risk propositions.

[0397] The combined BPA value m A,B (Z) reflects the comprehensive degree of support of evaluation factors A and B for the risk proposition Z. This value is obtained by considering all possible combinations of X and Y (whose intersection is equal to Z) and calculating the sum of the products of their BPA values. Finally, these combined BPA values will be used to calculate the belief function Bel and the plausibility function Pl, and then form the treatment factor Saving. Therefore, the connection among X, Y, and Z lies not only in the fact that they are subsets of the frame of discernment I, but more importantly, they are interrelated through the Dempster combination rule and jointly affect the final decision result.

[0398] Specifically, in step S502, evaluation and decision-making: calculate the belief function Bel and the plausibility function Pl to form the treatment factor S aving . The method is as follows:

[0399] (1) Construct the belief function Bel and the plausibility function Pl:

[0400] ;

[0401] ;

[0402] Among them, the symbol represents the empty set;

[0403] (2) According to the values of the belief function Bel and the plausibility function Pl, form the treatment factor S aving . The rule is that if the belief degree of the proposition of low risk LR (generally speaking, the optimistic expectation) is greater than the plausibility degree of the high-risk proposition HR (generally speaking, the pessimistic expectation), then it can be considered that the treatment effect is better, so the value of the treatment factor S aving will also be larger; the method process is as follows:

[0404] 1) Calculate the belief function Bel(LR): It represents the belief degree of the proposition Z, which is the sum of the BPA values of all subsets containing Z. For the low-risk proposition LR, its belief degree Bel(LR) is calculated as:

[0405] ;

[0406] Among them, m a,b (Y) is the basic probability assignment of the subset Y after the evaluation factors A and B are combined; the belief function Bel(Z) represents the belief degree of the proposition Z (for the scenario of this step, the proposition Z refers to the subset of low risk LR), which is the sum of the BPA values of all subsets containing Z.

[0407] 2) Calculate the plausibility function PI(HR): It represents the plausibility degree of the proposition Z, which is the sum of the BPA values of all subsets having an intersection with Z. For the high-risk proposition HR, its plausibility degree Pl(HR) is calculated as:

[0408] ;

[0409] Among them, it is necessary to consider all subsets having an intersection with {HR} and add their BPA values. This includes the BPA value directly assigned to {HR}, as well as the BPA values assigned to larger subsets containing {HR} (such as {MR, HR} or {LR, MR, HR}) (if any).

[0410] 3) Saving = Bel(LR) - Pl(HR);

[0411] where the treatment factor S aving The larger the value, the greater the degree of belief in the proposition of low risk LR than the likelihood of the high-risk proposition HR, that is, the higher the optimistic expectation of low risk, that is, the better the treatment effect, and the lower the overall risk after combining the basic disease assessment information and mortality assessment information of the patient.

[0412] It can be understood that through the D-S evidence theory algorithm, the evaluation factors A and B of different sources and natures are organically combined, improving the accuracy and comprehensiveness of hospital reputation assessment. The treatment factor S aving The introduction of provides an intuitive and comparable index for quantifying the treatment effect. The treatment factor S aving The value directly reflects the quality of the treatment effect, providing strong decision-making support for hospital managers and patients. By comparing the treatment factor S aving values of different hospitals or physicians, patients can more objectively select medical service providers. The application of the D-S evidence theory algorithm effectively solves the uncertainty and conflict problems in multi-source information fusion, improving the efficiency and accuracy of information fusion.

[0413] Furthermore, the python execution program for step S5 above is as follows:

[0414] # Step S500: Define the frame of discernment and assign the basic probability BPA

[0415] def define_framework_and_bpa():

[0416] # Define the frame of discernment

[0417] framework = {'LR': 'Low Risk', 'MR': 'Medium Risk', 'HR': 'HighRisk'}

[0418] # Assign the basic probability BPA to evaluation factors A and B

[0419] mA = {'LR': 0.6, 'MR': 0.2, 'HR': 0.2} # BPA of evaluation factor A

[0420] mB = {'LR': 0.5, 'MR': 0.3, 'HR': 0.2} # BPA of evaluation factor B

[0421] return framework, mA, mB

[0422] # Step S501: Combine Basic Probability Assignment (BPA)

[0423] def combine_bpa(mA, mB):

[0424] # Initialize the combined BPA

[0425] mA_B = {}

[0426] # Calculate the normalization constant D

[0427] D = sum(mA[x] * mB[y] for x in mA for y in mB if x == y) # Only consider the case of X = Y, i.e., direct assignment

[0428] # The combined BPA can be directly obtained by multiplying the corresponding terms of mA and mB and normalizing

[0429] for x in mA:

[0430] mA_B[x] = mA[x] * mB[x] / D # Consider the case of direct assignment

[0431] return mA_B

[0432] # Step S502: Evaluation and Decision

[0433] def calculate_bel_and_pl(mA_B):

[0434] # Calculate the belief function Bel and plausibility function Pl

[0435] bel = {'LR': 0, 'MR': 0, 'HR': 0}

[0436] pl = {'LR': 0, 'MR': 0, 'HR': 0}

[0437] # Calculate Bel(LR)

[0438] bel['LR'] = mA_B.get('LR', 0) # If 'LR' is not in mA_B, return 0

[0439] # For the plausibility function, we need to consider all subsets that intersect with 'HR'

[0440] pl['HR'] = mA_B.get('HR', 0) # If 'HR' is not in mA_B, return 0

[0441] # Calculate the treatment factor Saving

[0442] saving = bel['LR'] - pl['HR']

[0443] return bel, pl, saving

[0444] # Main program

[0445] def main():

[0446] # Execute step S500

[0447] framework, mA, mB = define_framework_and_bpa()

[0448] # Execute step S501

[0449] mA_B = combine_bpa(mA, mB)

[0450] # Execute step S502

[0451] bel, pl, saving = calculate_bel_and_pl(mA_B)

[0452] # Output the results

[0453] print("Combined BPA:", mA_B)

[0454] print("Belief function Bel:", bel)

[0455] print("Plausibility function Pl:", pl)

[0456] print("Treatment factor Saving:", saving)

[0457] # Run the main program

[0458] if __name__ == "__main__":

[0459] main()

[0460] In the above program, the define_framework_and_bpa function defines the recognition framework I (including 'LR', 'MR', 'HR') and the basic probability assignment (BPA) of evaluation factors A and B. The combine_bpa function is responsible for combining the BPA values of evaluation factors A and B. The combined BPA value is obtained by multiplying the corresponding terms and normalizing. The calculate_bel_and_pl function calculates the values of the belief function Bel and the plausibility function Pl, and calculates the treatment factor Saving based on these values. The main function calls the above functions in sequence and outputs the results of the combined BPA, the belief function Bel, the plausibility function Pl, and the treatment factor Saving.

[0461] In this embodiment, in order to help patients more objectively and comprehensively select hospitals or physicians, we design a comprehensive consideration factor CCF (Comprehensive Consideration Factor) as the decision-making basis. That is, step S6, the stage of selecting a hospital / physician: combine the treatment factor S aving and the comprehensive reputation characteristic value S into a comprehensive consideration factor C that can help patients select a hospital / physician. CF It specifically includes the following steps S600~S601. The aim is to quantify the comprehensive performance of hospitals or physicians in terms of medical technology, service quality, and reputation.

[0462] Specifically, in step S600, determine the weights: use the entropy weight method to determine the weight W5 of the treatment factor S aving and the weight W6 of the comprehensive reputation characteristic value S:

[0463] (1) Construct an evaluation matrix EM: Suppose there are n hospitals / physicians to be evaluated, and each hospital / physician has two evaluation indicators: the treatment factor S aving and the comprehensive reputation characteristic value S. Construct an n*2 evaluation matrix EM, where any element x ij in the evaluation matrix EM represents the value of the i-th hospital / physician on the j-th evaluation indicator.

[0464] (2) Data standardization: Since both the treatment factor S aving and the comprehensive reputation characteristic value S are positive indicators (the larger the value, the better), use the same standardization operation:

[0465] ;

[0466] where x ij ' is the value of the i-th hospital / physician on the j-th evaluation indicator after standardization.

[0467] (3) For each evaluation indicator j, calculate the entropy value ej :

[0468] ;

[0469] where p ij represents the proportion of the standardized value of the i-th hospital / physician on the j-th evaluation index ( ); k is a constant equal to , where n is the number of hospitals / physicians to be evaluated.

[0470] (4) For each evaluation index j, calculate its coefficient of variation g j : g j = 1 - e j ;

[0471] (5) For each evaluation index j, calculate the corresponding weight w j , to quantify the contribution of this index to the comprehensive consideration factor CCF: ;

[0472] Specifically, in step S601, calculate the comprehensive consideration factor: Combine the treatment factor S aving and the comprehensive reputation eigenvalue S, and calculate the comprehensive consideration factor C CF by weighted summation: CCF = W5 * Saving + W6 * S;

[0473] Patients can select a hospital / physician based on the value of the comprehensive consideration factor. The higher the value, the better the hospital / physician performs in terms of medical technology, service quality, reputation, etc.

[0474] It can be understood that determining the weight by the entropy weight method avoids the subjectivity of artificial weighting, making the comprehensive consideration factor CCF more objective. At the same time, integrating the treatment factor and the comprehensive reputation eigenvalue comprehensively evaluates the performance of hospitals / physicians from multiple dimensions. Through standardization operations and weighted summation, quantitative data from different sources and of different natures are organically combined to form a unified comprehensive consideration factor CCF, which is convenient for patients to compare different hospitals / physicians. Moreover, the comprehensive consideration factor CCF provides an intuitive decision-making basis for patients. The higher its value, the better the hospital / physician performs in terms of medical technology, service quality, reputation, etc., which helps patients make a more informed choice.

[0475] Further, the python execution program for step S6 is as follows:

[0476] import numpy as np

[0477] n = 5 # Number of hospitals / physicians

[0478] Saving = np.array([85, 90, 78, 88, 92]) # Example data of treatment factors

[0479] S = np.array([80, 85, 75, 82, 88]) # Example data of comprehensive reputation eigenvalues

[0480] # Step S600: Determine weights

[0481] def calculate_weights(em):

[0482] # Data standardization

[0483] min_vals = np.min(em, axis=0)

[0484] max_vals = np.max(em, axis=0)

[0485] em_standardized = (em - min_vals) / (max_vals - min_vals)

[0486] # Calculate entropy value

[0487] k = 1 / np.log(n)

[0488] p = em_standardized / np.sum(em_standardized, axis=0)

[0489] e = -k * np.sum(p * np.log(p + 1e-10), axis=0) # Add 1e-10 to prevent log(0) situation

[0490] # Calculate difference coefficient and weights

[0491] g = 1 - e

[0492] w = g / np.sum(g)

[0493] return w

[0494] # Construct evaluation matrix EM

[0495] EM = np.column_stack((Saving, S))

[0496] # Calculate weights

[0497] weights = calculate_weights(EM)

[0498] W5, W6 = weights[0], weights[1]

[0499] # Step S601: Calculate the comprehensive consideration factor CCF

[0500] CCF = W5 * Saving + W6 * S

[0501] # Output the result

[0502] print("The weight W5 of the treatment factor Saving:", W5)

[0503] print("The weight W6 of the comprehensive reputation eigenvalue S:", W6)

[0504] print("The comprehensive consideration factor CCF:", CCF)

[0505] In the above program, the np.column_stack function is used to combine Saving and S into an evaluation matrix EM with n rows and 2 columns. A calculate_weights function is defined that accepts the evaluation matrix em as input and returns the weights of each evaluation index. Inside the function, first, data standardization operations are performed to eliminate the influence of dimensions. The np.min and np.max functions are used to calculate the minimum and maximum values of each evaluation index, and then the original data is converted into standardized data using the standardization formula.

[0506] Next, the entropy value is calculated. The np.sum function is used to calculate the sum of the standardized data for each evaluation index, and then the proportion p of the standardized value of each hospital / physician for each evaluation index is calculated. Finally, the entropy value e of each evaluation index is calculated using the entropy value formula. The difference coefficient g and the weight w are calculated. The difference coefficient g is 1 minus the entropy value e, and the weight w is the difference coefficient g divided by the sum of the difference coefficients.

[0507] Using the calculated weights W5 and W6, as well as the original treatment factor Saving and the comprehensive reputation eigenvalue S, the comprehensive consideration factor CCF is calculated by weighted summation.

[0508] Example 2: Based on step S302 of Example 1, this example further provides a preferred technical solution.

[0509] Considering that in practice, different disease types De have obvious seasonal characteristics (for example, the incidence rate of cardiovascular diseases in winter is significantly higher than that in other seasons), and at the same time, different disease types De have the distinction between acute and chronic. Therefore, the objective function in step S302 of the first embodiment can be further optimized.

[0510] (1) Introduce a time variable: Let t i be the onset or diagnosis time of patient i. Through this time variable, the starting point of the death rate curve D vc is located. Denote the death rate curve D vc as a function of time D vc (t). Then for patient i, the starting point of the corresponding death rate curve is D vc (t i ).

[0511] (2) Consider seasonal factors: Let S(t) be a seasonal factor function (such as a weight function or a penalty function) that reflects the seasonal variation of the incidence or mortality rate of disease type De at time t. It can be constructed based on historical data or expert knowledge.

[0512] (3) Optimized objective function: Combining the above two factors, the original objective function is optimized to:

[0513] ;

[0514] where d i,obs represents the actual death time (or survival time) of patient i. d i,pred (t i , S(t i )) represents the predicted death time (or survival time) of patient i according to the model, which depends on the onset or diagnosis time t i and the seasonal factor S(t i ) with weight properties. σ i is the standard deviation of the death time (or survival time) of patient i, which is used to measure the prediction error. λ is the weight coefficient of the penalty term. Penalty(G k ) is the penalty factor for group G k , which is used to avoid the group being too large or too small.

[0515] It can be understood that the seasonal factor function S(t) reflects the seasonal variation of the incidence or mortality rate of disease type De at time t. When predicting the death time (or survival time) of a patient, the model not only considers the individual characteristics of the patient but also the seasonal factor S(t i corresponding to the onset or diagnosis time t i) In this way, the model can more accurately reflect the impact of seasonal changes on the survival status of patients with different disease types.

[0516] Acute patients usually have a sudden onset and a short course of disease, while chronic patients have a longer course of disease. Among patients with different onset or diagnosis times, their survival status and death rate curves will be different. Therefore, by introducing the time variable t i and the seasonal factor S(t i ), the differences between acute and chronic patients can be captured. In other words, the main purpose of the objective function is to minimize the sum of the prediction error and the penalty term. By incorporating the seasonal factor and considering the individual characteristics of patients (including the type of disease course), the differences between acute and chronic patients can be reflected to a certain extent. For example, if the model can more accurately predict the survival status of patients with different disease courses, the value of the objective function will be smaller.

[0517] Furthermore, let L(j,k) denote the minimum value of the objective function for dividing the first j patients into k groups. A (n + 1)*(K + 1) dynamic programming table can be constructed to store these values. For each j = 2, 3, …, n and k = 2, 3, …, K, we have:

[0518] ;

[0519] where Cost(i + 1,j) represents the value of the objective function for dividing patients from i + 1 to j into one group.

[0520] After the dynamic programming table is filled, we can start backtracking from L(n,K) to find the optimal splitting point. The specific steps are as follows:

[0521] P1. Initialize an empty list partitions to store the indices of patients in each group.

[0522] P2. Let j = n and k = K.

[0523] P3. When k > 1, find the i that minimizes L(j,k) (i.e., the optimal splitting point).

[0524] P4. Add the indices of patients from i + 1 to j to the empty list partitions.

[0525] P5. Update j = i and k = k - 1, and repeat steps P3 - P4 above.

[0526] P6. When k = 1, add the remaining patients (i.e., the first j patients) to a new group and add their indices to partitions.

[0527] The Python program to implement the above solution is as follows:

[0528] def backtrack_dp_table(dp_table, split_points):

[0529] partitions = []

[0530] j, k = len(dp_table) - 1, len(dp_table[0]) - 1

[0531] while k > 1:

[0532] i = split_points[j, k] # Find the optimal split point

[0533] partitions.append(list(range(i + 1, j + 1))) # Add patient indices to the group

[0534] j, k = i, k - 1

[0535] partitions.append(list(range(1, j + 1))) # Add the remaining patients to a new group

[0536] partitions.reverse() # Since it backtracks from the end, the list needs to be reversed

[0537] # The indices returned here are 1-based. Subtract 1 for 0-based indexing

[0538] return [[idx - 1 for idx in group] for group in partitions]

[0539] # The dp_table and split_points are filled by the dynamic programming algorithm (see Example 1 for details)

[0540] dp_table =...

[0541] split_points =...

[0542] # Obtain the optimal partition in the following way

[0543] optimal_partitions = backtrack_dp_table(dp_table, split_points)

[0544] Embodiment 3: This embodiment further provides a standard deviation weight assignment method for weights W1, W2, W3, and W4 in step S401 of embodiment 1. This method assigns weights according to the degree of dispersion (i.e., standard deviation) of each feature sequence data. Features with greater dispersion have a more important position in the comprehensive evaluation:

[0545] S4010, for the feature sequences R1, R2, R3 and R4, calculate their standard deviations σ1, σ2, σ3 and σ4 respectively.

[0546] S4011, normalized standard deviation: Since the standard deviations of different feature sequences may have different dimensions, normalization is required. Divide each standard deviation by the sum of all standard deviations to obtain the normalized standard deviation ratio: ;where i=1,2,3,4, σ i ' is the normalized σ i .

[0547] S4012, assignment:

[0548] W1=σ1 / [σ1+σ2+σ3+σ4];

[0549] W2=σ2 / [σ1+σ2+σ3+σ4];

[0550] W3=σ3 / [σ1+σ2+σ3+σ4];

[0551] W4=σ4 / [σ1+σ2+σ3+σ4];

[0552] Among them, all the above expressions represent the characteristic sequence R i The standard deviation σ i The ratio of the feature sequence to the sum of the standard deviations of all feature sequences.

[0553] The scheme of this embodiment allocates weights based on the variability of the data, and considers that features with large variability are more important in the comprehensive evaluation. This method avoids the subjectivity of artificial weighting and ensures the objectivity and rationality of the weights. The weight allocation is completely based on the characteristics of the data and is not interfered by human factors, which improves the objectivity of the evaluation. It can also capture small changes in the data, which has a positive impact on the accuracy of the evaluation results.

[0554] Embodiment 4: Figure 3 As shown, this embodiment discloses a patient condition assessment system based on a multi-source information fusion mechanism: the system is used to implement the patient condition assessment method based on a multi-source information fusion mechanism described in all the above embodiments, and the system includes:

[0555] (1)User side, including an interaction interface; patients can call the reputation information of hospitals / physicians through this interaction interface.

[0556] (2)Server side, including:

[0557] (2.1)Data collection module responsible for automatically collecting the basic disease index datasets of patients from the hospital information system: having functions of interface development, data scraping, data cleaning and preprocessing.

[0558] (2.2)Data analysis module for applying principal component analysis to extract the comprehensive indicators of patients' basic diseases and applying ordered clustering analysis to evaluate the mortality risk of patients: having functions of PCA algorithm implementation, ordered clustering algorithm implementation (Fisher optimal segmentation method), time series analysis and curve fitting.

[0559] (2.3)Feature quantization and weight assignment module responsible for implementing the entropy weight method: having functions of formulating quantization standards, calculating weights by the entropy weight method and weight assignment strategies.

[0560] (2.4)Fusion module responsible for implementing the D-S evidence theory algorithm: using the Dempster combination rule to fuse the treatment factors (evaluation factors A and B) based on patients' basic diseases and mortality risk with the comprehensive reputation eigenvalue of the hospital to form a comprehensive consideration factor. Including functions of D-S evidence theory algorithm implementation, basic probability assignment (BPA) calculation, belief function and plausibility function calculation.

[0561] (2.5)Result interpretation module for evaluating the reputation of hospitals or physicians: interpreting the fused consideration factor, and the result interpretation module numerically converts the reputation information of the hospital / physician calculated in step S6 (i.e., the comprehensive consideration factor C CF )and feeds it back to the interaction interface of the user side.

[0562] (3)Data side: including the existing hospital information system, and the data collection module of the server side is communicatively connected to it and extracts relevant data information into the server side.

[0563] All the above embodiments only express the implementation manners of the related practical applications of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent should be subject to the appended claims.

[0564] For those skilled in the art, it can be further realized that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0565] At the same time, those skilled in the art can understand that all or part of the processes of implementing the methods of all the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in the present application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

Claims

1. A patient condition assessment method based on a multi-source information fusion mechanism, characterized in that: When a patient calls the medical reputation of a hospital / physician, after collecting the basic disease indicator dataset D from the hospital information system, the following steps are performed: S1, construct the basic disease indicator dataset D into a data matrix d with rows representing different historical patients and columns representing different indicators m ; S2, the data matrix d m After being converted into a covariance matrix C that describes the strength and direction of the linear relationship between the indicators, it is decomposed to obtain the eigenvalue E and the corresponding eigenvector F; according to the size of the eigenvalue E, the principal component P whose cumulative contribution rate reaches the preset threshold is selected; then the original data is projected onto a new feature space F consisting of the selected principal component P n , calculate the new coordinates P of each data point c And perform normalization, and weighted sum according to contribution rate to obtain evaluation factor A; S3, set the time window T of the patient in the historical medical record data D' w , and count the month of death M d ; For each historical patient, analyze different diseases De and their corresponding death time d t The distribution of the death rate curve D was obtained by time series analysis and curve fitting using the exponential model. vc Fisher optimal segmentation method according to the historical medical record data D' and the death rate curve D vc Clustering the patients, dividing n patients into K ordered groups, maximizing the similarity within the ordered groups and maximizing the difference between the ordered groups; quantifying the evaluation index, and calculating the evaluation factor B; S4, collects data on social reputation, treatment level, service attitude, and patient satisfaction to form a comprehensive reputation characteristic value S; S5, using the DS evidence theory algorithm to combine the evaluation factor A and the evaluation factor B to form a treatment factor S with the patient's basic disease assessment information and the patient's mortality assessment information aving ; The implementation method of the DS evidence theory algorithm includes: S500, define the identification framework I = {LR, MR, HR}, where LR represents low risk, MR represents medium risk, and HR represents high risk; then assign basic probabilities BPA to evaluation factors A and B, denoted as m and A and m B ; S501, use Dempster's combination rule to combine the basic probability m A and the basic probability m B Combine and calculate the normalization constant D: ; Wherein, X and Y are subsets of the identification framework I; symbol represents the empty set; The combined basic probability BPA value m A,B Calculated as: ; Wherein, Z is also a subset of the identification framework I; S502, calculate the trust function Bel and the likelihood function Pl to form the rescue factor S aving : I. First calculate the trust function Bel and the likelihood function Pl: ; ; II. Then, based on the values ​​of the trust function Bel and the likelihood function Pl, the rescue factor S is formed. aving ; The rules are: S aving = Bel(LR)-Pl(HR); Among them, Bel(LR) represents the confidence of low-risk propositions, and Pl(HR) represents the likelihood of high-risk propositions; S6, the rescue factor S aving Combined with the comprehensive reputation characteristic value S, it becomes a comprehensive consideration factor C that can help patients choose hospitals / doctors CF .

2. The method for assessing a patient's condition according to claim 1, characterized in that: In S2, the covariance matrix C is: ; in, Denotes the data matrix d m The mean vector of Denotes the data matrix d m The i-th row in , n is the number of patients; T represents the transpose operation of the matrix; In S2, the decomposition method is: C=F*E*F T ; Where T still represents the transpose operation of the matrix; In S2, the methods of the selection, the cumulative contribution rate, the projection, the normalization and the weighted summation are respectively as follows: I. The selection: Assume that k principal components are selected so that the cumulative contribution rate R C Reaching a preset threshold; II. The cumulative contribution rate R c : ; Among them, E jj represents the j-th diagonal element of the diagonal matrix E; m is the number of indices j; III. The projection: P c =d m *F k ; Among them, F k is the matrix composed of the eigenvectors corresponding to the selected k principal components; P c is the principal component score; F n is the new feature space; V. Normalization: ; Among them, P cn Represents the normalized value of the jth principal component score; VI. The weighted summation: ;in, Represents the sum of the variance contributions of the first k principal components.

3. The method for assessing a patient's condition according to claim 2, wherein: In S3, the time series analysis and curve fitting methods are as follows: I. The index model: D vc (t)=a·exp(b·t)+c; where a, b and c are the parameters to be fitted; t is the time variable; II. Use the least squares method to minimize the sum of squared differences between the observed values ​​and the model's predicted values: ; where N is the number of observations; y i is the actual death data, Dvc(t i ) is the predicted value of the exponential model; solving the least squares problem to obtain the values ​​of parameters a, b and c.

4. The method for assessing a patient's condition according to claim 3, characterized in that: In S3, the method of the Fisher optimal segmentation method includes: I. Let G1, G2, …, G K It is a segmentation that divides n patients into K groups, and defines the objective function L(n,K) as: ; Among them, n i is the i-th group G i The number of patients in It is group G i The feature mean vector of the patients in , is the feature mean vector of all patients; II. Define L(j,i) as the objective function value of the optimal segmentation of the first j patients into i groups, and then obtain the following recursive relationship: ; in, is the feature mean vector from patient s+1 to patient j; represents the difference between the feature mean vector from patient s+1 to patient j and the feature mean vector of all patients; represents the transpose of the difference vector; s represents a limit of the patient number; L(s,i−1) represents the objective function value of the optimal segmentation of the first s patients into i−1 groups; III. Initialization: when i=1, L(j,1)=0; V. Starting from i=2, gradually increase the number of groups until the predetermined number of groups K is reached; VI. Obtain the optimal partitioning of n patients into K groups; VII. Quantification: ; in, It is group G i The proportion of patients in the total number of patients, R i is the group G i risk score.

5. The method for assessing a patient's condition according to any one of claims 1 to 4, characterized in that: The implementation method of S6 includes: S600, using entropy weight method to determine the treatment factor S aving The weight W5 of and the weight W6 of the comprehensive reputation feature value S; S601, combined with rescue factor S aving and the comprehensive reputation characteristic value S, and the comprehensive consideration factor C is calculated by weighted summation CF :CCF=W5*Saving+W6*S.

6. The method for assessing a patient's condition according to claim 5, characterized in that: In S6, the implementation method of the entropy weight method includes: I. There are n hospitals / doctors to be evaluated. Each hospital / doctor has two evaluation indicators, namely, the treatment factor S aving and the comprehensive reputation characteristic value S; any element x of the evaluation matrix EM of n*2 ij represents the value of the i-th hospital / physician on the j-th evaluation index; II. Standardized operation: ; where x ij ′ is the standardized value of the i-th hospital / physician on the j-th evaluation index; III. For each evaluation index j, calculate its entropy value e j : ; Among them, p ij represents the proportion of the standardized value of the i-th hospital / physician on the j-th evaluation index; k is a constant; IV. For each evaluation index j, calculate its difference coefficient g j :g j =1−e j ; V. For each evaluation index j, calculate the corresponding weight w j : .

7. A system for implementing the patient condition assessment method according to any one of claims 1 to 6, characterized in that: include: The data collection module is responsible for automatically collecting the patient's basic disease indicator data set from the hospital information system; A data analysis module that uses principal component analysis to extract comprehensive indicators of patients' underlying diseases and an ordered cluster analysis to assess patients' mortality risk; Responsible for the feature quantization and weight allocation module that implements the entropy weight method; Responsible for implementing the fusion module of DS evidence theory algorithm; Results interpretation module to evaluate the reputation of a hospital or physician.

Citation Information

Patent Citations

  • Quality of service (QoS) and reputation based method for evaluating service trust levels

    CN103412918A

  • Reputation risk management capability assessment method and device, equipment and storage medium

    CN113570182A

  • Reputation risk monitoring and quantitative evaluation method, electronic equipment and computer readable storage medium

    CN114661860A

  • V2X node message credibility evaluation method and device based on reputation

    CN114945149A

  • Enterprise credit assessment method and system based on big data analysis

    CN116664012A