A data analysis method and device based on multi-dimensional feature similarity
By calculating multidimensional feature similarity and utilizing patient characteristics and historical case databases, similarity and posterior probability are determined, which solves the problems of poor interpretability and insufficient prediction of rare diseases in clinical decision support systems, and achieves higher diagnostic accuracy and rare disease prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PEKING UNION MEDICAL COLLEGE HOSPITAL
- Filing Date
- 2026-06-03
- Publication Date
- 2026-07-03
AI Technical Summary
Existing clinical decision support systems suffer from poor clinical interpretability, difficulty in utilizing historical case similarities, and insufficient predictive ability for rare diseases.
By calculating multidimensional feature similarity, we obtain the patient's basic personal information and disease information features, use specificity coefficients and weights to determine the weights in the historical case set, and calculate similarity and posterior probability, matching probability to improve interpretability and rare disease prediction ability.
It improves the interpretability of clinical decision support systems and reduces the probability of misdiagnosis, especially in predicting rare diseases.
Smart Images

Figure CN122337567A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically, to a data analysis method and apparatus based on multidimensional feature similarity. Background Technology
[0002] Clinical decision support systems are an important application area of medical artificial intelligence. Existing clinical decision support technologies are mainly divided into two categories: rule engine-based expert systems and machine learning-based predictive models.
[0003] Rule-based systems rely on manually written medical rule bases and have low utilization of historical cases. Machine learning-based methods (such as deep learning and random forests) can automatically learn complex patterns from data, but the models have poor interpretability, are considered black boxes, are difficult for clinicians to trust, and require a large amount of labeled data for training, and have insufficient predictive ability for rare diseases. Summary of the Invention
[0004] In view of this, embodiments of this application provide a data analysis method and apparatus based on multidimensional feature similarity to solve the technical problems of poor clinical interpretability, difficulty in utilizing historical case similarity, and insufficient predictive ability for rare diseases in the prior art.
[0005] In a first aspect, embodiments of this application provide a data analysis method based on multidimensional feature similarity, the method comprising: Obtain the patient's basic personal information and disease information; For each symptom information feature, the weight of the symptom information relative to the historical case set containing the personal basic information feature is determined based on the specificity coefficient of the personal basic information feature and the symptom information feature. Based on the disease information features and corresponding weights, determine the similarity between each historical case in the historical case set and the disease obtained by the patient; Based on the similarity, a preset number of target historical cases with the highest similarity are determined; Based on the first prior probability of the patient's condition in the historical case set, the first likelihood probability of the patient's condition in the historical case set, the second prior probability of the patient's condition in the target historical cases, and the second likelihood probability of the patient's condition in the target historical cases, the posterior probability of the patient's condition and each of the target historical cases is determined. Based on the posterior probability, the matching probability between the patient's condition and each of the target historical cases is obtained.
[0006] Secondly, embodiments of this application provide a data analysis apparatus based on multidimensional feature similarity, the apparatus comprising: The acquisition unit is used to acquire the patient's basic personal information characteristics and disease information characteristics; The first determining unit is used to determine the weight of each symptom information feature relative to the historical case set in which the personal basic information feature is located, based on the personal basic information feature and the specificity coefficient of the symptom information feature. The second determining unit is used to determine the similarity between each historical case in the historical case set and the patient's disease based on the disease information features and corresponding weights. The third determining unit is used to determine a preset number of target historical cases with the highest similarity based on the similarity. The fourth determining unit is used to determine the posterior probability of the patient's symptom and each of the target historical cases based on the first prior probability of the patient's symptom in the historical case set, the first likelihood probability of the patient's symptom in the historical case set, the second prior probability of the patient's symptom in the target historical cases, and the second likelihood probability of the patient's symptom in the target historical cases. The processing unit is used to obtain the matching probability between the patient's symptom and each of the target historical cases based on the posterior probability.
[0007] The technical solution provided in this application includes, but is not limited to, the following beneficial effects: In this application, by calculating the weighted similarity of multidimensional features, the most similar historical cases to the current patient are found. Then, based on the first prior probability, the first likelihood probability, the second prior probability, and the second likelihood probability, the posterior probability of the patient's condition and each target historical case is predicted and determined, that is, the degree of similarity between the patient's condition and each target historical case. Thus, the matching probability between the patient's condition and each target historical case is determined. Through the above method, historical cases can be effectively utilized, that is, historical cases can be used as diagnostic references, which improves interpretability. At the same time, it is also beneficial for the prediction of rare diseases and reduces the probability of misdiagnosis.
[0008] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0009] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 A flowchart illustrating a data analysis method based on multidimensional feature similarity provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of a data analysis device based on multidimensional feature similarity provided in an embodiment of this application. Detailed Implementation
[0011] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0012] Figure 1 A flowchart illustrating a data analysis method based on multidimensional feature similarity provided in this application embodiment is shown below. Figure 1 As shown, the method includes the following steps: Step 101: Obtain the patient's basic personal information characteristics and disease information characteristics.
[0013] Step 102: For each symptom information feature, determine the weight of the symptom information relative to the historical case set in which the personal basic information feature is located, based on the specificity coefficient of the personal basic information feature and the symptom information feature.
[0014] Step 103: Based on the disease information features and corresponding weights, determine the similarity between each historical case in the historical case set and the disease obtained by the patient.
[0015] Step 104: Based on the similarity, determine a preset number of target historical cases with the highest similarity.
[0016] Step 105: Determine the posterior probability of the patient's condition relative to each of the target historical cases based on the first prior probability of the patient's condition in the historical case set, the first likelihood probability of the patient's condition in the historical case set, the second prior probability of the patient's condition in the target historical cases, and the second likelihood probability of the patient's condition in the target historical cases.
[0017] Step 106: Based on the posterior probability, obtain the matching probability between the patient's condition and each of the target historical cases.
[0018] Specifically, the patient's medical records are first obtained, and then the patient's basic personal information and disease information features are obtained through the medical records. The basic personal information includes gender, age, height, weight, and underlying disease information, while the disease information features include body temperature, blood pressure, and duration of disease. For numerical features in the basic personal information and disease information features, such as age, height, weight, and blood pressure, these features are standardized to eliminate differences in the dimensions of different features, map various features to comparable intervals, and improve the stability and accuracy of the matching algorithm. For categorical features in the basic personal information and disease information features, such as gender and presence of symptoms, binary encoding is performed, e.g., male = 0, female = 1, presence of symptoms = 1, absence of symptoms = 0, and uncertainty of symptoms = 0.5. For temporal features in the basic personal information and disease information features, such as duration of disease, segmented discretization is performed, e.g., acute (<7 days) = 0, subacute (7-30 days) = 0.5, and chronic (>30 days) = 1.
[0019] Because the physical conditions corresponding to individual basic information vary with age, gender, height, and weight, it is necessary to determine a suitable set of historical cases for each patient based on their basic information characteristics. For example, if a patient is a 40-year-old male, 170 cm tall, and weighs 140 catties, the determined set of historical cases would include all historical cases of males aged 38-42, with heights of 165-175 cm and weights of 130-150 catties. The specificity coefficient reflects the ability of a feature to distinguish diseases; that is, it is an indicator of the ability to identify non-cases in diagnostic tests. Furthermore, it reflects the ability to correctly exclude "disease-free" individuals. Through the specificity coefficient, different weights can be dynamically assigned to different features within the historical case set. These weights represent the degree of attention doctors pay to a particular feature in clinical practice.
[0020] Then, by using the multidimensional symptom information features of the patient and their corresponding weights, the similarity between each historical case in the historical case set and the patient's symptom is determined. In this way, target historical cases with the highest matching degree between a certain number of patients and each historical case are selected, thereby providing a reference for subsequent diagnosis.
[0021] By inferring the cause from the effect using conditional probabilities such as the first prior probability, the first likelihood probability, the second prior probability, and the second likelihood probability, the posterior probability of the disease obtained by each target historical case and the patient can be determined. By screening target historical cases, the amount of data processing can be effectively reduced, and the data can be more targeted.
[0022] After obtaining the posterior probability, the matching probability between the patient's condition and each of the target historical cases can be determined based on the candidate probability, thereby providing a reference for the doctor's subsequent diagnosis. The above method can effectively utilize historical cases, that is, use historical cases as a diagnostic reference, which improves interpretability and is also conducive to the prediction of rare diseases, reducing the probability of misdiagnosis.
[0023] In a feasible implementation, the disease information features include: physiological parameter features, disease classification features, and disease duration features.
[0024] It should be noted that the specific disease information features used can be set according to actual needs, and no specific restrictions are made here.
[0025] In one feasible implementation, when performing the step of determining the weight of the disease information relative to the historical case set containing the personal basic information feature based on the specificity coefficient of the personal basic information feature and the disease information feature, firstly, the historical case database to which the personal basic information feature belongs is determined based on the personal basic information feature; then, the specificity coefficient of the disease information feature is determined based on the ratio of the total number of cases in the historical case database containing the disease information feature to the total number of cases included in the historical case database; finally, the weight is calculated using the following formula: Formula 1 in, This is the specificity coefficient of the disease information features. The importance score for the information features of this disease is given, with a value range of [1, 5]. For the patient's first The specificity coefficient of each disease information feature relative to the historical case database. For the patient's first The importance score of individual disease information features The number of features of the disease information.
[0026] Specifically, in clinical practice, doctors pay different levels of attention to different symptoms and signs. This difference depends on two factors: 1. Statistical characteristics of the data: the distribution pattern of features in the disease; 2. Clinical knowledge: the actual value of features in diagnosis, treatment, and prognosis. Among these, the specificity coefficient represents data-driven, the importance score represents knowledge-driven, and the weight represents the diagnostic value of the feature. By multiplying and fusing the specificity coefficient and the importance score, prominent key features can be amplified, thereby assigning different weights to different features.
[0027] The ratio of the total number of cases in the historical case database that exhibit the symptom information feature to the total number of cases included in the historical case database represents the probability of the symptom information feature appearing in all diseases included in the historical case database. The difference between this ratio and the numerical value is then calculated and used as the specificity coefficient of the symptom information feature, i.e., the specificity coefficient of the symptom information feature relative to the historical case database. For example, if the historical case database includes 10,000 cases, and fever appears in 6,000 cases, then the specificity coefficient of the feature "fever" relative to the historical case database is 0.4, indicating a common symptom with low specificity. In extreme cases, such as when the number of occurrences of a certain feature is 0, the specificity coefficient is 1, indicating that the feature has complete specificity. In practical processing, a lower threshold can be set, such as a specificity coefficient ≥ 0.01, to prevent excessive amplification of noise. If the number of occurrences of a certain feature is 10,000, the specificity coefficient is 0, indicating that the feature has no distinguishing value. In practical processing, a specificity coefficient of 0.01 is set as the minimum value to avoid a weight of 0.
[0028] Taking features including fever, rusty sputum, and elevated white blood cell count as examples, the specificity coefficient for fever is 0.4, and the importance score is 2.2; the specificity coefficient for rusty sputum is 0.98, and the importance score is 4.2; the specificity coefficient for elevated white blood cell count is 0.70, and the importance score is 3.3. At this point, the product of the specificity coefficient and importance score for fever is 0.88; the product of the specificity coefficient and importance score for rusty sputum is 4.12; and the product of the specificity coefficient and importance score for elevated white blood cell count is 2.31. The sum of these three features is 0. 0.88 + 4.12 + 2.31 = 7.31, normalized to get the final weights: Fever: 0.88 / 7.31 = 0.12, Coughing up rusty sputum: 4.12 / 7.31 = 0.56, Elevated white blood cell count: 2.31 / 7.31 = 0.32. Coughing up rusty sputum has the highest weight (0.56) because of its high specificity (0.98) and high clinical importance (4.2). Fever has the lowest weight (0.12) because although its clinical importance is acceptable, its specificity is too low. Elevated white blood cell count has a medium weight (0.32), with both specificity and importance being in the middle range.
[0029] Regarding the specificity coefficient Update using the following formula: ; in, This represents the current specificity coefficient. This indicates the specificity of the recalculation based on newly added cases. This is a smoothing coefficient, typically set to 0.9-0.95.
[0030] It should be noted that the importance score of the disease information features can be set according to the importance score of clinical experts. The specific setting method can be set according to actual needs, and no specific limitation is made here.
[0031] In a feasible implementation, when performing the step of determining the similarity between each historical case in the historical case set and the patient's symptom based on the symptom information features and corresponding weights, the similarity can be determined according to the following formula: Formula 2 in, For the patient's first Individual disease information characteristics, For the historical case database, the first The first case in the Individual disease information characteristics, For the first A similarity calculation function for individual disease information features. For the first The weights of individual disease information features.
[0032] For example, for numerical characteristics such as age, body temperature, and blood pressure, as follows: ; in, The first in historical cases Individual disease information characteristics, For the first The range of values for each symptom information feature in the historical case database; Symptom Presence Similarity Function as follows: ; For pain intensity characteristics such as grade and severity as follows: ; in, The first in historical cases The maximum possible strength value of a disease information feature.
[0033] It should be noted that the specific similarity function used for different features can be set according to actual needs, and no specific restrictions are made here.
[0034] In a feasible implementation, when determining the posterior probability of the patient's condition relative to each of the target historical cases based on a first prior probability of the patient's condition in the historical case set, a first likelihood probability of the patient's condition in the historical case set, a second prior probability of the patient's condition in the target historical cases, and a second likelihood probability of the patient's condition in the target historical cases, the first prior probability of the patient's condition in the historical case set, the first likelihood probability of the patient's condition in the historical case set, and the second prior probability of the patient's condition in the target historical cases, and the second likelihood probability of the patient's condition in the target historical cases, can be input into a Bayesian formula to obtain the posterior probability of the patient's condition relative to each of the target historical cases. The specific methods for obtaining the prior probabilities and likelihood probabilities, and the calculation methods for the posterior probabilities, can be found in existing technologies and will not be elaborated upon here.
[0035] In a feasible implementation, when performing the step of obtaining the matching probability between the patient's condition and each of the target historical cases based on the posterior probability, the matching probability is obtained using the following Formula 3: Formula 3 in, , Let be the posterior probability. The slope parameter for logistic regression fitting of the logit model (used for probability-to-log odds conversion). The intercept parameter is used for logistic regression fitting of the logit model. Parameter a is the slope, controlling the discriminative power of the probability predictions; |a| < 1 indicates the original model is overconfident, and |a| > 1 indicates the original model is overly conservative. Parameter b is the intercept, controlling the overall bias of the probability predictions; b > 0 indicates the original probability is systematically underestimated, and b < 0 indicates the original probability is systematically overestimated.
[0036] Since the posterior probability has a systematic bias, Formula 3 is used to calibrate the posterior probability. The purpose of calibration is to make the output probability value reflect the true risk of disease. Calibration solves the problem of correcting the deviation between the predicted probability and the actual observed frequency.
[0037] In one feasible implementation, the method further includes: outputting target historical cases with a matching probability greater than a preset threshold as a reference.
[0038] Figure 2 A schematic diagram of the structure of a data analysis device based on multidimensional feature similarity provided in this application embodiment is shown below. Figure 2 As shown, the device includes: Acquisition unit 21 is used to acquire the patient's basic personal information features and disease information features; The first determining unit 22 is used to determine the weight of the disease information relative to the historical case set in which the personal basic information feature is located, based on the personal basic information feature and the specificity coefficient of the disease information feature for each disease information feature. The second determining unit 23 is used to determine the similarity between each historical case in the historical case set and the patient's disease based on the disease information features and corresponding weights. The third determining unit 24 is used to determine a preset number of target historical cases with the highest similarity based on the similarity. The fourth determining unit 25 is used to determine the posterior probability of the patient's symptom and each of the target historical cases based on the first prior probability of the patient's symptom in the historical case set, the first likelihood probability of the patient's symptom in the historical case set, the second prior probability of the patient's symptom in the target historical cases, and the second likelihood probability of the patient's symptom in the target historical cases. Processing unit 26 is used to obtain the matching probability between the patient's symptom and each of the target historical cases based on the posterior probability.
[0039] In a feasible implementation, the disease information features include: physiological parameter features, disease classification features, and disease duration features.
[0040] In one feasible implementation, when the first determining unit determines the weight of the disease information relative to the historical case set containing the personal basic information feature based on the specificity coefficient of the personal basic information feature and the disease information feature, the method includes: Based on the aforementioned basic personal information characteristics, determine the historical case database to which the aforementioned basic personal information characteristics belong; The specificity coefficient of the disease information feature is determined by the ratio of the total number of cases with the disease information feature in the historical case database to the total number of cases included in the historical case database. The weights are calculated using the following formula: Formula 1 in, This is the specificity coefficient of the disease information features. The importance score for the information features of this disease is given, with a value range of [1, 5]. For the patient's first The specificity coefficient of each disease information feature relative to the historical case database. For the patient's first The importance score of individual disease information features The number of features of the disease information.
[0041] In one feasible implementation, when the second determining unit determines the similarity between each historical case in the historical case set and the patient's symptom based on the symptom information features and corresponding weights, it includes: The similarity is determined according to the following formula: Formula 2 in, For the patient's first Individual disease information characteristics, For the historical case database, the first The first case in the Individual disease information characteristics, For the first A similarity calculation function for individual disease information features. For the first The weights of individual disease information features.
[0042] In one feasible implementation, the fourth determining unit, when determining the posterior probability of the patient's symptom relative to each of the target historical cases based on a first prior probability of the patient's symptom in the historical case set, a first likelihood probability of the patient's symptom in the historical case set, a second prior probability of the patient's symptom in the target historical cases, and a second likelihood probability of the patient's symptom in the target historical cases, includes: The first prior probability of the patient's condition in the historical case set, the first likelihood probability of the patient's condition in the historical case set, the second prior probability of the patient's condition in the target historical cases, and the second likelihood probability of the patient's condition in the target historical cases are input into the Bayesian formula to obtain the posterior probability of the patient's condition and each of the target historical cases.
[0043] In one feasible implementation, when the processing unit is used to obtain the matching probability between the patient's symptom and each of the target historical cases based on the posterior probability, it includes: The matching probability is obtained using the following formula three: Formula 3 in, , Let be the posterior probability. The slope parameter for logistic regression fitting of the logit model. The intercept parameter for logistic regression fitting of the logit model.
[0044] In one feasible implementation, the device further includes: The output unit is used to output target historical cases with a matching probability greater than a preset threshold as a reference.
[0045] about Figure 2 For detailed principles related to this content, please refer to [link / reference]. Figure 1 Detailed explanations of the relevant content will not be repeated here.
[0046] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0047] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0048] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0049] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0050] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0051] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.
Claims
1. A data analysis method based on multi-dimensional feature similarity, characterized in that, The method includes: Obtain the patient's basic personal information and disease information; For each symptom information feature, the weight of the symptom information relative to the historical case set containing the personal basic information feature is determined based on the specificity coefficient of the personal basic information feature and the symptom information feature. Based on the disease information features and corresponding weights, determine the similarity between each historical case in the historical case set and the disease obtained by the patient; Based on the similarity, a preset number of target historical cases with the highest similarity are determined; Based on the first prior probability of the patient's condition in the historical case set, the first likelihood probability of the patient's condition in the historical case set, the second prior probability of the patient's condition in the target historical cases, and the second likelihood probability of the patient's condition in the target historical cases, the posterior probability of the patient's condition and each of the target historical cases is determined. Based on the posterior probability, the matching probability between the patient's condition and each of the target historical cases is obtained.
2. The data analysis method of claim 1, wherein, The disease information features include: physiological parameter features, disease classification features, and disease duration features.
3. The data analysis method of claim 1, wherein, The step of determining the weight of the disease information relative to the historical case set containing the personal basic information features based on the specificity coefficients of the personal basic information features and the disease information features includes: Based on the aforementioned basic personal information characteristics, determine the historical case database to which the aforementioned basic personal information characteristics belong; The specificity coefficient of the disease information feature is determined by the ratio of the total number of cases with the disease information feature in the historical case database to the total number of cases included in the historical case database. The weights are calculated using the following formula: ; Equation One in, This is the specificity coefficient of the disease information features. The importance score for the information features of this disease is given, with a value range of [1, 5]. For the patient's first The specificity coefficient of each disease information feature relative to the historical case database. For the patient's first The importance score of individual disease information features The number of features of the disease information.
4. The data analysis method of claim 1, wherein, The step of determining the similarity between each historical case in the historical case set and the patient's condition based on the disease information features and corresponding weights includes: The similarity is determined according to the following formula: ; Equation Two wherein, is a jth feature of the jth medical condition information of the patient, is a jth feature of the jth medical condition information of the patient, is a jth feature of the jth medical condition information of the jth case in the historical case database, is a jth feature of the jth medical condition information of the jth case in the historical case database, is a jth feature of the jth medical condition information of the jth case in the historical case database, is a similarity calculation function of the jth medical condition information, is a similarity calculation function of the jth medical condition information, is a weight of the jth medical condition information, is a weight of the jth medical condition information.
5. The data analysis method as described in claim 1, characterized in that, The step of determining the posterior probability of the patient's condition relative to each of the target historical cases based on the first prior probability of the patient's condition in the historical case set, the first likelihood probability of the patient's condition in the historical case set, the second prior probability of the patient's condition in the target historical cases, and the second likelihood probability of the patient's condition in the target historical cases includes: The first prior probability of the patient's condition in the historical case set, the first likelihood probability of the patient's condition in the historical case set, the second prior probability of the patient's condition in the target historical cases, and the second likelihood probability of the patient's condition in the target historical cases are input into the Bayesian formula to obtain the posterior probability of the patient's condition and each of the target historical cases.
6. The data analysis method as described in claim 1, characterized in that, The step of obtaining the matching probability between the patient's condition and each of the target historical cases based on the posterior probability includes: The matching probability is obtained using the following formula three: Formula 3 in, , Let be the posterior probability. The slope parameter for logistic regression fitting of the logit model. The intercept parameter for logistic regression fitting of the logit model.
7. The data analysis method as described in claim 1, characterized in that, The method further includes: Output target historical cases with a matching probability greater than a preset threshold as a reference.
8. A data analysis device based on multidimensional feature similarity, characterized in that, The device includes: The acquisition unit is used to acquire the patient's basic personal information characteristics and disease information characteristics; The first determining unit is used to determine the weight of each symptom information feature relative to the historical case set in which the personal basic information feature is located, based on the personal basic information feature and the specificity coefficient of the symptom information feature. The second determining unit is used to determine the similarity between each historical case in the historical case set and the patient's disease based on the disease information features and corresponding weights. The third determining unit is used to determine a preset number of target historical cases with the highest similarity based on the similarity. The fourth determining unit is used to determine the posterior probability of the patient's symptom and each of the target historical cases based on the first prior probability of the patient's symptom in the historical case set, the first likelihood probability of the patient's symptom in the historical case set, the second prior probability of the patient's symptom in the target historical cases, and the second likelihood probability of the patient's symptom in the target historical cases. The processing unit is used to obtain the matching probability between the patient's symptom and each of the target historical cases based on the posterior probability.
9. The data analysis device as described in claim 8, characterized in that, The disease information features include: physiological parameter features, disease classification features, and disease duration features.
10. The data analysis apparatus as described in claim 8, characterized in that, When the first determining unit determines the weight of the disease information relative to the historical case set containing the personal basic information feature based on the specificity coefficient of the personal basic information feature and the disease information feature, the determination includes: Based on the aforementioned basic personal information characteristics, determine the historical case database to which the aforementioned basic personal information characteristics belong; The specificity coefficient of the disease information feature is determined by the ratio of the total number of cases with the disease information feature in the historical case database to the total number of cases included in the historical case database. The weights are calculated using the following formula: Formula 1 in, This is the specificity coefficient of the disease information features. The importance score for the information features of this disease is given, with a value range of [1, 5]. For the patient's first The specificity coefficient of each disease information feature relative to the historical case database. For the patient's first The importance score of individual disease information features The number of features of the disease information.