Interpretable ordered classification system for predicting prognosis of pulmonary embolism patient
By employing a multi-stage ordered classification modeling framework and interpretable design, the problems of transparency and class imbalance in the prognostic assessment of pulmonary embolism patients are resolved, achieving closed-loop support for predictive accuracy and individualized intervention, thereby enhancing the credibility and practicality of clinical applications.
Patent Information
- Application Number
- CN202511337371.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-12-05
AI Technical Summary
Existing technologies for prognostic assessment of pulmonary embolism patients suffer from opaque decision-making processes, class imbalances, and high-dimensional sparsity, making it difficult to form a closed-loop support system from risk identification to intervention implementation, thus affecting model trust and clinical acceptance.
A multi-stage integrated ordered classification modeling framework is adopted, including a data processing module, an interpretable ordered classifier module, and a clinical decision support system module. Through feature engineering, global and local interpretable modeling, and class imbalance handling, it generates visual explanations and individualized intervention suggestions.
It improves the accuracy and transparency of prognostic risk prediction, enhances the interpretability of the model, forms a closed loop of clinical decision-making integrating prediction and intervention, and improves the support capability for personalized diagnosis and treatment.
Smart Images

Figure CN121075619A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical data processing and clinical decision support, and in particular to an interpretable ordered classification system for predicting the prognosis of patients with pulmonary embolism. Background Technology
[0002] Pulmonary embolism, a common thrombotic disease, poses a significant threat to patients' lives due to its diverse clinical manifestations and rapid progression. Accurate prognosis assessment is crucial for guiding treatment strategies and optimizing the allocation of medical resources. Currently, various risk stratification tools (such as the PESI score and sPESI score) have been developed for prognostic evaluation in clinical practice. Furthermore, with the advancement of medical informatization, machine learning-based intelligent prediction tools are gradually being integrated into the field of clinical decision support. Utilizing multi-source clinical data to improve the automation and precision of risk assessment has become an important development direction in this field.
[0003] However, existing technologies still have significant limitations in practical applications: First, mainstream prediction models mostly rely on "black box" models such as deep neural networks and random forests, whose decision-making processes lack transparency, making it difficult for clinicians to trace the basis of risk assessments, thus limiting model trust and clinical acceptance; Second, clinical data from pulmonary embolism patients often suffer from class imbalance (such as a low proportion of severe prognostic events) and high-dimensional sparsity, which traditional models have limited ability to handle, easily leading to prediction bias or insufficient stability; Third, existing tools mostly remain at the "prognostic prediction" stage, failing to effectively connect the decision-making logic of "predictive results" and "clinical intervention," making it difficult to form a closed-loop support system from risk identification to intervention implementation, thus limiting their comprehensive value in clinical practice. Summary of the Invention
[0004] The purpose of this invention is to overcome one or more shortcomings of the prior art and provide an interpretable ordered classification system for predicting the prognosis of patients with pulmonary embolism.
[0005] The objective of this invention is achieved through the following technical solution:
[0006] An interpretable ordered classification system for predicting the prognosis of patients with pulmonary embolism includes:
[0007] Data processing module, interpretable ordered classifier module, and clinical decision support system module;
[0008] The data processing module is used to acquire multidimensional clinical data of patients with pulmonary embolism, and form a structured dataset suitable for interpretable modeling through feature engineering and data preprocessing.
[0009] The interpretable ordered classifier module is used to perform prognostic risk level prediction based on the structured dataset using a multi-stage fusion ordered classification modeling framework. The modeling framework includes an interpretable feature ranking and selection mechanism, a modeling method that combines global and local interpretability, and a class imbalance handling mechanism.
[0010] The clinical decision support system module is used to output the patient's prognostic risk level, provide a visual explanation of risk factors, and generate clinical intervention recommendations.
[0011] Furthermore, the multidimensional clinical data acquired by the data processing module comes from at least one of the following sources: the electronic health records of pulmonary embolism patients, regional clinical data resource platforms, or public machine learning databases, covering demographic characteristics, imaging data, laboratory test results, and medical history.
[0012] Furthermore, the interpretability feature ranking and selection mechanism is used for feature processing of the multi-category distribution of pulmonary embolism prognosis, which includes improvement and absorption, conversion to chronic thrombosis, and death.
[0013] Furthermore, the modeling approach that combines global and local interpretability includes interpretable structure modeling based on rule set mining.
[0014] Furthermore, the modeling approach that combines global and local interpretability includes an ordered label modeling strategy based on ranking learning.
[0015] Furthermore, the modeling approach that combines global and local interpretability includes integrating counterfactual analysis and attention mechanisms to enhance the model's ability to perceive and interpret key variables.
[0016] Furthermore, the class imbalance handling mechanism is an undersampling strategy that combines K-means clustering and hierarchical density clustering to improve the model's ability to identify high-risk classes.
[0017] Furthermore, the risk factor visualization interpretation provided by the clinical decision support system module includes at least one of SHAP value and counterfactual analysis path.
[0018] Furthermore, the clinical intervention recommendations generated by the clinical decision support system module include at least one of the following: adjustment of follow-up frequency and adjustment of anticoagulant drug strategy.
[0019] Furthermore, the clinical decision support system module is deployed in the hospital information system or regional health management platform and applied to scenarios such as inpatient admission assessment, monitoring of changes in condition, and discharge follow-up management.
[0020] The beneficial effects of this invention are:
[0021] (1) By standardizing and feature engineering multi-source clinical data, the adaptability of data to interpretability modeling can be improved;
[0022] (2) By using an interpretable ordered classification modeling framework (integrating rule set mining, ranking learning, counterfactual analysis and class imbalance handling), the goal of balancing the accuracy of prognostic risk prediction with the transparency of model decision-making is achieved.
[0023] (3) By designing a closed-loop clinical auxiliary decision-making system that integrates prediction and intervention, the system’s ability to support individualized diagnosis and treatment and its clinical applicability are improved. Attached Figure Description
[0024] Figure 1 This is a diagram illustrating the structure of an interpretable ordered classification system for predicting the prognosis of patients with pulmonary embolism, provided in an embodiment of the present invention. Detailed Implementation
[0025] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] Example 1:
[0027] An interpretable ordered classification system for predicting the prognosis of patients with pulmonary embolism is provided, comprising:
[0028] Data processing module, interpretable ordered classifier module, and clinical decision support system module;
[0029] The data processing module is used to acquire multidimensional clinical data of patients with pulmonary embolism, and form a structured dataset suitable for interpretable modeling through feature engineering and data preprocessing.
[0030] The interpretable ordered classifier module is used to perform prognostic risk level prediction based on the structured dataset using a multi-stage fusion ordered classification modeling framework. The modeling framework includes an interpretable feature ranking and selection mechanism, a modeling method that combines global and local interpretability, and a class imbalance handling mechanism.
[0031] The clinical decision support system module is used to output the patient's prognostic risk level, provide a visual explanation of risk factors, and generate clinical intervention recommendations.
[0032] The multidimensional clinical data acquired by the data processing module comes from at least one of the following sources: electronic health records of pulmonary embolism patients, regional clinical data resource platforms, or public machine learning databases. It covers demographic characteristics, imaging data, laboratory test results, and medical history.
[0033] The interpretable feature ranking and selection mechanism is used to perform feature processing on the multi-category distribution of pulmonary embolism prognosis, which includes improvement and absorption, conversion to chronic thrombosis, and death.
[0034] The modeling approach that combines global and local interpretability includes interpretable structure modeling based on rule set mining.
[0035] The modeling approach that combines global and local interpretability includes an ordered label modeling strategy based on ranking learning.
[0036] The modeling approach that combines global and local interpretability includes integrating counterfactual analysis and attention mechanisms to enhance the model's ability to perceive and interpret key variables.
[0037] The class imbalance handling mechanism is an undersampling strategy that combines K-means clustering and hierarchical density clustering, used to improve the model's ability to identify high-risk classes.
[0038] The clinical decision support system module provides a visual interpretation of risk factors, including at least one of the following: SHAP value and counterfactual analysis path.
[0039] The clinical intervention recommendations generated by the clinical decision support system module include at least one of the following: adjustment of follow-up frequency and adjustment of anticoagulant drug strategy.
[0040] The clinical decision support system module is deployed in the hospital information system or regional health management platform and is applied to scenarios such as inpatient admission assessment, monitoring of changes in condition, and discharge follow-up management.
[0041] Example 2:
[0042] This embodiment provides an interpretable ordered classification system for predicting the prognosis of patients with pulmonary embolism. It aims to achieve a closed loop of the entire process of "data processing - risk prediction - decision support" through modular design, taking into account both prediction accuracy and interpretability, and assisting clinical intelligent decision-making from "risk identification" to "personalized intervention".
[0043] See Figure 1 This system comprises three core modules: a data processing module, an interpretable ordered classifier module, and a clinical decision support system module. These modules form an organic whole through data flow, and the specific architecture is as follows: Figure 1As shown in the diagram. The data processing module provides standardized input to the entire system; the interpretable ordered classifier module performs prognostic risk prediction based on the input data; and the clinical decision support system module transforms the prediction results into clinically understandable information and suggestions. These three modules work together to achieve a complete transformation from raw data to clinical decision-making. The main modules include:
[0044] (a) Data Processing Module:
[0045] The data processing module is the system's "data entry point," responsible for transforming multi-source, heterogeneous clinical data into structured datasets suitable for modeling. It comprises three sub-units: a data acquisition unit, a feature engineering unit, and a data preprocessing unit. Specific functions are as follows:
[0046] Data Acquisition Unit: This unit is used to collect multidimensional clinical data from patients with pulmonary embolism, including but not limited to:
[0047] Patient electronic health records (such as inpatient medical records, outpatient records, nursing records, etc.);
[0048] Regional-level clinical data resource platform (a regional medical database that integrates diagnosis and treatment data from multiple institutions);
[0049] Publicly available machine learning databases (such as standardized research datasets containing clinical characteristics and prognostic outcomes of PE patients).
[0050] The collected data covers four main categories of features:
[0051] Demographic characteristics: age, sex, body mass index (BMI), smoking history, etc.
[0052] Imaging data: CT angiography of the pulmonary artery (CTPA) showing the location and extent of the thrombus, and echocardiographic parameters (such as right ventricular diameter and pulmonary artery systolic pressure);
[0053] Laboratory test results: D-dimer level, coagulation function indicators (PT, APTT), myocardial enzyme profile (troponin), liver and kidney function indicators, etc.
[0054] Medical history: underlying diseases (such as hypertension, diabetes, malignant tumors), history of thrombosis, history of anticoagulant drug use, etc.
[0055] The data acquisition process connects with various data sources through standardized interfaces (such as the HL7 FHIR protocol) to ensure the security and compatibility of data transmission, and strictly follows medical data privacy protection standards (such as de-identification processing).
[0056] Feature Engineering Unit: This unit extracts, transforms, and filters features from the raw data to uncover key information related to PE prognosis. Specific steps include:
[0057] Feature extraction: Extracting structured features from unstructured data (such as extracting discrete features like "history of deep vein thrombosis in the lower extremities" and "recent surgical history" from medical record texts; and extracting binary features like "thrombosis involving the main pulmonary artery" from imaging reports).
[0058] Feature transformation: binning is performed on continuous features (such as D-dimer levels and age) (e.g., age is divided into <40 years, 40-60 years, and >60 years), and trend features are extracted from time series data (such as multiple laboratory test results) (e.g., "D-dimer change rate within 72 hours").
[0059] Feature selection: Based on domain knowledge (such as PE risk factors defined in clinical guidelines) and statistical methods (such as variance inflation factor (VIF) to remove multicollinearity and mutual information value (MIV) to screen highly relevant features), features that are significantly related to prognosis are retained, and redundant information is reduced.
[0060] Data preprocessing unit: This unit cleans and standardizes the feature-engineered dataset to form a structured dataset suitable for interpretable modeling. Specific operations include:
[0061] Data cleaning: Remove duplicate records, correct logical errors (such as outliers like "age=-5"), and mark missing values (distinguishing between "no record" and "not detected");
[0062] Missing value handling: Differentiated strategies are adopted for different feature types (e.g., continuous features are filled with the median, discrete features are filled with the mode, and key clinical indicators are filled with multiple imputation to preserve data distribution characteristics).
[0063] Standardization processing: Z-score standardization is performed on continuous features (making the mean 0 and the standard deviation 1), and one-hot encoding is performed on discrete features (such as converting "hypertension" in "medical history" into a 0 / 1 binary feature) to ensure that features of different magnitudes are weighted fairly in modeling.
[0064] After the above processing, the data processing module outputs a structured dataset in the format of a two-dimensional table (rows represent patient samples, and columns represent preprocessed features), which provides input to the interpretable ordered classifier module.
[0065] (ii) Interpretable Ordered Classifier Module:
[0066] This module is the system's "core prediction engine." Based on the structured dataset output by the data processing module, it employs a multi-stage fusion ordered classification modeling framework to predict prognostic risk levels. It includes three sub-units: a feature ranking and selection unit, a global-local interpretable modeling unit, and a class imbalance handling unit. The specific implementation is as follows:
[0067] Feature sorting and selection unit:
[0068] This unit addresses the multi-category distribution of PE prognosis (including three ordered categories: "improvement and absorption," "conversion to chronic thrombosis," and "death") by constructing an interpretable feature ranking and selection mechanism. The specific process is as follows:
[0069] Prior weights for the importance of features are defined based on clinical domain knowledge (e.g., "elevated troponin" has a higher weight in PE prognosis than "gender").
[0070] Tree-based feature importance assessment (e.g., calculating the contribution of features to decision tree splitting) and statistical tests (e.g., ANOVA analysis of variance) are used to quantify the association between features and prognostic categories.
[0071] By combining prior weights and quantization results, a feature ranking list is generated, and the Top N key features (e.g., N=20) are selected. This reduces the modeling dimensionality while retaining variables that have a significant impact on the prognosis, ensuring the interpretability of subsequent modeling.
[0072] Global-Local Interpretable Modeling Unit: This unit adopts a modeling approach that combines "global interpretability" and "local interpretability," balancing the overall logical transparency of the model with the traceability of individual predictions. It comprises three sub-units:
[0073] (1) Rule set mining sub-unit: Based on the filtered features, an interpretable rule set is constructed, expressing the prognostic judgment logic in the form of "If-Then". For example:
[0074] Rule 1: If "pulmonary artery trunk thrombosis" = Yes, "systolic blood pressure < 90 mmHg" = Yes, and "troponin > 0.5 ng / mL" = Yes, then the predicted category is "death";
[0075] Rule 2: If "pulmonary thrombosis below the segment level only" = yes, "D-dimer < 500 ng / mL" = yes, and "no history of malignant tumors" = yes, then the predicted category is "improvement and absorption". The rule set is mined from the training data using a greedy algorithm to ensure coverage of the majority of samples and no conflicts between rules, allowing clinicians to directly understand the model's global decision-making logic.
[0076] (2) Ranking Learning Subunit: To address the ordered nature of PE prognostic categories ("improvement and absorption" < "transition to chronic thrombosis" < "death"), a ranking learning (Learning to Rank) strategy is employed for modeling. By defining a "risk score function," patient characteristics are mapped to continuous risk values, and then prognostic categories are divided based on risk value thresholds. For example:
[0077] Risk value <0.3 → Improvement and absorption;
[0078] 0.3 ≤ risk value < 0.7 → Turning into chronic thrombosis;
[0079] Risk value ≥ 0.7 → Death.
[0080] (3) Counterfactual analysis and attention sub-unit: Integrating counterfactual analysis and attention mechanisms enhances the model's ability to perceive key variables and provide local explanations.
[0081] The parameters of the risk score function are determined by optimizing the "ranking loss" (such as pairwise loss) to ensure the model's sensitivity to class order and improve the accuracy of ordered classification.
[0082] Attention mechanism: assign dynamic weights to each feature (e.g., "systolic blood pressure" has a higher weight in patients with low blood pressure), and use a visual heatmap to show the contribution of each feature to individual predictions;
[0083] Counterfactual analysis: Generate counterfactual samples for individual prediction results (e.g., "If a patient's systolic blood pressure rises to 100 mmHg, will the risk category change?"), trace the threshold effect of key features, and help doctors understand "how to adjust variables to improve prognosis".
[0084] Class Imbalance Handling Unit: Addressing the issue of insufficient samples (class imbalance) for high-risk categories such as "death" in clinical data, an undersampling strategy combining K-means clustering and Hierarchical Density Clustering (HDBSCAN) is employed. The specific steps are as follows:
[0085] Perform K-means clustering on the majority class samples (e.g., “improved absorption”), dividing the samples into K clusters (e.g., K=5), with each cluster representing a subset of the data distribution;
[0086] HDBSCAN was used to perform density analysis on each cluster, retaining core samples within the cluster (high density and strong representativeness) and removing edge samples (which may be noise or outliers).
[0087] By adjusting the sampling ratio, the proportion of samples in the three categories of "improvement and absorption", "chronic thrombosis" and "death" in the processed dataset is close to 1:1:1. This avoids the ethical risks of generating artificially synthesized samples through traditional oversampling (such as SMOTE) while ensuring the model's ability to learn high-risk categories.
[0088] (III) Clinical Decision Support System Module:
[0089] This module serves as the system's "clinical interface," transforming the classifier module's predictions into clinically actionable information. It comprises four sub-units: a risk level output unit, a visualization interpretation unit, an intervention suggestion generation unit, and a system deployment interface. Specific functions are as follows:
[0090] Risk level output unit: Receives the prediction results from the interpretable ordered classifier module and outputs the patient's prognostic risk level in an intuitive format. For example:
[0091] The patient's prognostic risk level is displayed in text form as "High risk (predicted outcome: death)"; it is also displayed in color-coded labels as "Red (high risk), Yellow (medium risk), Green (low risk)"; and it includes the following risk probabilities: "Probability of death: 85%, probability of developing into chronic thrombosis: 12%, probability of improvement and absorption: 3%".
[0092] Visual Explanation Unit: This unit presents the basis for risk prediction in a visual way, helping clinicians understand the model's decision-making logic. Specific formats include:
[0093] The contribution of each feature to the current prediction is shown by a bar chart of SHAP values (positive values are risk-promoting factors, and negative values are protective factors). For example, "Elevated troponin (SHAP value = 0.6) is the main cause of high risk".
[0094] Through counterfactual analysis, a flowchart is used to show "the trajectory of risk level changes if a certain feature (such as increasing the dose of anticoagulant drugs) is adjusted". For example, "If the dose of anticoagulant drugs is increased from 5 mg / day to 10 mg / day, the probability of death can be reduced from 85% to 40%".
[0095] Rule matching diagram: Highlights the key rules in the rule set that the current patient is matched with (such as "pulmonary artery trunk thrombosis + hypotension" in rule 1) to clarify the basis for prediction.
[0096] Intervention Recommendation Generation Unit: Based on risk level and explanatory information, and in conjunction with clinical guidelines, it generates individualized intervention recommendations, specifically including:
[0097] Follow-up frequency adjustment: For high-risk patients, it is recommended to "check D-dimer and myocardial enzyme profile every 6 hours", for medium-risk patients, it is recommended to "check daily", and for low-risk patients, it is recommended to "check every 3 days".
[0098] Anticoagulation strategy adjustment: Recommend drug type and dosage based on risk factors (such as renal function and history of bleeding). For example, "For high-risk patients with no history of bleeding, intravenous unfractionated heparin (initial dose 80 U / kg) is recommended" and "For intermediate-risk patients with renal insufficiency, low molecular weight heparin (1 mg / kg, once daily) is recommended".
[0099] Referral and monitoring recommendations: High-risk patients are recommended to be transferred to the ICU for monitoring; medium-risk patients are recommended to be monitored by a cardiology specialist; and low-risk patients are recommended to be observed in a general ward.
[0100] System Deployment Interface: Provides standardized interfaces to connect the system with clinical information systems, supporting multiple application scenarios; can be integrated into hospital information systems (HIS), electronic medical record systems (EMR), or regional health management platforms; application scenarios include automatically collecting data and outputting initial risk levels upon patient admission to assist in developing admission treatment plans; real-time synchronization of patient vital signs and examination results, dynamic updating of risk levels, and triggering early warnings (such as automatically reminding doctors when the risk level rises from low to high); generating follow-up plans based on the risk level at discharge (such as monthly follow-ups for high-risk patients and every 3 months for low-risk patients), and recording follow-up results for model iteration.
[0101] The system's workflow forms a closed loop of "data input - model prediction - decision output - clinical feedback," with the specific steps as follows:
[0102] 1) Data Acquisition and Preprocessing: The data processing module collects multidimensional clinical data from patients through the data acquisition unit, extracts key features through the feature engineering unit, and then cleans and standardizes the data through the data preprocessing unit to output a structured dataset.
[0103] 2) Risk Prediction: The interpretable ordered classifier module receives a structured dataset, filters key features through the feature sorting and selection unit, performs ordered classification by the global-local interpretable modeling unit (combining rule set, sorting learning, and counterfactual analysis), and outputs the patient's prognostic risk level and feature importance information after optimization by the class imbalance processing unit.
[0104] 3) Decision support output: The clinical decision support system module receives the risk level and feature importance, displays the results through the risk level output unit, provides the basis for prediction through the visualization interpretation unit, outputs individualized plans through the intervention suggestion generation unit, and finally pushes them to the clinical terminal through the system deployment interface;
[0105] 4) Clinical feedback and model iteration: Doctors evaluate the system's recommendations based on actual clinical outcomes (such as the patient's final prognosis) and provide feedback data (such as intervention effects and prediction bias) to the data processing module to update the training set and achieve continuous optimization of the model.
[0106] To verify the effectiveness of this system, a comparative trial was conducted using multicenter clinical data (including 5000 PE patients, of whom 3000 improved and resolved, 1500 had chronic thrombosis, and 500 died) with existing mainstream models. The results are as follows:
[0107] The system's multi-class classification accuracy is 89.2%, which is higher than XGBoost (82.5%), Random Forest (81.3%), and Logistic Regression (76.8%). The sensitivity for identifying the "death" category is 91.5%, which is significantly higher than traditional models (75%-82%), indicating that it has a better ability to identify high-risk patients.
[0108] Through a questionnaire survey of 100 attending physicians and above, the "decision comprehensibility" score of this system (4.7 / 5) was significantly higher than that of the "black box" model (2.3 / 5), and 87% of the doctors believed that the system's explanation could help them formulate intervention plans.
[0109] The accuracy fluctuation on data with different centers (where the data distribution is different) is less than 3%, while XGBoost and Random Forest fluctuate by 5%-8%, indicating that this system is more robust to changes in data distribution.
[0110] In pilot hospitals, the average length of hospital stay for patients using this system was shortened by 1.2 days, and the mortality rate of high-risk patients was reduced by 12.3%, validating its clinical value.
[0111] This embodiment discloses an interpretable ordered classification system for predicting the prognosis of pulmonary embolism patients. Through a data processing module, it standardizes multi-source data; an interpretable ordered classifier module balances prediction accuracy and transparency; and a clinical decision support system module connects prediction and intervention, forming a complete clinical decision-making loop. This system addresses the problems of insufficient interpretability, weak handling of class imbalance, and lack of intervention coordination in existing technologies. It provides a reliable tool for precise risk stratification and individualized diagnosis and treatment of pulmonary embolism patients, and has significant clinical application value.
[0112] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.
Claims
1. An interpretable ordered classification system for predicting the prognosis of patients with pulmonary embolism, characterized in that, include: Data processing module, interpretable ordered classifier module, and clinical decision support system module; The data processing module is used to acquire multidimensional clinical data of patients with pulmonary embolism, and form a structured dataset suitable for interpretable modeling through feature engineering and data preprocessing. The interpretable ordered classifier module is used to perform prognostic risk level prediction based on the structured dataset using a multi-stage fusion ordered classification modeling framework. The modeling framework includes an interpretable feature ranking and selection mechanism, a modeling method that combines global and local interpretability, and a class imbalance handling mechanism. The clinical decision support system module is used to output the patient's prognostic risk level, provide a visual explanation of risk factors, and generate clinical intervention recommendations.
2. The system according to claim 1, characterized in that, The multidimensional clinical data acquired by the data processing module comes from at least one of the following sources: electronic health records of pulmonary embolism patients, regional clinical data resource platforms, or public machine learning databases. It covers demographic characteristics, imaging data, laboratory test results, and medical history.
3. The system according to claim 1, characterized in that, The interpretable feature ranking and selection mechanism is used to perform feature processing on the multi-category distribution of pulmonary embolism prognosis, which includes improvement and absorption, conversion to chronic thrombosis, and death.
4. The system according to claim 1, characterized in that, The modeling approach that combines global and local interpretability includes interpretable structure modeling based on rule set mining.
5. The system according to claim 1, characterized in that, The modeling approach that combines global and local interpretability includes an ordered label modeling strategy based on ranking learning.
6. The system according to claim 1, characterized in that, The modeling approach that combines global and local interpretability includes integrating counterfactual analysis and attention mechanisms to enhance the model's ability to perceive and interpret key variables.
7. The system according to claim 1, characterized in that, The class imbalance handling mechanism is an undersampling strategy that combines K-means clustering and hierarchical density clustering, used to improve the model's ability to identify high-risk classes.
8. The system according to claim 1, characterized in that, The clinical decision support system module provides a visual interpretation of risk factors, including at least one of the following: SHAP value and counterfactual analysis path.
9. The system according to claim 1, characterized in that, The clinical intervention recommendations generated by the clinical decision support system module include at least one of the following: adjustment of follow-up frequency and adjustment of anticoagulant drug strategy.
10. The system according to claim 1, characterized in that, The clinical decision support system module is deployed in the hospital information system or regional health management platform and is applied to scenarios such as inpatient admission assessment, monitoring of changes in condition, and discharge follow-up management.