Configuration method and device of pulmonary embolism thrombolysis risk prediction model and processing equipment

Through the enhancement and labeling of deep learning models and electronic health file data, a targeted pulmonary embolism thrombolysis risk prediction model was trained, which solved the problem of unstable accuracy of thrombolysis risk prediction in the existing technology, achieved high-precision risk prediction, and provided high-quality support for clinical decision-making.

CN119993451APending Publication Date: 2025-05-13TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202411885316.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Prior art predicts unstable or limited accuracy when predicting thrombolysis risk in patients with hemodynamicly stable pulmonary embolism.

Method used

A targeted pulmonary embolism thrombolysis risk prediction model is trained through deep learning models, combining data augmentation and labeling of electronic health profiles (EHR) data. The model includes obtaining sample data, performing data augmentation and annotation, and training deep learning models to predict thrombolysis risk.

Benefits of technology

It has achieved a highly targeted and adaptive thrombolysis risk prediction in patients with hemodynamic stability, improved prediction accuracy, provided high-quality data assistance services for clinical decision-making, and reduced treatment risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993451A_ABST
    Figure CN119993451A_ABST
Patent Text Reader

Abstract

The invention provides a configuration method and device of a pulmonary embolism thrombolysis risk prediction model and processing equipment, and aims to provide a specific pulmonary embolism thrombolysis risk prediction model configuration scheme on the basis of deep learning. The pulmonary embolism thrombolysis risk prediction model obtained through configuration has high pertinence and adaptability for the pulmonary embolism patient with the stable hemodynamics, and compared with general model configuration logic, the prediction of the thrombolysis risk of the pulmonary embolism patient with the stable hemodynamics can be accurately achieved, and the prediction accuracy of the thrombolysis risk of the pulmonary embolism patient with the stable hemodynamics is improved. And high-quality data auxiliary service is provided for clinical decision making, so that the treatment risk is reduced, and high-quality medical service is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of medical technology, and specifically to a configuration method, device and processing equipment for a pulmonary embolism thrombolysis risk prediction model. Background Art

[0002] Pulmonary embolism (PE) refers to a clinical and pathophysiological syndrome of pulmonary circulation disorder caused by endogenous or exogenous thrombi (most commonly blood clots) blocking the pulmonary artery or its branch vessels. Systemic thrombolytic therapy is the standard treatment for patients with pulmonary embolism. This method dissolves the formed blood clots through intravenous injection of thrombolytic drugs, thereby rapidly restoring pulmonary artery blood flow, alleviating symptoms, reducing pulmonary artery pressure and increasing blood oxygen saturation.

[0003] However, in clinical practice, the benefits of systemic thrombolytic therapy for hemodynamically stable patients may be limited or not obvious, and it is accompanied by increased risks.

[0004] In this context, the inventors of the present application discovered that when using a corresponding deep learning model to predict the thrombolytic risk faced by patients with pulmonary embolism, for patients with hemodynamically stable pulmonary embolism, the thrombolytic risk prediction model configured using a general model configuration logic has the problem of unstable or limited accuracy in thrombolytic risk prediction for patients with hemodynamically stable pulmonary embolism. Summary of the invention

[0005] The present application provides a configuration method, device and processing equipment for a pulmonary embolism thrombolysis risk prediction model, which is used to provide a specific pulmonary embolism thrombolysis risk prediction model configuration scheme based on deep learning. The pulmonary embolism thrombolysis risk prediction model thus configured is highly targeted and adaptable to patients with hemodynamically stable pulmonary embolism. Compared with the general model configuration logic, it can accurately predict the thrombolysis risk of patients with hemodynamically stable pulmonary embolism, provide high-quality data-assisted services for clinical decision-making, and thereby promote the reduction of treatment risks and obtain high-quality medical services.

[0006] In a first aspect, the present application provides a method for configuring a pulmonary embolism thrombolysis risk prediction model, the method comprising:

[0007] Acquire first sample EHR data corresponding to a sample pulmonary embolism patient, wherein the sample pulmonary embolism patient is specifically a pulmonary embolism patient who has received systemic thrombolytic therapy;

[0008] Performing data enhancement operation on the first sample EHR data, and retaining the data involved in the target feature during the enhancement process, to obtain the second sample EHR data, wherein the target feature includes three aspects of features: basic information feature, medical history feature, and laboratory result feature;

[0009] Annotating the second sample EHR data, wherein the annotation result includes whether death occurred after systemic thrombolytic therapy was administered;

[0010] Based on the annotated second sample EHR data, a pulmonary embolism thrombolysis risk prediction model is trained, wherein the pulmonary embolism thrombolysis risk prediction model is a deep learning model, which is used to predict the death risk of the corresponding patient after systemic thrombolysis therapy based on the EHR data input into the model.

[0011] In a second aspect, the present application provides a configuration device for a pulmonary embolism thrombolysis risk prediction model, the device comprising:

[0012] an acquisition unit, configured to acquire first sample EHR data corresponding to a sample pulmonary embolism patient, wherein the sample pulmonary embolism patient is specifically a pulmonary embolism patient who has received systemic thrombolytic therapy;

[0013] An operation unit is used to perform a data enhancement operation on the first sample EHR data, and retain the data involved in the target feature during the enhancement process to obtain the second sample EHR data, wherein the target feature includes three aspects of features: basic information feature, medical history feature and laboratory result feature;

[0014] a labeling unit, configured to label the second sample EHR data, wherein the labeling result includes whether death occurs after systemic thrombolytic therapy is applied;

[0015] A training unit is used to train a pulmonary embolism thrombolysis risk prediction model based on the annotated second sample EHR data, wherein the pulmonary embolism thrombolysis risk prediction model is a deep learning model, and the pulmonary embolism thrombolysis risk prediction model is used to predict the death risk of the corresponding patient after systemic thrombolytic therapy based on the EHR data input by the model.

[0016] In a third aspect, the present application provides a processing device, including a processor and a memory, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the method provided in the first aspect of the present application or any possible implementation method of the first aspect of the present application is executed.

[0017] In a fourth aspect, the present application provides a computer-readable storage medium, which stores multiple instructions, and the instructions are suitable for a processor to load to execute the method provided by the first aspect of the present application or any possible implementation of the first aspect of the present application.

[0018] From the above content, it can be concluded that the present application has the following beneficial effects:

[0019] With the goal of predicting the thrombolytic risk of pulmonary embolism patients, this application provides a specific pulmonary embolism thrombolytic risk prediction model configuration scheme based on deep learning. The pulmonary embolism thrombolytic risk prediction model thus configured is highly targeted and adaptable to hemodynamically stable pulmonary embolism patients. Compared with the general model configuration logic, it can accurately predict the thrombolytic risk of hemodynamically stable pulmonary embolism patients, provide high-quality data-assisted services for clinical decision-making, and thereby reduce treatment risks and obtain high-quality medical services. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0021] Figure 1 A schematic diagram of a process for configuring a pulmonary embolism thrombolysis risk prediction model for this application;

[0022] Figure 2 A logical schematic diagram of the XGBoost model of this application;

[0023] Figure 3 A schematic diagram of an example of performance evaluation of the model of this application;

[0024] Figure 4 A schematic diagram of another example of the performance evaluation of the model of this application;

[0025] Figure 5 This is a schematic diagram of an example of the explanatory analysis based on SHAP in this application;

[0026] Figure 6 This is another example schematic diagram of the explanatory analysis based on SHAP in this application;

[0027] Figure 7 A schematic diagram of a configuration device for a pulmonary embolism thrombolysis risk prediction model of the present application;

[0028] Figure 8 A schematic diagram of the structure of the processing equipment of this application. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0030] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or inherent to these processes, methods, products or devices. The naming or numbering of steps in this application does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The process steps that have been named or numbered can change the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved.

[0031] The division of modules in this application is a logical division. There may be other division methods when it is implemented in actual applications. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection between modules can be electrical or other similar forms, which are not limited in this application. In addition, the modules or submodules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed in multiple circuit modules, and some or all of the modules may be selected according to actual needs to achieve the purpose of the present application.

[0032] Before introducing the configuration method of the pulmonary embolism thrombolysis risk prediction model provided by the present application, the background content involved in the present application is first introduced.

[0033] The configuration method, device and computer-readable storage medium of the pulmonary embolism thrombolysis risk prediction model provided in the present application can be applied to processing equipment to provide a specific pulmonary embolism thrombolysis risk prediction model configuration scheme based on deep learning. The pulmonary embolism thrombolysis risk prediction model thus configured is highly targeted and adaptable to patients with hemodynamically stable pulmonary embolism. Compared with the general model configuration logic, it can accurately predict the thrombolysis risk of patients with hemodynamically stable pulmonary embolism, provide high-quality data-assisted services for clinical decision-making, and thereby promote the reduction of treatment risks and obtain high-quality medical services.

[0034] The configuration method of the pulmonary embolism thrombolysis risk prediction model mentioned in this application can be executed by a configuration device of the pulmonary embolism thrombolysis risk prediction model, or a server, physical host or user equipment (UE) and other different types of processing devices that integrate the configuration device of the pulmonary embolism thrombolysis risk prediction model. Among them, the configuration device of the pulmonary embolism thrombolysis risk prediction model can be implemented in hardware or software, and the UE can be a terminal device such as a smart phone, tablet computer, laptop computer, desktop computer or personal digital assistant (PDA), and the processing device can be set in the form of a device cluster.

[0035] It is understandable that, considering that the present application scheme is mainly a data processing based on the existing electronic health record (EHR) data, the processing equipment that executes the configuration method of the pulmonary embolism thrombolysis risk prediction model of the present application, or the processing equipment equipped with the application service corresponding to the configuration method of the pulmonary embolism thrombolysis risk prediction model of the present application, in actual applications, only needs to have the required data processing capabilities, and can be any type of processing equipment such as a server, a physical host or even a UE.

[0036] If the specific collection of EHR data or the application of pulmonary embolism thrombolysis risk prediction model is also involved, the processing equipment needs to continue to be adaptively configured in terms of equipment type and deployment method.

[0037] As an example, the processing device may specifically include a first device part responsible for collecting EHR data, a second device part for background model training, and a third device part for on-site model application.

[0038] Next, the configuration method of the pulmonary embolism thrombolysis risk prediction model provided by the present application is introduced.

[0039] First, see Figure 1 , Figure 1A flow chart of a configuration method of a pulmonary embolism thrombolysis risk prediction model of the present application is shown. The configuration method of a pulmonary embolism thrombolysis risk prediction model provided by the present application may specifically include the following steps S101 to S104:

[0040] Step S101, obtaining first sample EHR data corresponding to a sample pulmonary embolism patient, wherein the sample pulmonary embolism patient is specifically a pulmonary embolism patient who has received systemic thrombolytic therapy;

[0041] Corresponding to the training requirements of the pulmonary embolism thrombolysis risk prediction model of this application, the early stage will involve the configuration of corresponding training samples. For this, the initial sample, that is, the first sample EHR data of the sample pulmonary embolism patient, can be obtained here.

[0042] For the convenience of explanation, this application records EHR data in different states as first sample EHR data, second sample EHR data and target EHR data.

[0043] It can be understood that the data source behind the first sample EHR data obtained here is relatively flexible. It can be obtained by manual entry, or it can be obtained locally on the device or retrieved from other devices.

[0044] In actual applications, hospitals will involve many different online business systems, and these different online business systems store the first sample EHR data that can be obtained here. Specifically, EHR data covers medical documents including medical record summaries, examination reports, test results, medication records, etc. These documents can be indexed by unique identifiers such as assigned patient identification codes and medical treatment identifiers. This ensures that the identity of each patient can be uniquely confirmed. For the original EHR data, keyword matching and other technologies can be used to extract information.

[0045] Step S102, performing a data enhancement operation on the first sample EHR data, and retaining the data involved in the target feature during the enhancement process, to obtain the second sample EHR data, wherein the target feature includes three aspects of features: basic information feature, medical history feature, and laboratory result feature;

[0046] It is understandable that the relatively original EHR data obtained earlier, that is, the first sample EHR data, cannot be directly put into the model training phase in the present application scheme. The data quality can be enhanced through the data enhancement phase here, thus laying a good foundation for subsequent more convenient and higher-quality model training.

[0047] Among them, while enhancing the data quality through a series of data enhancement operations, it can be noted that specific feature extraction processing or feature filtering processing is also involved here.

[0048] In the pulmonary embolism thrombolytic risk prediction logic designed in this application, the three characteristics of basic information characteristics, medical history characteristics and laboratory result characteristics of pulmonary embolism patients have made a more outstanding contribution to the risks associated with systemic thrombolytic therapy for pulmonary embolism patients.

[0049] In other words, when analyzing how to better quantify the risks associated with pulmonary embolism patients after systemic thrombolytic therapy, this application filters out key features based on a series of potential features and forms three aspects of features: basic information features, medical history features, and laboratory result features. This facilitates efficient and high-quality processing of features in collection, data retrieval, data management, etc.

[0050] In specific operations, if the feature extraction processing for these three aspects of features is independent of the data enhancement operation, or is considered to be two operations, the feature extraction processing can be performed first, and then the data enhancement operation, or the data enhancement operation can be performed first and then the feature extraction processing; if the feature extraction processing is also considered to be a data enhancement operation, the specific execution order between different data enhancement operations can be configured according to actual needs.

[0051] In data enhancement operations other than feature extraction processing, it is understandable that specific data enhancement has relatively mature related technologies, so existing technologies can be directly adopted. Of course, improved solutions of existing technologies or even self-developed novel solutions can also be adopted, which are all acceptable.

[0052] As an exemplary embodiment here, the data enhancement operation involved in this application may specifically include:

[0053] 1. The first operation of removing samples with a missing rate exceeding 10%;

[0054] It is understandable that the EHR data that has been carefully combed above may also have many missing items in actual situations. Therefore, for each sample, the missing items can be counted to determine the corresponding missing rate. When it exceeds 10, deletion measures can be taken to ensure the data integrity of the samples used in specific model training and to ensure the model training effect.

[0055] 2. A second operation of removing samples whose indicator values ​​deviate from the normal range by a greater degree than a preset degree;

[0056] It can be understood that the present application believes that in order to retain the valid information contained in the sample to the maximum extent, a prudent strategy can be adopted to deal with outliers in the sample.

[0057] Specifically, only data points that obviously deviate from the normal range in the indicators manually or systematically recorded by doctors are excluded. These abnormal data points may be caused by factors such as instrument failure, operational errors of medical staff or negligence during recording. The indicators involved may include systolic blood pressure, diastolic blood pressure, respiratory rate, pulse and body temperature. For data points that are higher / lower than the normal range but the degree of deviation is less than the preset degree (degree threshold), it can be considered that these data fluctuations actually reflect the normal changes in the patient's vital characteristics to a certain extent. Therefore, these data points can be retained, and the data points whose indicator values ​​deviate from the normal range by more than the preset degree are eliminated, which further enhances the details and effectiveness of the sample in terms of details.

[0058] 3. The third operation of unifying the numerical units involved in the sample;

[0059] It can be understood that the unification and standardization of numerical units will help with the management and use of data, and thus help ensure good model training results.

[0060] 4. The fourth operation uses both the SMOTENC sampling method and the ClusterCentroids sampling method to balance the sample distribution;

[0061] It is understandable that in actual situations, samples collected from real patients may have problems with unbalanced sample distribution. For example, the number of dead individuals in EHR data may be significantly less than that of non-dead individuals, such as a ratio of approximately 1:60. The imbalance in data distribution may cause the model to tend to ignore those small numbers of death samples during model training. In order to alleviate this unbalanced sample distribution, the present application may use the SMOTENC sampling method and the ClusterCentroids sampling method to sample the samples, so as to achieve the purpose of balancing the sample distribution and avoid the interference of unbalanced sample distribution on the model training effect.

[0062] Among them, SMOTENC, namely Synthetic Minority Over-sampling Technique with Neighborhood Components, is based on synthetic minority over-sampling of neighborhood components; ClusterCentroids, namely cluster centers.

[0063] The SMOTENC sampling method and the ClusterCentroids sampling method are two existing sampling methods. Considering that the algorithm itself is not the focus of this application solution, a detailed explanation is not given here.

[0064] 5. The fifth operation of filling the missing values ​​in the sample using the MissForest method.

[0065] It can be understood that for missing values ​​in samples, the present application believes that in the case of the first operation of removing samples with a missing rate of more than 10% introduced above, there will still be samples with a small amount of missing content remaining. Although this is within an acceptable range, for these samples, the filling method can continue to be used to supplement the missing content to avoid samples with missing content from bringing adverse effects to model training, which can be specifically accomplished by using the MissForest method.

[0066] Among them, the MissForest method is a data interpolation method based on the random forest algorithm. Considering that the algorithm itself is not the focus of this application, it will not be described in detail here.

[0067] As an example, training samples can be divided into training sets and test sets in some model training architectures. When using the MissForest algorithm to fill missing values, the data of the training set and the test set can be filled separately to avoid data leakage.

[0068] As for the three aspects of basic information characteristics, medical history characteristics and laboratory result characteristics, this application finally identified 71 characteristics for analysis, including 15 basic information characteristics, 10 medical history characteristics and 46 laboratory result characteristics.

[0069] Correspondingly, as an exemplary embodiment, the 15 basic information features may specifically include:

[0070] 1. Gender, 2. Age, 3. Smoking history, 4. Drinking history, 5. Blood transfusion history, 6. Allergy history, 7. Pregnancy history, 8. Surgery history, 9. Past medical history, 10. Infectious diseases, 11. Family clustering diseases, 12. Abortion history, 13. Chemotherapy history, 14. Targeted therapy history, 15. Endocrine therapy history; 11. Body temperature

[0071] The diseases involved in the 10 medical history characteristics may include:

[0072] 1. Coronary heart disease, 2. Diabetes, 3. Hypertension, 4. Hyperlipidemia, 5. Cerebrovascular disease, 6. Lower limb vascular disease, 7. Kidney disease, 8. Cancer, 9. Fracture or trauma, 10. Coagulation system disease;

[0073] The 46 laboratory result characteristics may specifically include:

[0074] 1. Systolic blood pressure (mmHg), 2. Diastolic blood pressure (mmHg), 3. Pulse (bpm), 4. Temperature (℃), 5. Respiratory rate (bpm), 6. Corrected calcium (mmol / L), 7. Alkaline phosphatase (IU / L), 8. Total cholesterol (mmol / L), 9. Albumin (g / L), 10. Chloride (mmoL / L), 11. Direct bilirubin (umol / L), 12. γ-glutamyl transpeptidase (IU / L), 1 3. Bicarbonate (mmoL / L), 14. Globulin (g / L), 15. Potassium (mmoL / L), 16. TBIL*0.8, 17. Sodium (mmol / L), 18. White / globulin ratio, 19. Calcium (mmoL / L), 20. Total bilirubin (umoL / L), 21. Creatinine (umoL / L), 22. Urea (mmol / L), 23. Total protein (g / L), 24. Uric acid (umoL / L), 2 5. Basophils (%), 26. Monocytes (%), 27. Platelet count (10^9 / L), 28. Monocytes (10^9 / L), 29. Average hemoglobin concentration (g / L), 30. Average hemoglobin content (pg), 31. Lymphocytes (10^9 / L), 32. Eosinophils (%), 33. Basophils (10^9 / L), 34. Hematocrit (%), 35. Hemoglobin (g / L), 36 .Lymphocytes (%), 37. Average RBC volume (fL), 38. WBC count (10^9 / L), 39. Eosinophils (10^9 / L), 40. Neutrophils (10^9 / L), 41. Average PLT volume (fL), 42. Neutrophils (%), 43. Red blood cell count ( / 10^12 / L), 44. Platelet volume (%), 45. PLT distribution width (FL), 46. TP*0.75.

[0075] It can be seen that the embodiments here are the three aspects of the features that can be involved in the pulmonary embolism thrombolysis risk prediction logic designed by the present application, and provide a specific implementation plan, namely the series of specific indicators listed above, which have better practical significance. Among them, the above series of indicators themselves are indicators that can be involved in the clinical work of the hospital, but it can be understood that they should not be directly understood as the stacking of existing indicators. Although there are a small number of indicators that doctors do consider when conducting thrombolysis risks in actual clinical practice, there are also a large number of indicators that are pre-mined by the present application through the configured data mining method. These indicators would actually be directly abandoned and ignored in actual clinical work. The present application captures and mines these indicators in response to the model processing requirements, thereby matching a more delicate and more targeted indicator system for the thrombolysis risk prediction target of the present application.

[0076] Among them, the data mining methods mentioned here may specifically involve the introduction of artificial intelligence (AI) technology, so as to realize deep-level indicator mining and processing through powerful model computing capabilities, so as to obtain corresponding indicators that can contribute to thrombolytic risk prediction at the subtle data processing level, and even in terms of details, it is possible to obtain an indicator set in which a single indicator itself will not make a significant contribution but will make a significant contribution in combination with other indicators. Therefore, it has significant sensitivity and pertinence for the thrombolytic risk prediction goal of this application.

[0077] Step S103, annotating the second sample EHR data, wherein the annotated result includes whether the person died after receiving systemic thrombolytic therapy;

[0078] After obtaining the simplified second sample EHR data, the labeling process can be carried out. It can be understood that the labeling can also be understood by the model prediction true value, that is, the prediction result under the theoretical state. In this way, during the model training process, the model training is guided by the labeling of the training samples.

[0079] Among them, it should be noted that the thrombolytic risk marked in the present application is specifically the risk of death, that is, the thrombolytic risk of pulmonary embolism made in the present application is specifically the risk of whether a patient with pulmonary embolism will die after systemic thrombolytic therapy is applied. This corresponds to the technical defect that the thrombolytic risk prediction model configured by the general model configuration logic in the prior art has unstable thrombolytic risk prediction accuracy for patients with hemodynamically stable pulmonary embolism.

[0080] In specific operations, the annotation process can be completed by manual annotation or by automated annotation tools, wherein the automated annotation tools need to be pre-configured with corresponding automated annotation logic / strategies.

[0081] In this way, after completing the labeling process for the second sample EHR data here, a training sample that can be used for specific model training is obtained.

[0082] As an example, in some model training architectures, training samples can be divided into training sets and test sets, and the ratio of the two can be 4:1, as shown in the following table:

[0083] Table 1 - Training sample examples

[0084] Initial death data Training set Test Set Training set (after sampling) Not dead 3098 2481 619 1240 die 52 40 12 1240

[0085] Step S104, based on the annotated second sample EHR data, training a pulmonary embolism thrombolysis risk prediction model, wherein the pulmonary embolism thrombolysis risk prediction model is a deep learning model, and the pulmonary embolism thrombolysis risk prediction model is used to predict the death risk of the corresponding patient after systemic thrombolytic therapy based on the EHR data input into the model.

[0086] For the pulmonary embolism thrombolysis risk prediction model, the thrombolysis risk prediction task based on EHR data can be regarded as a binary classification task.

[0087] The model input is EHR data, denoted as X input ={x1,x2,......,x n}, n represents the number of patients involved, x i ={x i1 ,x i2 ,.......,x if},i=1,2,...,n,x i is the eigenvector of the i-th individual, and f is the dimension of the eigenvector.

[0088] The goal of the thrombolytic risk prediction task is to analyze the probability of death caused by thrombolytic therapy in patients with pulmonary embolism and record it as Youtput, that is, to learn a classification function f H (·) to map the patient's feature vector to a death probability value between 0 and 1.

[0089] In addition, the model training process can generally include the following:

[0090] In each round of model training, a training sample (i.e., the second sample EHR data that has been labeled) is input into the model, so that the model can carry out the corresponding pulmonary embolism thrombolysis risk prediction processing to achieve forward propagation. Then, based on the pulmonary embolism thrombolysis risk prediction results promoted by the model, the loss function is calculated in combination with the annotations of the training samples, and the model parameters are optimized according to the calculation results of the loss function to achieve reverse propagation. In this way, when the preset model training requirements such as training duration, number of trainings or prediction accuracy are met, the model training can be completed, and a pulmonary embolism thrombolysis risk prediction model that can be put into practical use can be obtained.

[0091] It can be understood that the pulmonary embolism risk thrombolysis prediction model is specifically built under the conditions of AI technology based on the deep learning technology / model in machine learning. Therefore, it has strong learning ability and can be used as a carrier of the pulmonary embolism thrombolysis risk prediction logic designed in this application. It has high processing efficiency and prediction accuracy, which can well realize the risk of death for patients with pulmonary embolism after systemic thrombolytic therapy, especially for patients with pulmonary embolism with stable hemodynamics. It can also ensure a high degree of prediction accuracy.

[0092] Furthermore, for the pulmonary embolism risk thrombolysis prediction model, its specific model structure, model training architecture and loss function used in the model training process can adopt existing solutions, or make further optimization and improvements on the basis of existing solutions, or adopt novel self-developed solutions. This is all possible in actual situations and can be configured according to actual needs.

[0093] As an exemplary embodiment, the pulmonary embolism thrombolysis risk prediction model involved in the present application specifically adopts the XGBoost model (i.e., a deep neural network model based on the XGBoost integrated learning method). During the model training process, the XGBoost model is iteratively optimized based on regularization, and grid search is used to determine the optimal model parameters.

[0094] Specifically, the XGBoost model can involve the following configuration content:

[0095] The schematic diagram of XGBoost principle can be referred to Figure 2 A logical schematic diagram of the XGBoost model of the present application is shown. XGBoost is based on the Boosting integration concept and uses the prediction results of multiple base models to sum up to obtain the final output result:

[0096]

[0097] f(x)=w q (x),

[0098] Among them, f(x) represents the wth q The weight of (x) leaf nodes.

[0099] Objective Function of XGBoost The expression of is as follows, l is the loss function, and Ω is the regularization term, which is used to control the complexity of the model and prevent overfitting:

[0100]

[0101] Where T is the number of leaves and w is the weight vector of the leaves.

[0102] Assuming that the first t-1 base models have been solved, the next model to be solved is the tth model. According to the principle of Boosting:

[0103]

[0104] So the objective function Can be transformed into:

[0105]

[0106] Make a second-order Taylor expansion of the loss function l:

[0107]

[0108] Deleting the constant term does not affect the solution of the objective function:

[0109]

[0110] The regularization term Ω(f t ) Substituting into:

[0111]

[0112] Taking the derivative of the above formula and setting it to 0, we can get w j The optimal solution

[0113]

[0114] Will Substitute into Can be eliminated Variables, get:

[0115]

[0116] Feature partitioning criteria, Gain (information gain) can be used to guide decision tree splitting. In actual partitioning, XGBoost uses a greedy algorithm to perform partitioning based on the maximum Gain each time. When Gain cannot obtain a positive number, the node will no longer split.

[0117] The expression of Gain is:

[0118]

[0119] For determining the best model parameters by using grid search, specifically:

[0120] Use threshold adjustment to improve model performance. Threshold adjustment is a technique in classification tasks, which is mainly used to optimize the classification performance of the model, especially in the case of unbalanced data sets. In binary or multi-classification tasks, the model usually outputs the probability that each sample belongs to a certain category, but directly using 0.5 as the default threshold to make classification decisions (that is, samples with a probability greater than or equal to 0.5 are classified as positive, and samples with a probability less than 0.5 are classified as negative) is not always the best choice. Therefore, grid search can be used to determine the optimal parameters of the model involved to achieve the best model training effect.

[0121] In addition, after completing the model training, the model performance evaluation stage may also be involved.

[0122] In this regard, as an example, the present application may specifically use accuracy, specificity, sensitivity, positive predictive value (PPV), negative predictive value (NPV), F1 score (F1 Score), and AUC (Area Under the Curve) as the criteria for evaluating the model. For details, please refer to Figure 3 An example schematic diagram of the performance evaluation of the model of the present application is shown.

[0123] For specific evaluation indicators, please refer to Table 2 below:

[0124] Table 2 - Examples of evaluation indicators

[0125] ACC AUC SEN SPE PVV NPV F1 0.968 0.920 0.750 0.973 0.346 0.995 0.474

[0126] The confusion matrix in the evaluation index can be referred to Figure 4 A schematic diagram showing another example of the performance evaluation of the model of the present application is shown.

[0127] Performing multi-fold cross validation on the model, we can have:

[0128] Table 3 - Multi-fold cross validation example

[0129] ACC (95% CI) AUC (95% CI) SEN (95% CI) SPE (95% CI) 0.975(0.963,0.987) 0.925(0.872,0.978) 0.596(0.497,0.696) 0.981(0.968,0.994)

[0130] From the above series of evaluation indicator examples, it can be seen that under the pulmonary embolism thrombolysis risk prediction logic designed in this application, the trained pulmonary embolism thrombolysis risk prediction model has a better prediction effect on the thrombolysis risk of pulmonary embolism, especially in patients with hemodynamically stable pulmonary embolism.

[0131] Furthermore, for the trained pulmonary embolism thrombolysis risk prediction model, the present application can also use SHAP values ​​to perform explanatory analysis on the model to demonstrate the degree of influence of each feature on the prediction result.

[0132] Specifically, the SHAP score is an advanced model interpretability method based on the principles of game theory. It provides a quantitative measure of the role of each feature in model prediction. Unlike traditional global interpretation methods, the SHAP score considers all possible feature combinations and assigns a unique contribution value to each feature. This value reflects the average impact of the feature on the model prediction results. The core advantage of the SHAP score lies in its fairness and accuracy, which ensures that the contribution of each feature is fairly distributed, that is, no feature will be over- or underestimated due to its relative position in the model or its interaction with other features.

[0133] As an example, refer to Figure 5 An example schematic diagram of the SHAP-based explanatory analysis of the present application is shown, which lists the top 20 features determined based on the SHAP score.

[0134] In addition, you can continue to refer to Figure 6 Another example schematic diagram of the present application based on SHAP's explanatory analysis is shown. For each prediction processing of the model, SHAP's explanatory analysis can also be combined to display the contribution of different features to the pulmonary embolism thrombolysis risk prediction results through visual display, so that clinicians / systems can clearly determine the specific risk factors of each patient and make more reasonable treatment decisions.

[0135] Furthermore, considering that the pulmonary embolism thrombolysis risk prediction model can obtain and determine the relevant characteristics that cause the risk of pulmonary embolism thrombolysis during the prediction process, the pulmonary embolism thrombolysis risk prediction model can also generate new treatment plans or update the original treatment plans by predicting treatment plans (which may involve specific plans in terms of physical intervention measures, medication, psychological counseling, etc.) to provide data support for auxiliary decision-making in clinical work. The predicted treatment plan can effectively improve the corresponding characteristics of the patient, thereby achieving the goal of effectively reducing the risk of pulmonary embolism thrombolysis, which is also feasible in practical applications.

[0136] Correspondingly, as an exemplary embodiment, the pulmonary embolism thrombolysis risk prediction model of the present application can also be used to predict the appropriate treatment plan for the corresponding patient based on the EHR data input into the model, so as to reduce the risk of death after the application of systemic thrombolytic therapy.

[0137] In addition, it is understandable that, after determining a series of characteristics of the current patient, personalized or customized prediction processing can be performed in the process of predicting the treatment plan to reduce the risk of thrombolysis by improving specific characteristic indicators. Specifically, a highly adapted treatment plan prediction process can be performed based on the specific situation of a series of characteristics of the current patient (the corresponding indicators in the EHR data input by the model), rather than using a simple, general treatment plan prediction logic. The historical characteristics of the current patient (the corresponding indicators in the historical EHR data) can also be combined to more accurately quantify the dynamic changes of a series of characteristics of the current patient in the time dimension, thereby achieving a more accurate prediction of the treatment plan, which can provide a more delicate and effective data-assisted support effect in specific applications.

[0138] Furthermore, corresponding to the subsequent practical application of the model, the present application scheme may also involve the application of the pulmonary embolism thrombolysis risk prediction model.

[0139] Correspondingly, after training the pulmonary embolism thrombolysis risk prediction model based on the annotated second sample EHR data in step S104, the method of the present application may further include:

[0140] Obtain target EHR data;

[0141] Inputting the target EHR data into the pulmonary embolism thrombolysis risk prediction model so that the pulmonary embolism thrombolysis risk prediction model carries out corresponding pulmonary embolism thrombolysis risk prediction processing;

[0142] Extract the pulmonary embolism thrombolysis risk prediction results output by the pulmonary embolism thrombolysis risk prediction model.

[0143] It can be understood that in specific operations, the target EHR data input into the model in actual applications is usually configured according to the data format of the unlabeled second sample EHR data in the previous model training link, or, in some cases, it can also follow the data format of the first sample EHR data in the previous model training link, that is, the original EHR data that can be collected in various online business systems of the hospital under actual circumstances.

[0144] In this way, when there is a need for pulmonary embolism thrombolysis risk prediction, it can be triggered manually or automatically by the system to input the corresponding target EHR data into the model, so that the model can carry out the corresponding pulmonary embolism thrombolysis risk prediction processing to predict the corresponding patient's risk of death after systemic thrombolysis therapy. After completing the risk prediction, the model can output the corresponding pulmonary embolism thrombolysis risk prediction result. At this time, the pulmonary embolism thrombolysis risk prediction result output by the model can be extracted according to the preset extraction method.

[0145] In this case, further data application processing can be performed on the pulmonary embolism thrombolysis risk prediction result, such as local storage, remote storage, result display, output of prompts for completion of prediction processing, result push, further data analysis, etc. It is easy to understand that the specific data application processing content can be flexibly adjusted according to the pre-set and real-time data application strategies / rules, and this application again does not make specific limitations.

[0146] Corresponding to the data support work in auxiliary decision-making mentioned many times above, in practical application, this application may mainly involve the data application processing method of result display.

[0147] In this regard, as an exemplary embodiment, after extracting the pulmonary embolism thrombolysis risk prediction result output by the pulmonary embolism thrombolysis risk prediction model, the method of the present application may further include:

[0148] The content of the pulmonary embolism thrombolysis risk prediction results is output through a visual interface.

[0149] It can be understood that in specific operations, the visualization interface can be displayed through the device's own display screen (such as a touch screen), through an external display screen device, or through other related devices with a display screen. This is relatively flexible and corresponds to the flexible and changeable application environment in actual clinical work.

[0150] In the visualization interface, in addition to displaying the specific predicted probability value of pulmonary embolism thrombolysis risk (in numerical form), the content of the target EHR data can also be displayed, the specific features that make corresponding contributions to the probability can also be displayed, and the corresponding predicted treatment plan content can also be displayed. This information can also be sorted according to the sorting strategy (which can be adjusted in real time) based on the sorting of the contribution made by the features. For example, the corresponding detailed content of the features that can be significantly improved in the predicted treatment plan to significantly reduce the risk can be displayed in the first place (the first place can be the first item in the list, the first key content display area, etc., or a pop-up window display position. The specific form of the first place can be adjusted according to the specific content display strategy in the interface) to further improve the readability of the displayed content and achieve a more intuitive display effect.

[0151] Finally, in general, for the above scheme contents, this application provides a specific pulmonary embolism thrombolysis risk prediction model configuration scheme based on deep learning under the goal of predicting thrombolysis risk in patients with pulmonary embolism. The pulmonary embolism thrombolysis risk prediction model thus configured is highly targeted and adaptable for patients with hemodynamically stable pulmonary embolism. Compared with the general model configuration logic, it can accurately predict the thrombolysis risk for patients with hemodynamically stable pulmonary embolism, provide high-quality data-assisted services for clinical decision-making, and thus promote the reduction of treatment risks and obtain high-quality medical services.

[0152] The above is an introduction to the configuration method of the pulmonary embolism thrombolysis risk prediction model provided by the present application. In order to facilitate better implementation of the configuration method of the pulmonary embolism thrombolysis risk prediction model provided by the present application, the present application also provides a configuration device for a pulmonary embolism thrombolysis risk prediction model from the perspective of functional modules.

[0153] See also Figure 7 , Figure 7 This is a schematic diagram of a configuration device for a pulmonary embolism thrombolysis risk prediction model of the present application. In the present application, the configuration device 700 for a pulmonary embolism thrombolysis risk prediction model may specifically include the following structure:

[0154] An acquisition unit 701 is used to acquire first sample EHR data corresponding to a sample pulmonary embolism patient, wherein the sample pulmonary embolism patient is specifically a pulmonary embolism patient who has received systemic thrombolytic therapy;

[0155] The operation unit 702 is used to perform a data enhancement operation on the first sample EHR data, and retain the data involved in the target feature during the enhancement process to obtain the second sample EHR data, wherein the target feature includes three aspects of features: basic information feature, medical history feature and laboratory result feature;

[0156] A labeling unit 703 is used to label the second sample EHR data, wherein the labeling result includes whether the patient died after the systemic thrombolytic therapy was applied;

[0157] The training unit 704 is used to train a pulmonary embolism thrombolysis risk prediction model based on the annotated second sample EHR data, wherein the pulmonary embolism thrombolysis risk prediction model is a deep learning model, and the pulmonary embolism thrombolysis risk prediction model is used to predict the death risk of the corresponding patient after systemic thrombolytic therapy based on the EHR data input by the model.

[0158] In an exemplary embodiment, the data enhancement operation specifically includes:

[0159] The first operation is to remove samples with a missing rate of more than 10%;

[0160] A second operation of removing samples whose indicator values ​​deviate from the normal range by a degree greater than a preset degree;

[0161] The third operation of unifying the numerical units involved in the sample;

[0162] The fourth operation is to balance the sample distribution by choosing between the SMOTENC sampling method and the ClusterCentroids sampling method;

[0163] The fifth operation uses the MissForest method to fill in the missing values ​​in the sample.

[0164] In another exemplary embodiment, the 15 basic information features specifically include:

[0165] Gender, age, smoking history, drinking history, blood transfusion history, allergy history, pregnancy history, surgical history, previous medical history, infectious diseases, family clustering diseases, miscarriage history, chemotherapy history, targeted therapy history, and endocrine therapy history;

[0166] The diseases involved in the 10 medical history characteristics include:

[0167] Coronary heart disease, diabetes, hypertension, hyperlipidemia, cerebrovascular disease, lower limb vascular disease, kidney disease, cancer, fracture or trauma, coagulation system disease;

[0168] The 46 laboratory result characteristics specifically include:

[0169] Systolic blood pressure, diastolic blood pressure, pulse, temperature, respiratory rate, corrected calcium, alkaline phosphatase, total cholesterol, albumin, chloride, direct bilirubin, gamma-glutamyl transpeptidase, bicarbonate, globulin, potassium, TBIL*0.8, sodium, W / B ratio, calcium, total bilirubin, creatinine, urea, total protein, uric acid, basophils, monocytes, platelet count, monocytes, mean hemoglobin concentration, mean hemoglobin content, lymphocytes, eosinophils, basophils, hematocrit, hemoglobin, lymphocytes, mean RBC volume, white blood cell count, eosinophils, neutrophils, mean PLT volume, neutrophils, red blood cell count, platelet hematocrit, PLT distribution width, TP*0.75.

[0170] In another exemplary embodiment, the pulmonary embolism thrombolysis risk prediction model specifically adopts the XGBoost model. During the model training process, the XGBoost model is iteratively optimized based on regularization, and grid search is used to determine the optimal model parameters.

[0171] In yet another exemplary embodiment, the pulmonary embolism thrombolytic risk prediction model is also used to predict an adaptive treatment plan for the corresponding patient based on EHR data input into the model, so as to reduce the risk of death after the application of systemic thrombolytic therapy.

[0172] In yet another exemplary embodiment, the apparatus further includes an application unit 705, configured to:

[0173] Obtain target EHR data;

[0174] Inputting the target EHR data into the pulmonary embolism thrombolysis risk prediction model so that the pulmonary embolism thrombolysis risk prediction model carries out corresponding pulmonary embolism thrombolysis risk prediction processing;

[0175] Extract the pulmonary embolism thrombolysis risk prediction results output by the pulmonary embolism thrombolysis risk prediction model.

[0176] In yet another exemplary embodiment, the apparatus further includes an output unit 706, configured to:

[0177] The content of the pulmonary embolism thrombolysis risk prediction results is output through a visual interface.

[0178] This application also provides a processing device from the perspective of hardware structure, see Figure 8 , Figure 8801, a memory 802, and an input / output device 803. The processor 801 is used to execute the computer program stored in the memory 802 to implement the following Figure 1 The steps of the configuration method of the pulmonary embolism thrombolysis risk prediction model in the corresponding embodiment; or, the processor 801 is used to execute the computer program stored in the memory 802 to implement the following Figure 7 Corresponding to the functions of each unit in the embodiment, the memory 802 is used to store the processor 801 executing the above Figure 1 A computer program required for the configuration method of the pulmonary embolism thrombolysis risk prediction model in the corresponding embodiment.

[0179] Exemplarily, the computer program may be divided into one or more modules / units, one or more modules / units are stored in the memory 802, and executed by the processor 801 to complete the present application. One or more modules / units may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the computer device.

[0180] The processing device may include, but is not limited to, a processor 801, a memory 802, and an input / output device 803. Those skilled in the art will appreciate that the illustration is merely an example of a processing device and does not constitute a limitation on the processing device, and may include more or fewer components than shown in the illustration, or a combination of certain components, or different components. For example, the processing device may also include a network access device, a bus, etc., and the processor 801, the memory 802, the input / output device 803, etc. are connected via a bus.

[0181] The processor 801 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the processing device, and uses various interfaces and lines to connect various parts of the entire device.

[0182] The memory 802 can be used to store computer programs and / or modules. The processor 801 implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory 802 and calling the data stored in the memory 802. The memory 802 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the processing device, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0183] When the processor 801 is used to execute the computer program stored in the memory 802, the following functions can be implemented:

[0184] Acquire first sample EHR data corresponding to a sample pulmonary embolism patient, wherein the sample pulmonary embolism patient is specifically a pulmonary embolism patient who has received systemic thrombolytic therapy;

[0185] Performing data enhancement operation on the first sample EHR data, and retaining the data involved in the target feature during the enhancement process, to obtain the second sample EHR data, wherein the target feature includes three aspects of features: basic information feature, medical history feature, and laboratory result feature;

[0186] Annotating the second sample EHR data, wherein the annotation result includes whether death occurred after systemic thrombolytic therapy was administered;

[0187] Based on the annotated second sample EHR data, a pulmonary embolism thrombolysis risk prediction model is trained, wherein the pulmonary embolism thrombolysis risk prediction model is a deep learning model, which is used to predict the death risk of the corresponding patient after systemic thrombolysis therapy based on the EHR data input into the model.

[0188] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the configuration device, processing device and corresponding units of the pulmonary embolism thrombolysis risk prediction model described above can refer to the following. Figure 1 The description of the configuration method of the pulmonary embolism thrombolysis risk prediction model in the corresponding embodiment will not be repeated here.

[0189] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0190] To this end, the present application provides a computer-readable storage medium, in which a plurality of instructions are stored, and the instructions can be loaded by a processor to execute the present application as follows: Figure 1 The steps of the configuration method of the pulmonary embolism thrombolysis risk prediction model in the corresponding embodiment, the specific operation can refer to the following Figure 1 The description of the configuration method of the pulmonary embolism thrombolysis risk prediction model in the corresponding embodiment will not be repeated here.

[0191] The computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0192] Due to the instructions stored in the computer-readable storage medium, the present application can be executed. Figure 1 The steps of the configuration method of the pulmonary embolism thrombolysis risk prediction model in the corresponding embodiment, therefore, the present application can be implemented as follows Figure 1 The beneficial effects that can be achieved by the configuration method of the pulmonary embolism thrombolysis risk prediction model in the corresponding embodiment are detailed in the previous description and will not be repeated here.

[0193] The above is a detailed introduction to the configuration method, device, processing equipment and computer-readable storage medium of the pulmonary embolism thrombolysis risk prediction model provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the core idea of ​​the present application; at the same time, for technical personnel in this field, according to the idea of ​​the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A method for configuring a pulmonary embolism thrombolysis risk prediction model, characterized in that: The method comprises: Acquire first sample EHR data corresponding to a sample pulmonary embolism patient, wherein the sample pulmonary embolism patient is specifically a pulmonary embolism patient who has received systemic thrombolytic therapy; Performing data enhancement operation on the first sample EHR data, and retaining data related to target features during the enhancement process, to obtain second sample EHR data, wherein the target features include three aspects of features: basic information features, medical history features, and laboratory result features; annotating the second sample EHR data, wherein the annotated result includes whether death occurs after the systemic thrombolytic therapy is applied; Based on the annotated second sample EHR data, a pulmonary embolism thrombolysis risk prediction model is trained, wherein the pulmonary embolism thrombolysis risk prediction model is a deep learning model, and the pulmonary embolism thrombolysis risk prediction model is used to predict the death risk of the corresponding patient after the systemic thrombolytic therapy is applied based on the EHR data input by the model.

2. The method according to claim 1, characterized in that The data enhancement operation specifically includes: The first operation is to remove samples with a missing rate of more than 10%; A second operation of removing samples whose indicator values ​​deviate from the normal range by a degree greater than a preset degree; The third operation of unifying the numerical units involved in the sample; The fourth operation uses both the SMOTENC sampling method and the ClusterCentroids sampling method to balance the sample distribution; The fifth operation uses the MissForest method to fill in the missing values ​​in the sample.

3. The method according to claim 1, characterized in that The 15 basic information features specifically include: Gender, age, smoking history, drinking history, blood transfusion history, allergy history, pregnancy history, surgical history, previous medical history, infectious diseases, family clustering diseases, miscarriage history, chemotherapy history, targeted therapy history, and endocrine therapy history; The diseases involved in the 10 medical history characteristics specifically include: Coronary heart disease, diabetes, hypertension, hyperlipidemia, cerebrovascular disease, lower limb vascular disease, kidney disease, cancer, fracture or trauma, coagulation system disease; The 46 laboratory result characteristics specifically include: Systolic blood pressure, diastolic blood pressure, pulse, temperature, respiratory rate, corrected calcium, alkaline phosphatase, total cholesterol, albumin, chloride, direct bilirubin, gamma-glutamyl transpeptidase, bicarbonate, globulin, potassium, TBIL*0.8, sodium, W / B ratio, calcium, total bilirubin, creatinine, urea, total protein, uric acid, basophils, monocytes, platelet count, monocytes, mean hemoglobin concentration, mean hemoglobin content, lymphocytes, eosinophils, basophils, hematocrit, hemoglobin, lymphocytes, mean RBC volume, white blood cell count, eosinophils, neutrophils, mean PLT volume, neutrophils, red blood cell count, platelet hematocrit, PLT distribution width, TP*0.

75.

4. The method according to claim 1, characterized in that: The pulmonary embolism thrombolysis risk prediction model specifically adopts the XGBoost model. During the model training process, the XGBoost model is iteratively optimized based on regularization, and grid search is used to determine the optimal model parameters.

5. The method according to claim 1, characterized in that: The pulmonary embolism thrombolysis risk prediction model is also used to predict the adaptive treatment plan for the corresponding patient based on the EHR data input into the model, so as to reduce the risk of death after the application of the systemic thrombolytic therapy.

6. The method according to claim 1, characterized in that After training the pulmonary embolism thrombolysis risk prediction model based on the annotated second sample EHR data, the method further includes: Obtain target EHR data; Inputting the target EHR data into the pulmonary embolism thrombolysis risk prediction model so that the pulmonary embolism thrombolysis risk prediction model carries out corresponding pulmonary embolism thrombolysis risk prediction processing; The pulmonary embolism thrombolysis risk prediction result output by the pulmonary embolism thrombolysis risk prediction model is extracted.

7. The method according to claim 6, characterized in that After extracting the pulmonary embolism thrombolysis risk prediction result output by the pulmonary embolism thrombolysis risk prediction model, the method further includes: The content of the pulmonary embolism thrombolysis risk prediction result is output through a visual interface.

8. A device for configuring a pulmonary embolism thrombolysis risk prediction model, characterized in that: The device comprises: an acquisition unit, configured to acquire first sample EHR data corresponding to a sample pulmonary embolism patient, wherein the sample pulmonary embolism patient is specifically a pulmonary embolism patient who has received systemic thrombolytic therapy; An operation unit, configured to perform a data enhancement operation on the first sample EHR data, and retain the data involved in the target feature during the enhancement process, to obtain the second sample EHR data, wherein the target feature includes three aspects of features: basic information feature, medical history feature, and laboratory result feature; a labeling unit, configured to label the second sample EHR data, wherein the labeling result includes whether death occurs after the systemic thrombolytic therapy is applied; A training unit is used to train a pulmonary embolism thrombolysis risk prediction model based on the labeled second sample EHR data, wherein the pulmonary embolism thrombolysis risk prediction model is a deep learning model, and the pulmonary embolism thrombolysis risk prediction model is used to predict the death risk of the corresponding patient after the application of the systemic thrombolytic therapy based on the EHR data input by the model.

9. A processing device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program, and the processor executes the method according to any one of claims 1 to 7 when calling the computer program in the memory.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Cerebral stroke END risk prediction model establishment method and device, END risk prediction system, electronic equipment and medium

    CN116403714A

  • Thrombolysis risk prediction system and processing method thereof

    CN117457194A

  • Processing method, device and processing equipment for thrombolysis risk dynamic evaluation model

    CN118262910A

  • Thrombolysis risk prediction system and method based on nano-robot

    CN118571468A

  • Computed tomography pulmonary angiography examination initiating method and device

    CN118888164A