Patient death rate prediction method based on self-attention mechanism and generative adversarial network
By combining self-attention mechanisms with generative adversarial networks, the problems of missing data and high-dimensional nonlinear modeling in heart failure patients are solved, improving the quality of data imputation and prediction accuracy, providing clinical interpretability, and supporting individualized intervention programs.
Patent Information
- Application Number
- CN202510690947.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies have data missing problems in predicting mortality in heart failure patients, resulting in limited prediction accuracy. The value distribution generated by traditional interpolation methods deviates from the real data. The model is insufficient in modeling high-dimensional nonlinear relationships and lacks clinical interpretability, making it difficult to support individualized intervention plans.
By combining self-attention mechanism with generative adversarial network, the feature space is optimized through dynamic weighted fusion of Monte Carlo sampling results. The adversarial training framework is used to improve data imputation quality and nonlinear modeling capability, and SHAP analysis is used to provide clinical interpretability.
It significantly improved the diversity and accuracy of data imputation, enhanced the model's ability to capture high-dimensional nonlinear relationships, improved mortality prediction accuracy by about 15%, and clarified the contribution direction of key biomarkers through SHAP analysis, supporting individualized intervention decisions.
Smart Images

Figure CN120809180A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical prediction, in particular to a patient mortality prediction method based on a self-attention mechanism and a generative adversarial network. BACKGROUND
[0002] Heart failure (HF) as the end stage of CVD, its disease burden is particularly severe, epidemiological studies have shown that in the past 15 years, the prevalence of heart failure in China has increased by 44%, with an additional 3.8 million newly diagnosed patients. Due to differences in patient characteristics, comorbidity spectrum and medical accessibility, the clinical prognosis of heart failure patients presents high heterogeneity, with a mortality rate of 5% to 75%. In this context, how to construct a precise prognosis model based on the heterogeneity of patient characteristics, and then develop individualized intervention programs and optimize medical resource allocation, has become a key breakthrough to improve prognosis and reduce social and economic burden.
[0003] The prior art has the following key defects in practical application: data missing problem limits the prediction accuracy, electronic medical record data generally has information missing (such as laboratory indicators not detected, incomplete medical history records), and traditional interpolation methods are difficult to effectively solve: single interpolation method (mean interpolation, K-neighbor interpolation): although efficient in calculation, the generated value distribution is significantly deviated from the real data generation mechanism. For example, mean interpolation ignores the nonlinear correlation between features, resulting in an excessively smooth data distribution after interpolation, which cannot reflect the complexity of the real clinical scenario. Multiple interpolation methods (interpolation based on VAE): the traditional variational autoencoder (VAE) is over-regularized due to the KL divergence constraint, and the representation learning ability is limited, and the diversity of generated samples is insufficient. The interpolation results are biased towards the mean value region of the data distribution, and it is difficult to cover the extreme values or rare modes of real cases. The model lacks the ability to model high-dimensional nonlinear relationships, and the clinical characteristics of heart failure patients (such as blood routine, biochemical indicators, and imaging parameters) have high dimensions and complex nonlinear interactions. Existing models (such as simple multilayer perceptron) are difficult to fully capture the deep relationships between features due to their single structure and improper selection of activation functions. For example, traditional linear encoders cannot effectively integrate multi-modal data, resulting in the loss of key prognostic information and limited prediction accuracy. Lack of clinical interpretability hinders practical application, existing prediction models are mostly "black box" designs, and the contribution of each clinical feature to the risk of death cannot be quantified, making it difficult for doctors to trust the model results. For example, although key biomarkers such as D-dimer have been confirmed by literature to be related to prognosis, traditional models cannot clearly determine their specific role direction (such as increasing or decreasing risk), limiting their application value in developing individualized intervention plans. The application of generative adversarial networks (GAN) in data interpolation is limited, although GAN performs well in image generation, but its direct application in electronic medical record interpolation has the following problems: the discriminator design is simple (such as a single-layer network), which cannot effectively distinguish the real distribution of high-dimensional clinical data from the generated distribution, resulting in unstable quality of interpolated data. Lack of synergistic optimization with self-attention mechanism, the interpolation process ignores the dynamic correlation between features, for example, it fails to dynamically adjust the interpolation weight according to patient-specific features (such as kidney function indicators). SUMMARY
[0004] The application provides a patient mortality prediction method based on a self-attention mechanism and a generative adversarial network. The application dynamically weights and fuses Monte Carlo sampling results through a self-attention mechanism, optimizes a feature space in combination with an adversarial training framework, significantly reduces data interpolation bias (increases sample diversity by 15%-20% compared with a traditional VAE method), and enhances the modeling capability of the model for high-dimensional nonlinear relationships, so that the mortality prediction accuracy is increased by about 15%. Real-world data of 15,773 heart failure patients are used, and three-level cleaning and five-fold cross-validation are performed to ensure the generalization and stability of the model. The SHAP analysis module clearly shows the contribution direction of key biomarkers such as D-dimer and NT-ProBNP (blue / red feature distinguishes risk), fills the explanatory gap of traditional black box models, and supports clinical precision decision-making. In addition, the model can be extended to other cardiovascular disease prognosis prediction by adjusting parameters, and has the potential for cross-disease migration. The technical scheme breaks through the bottleneck of existing methods in data quality, nonlinear modeling and explainability, and provides a reliable tool for optimizing medical resource allocation.
[0005] In order to realize the above effect, the application provides the following technical scheme: a patient mortality prediction method based on a self-attention mechanism and a generative adversarial network, comprising the following steps:
[0006] S1, data acquisition and preprocessing: collecting electronic medical record data of heart failure patients, constructing a clinical feature data set containing demographic characteristics, medical history records, laboratory test indicators and imaging examination results, and performing anonymization processing and three-level cleaning on the data.
[0007] S2, data interpolation based on an attention mechanism: encoding missing data using an improved variational autoencoder (VAE), dynamically weighting and fusing multiple Monte Carlo sampling results through a self-attention mechanism, generating latent feature representation, and optimizing feature space discriminativeness through an adversarial training framework.
[0008] S3, mortality prediction: inputting the complete data set after interpolation into a multilayer perceptron (MLP) model for training, and using a cross-entropy loss function and an Adam optimizer to complete a binary classification prediction task.
[0009] S4, important feature analysis: quantifying the contribution of each feature to the prediction result through SHAP value, and verifying the clinical rationality of the model decision mechanism.
[0010] Further, according to the operation steps in S1,
[0011] S101, mark Null values, zero values and negative values as missing values.
[0012] S102, set the abnormal values of continuous features as missing.
[0013] S103, the similar categories of discrete features are integrated, and finally 105 clinical features are retained.
[0014] Further, according to the operation steps in S1, in the three-level data cleaning process, the determination standard of continuous feature abnormal value is exceeding the range of mean value ± 3 times standard deviation.
[0015] Further, according to the operation steps in S2, the calculation method of the self-attention mechanism is: randomly sampling multiple potential samples from a normal distribution, calculating the weight of each sample by cosine similarity, and generating the final potential variable by weighted summation, the formula is:
[0016] Among them, the weight Obtained by normalizing the cosine similarity.
[0017] Further, according to the operation steps in S2, the discriminator network structure of the adversarial training framework is [105, 32, 1], the activation function adopts LeakyReLU, and the objective function is:
[0018]
[0019] Wherein, the generator and the discriminator are optimized to generate data quality through dynamic game.
[0020] Further, according to the operation steps in S3, the network structure of the multi-layer perception (MLP) model is [105, 512, 128, 16, 2], the activation function adopts ReLU, the learning rate is set to 0.001, the training period is 450 times, and the loss function is cross entropy loss.
[0021] Further, according to the operation steps in S3, the total loss function of the cross entropy loss function is composed of data interpolation loss L MI And mortality prediction loss L MP The weights are 1.0 and 1.5 respectively, and the expression is:
[0022] L MI =L VAE +1.5L GAN
[0023] Further, according to the operation steps in S4, the average absolute SHAP value is calculated to determine the feature importance, wherein D-dimer (D-Dimer), N-terminal brain natriuretic peptide (NT-Pro BNP) and creatinine (Crea) are key prognostic biomarkers.
[0024] Further, according to the operation steps in S3, the decoder of the variational autoencoder is composed of a multi-layer perception mechanism, and the generated data after adversarial training is merged with the real data, the formula is:
[0025]
[0026] wherein, is the imputed data generated by the GAN.
[0027] Further, according to the operation steps in S3, the method adopts five-fold cross-validation to evaluate the model performance, ensuring the generalization and stability of the prediction results.
[0028] The application provides a patient mortality prediction method based on a self-attention mechanism and a generative adversarial network, which has the following beneficial effects:
[0029] (1) The application improves the data imputation quality and reduces the imputation bias. The traditional VAE method causes over-regularization of the latent space in data imputation due to the KL divergence constraint, and the generated sample diversity is insufficient. The application introduces a self-attention mechanism to dynamically weight and fuse multiple sets of Monte Carlo sampling results, calculate the sample weight through cosine similarity, and generate latent variables by weighted summation. This method effectively reduces the deviation of imputed values from the real data distribution, significantly improves the accuracy and diversity of missing data imputation.
[0030] (2) The application enhances the discriminability of the feature space and optimizes the nonlinear modeling capability of the model. By embedding an adversarial training framework in the model, the generator and the discriminator form a dynamic game mechanism, and the generated data and the real data distribution are dynamically aligned through the discriminator. The application of the discriminator network structure and the LeakyReLU activation function further optimizes the discriminability of the feature space, enhances the model's ability to capture high-dimensional nonlinear clinical relationships, and thus improves the accuracy of mortality prediction.
[0031] (3) The application improves the generalization performance and stability of the prediction model. The application uses 15,773 electronic medical record data (EMR-HF data set) of heart failure patients collected in the real world, through three-level data cleaning (labeling missing values, removing outliers, and integrating discrete features) and anonymization processing, to ensure data quality and privacy security. Combined with five-fold cross-validation and specific parameters of multilayer perceptron, the generalization ability and stability of the model on unknown data are significantly improved.
[0032] (4) The application provides clinical interpretability and assists in precision medical decision-making. The SHAP value analysis quantifies the contribution of each feature to the prediction result, and clearly identifies D-dimer, N-terminal pro-BNP (NT-ProBNP), and creatinine (Crea) as key prognostic biomarkers. The local explanation module visualizes the feature contribution direction (blue reduces risk, red increases risk), providing individualized prediction basis for clinicians and supporting treatment plan optimization and rational allocation of medical resources.
[0033] (5) The present application supports cross-disease expansion application. By adjusting the clinical feature data set and the model parameter, the present application can be expanded to be applied to the prognosis prediction of other cardiovascular diseases (such as coronary heart disease and myocardial infarction), has good migration and adaptability, and provides a technical basis for precise prognosis modeling of multiple diseases. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 SHAP diagram of the patient mortality prediction method based on the self-attention mechanism and the generative adversarial network according to the present application;
[0035] Figure 2 Feature explainability diagram of the patient mortality prediction method based on the self-attention mechanism and the generative adversarial network according to the present application;
[0036] Figure 3 Flowchart of the patient mortality prediction method based on the self-attention mechanism and the generative adversarial network according to the present application. DETAILED DESCRIPTION
[0037] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to specific embodiments.
[0038] In Example 1, the present application provides a technical solution: please refer to Figures 1-3 The patient mortality prediction method based on the self-attention mechanism and the generative adversarial network, characterized in that the method comprises the following steps:
[0039] S1, data collection and preprocessing: The data were collected from the Department of Cardiology of a certain third-grade class-A hospital. The electronic medical records of 15773 patients aged 18 years and discharged with heart failure as the first diagnosis from May 2000 to September 2023 were collected to construct the EMR-HF master dataset, including demographic characteristics, medical history records, laboratory test indicators, and imaging examination results during hospitalization. To ensure data security and privacy, the data were aligned based on the unique identifier of the patient, and sensitive information was anonymized. A three-level data cleaning process was performed: (a) Null values, zero values, and negative values in the variables were marked as missing values; (b) Abnormal values of continuous features were set to missing; (c) Similar categories of discrete features were integrated. After preprocessing, a total of 105 features were included (15 discrete features and 90 continuous features), including 6 demographic indicators (gender, ethnicity, age, marital status, education level, and cause of death), 22 indicators of blood routine examination (white blood cell count and mean corpuscular volume) (blood routine white blood cell count (WBC, 10^9 / L), blood routine red blood cell count (RBC, 10^12 / L), blood routine mean corpuscular volume (MCV, fL), blood routine mean corpuscular hemoglobin concentration (MCHC, g / L), blood routine mean corpuscular hemoglobin content (MCH, pg), blood routine red blood cell volume distribution width CV (RDW-CV, %), blood routine lymphocyte percentage (Lymph, %), blood routine lymphocyte count (Lymph, 10^9 / L), blood routine monocyte percentage (Mono, %), blood routine monocyte count (Mono, 10^9 / L), blood routine neutrophil percentage (Neut, %), blood routine neutrophil count (Neut, 10^9 / L), blood routine hematocrit (Hct, %) - venous blood test quantitative results, blood routine eosinophil percentage (Eos, %), blood routine basophil percentage (Baso, %), blood routine eosinophil count (Eos, 10^9 / L), blood routine basophil count (Baso, 10^9 / L), blood routine hemoglobin (Hb, g / L), blood routine platelet count (PLT, 10^9 / L), blood routine mean platelet volume (MPV, fL), blood routine platelet distribution width (PDW, ratio), blood routine platelet hematocrit, %), 27 indicators of biochemical examination (including total bilirubin and albumin) (biochemical examination alanine aminotransferase (ALT), biochemical examination aspartate aminotransferase (AST), biochemical examination gamma-glutamyl transferase (GGT), biochemical examination total bilirubin (TBIL), biochemical examination direct bilirubin (DBIL), biochemical examination indirect bilirubin (IBIL), biochemical examination albumin (ALB), biochemical examination globulin (GLO),Biochemical test_Total protein (TP), Biochemical test_Albumin / globulin ratio (ALB / GLO), Biochemical test_Total bile acid (TBA), Biochemical test_Creatinine (Crea), Biochemical test_Urea (Urea), Biochemical test_Uric acid (UA), Biochemical test_Total cholesterol (TC), Biochemical test_Triglycerides (TG), Biochemical test_High-density lipoprotein cholesterol (HDL-C), Biochemical test_Low-density lipoprotein cholesterol (LDL-C), Biochemical test_Potassium ion (K+), Biochemical test_Sodium ion (+), Biochemical test_Chloride ion (Cl-), Biochemical test_Total calcium (Ca), Biochemical test_Phosphorus (P), Biochemical test_Magnesium ion (Mg2+), Biochemical test_Glucose (Glu ), biochemical test_lactate dehydrogenase (LDH), biochemical test_alkaline phosphatase (ALP)), coagulation test including 9 indicators including thrombin (coagulation test_thrombin time (TT), coagulation test_prothrombin time (PT), coagulation test_activated partial thromboplastin time (APTT), coagulation test_activated partial thromboplastin time ratio (APTTR), coagulation test_prothrombin international normalized ratio (PT-INR), coagulation test_D-dimer (D-Dimer), coagulation test_fibrinogen degradation product (FDP), coagulation test_fibrinogen (Fbg), coagulation test_prothrombin time activity (PTA)), thyroid test 3 indicators (thyroid function test_ Thyroid stimulating hormone (TSH), thyroid function test_free thyroxine (FT4), thyroid function test_free triiodothyronine (FT3)), 2 indicators of glucose metabolism measurement (glycated serum protein (GSP), glucose metabolism measurement_glycated hemoglobin A1c (HbA1c)), 2 indicators of heart failure and myocardial injury markers (heart failure and myocardial injury marker_N-terminal pro-brain natriuretic peptide (NT-ProBNP), heart failure and myocardial injury marker_creatine kinase (CK)), 27 basic examination indicators upon admission (creatinine clearance upon admission (C_G formula) (ml / min), glomerular filtration rate upon admission (MDRD2006 formula) (ml / min), lower limb edema, heart rate (beats / min) ), body temperature (degrees Celsius), pulse (beats / min), respiratory rate (beats / min), diastolic blood pressure (mmHg), systolic blood pressure (mmHg), height (cm), weight (kg), aorta (sinus), aorta (annulus), right ventricular outflow tract, pulmonary artery, ventricular wall motion score, FS (left ventricular fractional shortening), SV (stroke volume), CO (stroke volume), left atrial diameter (mm), left ventricular end-systolic diameter (mm), left ventricular end-diastolic diameter (mm), interventricular septum thickness (mm), left ventricular posterior wall thickness (mm), right atrial diameter (mm), right ventricular internal diameter (mm), left ventricular ejection fraction (%)) and 7 indicators of basic medical history (hypertension, diabetes, coronary heart disease, history of cerebral infarction, history of surgery,Smoking, drinking)
[0040] S2, data imputation based on attention mechanism: Next, the improved variational autoencoder is used to encode the preprocessed heart failure patient characteristics. In order to improve the influence of missing electronic medical record data on feature encoding, a data imputation module based on attention mechanism is constructed.
[0041] The specific steps are as follows:
[0042] S201, define the preprocessed heart failure patient feature sample set as X={X o ,X m}, wherein X o is a sample feature set without missing information, and X m is a sample feature set with missing information.
[0043] S202, encode the sample feature set without missing information, and encode each sample in X o using a linear encoder, specifically: Z=Wx o +b, wherein W is a weight matrix and b is a bias vector.
[0044] S203, construct a resampling method that integrates attention mechanism to perform data imputation on the sample feature set with missing information. Subsequently, by adding a random variable reparameterization method, randomly sample latent samples from the sample space X o without missing information, and integrate them according to the attention mechanism as the latent feature representation of the sample x i that needs to be imputed, specifically: randomly collect h latent samples from the normal distribution (wherein, ) and represent them as In order to effectively integrate multiple latent samples, a self-attention mechanism is introduced to weight and sum the h latent samples, is composed of key and value data pairs, and the query vector in the self-attention mechanism is transformed by the characteristics of the Gaussian distribution to generate the final latent variable The calculation formula is as follows:
[0045]
[0046] wherein, is the normalization of cosine similarity, which can be calculated by the following formula:
[0047]
[0048] wherein, S j is the cosine similarity of q and , and its formula expression is as follows:
[0049]
[0050] Where: q and each It is considered as the concatenation of the mean vector of the Gaussian distribution and the diagonal elements of the covariance matrix.
[0051] S204, the latent variables in step 3 Through the decoder composed of multi-layer perceptron (MLP), in order to obtain the sample set X m The optimal feature representation of , further constructs a training module based on the adversarial network, specifically: from the noise distribution p z Generated and real data distribution p data It is difficult to distinguish samples, and D optimizes the discrimination criterion by distinguishing real data from generated data. Its objective function is:
[0052]
[0053] Where x is the real data, collected from the real data distribution r(·); z is the random noise vector, sampled from the prior distribution f(·); E(·) is the computational expectation, and the adversarial training mechanism promotes high-quality data samples.
[0054] In the adversarial training framework, the discriminator and the decoder form a dynamic game mechanism, whose network structure is [105, 32, 1], and the activation function uses LeakyReLU.
[0055] S205, after training the generative adversarial network, the The values at the corresponding positions in the data set are interpolated to X m In, with X o Combined, the formula is as follows:
[0056]
[0057] S3. Mortality Prediction: The mortality prediction task is treated as a binary classification problem (dead / not dead). The MLP is trained using 5-fold cross-validation with a network structure of [105, 512, 128, 16, 2]. The activation function is ReLU, the learning rate is set to 0.001, and the Adam optimizer is used for 450 epochs. The cross entropy loss function is used as the binary classification loss function. The formula is as follows:
[0058]
[0059] Where y i and are the true labels and predicted labels respectively.
[0060] In the process of training the model, the loss function of SGMP consists of two parts, including data imputation loss L MI and mortality prediction loss L MP Since data imputation focuses on adversarial training, its weights are 1.0 and 1.5, respectively, so its expression is:
[0061] L MI = L VAE + 1.5L GAN
[0062] Where L VAE is the loss of the variational autoencoder, and L GAN is the adversarial network loss.
[0063] S4, important feature analysis: by calculating the average absolute SHAP value of the feature, the overall contribution of each variable to the prediction result is quantified, as shown in Figure 1 D-dimer (D-Dimer), N-terminal brain natriuretic peptide (NT-ProBNP) and creatinine (Crea) are among the top three feature importance, which is highly consistent with the core prognostic biomarkers listed in the literature, verifying the clinical rationality of the model decision mechanism. Local interpretation based on SHAP analysis quantifies the contribution direction and strength of features to prediction in individual samples, as shown in Figure 2 Blue features in the feature force diagram promote "non-death", and red features promote "death". D-Dimer has a negative SHAP value and a large absolute value in true negative and false positive samples, increasing the risk of positive prediction; in false negative and true positive samples, the SHAP value is positive, reducing the likelihood of positive prediction. NT-ProBNP has a high value in true negative samples, supporting correct prediction; in true positive samples, the SHAP value is highly positive, consistent with heart failure diagnosis. Crea has a significantly negative SHAP value in true negative samples, indicating that impaired renal function dominates negative prediction; in true positive samples, the value is normal, interfering with positive determination; in false negative and false positive samples, the SHAP value is neutral, providing no effective information leading to misjudgment.
[0064] Although embodiments of the present application have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and alterations can be made without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A patient mortality prediction method based on self-attention mechanism and generative adversarial network, characterized by: The following steps are involved: S1. Data collection and preprocessing: Collect electronic medical records of heart failure patients to construct a clinical characteristic dataset including demographic characteristics, medical history records, laboratory test indicators and imaging examination results. The data are anonymized and cleaned three times. S2. Data interpolation based on attention mechanism: An improved variational autoencoder (VAE) is used to encode missing data. The self-attention mechanism is combined with dynamic weighted fusion of multiple sets of Monte Carlo sampling results to generate latent feature representations. The discriminability of the feature space is optimized through an adversarial training framework. S3. Mortality prediction: The complete dataset after interpolation is input into the multi-layer perceptron (MLP) model for training, using the cross entropy loss function and Adam optimizer to complete the binary classification prediction task; S4. Important feature analysis: SHAP values were used to quantify the contribution of each feature to the prediction results and verify the clinical rationality of the model decision-making mechanism.
2. The patient mortality prediction method based on self-attention mechanism and generative adversarial network according to claim 1 is characterized in that: The following steps are involved: According to the steps in S1, S101. Mark Null values, zero values, and negative values as missing values; S102. Set outliers of continuous features as missing; S103. Similar categories of discrete features were integrated, and 105 clinical features were finally retained.
3. The patient mortality prediction method based on self-attention mechanism and generative adversarial network according to claim 1 is characterized in that: The following steps are involved: According to the operation steps in S1, in the three-level data cleaning process, the judgment standard for continuous feature outliers is beyond the range of ±3 times the standard deviation of the mean.
4. The patient mortality prediction method based on self-attention mechanism and generative adversarial network according to claim 1 is characterized in that: The following steps are involved: According to the operation steps in S2, the calculation method of the self-attention mechanism is: randomly sample multiple potential samples from the normal distribution, calculate the weight of each sample by cosine similarity, and generate the final potential variable by weighted summation. The formula is: Among them, the weight Obtained by normalized cosine similarity.
5. The patient mortality prediction method based on self-attention mechanism and generative adversarial network according to claim 1 is characterized in that: The following steps are involved: According to the operation steps in S2, the discriminator network structure of the adversarial training framework is [105, 32, 1], the activation function uses LeakyReLU, and the objective function is: Among them, the generator and the discriminator optimize the quality of generated data through dynamic game.
6. The patient mortality prediction method based on self-attention mechanism and generative adversarial network according to claim 1 is characterized in that: The following steps are involved: According to the operation steps in S3, the network structure of the multi-layer perceptron (MLP) model is [105, 512, 128, 16, 2], the activation function is ReLU, the learning rate is set to 0.001, the training cycle is 450 times, and the loss function is cross entropy loss.
7. The patient mortality prediction method based on self-attention mechanism and generative adversarial network according to claim 1 is characterized in that: The following steps are involved: According to the operation steps in S3, the total loss function of the cross entropy loss function is composed of the data interpolation loss L MI and mortality rate predicted loss L MP The weights are 1.0 and 1.5 respectively, and the expression is: L MI =L VAE +1.5L GAN 8. The patient mortality prediction method based on self-attention mechanism and generative adversarial network according to claim 1 is characterized in that: The following steps are involved: According to the operation steps in S4, the calculated mean absolute SHAP value was used to determine the feature importance, among which D-dimer (D-Dimer), N-terminal pro-brain natriuretic peptide (NT-ProBNP) and creatinine (Crea) were key prognostic biomarkers.
9. The patient mortality prediction method based on self-attention mechanism and generative adversarial network according to claim 1, characterized in that: The following steps are involved: According to the operation steps in S3, the decoder of the variational autoencoder is composed of a multi-layer perceptron. After adversarial training, the generated data is merged with the real data. The formula is: in, Imputed data generated for adversarial networks.
10. The patient mortality prediction method based on self-attention mechanism and generative adversarial network according to claim 1, characterized in that: The following steps are involved: According to the operation steps in S3, the method uses five-fold cross-validation to evaluate the model performance to ensure the generalization and stability of the prediction results.