Prediction method for heart failure after myocardial infarction based on coronary microcirculation and renal function stratification

By employing a stratified method for predicting heart failure after myocardial infarction, combined with coronary microcirculation and renal function indicators, and utilizing multi-factor models and multi-source data feature fusion technology, the accuracy of existing models has been improved, achieving high-precision prediction and individualized assessment of heart failure risk.

CN121483558APending Publication Date: 2026-02-06PEOPLES HOSPITAL PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511640075.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing models for predicting the risk of heart failure after myocardial infarction lack precision, fail to effectively integrate coronary microcirculation resistance index and renal function indicators, resulting in insufficient predictive accuracy for specific subgroups of people, and fail to reveal the complex interactions between key pathophysiological indicators.

Method used

A method for predicting heart failure after myocardial infarction based on coronary microcirculation and renal function stratification was adopted. Feature extraction and model training were performed by combining multi-factor Cox proportional hazards model and simplified multi-factor Cox regression model with multi-source heterogeneous reporting data. Feature fusion was performed using one-dimensional convolutional neural network and random forest model to generate individualized heart failure risk levels.

Benefits of technology

It significantly improves prediction accuracy, particularly in identifying extremely high heart failure risk in patients with normal renal function but impaired coronary microcirculation, and provides simplified bedside scoring rules to facilitate precise clinical intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121483558A_ABST
    Figure CN121483558A_ABST
Patent Text Reader

Abstract

The invention provides a post-myocardial infarction heart failure prediction method based on coronary artery microcirculation and renal function stratification, and belongs to the technical field of intelligent medical prediction. The post-myocardial infarction heart failure prediction method comprises the steps that patients are divided into a normal renal function group and an abnormal renal function group; adopting a first prediction strategy for the normal renal function group to obtain a first total risk score, and adopting a second prediction strategy for the abnormal renal function group to obtain a second total risk score; according to the total risk score corresponding to the group to which the patient belongs, outputting an individualized first heart failure risk level by referring to a preset risk level threshold comparison table; performing result feature extraction on the multi-source heterogeneous report data according to a preset multi-dimensional index to obtain feature vectors, and performing preprocessing and model training on each feature vector in sequence; and obtaining a current variable set of the current patient and combining the first heart failure risk level to obtain a final prediction result and formulating an auxiliary scheme for output. And the prediction precision and the accuracy of the prediction result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent medical prediction, and in particular relates to a post-myocardial infarction heart failure prediction method based on coronary microcirculation and kidney function stratification. BACKGROUND

[0002] Heart failure (HF) after acute myocardial infarction (AMI) is a major cause of rehospitalization and death. Accurate prediction of heart failure risk is crucial for guiding individualized treatment and improving prognosis. Current clinical risk prediction models rely on clinical characteristics (such as age), vital signs, serological markers (such as creatinine, white blood cell count), and imaging indicators (such as left ventricular ejection fraction).

[0003] In recent years, coronary microvascular dysfunction (CMVD) has been recognized as playing a key role in myocardial infarction prognosis. Microcirculatory resistance index based on coronary angiography is an emerging technique for non-invasive assessment of CMVD. However, the existing technology has the following defects: Model homogeneity, lack of precision: existing models are mostly single models applied to all patients, ignoring the significant differences in heart failure risk driving factors under different physiological conditions (such as kidney function). This one-size-fits-all approach leads to insufficient prediction accuracy for specific subgroups.

[0004] Failure to effectively integrate key pathophysiological indicators: existing heart failure risk prediction models fail to systematically incorporate key indicators of coronary microcirculation resistance index reflecting coronary microcirculation status, and fail to reveal complex interactions between other variables (such as age, kidney function), thus limiting further improvement of model prediction performance.

[0005] Therefore, the present application proposes a post-myocardial infarction heart failure prediction method based on coronary microcirculation and kidney function stratification. SUMMARY

[0006] The present application provides a post-myocardial infarction heart failure prediction method based on coronary microcirculation and kidney function stratification to solve the above technical problems.

[0007] The present application provides a post-myocardial infarction heart failure prediction method based on coronary microcirculation and kidney function stratification, comprising the following steps: Step 1: Obtain the baseline serum creatinine level of different acute myocardial infarction patients, and divide the patients into a normal kidney function group and an abnormal kidney function group using a preset critical value; Step 2: obtaining a first total risk score by using a first prediction strategy for the normal renal function group, and obtaining a second total risk score by using a second prediction strategy for the abnormal renal function group, wherein the first prediction strategy comprises a multi-factor Cox proportional hazards model and a first simplified integer scoring system, and the second prediction strategy comprises a simplified multi-factor Cox regression model and a second simplified integer scoring system; Step 3: outputting an individualized first heart failure risk level according to the total risk score corresponding to the group to which the patient belongs and referring to a preset risk level threshold table; Step 4: obtaining multi-source heterogeneous report data based on coronary microcirculation and renal function stratification of each historical patient, and performing result feature extraction on the multi-source heterogeneous report data according to a preset multi-dimensional index to obtain a feature vector, and sequentially performing preprocessing and model training on each feature vector; Step 5: obtaining a current variable set of the current patient, inputting the current variable set into the trained model to obtain a second heart failure risk level, and combining the first heart failure risk level to obtain a final prediction result and formulate an auxiliary scheme output.

[0008] Preferably, the variable set of the multi-factor Cox proportional hazards model comprises age, white blood cell count, history of previous stroke, percutaneous coronary angioplasty, coronary microcirculation dysfunction, and an interaction term of age and coronary microcirculation dysfunction, and the coronary microcirculation dysfunction is defined as a coronary microcirculation resistance index greater than or equal to 25. The variable set of the simplified multi-factor Cox regression model comprises age and white blood cell count.

[0009] Preferably, the multi-factor Cox proportional hazards model is: wherein, wi is a weight of the ith term; xi is a value of the ith term; h0(t) is a baseline risk function; h1(t) is a first risk function at time t.

[0010] Preferably, the simplified multi-factor Cox regression model is: wherein, h2(t) is a second risk function at time t; wj is a weight of the jth term; xj is a value of the jth term.

[0011] Preferably, the multi-source heterogeneous report data comprises clinical electronic medical record reports, coronary angiography reports, laboratory test reports, and dynamic monitoring reports. The feature vector includes clinical semantic correlation, imaging pathology feature values, dynamic changes in laboratory indicators, dynamic monitoring indicators, and time-series risk trend change rates.

[0012] Preferably, each feature vector is preprocessed and the model is trained sequentially, including: Determine the frequency density of the feature vector in the historical dataset. If the frequency density is greater than a first preset density, the feature vector is considered a reasonable vector and retained. If the frequency density is less than a second preset density, the feature vector is considered an unreasonable vector and removed. Otherwise, it is considered a fuzzy vector to be judged. Calculate the deviation coefficient between each feature value in the fuzzy vector to be determined and the reference interval of the corresponding feature dimension. And use piecewise linear functions Determine the reasonable interval membership degree for the corresponding eigenvalue, where x is the eigenvalue. This represents the mean of all historical feature values ​​within the reference interval for the corresponding feature dimension. This represents the standard deviation of all historical feature values ​​within the reference interval for the corresponding feature dimension. It is the minimum value, and takes the value of ; This is the midpoint value of the standardized reference interval; This is the tolerance threshold, used to define the acceptable range of deviation. At the same time, a cubic function is used. Calculate the reasonable distribution membership degree of the fuzzy vector to be determined; Random noise is added to the fuzzy vector to be determined to determine the prediction error volatility, and a hyperbolic tangent function is used. Calculate the robust interval membership degree; For the reasonable distribution membership degree, robust interval membership degree, and all reasonable interval membership degrees, determine the comprehensive evaluation coefficient of the fuzzy vector to be decided; If the comprehensive evaluation coefficient is greater than the preset coefficient, the fuzzy vector to be determined is classified as a reasonable vector and retained; otherwise, it is classified as an unreasonable vector and removed. The retained vectors are used to train the model, resulting in a trained model.

[0013] Preferably, the retained vectors are used to train a model to obtain a trained model, including: The retained feature vectors were divided into a feature subset for the normal renal function group and a feature subset for the abnormal renal function group according to the patient's baseline serum creatinine level; For each feature subset, a one-dimensional convolutional neural network is used to extract features in the temporal dimension, generating intramodal temporal attention weights and obtaining modal temporal fusion features; dividing the modality time sequence fusion features into time sequence risk windows, calculating a weighted center of each time sequence risk window, and generating multi-scale clinical perception weights according to distances between the weighted window center and the time sequence risk window and the distance between the weighted window center ; wherein T is a time sequence length, b0 is a window size; c0 is a sliding step; u0 is a feature modality category; is a clinical sensitivity adjustment parameter; is a time sequence window index; performing cross-modality weighted aggregation on the modality time sequence fusion features according to the multi-scale clinical perception weights to obtain final fusion features; initializing network parameters of the final fusion features of the feature subset of the normal renal function group based on a random forest model, minimizing an F1 loss function by a self-adaptive optimization algorithm, and iteratively training until the AUC and the accuracy of the verification set simultaneously converge; the final fusion features of the feature subset of the abnormal renal function group, using the same model structure as the normal renal function group, introducing an interactive confidence penalty term of serum creatinine and age in the loss function, enhancing the recognition sensitivity of the model to high-risk features of the abnormal renal function subgroup, and outputting a trained model.

[0014] Preferably, based on the second heart failure risk level and in combination with the first heart failure risk level, a final prediction result is obtained, including: determining a first confidence of the first heart failure risk level and a second confidence of the second heart failure risk level; determining a correlation coefficient of each second feature in the second heart failure risk level and each first feature in the first heart failure risk level, generating an initial fusion weight, and obtaining a feature fusion weight in combination with the first confidence and the second confidence to obtain the final prediction result.

[0015] Compared with the prior art, the application has the following beneficial effects: Significant improvement in prediction accuracy: By stratifying renal function and customizing the model, the problem of model homogeneity is solved. Experimental data show that in the normal renal function group, the model discriminant C-index is as high as 0.865 after introducing CMVD and the interaction term, and CMVD is a very strong predictor (risk ratio HR = 836.45, P = 0.014) of this group.

[0016] Pioneering identification of high-risk phenotypes: This application first reveals that in patients with normal renal function but coronary microcirculation dysfunction, there is a very high risk of heart failure, providing a new target for clinical precision intervention.

[0017] ​Clinical practicability is extremely strong: the provided simplified integer scoring system successfully converts the complex statistical interaction (age x CMVD) into a simple bedside scoring rule for the first time, enabling high-precision prediction models to be quickly applied without relying on computers, greatly facilitating clinical translation. Additional features and advantages of the present application will be set forth in the description that follows, and in part will be apparent from the description, or can be learned by practice of the present application. The objectives and other advantages of the present application will be realized and attained by the structure particularly pointed out in the written description and claims thereof as well as the appended drawings.

[0018] The technical solutions of the present application are described in further detail below by means of the accompanying drawings and examples. BRIEF DESCRIPTION OF DRAWINGS

[0019] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, illustrate embodiments of the present application and explain the present application, and do not constitute a limitation of the present application. In the drawings: Figure 1 The flow chart of a post-infarction heart failure prediction method based on coronary microcirculation and kidney function stratification in an embodiment of the present application. DETAILED DESCRIPTION

[0020] The preferred embodiments of the present application are described below in conjunction with the accompanying drawings, and it should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and do not limit the present application.

[0021] The post-infarction heart failure prediction method based on coronary microcirculation and kidney function stratification of the present application comprises the following steps: Step 1: Obtain the baseline serum creatinine level of different acute myocardial infarction patients, and divide the patients into a normal kidney function group and an abnormal kidney function group according to a preset critical value; Step 2: For the normal kidney function group, a first total risk score is obtained by using a first prediction strategy, and for the abnormal kidney function group, a second total risk score is obtained by using a second prediction strategy, wherein the first prediction strategy comprises a multi-factor Cox proportional hazards model and a first simplified integer scoring system, and the second prediction strategy comprises a simplified multi-factor Cox regression model and a second simplified integer scoring system; Step 3: According to the total risk score corresponding to the group to which the patient belongs, and referring to a preset risk level threshold table, an individualized first heart failure risk level is output; Step 4: Obtain multi-source heterogeneous report data based on coronary microcirculation and kidney function stratification of each historical patient, and perform result feature extraction on the multi-source heterogeneous report data according to a preset multi-dimensional index to obtain a feature vector, and sequentially pre-process and model train each feature vector; Step 5: Obtain the current variable set of the current patient, and input it into the trained model to obtain the second heart failure risk level, and combine the first heart failure risk level to obtain the final prediction result and formulate the auxiliary scheme output.

[0022] Preferably, the variable set of the multi-factor Cox proportional hazards model comprises: age, white blood cell count, history of previous stroke, percutaneous coronary angioplasty, coronary microcirculation dysfunction, and the interaction term of age and coronary microcirculation dysfunction, and the coronary microcirculation dysfunction is defined as coronary microcirculation resistance index ≥ 25. The variable set of the simplified multi-factor Cox regression model comprises: age and white blood cell count.

[0023] Preferably, the multi-factor Cox proportional hazards model is: , wherein, is the weight of the ith term; is the value of the ith term; is the baseline risk function; is the first risk function at time t.

[0024] Preferably, the simplified multi-factor Cox regression model is: , wherein, is the second risk function at time t; is the weight of the jth term; is the value of the jth term.

[0025] Preferably, the multi-source heterogeneous report data comprises: clinical electronic medical record reports, coronary imaging reports, laboratory test reports, and dynamic monitoring reports. The feature vector comprises clinical semantic correlation degree, imaging and pathological feature value, laboratory index dynamic change amount, dynamic monitoring index, and time series risk trend change rate.

[0026] In this embodiment, the value of the preset threshold is 107.8 μmol / L, which is determined based on the ROC curve analysis of 2-year heart failure follow-up data of 1200 AMI patients (320 cases in the heart failure occurrence group and 880 cases in the non-occurrence group). The area under the ROC curve (AUC) is 0.82, the sensitivity is 0.78, and the specificity is 0.75, which is the optimal stratification threshold. The patients are divided into normal renal function group (creatinine < 107.8 μmol / L) and abnormal renal function group (creatinine ≥ 107.8 μmol / L).

[0027] In this embodiment, the variable coefficients of the first prediction strategy are as follows: White blood cell count (×109 / L): ; Age (years): ; Previous stroke (yes = 1, no = 0): ; PTCA (Yes = 1, No = 0): ; CMVD (Yes = 1, No = 1): ; Age × CMVD Interaction Items: ; As a key optimization, a simplified integer scoring system is further included. The derivation of the scoring rules is as follows: based on the regression coefficient (β) in the Cox model, the integral value is rounded down by a factor of 10. Continuous variables are assigned values ​​according to clinically interpretable intervals: approximately 7 points are awarded for every 5 years of age increase. One point is awarded for every 1×10⁹ / L increase in white blood cell count. Categorical variables are directly converted based on regression coefficients: Previous stroke history (18 points) PTCA scores 10. When a patient has CMVD, the score is -5 points for every 5 years of age increase. The scores of each patient's indicators were summed to obtain the first overall risk score. The scoring rules for the group with normal renal function (creatinine <107.8 μmol / L) are shown in Table 1.

[0028] Table 1 Scoring rules for the group with normal renal function (creatinine <107.8 μmol / L) Note: PTCA, percutaneous endovascular coronary angioplasty, refers to percutaneous coronary intervention (PCI) without the use of a stent. CMVD, coronary microcirculatory dysfunction, specifically refers to a microcirculatory resistance index (caIMR) ≥25 based on coronary angiography.

[0029] In this embodiment, the variable coefficients of the second prediction strategy are as follows: White blood cell count (×10⁹ / L): ; Age (years): ; Similarly, this strategy includes a simplified integer scoring system. The derivation of the scoring rules is as follows: the regression coefficients are proportionally multiplied by a factor of 10 to form the integral value. For the age variable, approximately 7 points are awarded for every 5 years of age increase. One point is awarded for every 1×10⁹ / L increase in white blood cell count. The scores of the various indicators of the patient are added to obtain a second total risk score. The scoring rules for the abnormal kidney function (creatinine ≥ 107.8 μmol / L) group are shown in Table 2.

[0030] Table 2 Scoring rules for the abnormal kidney function (creatinine ≥ 107.8 μmol / L) group The total score of the final model is the sum of the scores of the various variables, representing the linear prediction value of the individual. Then, the risk of 2-year major adverse cardiovascular events (MACE, death / heart failure rehospitalization / cardiac function III / IV) corresponding to different total scores is calculated according to the model benchmark survival rate function, to form a score-risk mapping table, which is used for clinical risk stratification and individualized prediction.

[0031] The specific operation method is to refer to the following preset risk level threshold table according to the group to which the patient belongs and the total risk score corresponding thereto, to output the individualized heart failure risk level. For the normal kidney function (creatinine < 107.8 μmol / L), refer to Table 3, and for the abnormal kidney function (creatinine ≥ 107.8 μmol / L), refer to Table 4.

[0032] Table 3 Risk level threshold (corresponding to the first total risk score) for the normal kidney function (creatinine < 107.8 μmol / L) group Table 4 Risk level threshold (corresponding to the second total risk score) for the abnormal kidney function (creatinine ≥ 107.8 μmol / L) group In this embodiment, the baseline serum creatinine level refers to the serum creatinine concentration detected by the first collection of venous blood (more than 8 hours of fasting) within 24 hours after the patient is diagnosed with acute myocardial infarction by electrocardiogram and myocardial enzyme detection, reflecting the baseline kidney function status of the patient. The enzyme detection reagent is a creatinine detection reagent kit matched with the instrument, and the sample processing procedure is as follows: centrifuge 3 mL of venous blood (3000 rpm, 10 minutes) to obtain serum, load the sample and reagent according to the kit instructions, incubate at 37°C for 5 minutes, then detect the absorbance, and calculate the creatinine concentration.

[0033] In this embodiment, is the baseline risk function, which is fitted by the Breslow method based on the 2-year follow-up data of 1000 normal group patients = 4%, at this time, the first preset strategy is targeted.

[0034] In this embodiment, is fitted based on the data of 800 abnormal group patients = 12% (higher than that of the normal group), at this time, the second preset strategy is targeted.

[0035] In this embodiment, the multi-source heterogeneous report data contains four types of data: clinical electronic medical record (medical history, PTCA record); coronary artery imaging (caIMR value, coronary stenosis degree); laboratory report (baseline and 72-hour values of creatinine, white blood cell, and B-type natriuretic peptide (BNP)); dynamic monitoring report (24-hour heart rate, blood pressure fluctuation). For example, historical patient C (male, 62 years old, underwent percutaneous coronary intervention (PCI, no stent, only PTCA) 6 hours after myocardial infarction, caIMR = 26, creatinine , white blood cell , 24-hour average heart rate 75 times per minute).

[0036] In this embodiment, when the current patient D (creatinine , normal group), the variable set is input into the model to obtain a second risk level of high risk (confidence 0.88); combined with the first risk level of high risk, the final output is high risk and an auxiliary scheme, such as daily monitoring of vital signs, LVEF, and intensified anti-heart failure drug treatment.

[0037] In this embodiment, the clinical electronic medical record report contains text and structured data of the whole cycle of patient diagnosis and treatment, and the core information includes: basic information (age, gender); medical history (history of hypertension / diabetes / stroke, disease duration); treatment record (PTCA implementation time, PCI intraoperative situation, postoperative medication such as aspirin); chief complaint and history of present illness (time of myocardial infarction, duration of chest pain). For example, the medical record report of patient K: female, 63 years old, admitted to hospital due to chest pain for 4 hours (diagnosed as myocardial infarction), with a history of hypertension for 8 years (taking amlodipine), underwent PCI (no stent, PTCA) 1 hour after admission, and took aspirin and clopidogrel after surgery. Implementation means: structured data is extracted from the hospital HIS system, unstructured text (such as chief complaint) is segmented by Pythonjieba library, and key terms (such as chest pain for 4 hours, PTCA) are extracted.

[0038] In this embodiment, the coronary artery imaging report contains coronary artery anatomy and microcirculation function data: coronary angiography results (stenosis blood vessels (such as left anterior descending branch), stenosis degree (such as 85%)); caIMR value (such as 29); left ventricular ejection fraction (LVEF, echocardiography detection, such as 42%). For example, the imaging report of patient K: left anterior descending branch stenosis 85%, caIMR = 29, LVEF = 42%. After the image is extracted from the PACS system, caIMR is automatically calculated by a coronary function detector, and is manually reviewed and recorded.

[0039] In this embodiment, the laboratory test report contains baseline and dynamic detection indexes: baseline indexes (creatinine, white blood cell count, BNP, troponin I); 72-hour recheck indexes (creatinine, white blood cell, BNP); reference range (such as normal range of creatinine ). For example, the test report of patient K: baseline creatinine (Abnormal), white blood cells BNP 650 pg / mL; 72-hour creatinine White blood cell count 8.5 BNP 520 pg / mL. Data were extracted from the LIS system and organized along a baseline-72-hour timeline. Missing values ​​were filled with the mean of the same group (e.g., mean 72-hour creatinine value for the abnormal group). ).

[0040] In this embodiment, the dynamic monitoring report includes time-series data of physiological indicators: 24-hour Holter monitoring (mean heart rate, highest / lowest heart rate, number of premature ventricular contractions); 24-hour Holter monitoring (mean systolic / diastolic blood pressure, fluctuation range); and daily bedside LVEF (e.g., LVEF changes from postoperative days 1-7). For example, the dynamic report for patient K shows: 24-hour mean heart rate 73 bpm (highest 88, lowest 62), mean systolic blood pressure 135 mmHg (fluctuation range 25 mmHg), LVEF 40% on postoperative day 1, and 42% on postoperative day 3. Time-series data are exported from the ECG monitoring system and bedside ultrasound and organized into a time series by hour / day.

[0041] In this embodiment, the co-occurrence strength of key terms related to myocardial infarction is quantified using clinical semantic association. The TF-IDF weighted co-occurrence method is employed to extract four terms: myocardial infarction, PTCA, CMVD, and hypertension. The co-occurrence frequency is calculated and normalized to (0-1). TF-IDF values ​​are calculated using the Python sklearn library, and a 4×4 co-occurrence matrix is ​​constructed. Association degree = co-occurrence frequency × sum of TF-IDF values ​​ / maximum possible value; Patient K contains myocardial infarction, PTCA, CMVD, and hypertension → co-occurrence frequency 6 → association degree = 6 / 6 = 1.0.

[0042] In this embodiment, the standardized values ​​(dimension-free) of the imaging pathological feature values ​​are standardized using min-max: (actual value - minimum value) / (maximum value - minimum value), with a value ranging from 0 to 1. The reference ranges for the indicators are determined (caIMR 20-40, LVEF 30-60%); for patients, KcaIMR = 29 → (29-20) / (40-20) = 0.45, LVEF = 42% → (42-30) / (60-30) = 0.4.

[0043] In this embodiment, the difference between the baseline and 72-hour period for the dynamic change of laboratory indicators reflects the trend of indicator change: Change = 72-hour value - baseline value (a negative value indicates a decrease, and a reduced risk). The difference is calculated directly; the change in the patient's creatinine K-cell ratio = 108 - 112 = -4. The change in BNP was 520 - 650 = -130 pg / mL.

[0044] In this embodiment, the time series slope of the time series risk trend change rate dynamic indicator reflects the physiological state stability, and a linear regression slope method is adopted, that is, time (days) is used as the x axis, and the indicator value is used as the y axis, y=kx+b is fitted, and k is the change rate. The Python numpy library is used to fit the linear regression;Patient K's LVEF value (40%, 42%) 1-3 days after operation → k=(42-40) / (3-1)=1% / day, in the rising trend, the risk is reduced.

[0045] In this embodiment, the feature vector is, for example, the feature vector of patient K (8 dimensions): {1.0 (clinical correlation degree), 0.45 (caIMR), 0.4 (LVEF), -4 (creatinine change), -130 (BNP change), 73 (mean heart rate), 25 (blood pressure fluctuation), 1 (LVEF change rate)}.

[0046] In this embodiment, multi-source data covers clinical, imaging, laboratory, and dynamic full dimensions, solving the defects of existing models using only static data;The feature vector is converted into a quantifiable feature that can be modeled by standardization and dynamic trend extraction;Clinical semantic correlation degree, time series change rate and other features accurately capture the individual risk differences of patients;Compared with a single type of data model, the multi-source feature model improves the AUC by 10%-15%, providing a comprehensive and high-quality data basis for subsequent model training.

[0047] The beneficial effects of the above technical solutions are: the prediction accuracy is significantly improved: through kidney function stratification and customized model, the model homogenization problem is solved. Experimental data show that in the normal kidney function group, after introducing CMVD and the interaction term, the model discrimination C-index is as high as 0.865, and CMVD is a very strong predictor (risk ratio HR=836.45, P=0.014) of this group.

[0048] It creates a high-risk phenotype: for the first time, it reveals that in patients with normal renal function but coronary microcirculation dysfunction, there is a very high risk of heart failure, providing a new target for clinical precision intervention.

[0049] It has strong clinical practicability: the simplified integer scoring system provided successfully converts the complex statistical interaction (age x CMVD) into a simple bedside scoring rule for the first time, enabling high-precision prediction models to be quickly applied without relying on computers, greatly promoting clinical translation.

[0050] The present application discloses a heart failure prediction method based on coronary microcirculation and kidney function stratification, which sequentially pre-processes and trains each feature vector, including: determining a frequency density of the feature vector in the historical data set, regarding the feature vector as a reasonable vector and retaining if the frequency density is greater than a first preset density, regarding the feature vector as an unreasonable vector and eliminating if the frequency density is less than a second preset density, otherwise, regarding the feature vector as a fuzzy vector to be judged; calculating a deviation coefficient of each feature value in the fuzzy vector to be judged and a reference interval of a corresponding feature dimension , and using a piecewise linear function determining a reasonable interval membership degree of the corresponding feature value, wherein x is the feature value, is a mean value of all historical feature values in the reference interval under the corresponding feature dimension, is a standard deviation of all historical feature values in the reference interval under the corresponding feature dimension; is a minimum value, and the value is ; is a midpoint value of the normalized reference interval; is a tolerance threshold value, used to define a tolerance range of the deviation degree; At the same time, a cubic function is used to calculate a reasonable distribution membership degree of the fuzzy vector to be judged; a random noise is added to the fuzzy vector to be judged to determine a prediction error fluctuation rate, and a hyperbolic tangent function is used to calculate a robust interval membership degree; the reasonable distribution membership degree, the robust interval membership degree and all reasonable interval membership degrees are used to determine a comprehensive evaluation coefficient of the fuzzy vector to be judged; if the comprehensive evaluation coefficient is greater than a preset coefficient, the fuzzy vector to be judged is attributed to a reasonable vector and retained, otherwise, the fuzzy vector to be judged is attributed to an unreasonable vector and eliminated; the retained vector is subjected to model training to obtain a trained model.

[0051] In this embodiment, In this embodiment, for example, the creatinine change amount , the creatinine change amount x=-4, at this time, .

[0052] In this embodiment, the feature vector is as shown in Table 5.

[0053] Table 5 Feature vector contains feature dimension, calculation method and example For example, the feature vector of the patient K is [1.0, 0.45, 0.4, -4, -130, 73, 25, 1].

[0054] In this embodiment, the frequency density is the frequency of a certain feature vector in the historical data set, which is calculated as: the number of occurrences of the vector / the total number of samples, which is valued between 0 and 1, reflecting the clinical commonality of the vector (the higher the frequency, the more reliable the data). The historical data set has 1200 cases (800 normal cases and 400 abnormal cases), and the Python pandas library is used to count the number of occurrences of each vector; the first preset density D1=0.06 (common vector threshold) and the second preset density D2=0.02 (abnormal vector threshold) are set by K-means clustering (k=3). For example, vector L occurs 72 times → frequency density = 72 / 1200 = 0.06 = D1 → retained (reasonable vector); vector M occurs 18 times → frequency density = 0.015 < D2 → rejected (unreasonable vector); vector N occurs 36 times → frequency density = 0.03 (D2 < 0.03 < D1) → fuzzy vector to be judged.

[0055] In this embodiment, the deviation coefficient Cmed quantifies the deviation of the feature value from the historical feature values in the reference interval.

[0056] In this embodiment, the reasonable interval membership degree The segmented linear function quantifies the reasonableness of the feature value in the reference interval, with a value between 0 and 1, The reference interval midpoint is, for example, the BNP change amount = 50, and = 0.5 is the tolerance threshold, which is determined by grid search, and the value of the model validation set AUC is optimal = 0.86.

[0057] In this embodiment, due to the original interval [-20, 10], the standardized reference interval is [-1, 0.5], the midpoint is -0.25, In this embodiment, The cubic function quantifies the reasonableness of the feature value in the historical distribution.

[0058] In this embodiment, the reasonable interval membership degree = the average cosine similarity of the vector and all reasonable vectors.

[0059] In this embodiment, the hierarchical model training is divided into The random forest model is trained in groups: the Python Scikit-learn library is used to build a random forest (100 trees, maximum depth 10); the data is divided into 7:3 (840 training set, 360 validation set); iterative training is performed, and the training is stopped when the validation set AUC is greater than or equal to 0.85 for 5 consecutive rounds; the normal group AUC = 0.88, and the abnormal group AUC = 0.86.

[0060] In this embodiment, the Gaussian noise is added as follows: mean 0, variance 0.01.

[0061] In this embodiment, the preset coefficient is 0.5, which is determined based on minimizing the false positive rate of the validation set.

[0062] In this embodiment, the comprehensive evaluation coefficient all The mean × 0.3.

[0063] The beneficial effects of the above technical solution are as follows: by filtering through frequency density and multiple membership, more than 30% of abnormal data are eliminated, significantly improving the quality of training data; hierarchical training adapts to differences in kidney function, avoiding subgroup bias caused by mixed modeling; the random forest model has strong generalization ability, with a validation set AUC≥0.85, ensuring prediction reliability; compared with the model without preprocessing, this solution improves the prediction accuracy by 8%-12%, providing a high-quality model foundation for two-level fusion.

[0064] This invention discloses a method for predicting post-myocardial infarction heart failure based on coronary microcirculation and renal function stratification. The method involves training a model using retained vectors to obtain a trained model, comprising: The retained feature vectors were divided into a feature subset for the normal renal function group and a feature subset for the abnormal renal function group according to the patient's baseline serum creatinine level; For each feature subset, a one-dimensional convolutional neural network is used to extract features in the temporal dimension, generating intramodal temporal attention weights and obtaining modal temporal fusion features; The temporal fusion features of each modality are divided into temporal risk windows. The weighting center of each temporal risk window is calculated, and then calculated based on clinical risk factors. With weighted window center Distance, generating multi-scale clinical perception weights ; Where T is the time series length, b0 is the window size, c0 is the sliding step size, and u0 is the feature mode category; For clinical sensitivity adjustment parameters; For time-series window index; The modal temporal fusion features are then weighted and aggregated across modalities according to multi-scale clinical perception weights to obtain the final fusion features. The final fusion features of the feature subset of the normal renal function group are initialized based on the network parameters of the random forest model, the adaptive optimization algorithm minimizes the F1 loss function, and iterative training is performed until the AUC and accuracy of the validation set converge simultaneously. For the final fusion features of the feature subset of the renal dysfunction group, the same model structure as the renal function normal group is adopted. An interactive confidence penalty term of serum creatinine and age is introduced into the loss function to enhance the model's sensitivity to identifying high-risk features of the renal dysfunction subgroup, and the trained model is output.

[0065] In this embodiment, the network structure of the one-dimensional convolutional neural network is shown in Table 6.

[0066] Table 6 Network structure of one-dimensional convolutional neural network In this embodiment, the dynamic feature time series of a patient in the normal renal function group is {1.0, 0.3, 0.6, 0.7, -5, -9.1, -150, -37.5, 0.83, -1.25}, which is processed by 1D-CNN to extract 64-dimensional local time sequence features, and then output 128-dimensional intra-modal features through the full connection layer.

[0067] In this embodiment, the intra-modal time sequence attention weight is used to quantify the contribution degree of different time step features to risk prediction, which is obtained by normalizing the importance score of each time step feature, with a value range of 0-1 and a total sum of 1. It is realized by the tf.keras.layers.Attention layer in TensorFlow, and the input is the time sequence feature sequence output by 1D-CNN. During the training process, the weight distribution is automatically adjusted according to the relevance of the feature and the prediction target, so that the key time step feature obtains a higher weight proportion. For example, the time sequence attention weight of a certain feature vector is {0.1, 0.05, 0.1, 0.05, 0.2, 0.1, 0.1, 0.05, 0.1, 0.1}, among which the weight of the 5th time step (related to creatinine change) is the highest, indicating that the feature at this time point has the greatest impact on the prediction result.

[0068] In this embodiment, the modal time sequence fusion feature is obtained by weighted calculation of the time sequence feature x the corresponding attention weight, which integrates the key time sequence information feature set, and the dimension remains the same as the original time sequence feature. For example, if the time sequence feature value of a certain time step is -5 and the corresponding attention weight is 0.05, then the fusion feature value of this time step is -5x0.05=-0.25; the complete modal time sequence fusion feature is {0.1, 0.015, 0.06, 0.035, -1.0, -0.91, -15.0, -1.875, 0.083, -0.125}. The weighted fusion of time sequence features and attention weights is realized by matrix multiplication operation, and Python numpy library is used to complete the calculation to ensure the dimension consistency of the fused features.

[0069] In this embodiment, the clinical risk factor is the weighted window center : The clinical risk factor is a feature risk benchmark value set based on the clinical diagnosis and treatment guidelines, which is used to define the risk level of the feature; the weighted window center is the weighted average value of the feature value in the time sequence window, and the weight is the intra-modal time sequence attention weight. For example, the risk benchmark value of BNP change in the laboratory mode is The concentration was set to -75 pg / mL; a sliding window method was used to divide the time series windows, with a window size of b0=3 and a sliding step size of c0=1 (based on the temporal resolution of clinical monitoring data, 3 data points constitute one risk assessment window), and a weighted average formula was used. Calculate the center of the window, where, The attention weights are for time step t. Let be the feature values ​​of time step t. For example, in a laboratory modality (u0=1), a certain time window contains 3 time steps (t=3,4,5), with feature values ​​[-5,-9.1,-150] and attention weights [0.05,0.2,0.1]. .

[0070] In this embodiment, multi-scale clinical perception weights This parameter measures the degree of fit between the characteristics of each time window and the clinical risk baseline. It is calculated using a normalized tanh function value, ranging from 0 to 1. A clinical sensitivity adjustment parameter is then set. It is determined based on grid search, balancing clinical sensitivity and model stability.

[0071] In this embodiment, for example, the window k0=2 for the laboratory mode u0=1, Based on risk benchmarks set by clinical guidelines.

[0072] In this embodiment, cross-modal weighted aggregation and final fusion features are obtained by weighting and summing the temporal fusion features of different modalities (clinical, imaging, laboratory, dynamic) according to multi-scale clinical perception weights, forming a final feature set that integrates multi-source information. By layer-by-layer concatenation of the features of each modality and then multiplying them with the multi-scale clinical perception weight matrix, weighted aggregation is achieved. The dimension of the final fusion feature is the sum of the dimensions of the features of each modality.

[0073] In this embodiment, the random forest model initialization and adaptive optimization algorithm are as follows: The random forest model initialization uses the output weights of the random forest as the initial parameters of the 1D-CNN to accelerate model convergence; the adaptive optimization algorithm uses the Adam algorithm to minimize the model's prediction error. A random forest model with 100 trees and a maximum depth of 10 is constructed, and its output weights are used as the initial weights of the 1D-CNN; the learning rate of the Adam algorithm is set to 0.001. The loss function used is Binary Crosssentropy, and a custom F1 score is used for model monitoring. During training, the batch size is set to 32, and the iteration is 50 rounds. Training stops when the validation set score shows no improvement for 5 consecutive rounds.

[0074] In this embodiment, the interaction confidence penalty term is a feature enhancement term designed for the abnormal kidney function group, used to improve the model's ability to identify high-risk features of creatinine and age interaction, and is achieved by quantifying the deviation of the interaction feature and the risk threshold. , and the standardization adopts min-max standardization. =0.01, the threshold is set as the median of the high-risk creatinine-age combination of the subgroup, and the threshold is 0.2 after standardization; and the total loss function is .

[0075] In this embodiment, when the model of the normal kidney function group is trained to 30 rounds, the AUC of the validation set reaches 0.88, and the accuracy is 0.85, which meets the convergence condition; after the model of the abnormal kidney function group introduces the penalty term, it can converge after 25 iterations, the AUC of the validation set is 0.86, and the accuracy is 0.83; and the model file is output after training for subsequent prediction.

[0076] The beneficial effects of the above technical solutions are: 1D-CNN extracts time sequence features and combines attention mechanism to accurately capture the dynamic change law of multi-source data; multi-scale clinical perception weight integrates clinical risk benchmark and data distribution characteristics to improve the clinical adaptability of cross-modal features; the interaction confidence penalty term for the abnormal kidney function group strengthens the identification ability of the high-risk features of the subgroup, and compared with the model without hierarchical training, the prediction AUC of the abnormal kidney function group is improved by 10%-15%, realizing accurate risk prediction at the subgroup level.

[0077] The present application discloses a kind of based on coronary microcirculation and kidney function stratification post-infarction heart failure prediction method, based on second heart failure risk level and in combination with first heart failure risk level, obtain final prediction result, including: Determine the first confidence of the first heart failure risk level and the second confidence of the second heart failure risk level; Determine the correlation coefficient of each second feature in the second heart failure risk level and each first feature in the first heart failure risk level, generate initial fusion weight, and obtain feature fusion weight in combination with first confidence, second confidence, obtain final prediction result.

[0078] In this embodiment, when the current variable set and the second heart failure risk level are that the current variable set is the real-time acquisition of the patient's multi-source clinical data, and the format is consistent with the feature vector used for model training; the second heart failure risk level is the risk level division (low / medium / high) obtained by inputting the variable set into the trained model. The real-time data of the current patient is converted into a fixed-dimensional feature vector, which is input into the trained 1D-CNN fusion model. After the model outputs the risk probability, the risk level is determined according to the preset threshold (<10%=low risk, 10%-25%=medium risk, >25%=high risk). For example, the current variable set of patient D is processed to obtain a feature vector such as patient D: 62 years old, creatinine (normal group), the feature vector [0.9, 0.5, 0.45, -3, -120, 75, 22, 0.9]), and the output risk probability is 28% after inputting the model, which is determined as the second heart failure risk level high risk.

[0079] In this embodiment, the first confidence degree = 1- (risk upper limit-risk lower limit) / (2x risk mean), for example, the normal group score is 130 points, that is, the risk is 48%, the range is (40%, 55%), at this time, the first confidence degree = 1- (55-40) / (2x48) = 0.84.

[0080] In this embodiment, the second confidence degree is used to measure the reliability of the prediction result of the second heart failure risk level, which is calculated by the complementary value of probability entropy, and the value range is 0-1. The closer the value is to 1, the more reliable the prediction result is. The entropy calculation formula is adopted , wherein p is the risk probability distribution, , n is the number of risk levels, and the calculation is realized through the Python scipy library.

[0081] In this embodiment, the correlation coefficient is used to quantify the correlation degree of the first heart failure risk level and the real heart failure outcome of the patient, which is calculated by the Pearson correlation coefficient, and the value range is -1 to 1. The closer the value is to 1, the stronger the correlation is. Based on the historical verification set (500 patients), the first risk level (low / medium / high are coded as 1 / 2 / 3) and the real outcome (1=heart failure occurs, 0=heart failure does not occur) are substituted into the Pearson correlation coefficient formula for calculation, which is realized through the Python pandas library. For example, after calculation, the correlation coefficient of the first risk level and the real outcome is 0.75, indicating that the first risk level has a strong indication on the real outcome.

[0082] In this embodiment, the fusion rule is a preset logic for calculating the weight of each risk level based on the confidence and correlation coefficient, and the fusion weight includes the second risk level weight W1 and the first risk level weight W2 (W2=1-W1). The weight coefficient is set to 0.6 and the grid search result based on the AUC maximization of the verification set, and the formula W1=weight coefficient x second confidence+(1-weight coefficient) x correlation coefficient is used for calculation, which ensures that the weight distribution takes into account the real-time prediction reliability and historical data relevance of the model.

[0083] In this embodiment, the first risk level and the second risk level are weighted and summed according to the fusion weight, and the final heart failure risk level is obtained after rounding. The complementary integration of the two risk assessment results is realized. The first risk level code value L1 and the second risk level code value L2 are substituted into the formula Final_L=W1xL2+W2xL1, and the calculation result is rounded after retaining two decimal places to obtain the final level. For example, the first risk level L1 of patient D is 2 (medium risk), and the second risk level L2 is 3 (high risk). Final_L=0.81x3+0.19x2=3, and the final risk level is high risk.

[0084] In this embodiment, the auxiliary scheme is a set of clinical intervention suggestions based on the final risk level, which covers monitoring frequency, drug adjustment, lifestyle guidance and other contents, and matches the severity of the risk level. A risk level-auxiliary scheme mapping table is established in advance, as shown in Table 7.

[0085] Table 7 Mapping table The beneficial effects of the above technical solutions are: through the fusion weight calculation of confidence and correlation coefficient, the advantages of real-time model prediction and historical risk assessment are fully integrated, and the deviation of single evaluation method is avoided; the double risk level fusion makes the final prediction accuracy increase by 8%-12% compared with single level; the auxiliary scheme is accurately matched with the risk level, which directly connects with the clinical intervention scene and can be quickly converted into individualized treatment decision, providing clear action guidance for clinicians, which helps to reduce the incidence of heart failure and rehospitalization rate of patients.

[0086] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.

Claims

1. A method for predicting heart failure after myocardial infarction based on coronary microcirculation and renal function stratification, characterized in that, Includes the following steps: Step 1: Obtain the baseline serum creatinine levels of different patients with acute myocardial infarction, and divide the patients into a normal renal function group and a renal function abnormal group based on a preset threshold. Step 2: For the group with normal renal function, the first prediction strategy is used to obtain the first total risk score. At the same time, for the group with abnormal renal function, the second prediction strategy is used to obtain the second total risk score. The first prediction strategy includes a multivariate Cox proportional hazards model and a first simplified integer scoring system. The second prediction strategy includes a simplified multivariate Cox regression model and a second simplified integer scoring system. Step 3: Based on the total risk score corresponding to the patient's group and referring to the preset risk level threshold comparison table, output the individualized first heart failure risk level; Step 4: Obtain multi-source heterogeneous report data based on the coronary microcirculation and renal function stratification of each historical patient, and extract feature vectors from the multi-source heterogeneous report data according to preset multi-dimensional indicators, and perform preprocessing and model training on each feature vector in sequence; Step 5: Obtain the current set of variables for the patient and input it into the trained model to obtain the second heart failure risk level. Combine this with the first heart failure risk level to obtain the final prediction result and formulate an auxiliary plan for output.

2. The method for predicting post-myocardial infarction heart failure based on coronary microcirculation and renal function stratification according to claim 1, characterized in that, The variable set of the multivariate Cox proportional hazards model includes: age, white blood cell count, history of stroke, percutaneous coronary angioplasty, coronary microcirculation dysfunction, and the interaction term between age and coronary microcirculation dysfunction, and coronary microcirculation dysfunction is defined as coronary microcirculation resistance index ≥25. The set of variables in the simplified multivariate Cox regression model includes age and white blood cell count.

3. The method for predicting post-myocardial infarction heart failure based on coronary microcirculation and renal function stratification according to claim 1, characterized in that, The multifactor Cox proportional hazards model is as follows: ,in, Let be the weight of the i-th term; Let be the value of the i-th term; Baseline risk function; Let be the first risk function at time t.

4. The method for predicting post-myocardial infarction heart failure based on coronary microcirculation and renal function stratification according to claim 3, characterized in that, The simplified multifactor Cox regression model is as follows: ,in, This is the second risk function at time t; Let be the weight of the j-th term; Let be the value of the j-th term.

5. The method for predicting post-myocardial infarction heart failure based on coronary microcirculation and renal function stratification according to claim 1, characterized in that, The multi-source heterogeneous reporting data includes: clinical electronic medical record reports, coronary artery imaging reports, laboratory test reports, and dynamic monitoring reports; The feature vector includes clinical semantic correlation, imaging pathology feature values, dynamic changes in laboratory indicators, dynamic monitoring indicators, and time-series risk trend change rates.

6. The method for predicting post-myocardial infarction heart failure based on coronary microcirculation and renal function stratification according to claim 1, characterized in that, Each feature vector is preprocessed and the model is trained sequentially, including: Determine the frequency density of the feature vector in the historical dataset. If the frequency density is greater than a first preset density, the feature vector is considered a reasonable vector and retained. If the frequency density is less than a second preset density, the feature vector is considered an unreasonable vector and removed. Otherwise, it is considered a fuzzy vector to be judged. Calculate the deviation coefficient between each feature value in the fuzzy vector to be determined and the reference interval of the corresponding feature dimension. And use piecewise linear functions Determine the reasonable interval membership degree for the corresponding eigenvalue, where x is the eigenvalue. This represents the mean of all historical feature values ​​within the reference interval for the corresponding feature dimension. This represents the standard deviation of all historical feature values ​​within the reference interval for the corresponding feature dimension. It is the minimum value, and takes the value of ; This is the midpoint value of the standardized reference interval; This is the tolerance threshold, used to define the acceptable range of deviation. At the same time, a cubic function is used Calculate the reasonable distribution membership degree of the fuzzy vector to be determined; Random noise is added to the fuzzy vector to be determined to determine the prediction error volatility, and a hyperbolic tangent function is used. Calculate the robust interval membership degree; For the reasonable distribution membership degree, robust interval membership degree, and all reasonable interval membership degrees, determine the comprehensive evaluation coefficient of the fuzzy vector to be decided; If the comprehensive evaluation coefficient is greater than the preset coefficient, the fuzzy vector to be determined is classified as a reasonable vector and retained; otherwise, it is classified as an unreasonable vector and removed. The retained vectors are used to train the model, resulting in a trained model.

7. The method for predicting post-myocardial infarction heart failure based on coronary microcirculation and renal function stratification according to claim 6, characterized in that, The retained vectors are used to train the model, resulting in a trained model, including: The retained feature vectors were divided into a feature subset for the normal renal function group and a feature subset for the abnormal renal function group according to the patient's baseline serum creatinine level; For each feature subset, a one-dimensional convolutional neural network is used to extract features in the temporal dimension, generating intramodal temporal attention weights and obtaining modal temporal fusion features; The temporal fusion features of each modality are divided into temporal risk windows. The weighting center of each temporal risk window is calculated, and then calculated based on clinical risk factors. With weighted window center Distance, generating multi-scale clinical perception weights ; Where T is the time series length, b0 is the window size, c0 is the sliding step size, and u0 is the feature mode category; For clinical sensitivity adjustment parameters; For time-series window index; The modal temporal fusion features are then weighted and aggregated across modalities according to multi-scale clinical perception weights to obtain the final fusion features. The final fusion features of the feature subset of the normal renal function group are initialized based on the network parameters of the random forest model, the adaptive optimization algorithm minimizes the F1 loss function, and iterative training is performed until the AUC and accuracy of the validation set converge simultaneously. For the final fusion features of the feature subset of the renal dysfunction group, the same model structure as the renal function normal group is adopted. An interactive confidence penalty term of serum creatinine and age is introduced into the loss function to enhance the model's sensitivity to identifying high-risk features of the renal dysfunction subgroup, and the trained model is output.

8. The method for predicting post-myocardial infarction heart failure based on coronary microcirculation and renal function stratification according to claim 1, characterized in that, Based on the second heart failure risk level and combined with the first heart failure risk level, the final prediction results are obtained, including: The first confidence level for determining the first risk level of heart failure and the second confidence level for determining the second risk level of heart failure; The correlation coefficients between each second feature in the second heart failure risk level and each first feature in the first heart failure risk level are determined, initial fusion weights are generated, and feature fusion weights are obtained by combining the first confidence level and the second confidence level to obtain the final prediction result.