Method and system for predicting cardiovascular event risk in copd patients based on multi-modal data

CN122531752APending Publication Date: 2026-08-07GUANGZHOU MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU MEDICAL UNIV
Filing Date
2026-07-04
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

现有技术在聚类建模时纳入的变量几乎完全局限于呼吸系统本身的临床表现和全身性炎症的血液替代指标,未能系统性地整合能够直接、客观评价心脏结构和功能的心脏超声影像学标志物,导致所识别出的亚型无法系统性地反映心肺共病这一核心病理轴对患者异质性的贡献,所以在预测目标不良心血管事件这一特定心肺复合终点时,现有技术的精确性和灵敏度低

Benefits of technology

[0015]The aforementioned method and system for predicting cardiovascular event risk in COPD patients based on multimodal data provides a solution that simultaneously acquires demographic data, blood biomarker data, and echocardiographic data from the electronic medical record system. It incorporates qualitative indicators of cardiac structural abnormalities and quantitative parameters of cardiac function and chambers into the same binary coding framework as demographic characteristics and blood biomarkers. This introduces echocardiographic parameters, which can directly and objectively reflect the state of cardiac structural remodeling and functional impairment, into the patient heterogeneity characterization system. This allows the identified patient subtypes to cover the dimension of cardiac target organ involvement on a pathophysiological basis, thus compensating for the structural lack of cardiopulmonary comorbidity information in existing classification schemes. This invention employs a three-stage cascaded screening process: sequentially performing univariate difference significance testing, lasso regression compression screening, and multivariate logistic regression confirmation of independent effects on multimodal binary feature vectors. This process eliminates variables with weak predictive contributions or those exhibiting collinearity with other variables during the progressive compression process. Ultimately, the core feature set is streamlined in terms of the number of variables, while retaining variable combinations with independent predictive effects on target adverse cardiovascular events. This reduces the complexity and fitting difficulty of the indicator variable set for subsequent latent class analysis models while maintaining the integrity of predictive information. This invention comprehensively determines the optimal number of categories in a latent category analysis model by combining multiple goodness-of-fit indices, including the Bayesian information criterion, the consistent Akaike information criterion, the sample size-corrected Bayesian information criterion, entropy, bootstrap likelihood ratio test, and the Rohman-Rubin likelihood ratio test, with model comparison tests. This avoids the bias that may arise from determining the number of categories based on a single indicator. Furthermore, it assigns corresponding clinical labels based on the differences in the conditional response probability distribution of each subtype on the corresponding indicator variables in the core feature set. This results in the final subtypes having clear distinguishability and clinical interpretability in pathological features such as cardiopulmonary structural remodeling, cardiorenal function, and valvular degeneration. This invention calculates the association strength between subtypes and target adverse cardiovascular events and specific cardiovascular diseases by using both logistic regression and Bolke-Kroll-Hageners models in parallel based on posterior probability vectors. The logistic regression model calculates association strength based on modal classification results, while the Bolke-Kroll-Hageners model constructs classification error correction weights based on posterior probability vectors before performing distant outcome analysis. The association strength results obtained from the two approaches are compared and verified, ensuring that the final cardiovascular event risk stratification results, after correcting for effect attenuation bias that may be introduced by latent category classification uncertainty, maintain consistency in the direction and value of association strength. This improves the robustness and reliability of subtype-based cardiovascular event risk stratification results, providing clinicians with a methodologically sound decision-making basis for developing individualized cardiopulmonary comorbidity prevention and treatment strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531752A_ABST
    Figure CN122531752A_ABST
Patent Text Reader

Abstract

The application is suitable for the technical field of multi-modal data prediction, and provides a COPD patient cardiovascular event risk prediction method and system based on multi-modal data.The method comprises the following steps: acquiring demographic characteristic data, blood biomarker data and echocardiographic index data of ECOPD patients, and constructing a multi-modal binary feature vector; performing single-factor difference significance test, lasso regression compression screening and multi-factor logistic regression independent effect confirmation on the multi-modal binary feature vector to obtain a core feature set; determining a posterior probability vector of each patient belonging to each subtype according to the core feature set, and calculating a cardiovascular event risk stratification result corresponding to each subtype based on the posterior probability vector.The application improves the robustness and reliability of the cardiovascular event risk stratification result based on the subtype, and provides a methodologically guaranteed decision basis for a clinician to formulate an individualized cardiopulmonary comorbidity prevention and treatment strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multimodal data prediction technology, and in particular to a method and system for predicting cardiovascular event risk in COPD patients based on multimodal data. Background Technology

[0002] Chronic obstructive pulmonary disease (COPD) and cardiovascular disease (CVD) are both high-burden epidemics worldwide. CVD is reported as the most important and common non-respiratory comorbidity of COPD, significantly increasing the hospitalization risk and worsening long-term prognosis in COPD patients. Acute exacerbations of COPD (ECOPD), a key exacerbation event in the course of COPD, not only accelerate the irreversible decline in lung function but also induce or exacerbate the risk of target adverse cardiovascular events (MACEs, such as cardiovascular death, non-fatal myocardial infarction, and stroke). Due to the high heterogeneity of the ECOPD patient population in terms of demographic characteristics, inflammation levels, respiratory function, and the degree of cardiac structural and functional impairment, single-dimensional biomarkers are insufficient to comprehensively capture this complex comorbidity and have limited predictive power for MACE.

[0003] To address the heterogeneity of ECOPD patients, researchers have attempted latent clustering schemes based on clinical characteristics and blood biomarkers. These schemes collect patient demographic and baseline characteristics, dyspnea severity scores, COPD duration, cardiovascular history, and peripheral blood inflammatory markers such as white blood cell count, eosinophil count, C-reactive protein, and fibrinogen from hospital information systems. Continuous variables are converted to categorical variables based on clinical cutoff values ​​and then input into a latent clustering model. The optimal number of clusters is determined using a goodness-of-fit index, dividing the study population into several subtypes with different clinical characteristics. Prospective follow-up studies are used to preliminarily validate differences in prognostic endpoints among these subtypes. However, current techniques almost entirely limit the variables included in clustering modeling to clinical manifestations of the respiratory system and blood substitutes for systemic inflammation. They fail to systematically integrate echocardiographic biomarkers that can directly and objectively evaluate cardiac structure and function. Consequently, the identified subtypes cannot systematically reflect the contribution of the core pathological axis of cardiopulmonary comorbidity to patient heterogeneity. Therefore, current techniques exhibit low accuracy and sensitivity in predicting the specific cardiopulmonary composite endpoint of adverse cardiovascular events. Summary of the Invention

[0004] This invention provides a method and system for predicting cardiovascular event risk in COPD patients based on multimodal data. This invention improves the robustness and reliability of subtype-based cardiovascular event risk stratification results, and provides clinicians with a methodologically sound decision-making basis for developing individualized cardiopulmonary comorbidity prevention and treatment strategies.

[0005] In a first aspect, embodiments of this application provide a method for predicting the risk of cardiovascular events in COPD patients based on multimodal data, including: Acquire demographic data, blood biomarker data, and echocardiographic data of ECOPD patients, and construct a multimodal binary feature vector; The multimodal binary feature vectors were subjected to single-factor difference significance test, lasso regression compression screening, and multi-factor logistic regression independent effect confirmation to obtain the core feature set; Based on the core feature set, the posterior probability vector of each patient belonging to each subtype is determined, and the cardiovascular event risk stratification result corresponding to each subtype is calculated based on the posterior probability vector.

[0006] Optionally, in a first implementation of the first aspect of the present invention, the step of acquiring demographic data, blood biomarker data, and echocardiographic data of ECOPD patients, and constructing a multimodal binary feature vector, includes: Demographic data, blood biomarker data, and echocardiogram data of ECOPD patients were obtained from the electronic medical record system. The demographic data, blood biomarker data, and echocardiogram data are divided into a set of continuous variables and a set of qualitative variables. The set of continuous variables is subjected to threshold binary encoding to obtain encoded continuous variables, and a multimodal binary feature vector is constructed based on the encoded continuous variables and the set of qualitative variables.

[0007] Optionally, in a second implementation of the first aspect of the present invention, the step of performing a single-factor significance test, lasso regression compression screening, and multi-factor logistic regression independent effect confirmation on the multimodal binary feature vector to obtain a core feature set includes: The ECOPD patients were divided into a positive group and a negative group, and a one-way significance test was performed on the multimodal binary feature vectors between the positive group and the negative group to obtain a set of significantly different variables. The set of significantly different variables is input into the lasso regression model for coefficient compression to obtain the initial screening feature set. The initial screening feature set is input into a multivariate logistic regression model for independent prediction effect confirmation to obtain the core feature set, which includes age, blood glucose, eosinophil percentage, serum creatinine, left atrial diameter, pulmonary artery diameter, fractional shortening rate, left atrial enlargement, and aortic valve degeneration.

[0008] Optionally, in a third implementation of the first aspect of the present invention, the step of determining the posterior probability vector of each patient belonging to each subtype based on the core feature set, and calculating the cardiovascular event risk stratification result corresponding to each subtype based on the posterior probability vector, includes: The core feature set is used as an indicator variable and input into latent class analysis models with different numbers of categories. The Bayesian information criterion, the consistent Akaike information criterion, the Bayesian information criterion corrected by the sample size, and the entropy value corresponding to each latent class analysis model are calculated to obtain the set of fitting indices for each latent class analysis model. Based on the set of fitting indices, perform bootstrap likelihood ratio test and Roe-Mendel-Rubin likelihood ratio test on the latent class analysis models with adjacent number of classes to determine the latent class analysis model corresponding to the optimal number of classes. The core feature set is input into the latent class analysis model corresponding to the optimal number of categories to calculate the posterior probability vector of each patient belonging to each subtype; Based on the comparison and verification of the posterior probability vector, the cardiovascular event risk stratification results corresponding to each subtype are obtained.

[0009] Optionally, in a fourth implementation of the first aspect of the present invention, the step of inputting the core feature set into the latent class analysis model corresponding to the optimal number of categories, and calculating the posterior probability vector of each patient belonging to each subtype, includes: The conditional response probability distributions of each subtype in the latent category analysis model corresponding to the optimal number of categories on each indicator variable of the core feature set are compared to obtain the feature patterns of each subtype. Based on the aforementioned characteristic patterns, the subtype with the highest probability of responding to abnormal levels of left atrial diameter, pulmonary artery diameter, and left atrial enlargement is identified as the cardiopulmonary structural remodeling with metabolic abnormalities subtype; the subtype with the lowest probability of responding to advanced age is identified as the relatively young subtype with normal cardiac and renal function; and the subtype with the highest probability of responding to aortic valve degeneration is identified as the cardiac aging subtype characterized by valvular degeneration. The core feature set of each patient is substituted into the latent class analysis model corresponding to the optimal number of categories to calculate the posterior probability, thereby obtaining the posterior probability vector of each patient belonging to each subtype.

[0010] Optionally, in a fifth implementation of the first aspect of the present invention, the step of performing comparison and verification based on the posterior probability vector to obtain the cardiovascular event risk stratification results corresponding to each subtype includes: The subtype corresponding to the maximum posterior probability in the posterior probability vector is determined as the modality classification result for each patient. The modality classification results are input into a logistic regression model to calculate the first association strength between each subtype and the target adverse cardiovascular events and cardiovascular diseases. Based on the posterior probability vector, classification error correction weights are constructed, and based on the classification error correction weights, remote outcome analysis is performed on the modality classification results, the target adverse cardiovascular events, and the cardiovascular diseases to obtain the second association strength result. The first correlation strength result and the second correlation strength result are compared and verified to obtain the cardiovascular event risk stratification results corresponding to each subtype.

[0011] Secondly, embodiments of this application provide a cardiovascular event risk prediction system for COPD patients based on multimodal data, including: The data acquisition module is used to acquire demographic data, blood biomarker data, and echocardiographic data of ECOPD patients, and to construct a multimodal binary feature vector. The feature analysis module is used to perform single-factor difference significance test, lasso regression compression screening, and multi-factor logistic regression independent effect confirmation on the multimodal binary feature vector to obtain the core feature set; The risk stratification module is used to determine the posterior probability vector of each patient belonging to each subtype based on the core feature set, and to calculate the cardiovascular event risk stratification result corresponding to each subtype based on the posterior probability vector.

[0012] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the above-described method for predicting the cardiovascular event risk of COPD patients based on multimodal data.

[0013] Fourthly, embodiments of this application provide a readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method for predicting cardiovascular event risk in COPD patients based on multimodal data.

[0014] Fifthly, embodiments of this application provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the above-described method for predicting cardiovascular event risk in COPD patients based on multimodal data.

[0015] The aforementioned method and system for predicting cardiovascular event risk in COPD patients based on multimodal data provides a solution that simultaneously acquires demographic data, blood biomarker data, and echocardiographic data from the electronic medical record system. It incorporates qualitative indicators of cardiac structural abnormalities and quantitative parameters of cardiac function and chambers into the same binary coding framework as demographic characteristics and blood biomarkers. This introduces echocardiographic parameters, which can directly and objectively reflect the state of cardiac structural remodeling and functional impairment, into the patient heterogeneity characterization system. This allows the identified patient subtypes to cover the dimension of cardiac target organ involvement on a pathophysiological basis, thus compensating for the structural lack of cardiopulmonary comorbidity information in existing classification schemes. This invention employs a three-stage cascaded screening process: sequentially performing univariate difference significance testing, lasso regression compression screening, and multivariate logistic regression confirmation of independent effects on multimodal binary feature vectors. This process eliminates variables with weak predictive contributions or those exhibiting collinearity with other variables during the progressive compression process. Ultimately, the core feature set is streamlined in terms of the number of variables, while retaining variable combinations with independent predictive effects on target adverse cardiovascular events. This reduces the complexity and fitting difficulty of the indicator variable set for subsequent latent class analysis models while maintaining the integrity of predictive information. This invention comprehensively determines the optimal number of categories in a latent category analysis model by combining multiple goodness-of-fit indices, including the Bayesian information criterion, the consistent Akaike information criterion, the sample size-corrected Bayesian information criterion, entropy, bootstrap likelihood ratio test, and the Rohman-Rubin likelihood ratio test, with model comparison tests. This avoids the bias that may arise from determining the number of categories based on a single indicator. Furthermore, it assigns corresponding clinical labels based on the differences in the conditional response probability distribution of each subtype on the corresponding indicator variables in the core feature set. This results in the final subtypes having clear distinguishability and clinical interpretability in pathological features such as cardiopulmonary structural remodeling, cardiorenal function, and valvular degeneration. This invention calculates the association strength between subtypes and target adverse cardiovascular events and specific cardiovascular diseases by using both logistic regression and Bolke-Kroll-Hageners models in parallel based on posterior probability vectors. The logistic regression model calculates association strength based on modal classification results, while the Bolke-Kroll-Hageners model constructs classification error correction weights based on posterior probability vectors before performing distant outcome analysis. The association strength results obtained from the two approaches are compared and verified, ensuring that the final cardiovascular event risk stratification results, after correcting for effect attenuation bias that may be introduced by latent category classification uncertainty, maintain consistency in the direction and value of association strength. This improves the robustness and reliability of subtype-based cardiovascular event risk stratification results, providing clinicians with a methodologically sound decision-making basis for developing individualized cardiopulmonary comorbidity prevention and treatment strategies. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of a COPD patient cardiovascular event risk prediction device based on multimodal data in one embodiment of the present invention; Figure 2 This is a flowchart illustrating a method for predicting cardiovascular event risk in COPD patients based on multimodal data, according to one embodiment of the present invention. Figure 3 yes Figure 2 A schematic diagram of the implementation process of step S10; Figure 4 yes Figure 2 A schematic diagram of the implementation process of step S20; Figure 5 yes Figure 2 A schematic diagram of the implementation process of step S30; Figure 6 yes Figure 5 A schematic diagram of the implementation process of step S33; Figure 7 This is a schematic diagram of a cardiovascular event risk prediction system for COPD patients based on multimodal data, according to one embodiment of the present invention. Figure 8 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof. It should also be understood that, as used in this specification and the appended claims, the term "and / or" refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0020] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0021] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0022] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0023] To illustrate the technical solution of the present invention, specific embodiments are described below.

[0024] To address the problems mentioned above, this application proposes a method and system for predicting cardiovascular event risk in COPD patients based on multimodal data. The method for predicting cardiovascular event risk in COPD patients based on multimodal data provided by this invention can be applied to, for example... Figure 1 The COPD patient cardiovascular event risk prediction device shown includes a client and a server.

[0025] In one embodiment, such as Figure 2 As shown, a method for predicting cardiovascular event risk in COPD patients based on multimodal data is provided, and this method is applied to... Figure 1 The following is an example of a COPD patient cardiovascular event risk prediction device based on multimodal data, which includes the following steps: S10: Obtain demographic data, blood biomarker data, and echocardiographic data of ECOPD patients, and construct a multimodal binary feature vector; S20: Perform single-factor significance test, lasso regression compression screening, and multi-factor logistic regression to confirm the independent effects of the multimodal binary feature vectors to obtain the core feature set; S30: Determine the posterior probability vector of each patient’s subtype based on the core feature set, and calculate the cardiovascular event risk stratification results corresponding to each subtype based on the posterior probability vector.

[0026] In this embodiment, demographic data, blood biomarker data, and echocardiogram data are simultaneously acquired from the electronic medical record system. Qualitative indicators of cardiac structural abnormalities and quantitative parameters of cardiac function and chambers are incorporated into the same binary coding framework as demographic characteristics and blood biomarkers. Echocardiogram parameters, which directly and objectively reflect the state of cardiac structural remodeling and functional impairment, are introduced into the patient heterogeneity characterization system. This allows the identified patient subtypes to cover the dimension of cardiac target organ involvement on a pathophysiological basis, compensating for the structural deficiency of cardiopulmonary comorbidity information in existing classification schemes. This invention employs a three-stage cascaded screening process on multimodal binary feature vectors, sequentially performing univariate difference significance testing, lasso regression compression screening, and multivariate logistic regression independent effect confirmation. Variables with weak predictive contributions or collinearity with other variables are eliminated during the progressive compression process. The final core feature set is streamlined in terms of the number of variables, while retaining variable combinations with independent predictive effects on target adverse cardiovascular events. This reduces the model complexity and fitting difficulty of the indicator variable set in subsequent latent class analysis models while maintaining the integrity of predictive information. This invention comprehensively determines the optimal number of categories in a latent category analysis model by combining multiple goodness-of-fit indices, including the Bayesian information criterion, the consistent Akaike information criterion, the sample size-corrected Bayesian information criterion, entropy, bootstrap likelihood ratio test, and the Rohman-Rubin likelihood ratio test, with model comparison tests. This avoids the bias that may arise from determining the number of categories based on a single indicator. Furthermore, it assigns corresponding clinical labels based on the differences in the conditional response probability distribution of each subtype on the corresponding indicator variables in the core feature set. This results in the final subtypes having clear distinguishability and clinical interpretability in pathological features such as cardiopulmonary structural remodeling, cardiorenal function, and valvular degeneration. This invention calculates the association strength between subtypes and target adverse cardiovascular events and specific cardiovascular diseases by using both logistic regression and Bolke-Kroll-Hageners models in parallel based on posterior probability vectors. The logistic regression model calculates association strength based on modal classification results, while the Bolke-Kroll-Hageners model constructs classification error correction weights based on posterior probability vectors before performing distant outcome analysis. The association strength results obtained from the two approaches are compared and verified, ensuring that the final cardiovascular event risk stratification results, after correcting for effect attenuation bias that may be introduced by latent category classification uncertainty, maintain consistency in the direction and value of association strength. This improves the robustness and reliability of subtype-based cardiovascular event risk stratification results, providing clinicians with a methodologically sound decision-making basis for developing individualized cardiopulmonary comorbidity prevention and treatment strategies.

[0027] In one embodiment, such as Figure 3 As shown, step S10 specifically includes the following steps: S11: Obtain demographic data, blood biomarker data, and echocardiogram data of ECOPD patients from the electronic medical record system; S12: Divide demographic data, blood biomarker data, and echocardiography data into a set of continuous variables and a set of qualitative variables; S13: Perform threshold binary encoding on the set of continuous variables to obtain the encoded continuous variables, and construct a multimodal binary feature vector based on the encoded continuous variables and the set of qualitative variables.

[0028] In this embodiment, a standardized data extraction interface for ECOPD patients is established in the electronic medical record system. The patient's unique visit number or hospitalization number is used as the main index to associate the admission diagnosis, discharge diagnosis, examination time, laboratory time, and echocardiogram report time. This ensures that demographic data, blood biomarker data, and echocardiogram index data all correspond to the same hospitalization cycle or the same acute exacerbation event. Target patients are screened based on diagnostic codes, diagnostic names, or acute exacerbation markers of chronic obstructive pulmonary disease in admission records. Demographic and clinical information, such as age, gender, department of visit, treatment method, medical insurance type, and comorbidity records, are then retrieved from the patient's medical record cover sheet, admission record, and medical order records. The first blood biomarker results after admission are retrieved from the laboratory system. Blood biomarkers can cover types such as complete blood count components, liver function, kidney function, myocardial function, inflammatory markers, arterial blood gas analysis, glucose metabolism markers, and exhaled nitric oxide levels. Up to 31 blood biomarkers can be retained as candidate inputs. At the same time, qualitative results of structural abnormalities and cardiac function parameters are retrieved from the cardiac ultrasound system. Up to 21 qualitative indicators of cardiac structural abnormalities can be retained, such as left atrial enlargement, right atrial enlargement, aortic valve degeneration, pulmonary hypertension, pericardial effusion, and valvular regurgitation. Quantifiable ultrasound parameters such as left atrial diameter, pulmonary artery diameter, and short axis shortening are extracted simultaneously.

[0029] The data were uniformly categorized according to variable attributes. Fields with continuous numerical meaning and comparable magnitudes, such as age, blood glucose, eosinophil percentage, serum creatinine, left atrial diameter, pulmonary artery diameter, and fractional shortening, were classified into the continuous variable set. Fields that naturally express "present / absent," "yes / no," or category labels, such as gender, whether inhaled corticosteroids were used, presence of left atrial enlargement, presence of aortic valve degeneration, presence of pulmonary hypertension, and presence of pericardial effusion, were classified into the qualitative variable set. Before classification, units, missing values, and outliers were standardized. For example, blood glucose was standardized to mmol / L, serum creatinine to μmol / L, and left atrial diameter and pulmonary artery diameter to mm. Data with unidentifiable units or that did not match the examination time were marked as records to be reviewed.

[0030] Threshold binary coding is applied to the set of continuous variables. The coding principle is to compare each continuous variable with a preset threshold or normal reference range. If the variable reaches an abnormal, high-risk, or elevated state, it is coded as 1; otherwise, it is coded as 0. When applying threshold binary coding to the set of continuous variables, each continuous variable is compared with a preset reference range, examination standard, or institutional configuration threshold. The abnormal judgment thresholds for variables such as age, blood glucose, eosinophil percentage, serum creatinine, left atrial diameter, pulmonary artery diameter, and fractional shortening can be determined based on patient gender, body surface area, laboratory reference intervals, ultrasound diagnostic standards, and medical institution configuration tables. If a variable is higher than the upper limit, lower than the lower limit, or falls into the abnormal range, the threshold is set accordingly. If the threshold is 1, it is coded as 1; otherwise, it is coded as 0. These thresholds can be configured in different hospital versions by the test reference interval and ultrasound diagnostic specifications to adapt the system to different testing platforms. Qualitative variables are no longer repeatedly binarized. Instead, "existence, positive, abnormal, present" are uniformly mapped to 1, and "absence, negative, normal, none" are uniformly mapped to 0. For multi-category variables, they can be split into several binary dummy variables according to whether they are related to cardiovascular event risk. The coded continuous variables and qualitative variables are concatenated in a fixed field order to form a multimodal binary feature vector for each patient, so that the same patient carries demographic status, blood metabolic inflammation status, and cardiac structural and functional status in one vector.

[0031] In one embodiment, such as Figure 4 As shown, step S20 specifically includes the following steps: S21: Divide ECOPD patients into positive and negative groups, and perform a one-way significance test on the multimodal binary eigenvectors between the positive and negative groups to obtain a set of significantly different variables; S22: Input the set of significantly different variables into the lasso regression model and compress the coefficients to obtain the initial screening feature set; S23: Input the initial screening feature set into the multivariate logistic regression model to confirm the independent predictive effect and obtain the core feature set. The core feature set includes age, blood glucose, eosinophil percentage, serum creatinine, left atrial diameter, pulmonary artery diameter, short axis shortening rate, left atrial enlargement, and aortic valve degeneration.

[0032] In this embodiment, ECOPD patients are grouped using the target adverse cardiovascular event as the outcome label. Patients with at least one of the six cardiovascular diseases—coronary artery disease, heart failure, arrhythmia, pulmonary heart disease, cerebral infarction, and hypertension—are labeled as the positive group, while patients without the above outcomes are labeled as the negative group. The multimodal binary feature vectors are mapped one-to-one with the outcome label according to the patient number, so that each row represents one patient and each column represents a binary candidate variable. A one-way significance test is performed column-by-column between the positive and negative groups. Since the input variables have been processed into binary variables, a chi-square test is used to determine whether there is a difference in the positive rate of a certain variable between the two groups. When the theoretical frequency of a certain variable is low, Fisher's exact test is used instead, with a preset significance level of 0.05. Candidate variables with a p-value less than 0.05 in the test results are included in the set of significantly different variables. Simultaneously, each... The proportion, direction of difference, and test statistics of variables in the positive and negative groups; for example, age, sex, left atrial enlargement, right atrial enlargement, aortic valve degeneration, pulmonary hypertension, pericardial effusion, lactate dehydrogenase, blood glucose, hemoglobin, red blood cell count, red blood cell distribution width, basophil count and percentage, eosinophil count and percentage, lymphocyte count and percentage, albumin, serum creatinine, etc. can be used as differential screening objects. At the same time, cardiac ultrasound parameters such as aortic root diameter, left atrial diameter, interventricular septum thickness, left ventricular diameter, left ventricular posterior wall thickness, pulmonary artery diameter, and aortic flow velocity propagation may show higher levels in the positive group, while the short axis shortening rate may show higher levels in the negative group. This shows that the univariate stage is not only used to narrow down the candidate range, but also to initially identify signals related to cardiovascular events such as cardiac structural remodeling, decreased cardiac function, metabolic abnormalities, renal dysfunction, and changes in inflammatory status.

[0033] When inputting a set of significantly different variables into a lasso regression model for coefficient compression, the target adverse cardiovascular event can be used as the dependent variable, and the significantly different variables as the independent variables. The L1 regularization penalty mechanism is used to compress the coefficients of some variables with weak contributions, redundant information, or high correlation with other variables to zero, thereby reducing the number of candidate variables while preserving predictive ability. The regularization parameter of lasso regression can be determined by cross-validation, and the parameter with the smallest cross-validation error or the error within an acceptable range is used as the final penalty strength. After this stage, several initial screening features with non-zero coefficients and stable contributions are selected to form an initial screening feature set.

[0034] The 15 initial screening features are input into a multivariate logistic regression model, or an independent predictive effect is confirmed using a stepwise approach. This ensures that each variable is tested while simultaneously adjusting for the influence of other candidate variables. If a variable remains statistically significant in the multivariate model and its regression direction is consistent with the univariate stage and clinical interpretation, it is retained as a core feature. If a variable is significant only in the univariate stage but loses its independent effect in the multivariate model due to collinearity or interpretation by other variables, it is removed. The final model consists of age, blood glucose, eosinophil count, and other predictive features. The study used a set of nine core features, including percentage, serum creatinine, left atrial diameter, pulmonary artery diameter, fractional shortening, left atrial enlargement, and aortic valve degeneration. Among these features, age reflects basic cardiovascular vulnerability, blood glucose reflects metabolic load, eosinophil percentage reflects differences in inflammatory phenotype, serum creatinine reflects renal function and circulatory load status, left atrial diameter and left atrial enlargement together characterize left ventricular structural remodeling, pulmonary artery diameter characterizes pulmonary circulation pressure-related changes, fractional shortening characterizes ventricular systolic function status, and aortic valve degeneration characterizes cardiac aging and valvular degeneration status.

[0035] In one embodiment, such as Figure 5 As shown, step S30 specifically includes the following steps: S31: Input the core feature set as an indicator variable into the latent class analysis model with different numbers of categories, calculate the Bayesian information criterion, the consistent Akaike information criterion, the Bayesian information criterion corrected by the sample size, and the entropy value corresponding to each latent class analysis model, and obtain the set of fitting indices for each latent class analysis model. S32: Based on the set of fitting indices, perform bootstrap likelihood ratio test and Rohman-Mendel-Rubin likelihood ratio test on latent class analysis models with adjacent number of classes to determine the latent class analysis model corresponding to the optimal number of classes. S33: Input the core feature set into the latent class analysis model corresponding to the optimal number of categories, and calculate the posterior probability vector of each patient belonging to each subtype; S34: Based on the posterior probability vector, the cardiovascular event risk stratification results corresponding to each subtype are obtained through comparison and verification.

[0036] In this embodiment, nine core features are uniformly used as indicator variables for latent class analysis. These include age, blood glucose, eosinophil percentage, serum creatinine, left atrial diameter, pulmonary artery diameter, fractional shortening, left atrial enlargement, and aortic valve degeneration, forming the patient's clustering input matrix. Since latent class analysis is suitable for handling categorical indicator variables, each patient corresponds to a binary state for these indicators, such as normal or abnormal, present or absent, high or low level. Before modeling, the system performs a completeness check on patient records, marking records with excessively high missing proportions or whose binary states cannot be confirmed as pending review. Simultaneously, it ensures that the encoding meaning of each indicator variable is consistent across different patients. Then, the number of candidate categories is preset to 1, 2, 3, and 4, and corresponding latent class analysis models are established for each. The model iteratively estimates the population proportion of each category and the conditional response probability of each indicator variable within each category using the expectation-maximization algorithm, enabling the model to automatically identify potential subtypes from the patient's multimodal binary response patterns. After fitting the model for each number of categories, the model simultaneously outputs the Bayesian information criterion, the consistent Akaike information criterion, the sample size-corrected Bayesian information criterion, and the entropy value, forming a set of fitting indices. The Bayesian information criterion, the consistent Akaike information criterion, and the sample size-corrected Bayesian information criterion are used to measure the balance between the model's fit and complexity. The lower the value, the better the model's interpretability without excessively increasing the number of categories. The entropy value is used to measure the classification clarity, and its value ranges from 0 to 1. The closer it is to 1, the higher the certainty that the patient has been assigned to the corresponding potential category.

[0037] By combining a set of fitting indices, a stepwise comparison is made between latent class analysis models with adjacent number of categories. That is, as the number of categories gradually increases, it is determined whether the addition of new categories truly improves the model's interpretability. Simultaneously, bootstrap likelihood ratio tests and Rohmandel-Rubin likelihood ratio tests are performed. If the comparison results of adjacent models show that the addition of new categories brings stable and statistically significant improvement in fit, and each category has sufficiently clear characteristic patterns and interpretable clinical meaning, then increasing the number of categories can be considered reasonable. Conversely, if the decrease in information criteria is limited, the entropy value decreases, the test results are unstable, or the addition of new categories only shows the splitting of a few abnormal samples without clear cardiopulmonary comorbidity characteristics, then the number of categories should not be increased further to avoid overfitting. Therefore, the latent class analysis model corresponding to the optimal number of categories should simultaneously meet the conditions of superior fitting indices, reliable test results of adjacent models, clear category boundaries, and subtype interpretation that conforms to clinical pathological logic.

[0038] After inputting the core features of each patient into the final latent category analysis model, the model calculates the probability of each patient belonging to each potential subtype based on the combination pattern of age, metabolic status, inflammatory characteristics, renal function status, and cardiac structural and functional abnormalities, and forms a posterior probability vector. The posterior probability vector preserves the possibility that the patient is close to different subtypes at the same time, thus reflecting the continuity and uncertainty of the cardiopulmonary comorbidity status of ECOPD patients.

[0039] The patient's initial subtype is determined based on the posterior probability vector. Then, the different subtypes are correlated with target adverse cardiovascular events, cardiovascular disease occurrence, and related outcomes. The bias caused by hard classification is reduced by classification error correction. Subsequently, the differences between the subtypes in terms of cardiac structural remodeling, pulmonary circulation load, cardiorenal dysfunction, valvular degenerative changes, and metabolic abnormalities are compared. This determines the risk level corresponding to different subtypes, so that the output results not only include which subtype the patient belongs to, but also further reflect the level of cardiovascular event risk corresponding to that subtype.

[0040] In one embodiment, such as Figure 6 As shown, step S33 specifically includes the following steps: S331: Compare the conditional response probability distributions of each subtype in the latent category analysis model corresponding to the optimal number of categories on each indicator variable in the core feature set to obtain the feature patterns of each subtype; S332: Based on the characteristic patterns, the subtype with the highest probability of conditioned response to abnormal levels of left atrial diameter, pulmonary artery diameter, and left atrial enlargement is identified as the cardiopulmonary structural remodeling with metabolic abnormalities subtype; the subtype with the lowest probability of conditioned response to advanced age is identified as the relatively young subtype with normal cardiac and renal function; and the subtype with the highest probability of conditioned response to aortic valve degeneration is identified as the cardiac aging subtype characterized by valvular degeneration. S333: Substitute the core feature set of each patient into the latent class analysis model corresponding to the optimal number of classes to calculate the posterior probability, and obtain the posterior probability vector of each patient belonging to each subtype.

[0041] In this embodiment, the output results of the latent class analysis model corresponding to the optimal number of categories are read, and the conditional response probabilities of each subtype on each indicator variable of the core feature set are organized into a probability distribution matrix of the same dimension. Rows correspond to different subtypes, and columns correspond to core indicator variables such as age, blood glucose, eosinophil percentage, serum creatinine, left atrial diameter, pulmonary artery diameter, fractional shortening, left atrial enlargement, and aortic valve degeneration. The values ​​in the matrix represent the probability that a patient of a certain subtype exhibits abnormality, high risk, or a pre-existing state on the corresponding variable. The conditional response probabilities of different subtypes are compared column by column, and the subtype characteristic patterns are identified by combining the pathological meaning of the variables. For example, the conditional response probabilities of a certain subtype on abnormal left atrial diameter, abnormal pulmonary artery diameter, and left atrial enlargement are compared. When the values ​​are all higher than other subtypes, and also show a relatively high trend in metabolic-related variables such as abnormal blood glucose, it can be said that this subtype mainly reflects a pattern of coexisting cardiac chamber enlargement, increased pulmonary circulation load, and metabolic abnormalities. When a subtype has a generally low probability of conditional response to indicators such as abnormal age, hyperglycemia, abnormal serum creatinine, left atrial enlargement, and abnormal pulmonary artery diameter, and the probability of reduced short-axis shortening is not prominent, it can be said that this subtype is closer to a pattern of relatively young age and relatively stable cardiac structure and renal function. When a subtype has the highest probability of conditional response to the presence of aortic valve degeneration, and a high probability of age-related abnormalities, while cardiac chamber enlargement or pulmonary artery changes are not the most prominent dominant features, it can be said that this subtype is closer to a pattern related to valvular degeneration and cardiac aging.

[0042] Based on the aforementioned characteristic patterns, semantic labeling was performed on each subtype. The subtype with the highest probability of conditional response to abnormal levels of left atrial diameter, pulmonary artery diameter, and left atrial enlargement was labeled as the cardiopulmonary structural remodeling with metabolic abnormalities subtype. The subtype with the lowest probability of conditional response in the elderly and relatively weak signals of cardiorenal dysfunction was labeled as the relatively young subtype with normal cardiorenal function. The subtype with the highest probability of conditional response to aortic valve degeneration was labeled as the cardiac aging subtype characterized by valvular degeneration. Furthermore, the labeling process was confirmed by combining the response probabilities of multiple core features to avoid subtype interpretation bias caused by fluctuations in a single test.

[0043] The core feature set of each patient is input into the identified and calibrated latent class analysis model. The model matches the patient's binary response status on various indicator variables with the conditional response probability distribution of each subtype, calculating the posterior probability of belonging to the cardiopulmonary structural remodeling with metabolic abnormalities subtype, the relatively young and normal cardiorenal function subtype, and the cardiac aging subtype characterized by valvular degeneration. The posterior probability vectors are then constructed according to a fixed subtype order. If a patient shows abnormalities in structural remodeling indicators such as left atrial diameter, pulmonary artery diameter, and left atrial enlargement, the posterior probability of belonging to the cardiopulmonary structural remodeling with metabolic abnormalities subtype is higher. If the patient has fewer abnormal indicators and no significant age-related risk, the posterior probability of belonging to the relatively young and normal cardiorenal function subtype is higher. If the patient has significant aortic valve degeneration, the posterior probability of belonging to the cardiac aging subtype is higher. This yields a patient-level subtype probability output that retains the classification results while reflecting classification uncertainty.

[0044] In one embodiment, step S34 specifically includes the following steps: The subtype corresponding to the maximum posterior probability in the posterior probability vector is determined as the modality classification result for each patient. The modality classification results are input into the logistic regression model to calculate the first association strength between each subtype and the target adverse cardiovascular events and cardiovascular diseases. Classification error correction weights are constructed based on posterior probability vectors, and the modality classification results are analyzed with the target adverse cardiovascular events and cardiovascular diseases based on the classification error correction weights to obtain the second association strength results. The results of the first and second association strengths were compared and verified to obtain the cardiovascular event risk stratification results for each subtype.

[0045] In this embodiment, the posterior probability vector of each patient is stored in a fixed subtype order, so that each component in the vector corresponds to the subtype of cardiopulmonary structural remodeling with metabolic abnormalities, the subtype of relatively young age with normal cardiorenal function, and the subtype of cardiac aging characterized by valvular degeneration. Then, the magnitude of each component in the vector is compared case by case, and the subtype corresponding to the maximum posterior probability is written into the patient's modality classification field, thereby obtaining a deterministic subtype label for conventional statistical modeling. If the maximum posterior probability of a patient is small compared with the probabilities of other subtypes, the low confidence label can be retained simultaneously without changing its modality classification result, so as to perform classification uncertainty correction.

[0046] Modal classification results were used as categorical independent variables and input into a logistic regression model. Target adverse cardiovascular events and specific cardiovascular diseases such as coronary heart disease, heart failure, arrhythmia, pulmonary heart disease, cerebral infarction, and hypertension were used as outcome variables. A relatively young subtype with normal cardiac and renal function was selected as the reference group. Odds ratios, confidence intervals, and significance results of other subtypes relative to the reference group were calculated to obtain the first association strength result. The first association strength result reflects the direction and strength of the association between different subtypes and cardiovascular outcomes when potential subtypes are directly regarded as determining grouping variables.

[0047] Classification error correction weights are constructed based on posterior probability vectors. The classification error of the latent class is corrected by the probability of each patient being assigned to each subtype, so that the uncertainty of subtype assignment can be included in the remote outcome analysis process. When performing remote outcome analysis, the target adverse cardiovascular event and each specific cardiovascular disease can be used as remote outcome variables, and the association strength between each subtype and the outcome is re-estimated according to the classification error correction weights to obtain the second association strength result.

[0048] The results of the first and second association strengths are compared. If the two results are consistent in the direction of association and the ranking of the strengths among high-risk, intermediate-risk, and low-risk subtypes is basically stable, it indicates that the risk stratification results do not rely on a single hard classification path and can resist the bias caused by the uncertainty of posterior probability. If a subtype is high-risk in logistic regression, but the association is significantly weakened or the direction changes after classification error correction, it is necessary to backtrack the posterior probability distribution and conditional response probability pattern of the subtype to determine whether there are unclear class boundaries or sample confounding issues. Finally, the subtypes that maintain stable high association after verification by both paths are identified as the high-risk layer, the subtypes with weaker association strength and stable results are identified as the low-risk layer, and the subtypes that are between the two or only show significant association for some cardiovascular diseases are identified as the intermediate-risk layer, thus outputting the cardiovascular event risk stratification results corresponding to each subtype.

[0049] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0050] In one embodiment, a cardiovascular event risk prediction system for COPD patients based on multimodal data is provided. This system corresponds one-to-one with the multimodal data-based cardiovascular event risk prediction method for COPD patients described in the above embodiments. Figure 7 As shown, this cardiovascular event risk prediction system for COPD patients based on multimodal data includes: The data acquisition module 701 is used to acquire demographic data, blood biomarker data and cardiac ultrasound index data of ECOPD patients, and construct a multimodal binary feature vector. The feature analysis module 702 is used to perform single-factor difference significance test, lasso regression compression screening, and multi-factor logistic regression independent effect confirmation on the multimodal binary feature vectors to obtain the core feature set; The risk stratification module 703 is used to determine the posterior probability vector of each patient to each subtype based on the core feature set, and to calculate the cardiovascular event risk stratification results corresponding to each subtype based on the posterior probability vector.

[0051] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0052] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0053] This application also provides a computer device, such as... Figure 8 As shown, the computer device includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, it implements the steps in any of the above method embodiments, or when the processor executes the computer program, it implements the functions of each module / unit in the above device embodiments.

[0054] For example, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete this application. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the computer device.

[0055] Those skilled in the art will understand that Figure 8 The computer device described is merely an example and does not constitute a limitation on the computer device. It may include more or fewer components than shown, or combine certain components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.

[0056] The aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0057] The memory can be an internal storage unit of the computer device, such as a hard drive or RAM. The memory can also be an external storage device of the computer device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory can include both internal and external storage units of the computer device.

[0058] This application also provides a readable storage medium storing a computer program, which, when executed by a processor, can implement the steps in the above-described method embodiments.

[0059] This application provides a computer program product that, when run on an electronic device, enables the electronic device to implement the steps described in the various method embodiments above.

[0060] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0061] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0062] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0063] In the embodiments provided in this application, it should be understood that the disclosed apparatus / devices and methods can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0064] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0065] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for predicting cardiovascular event risk in COPD patients based on multimodal data, characterized in that, include: Acquire demographic data, blood biomarker data, and echocardiographic data of ECOPD patients, and construct a multimodal binary feature vector; The multimodal binary feature vectors were subjected to single-factor difference significance test, lasso regression compression screening, and multi-factor logistic regression independent effect confirmation to obtain the core feature set; Based on the core feature set, the posterior probability vector of each patient belonging to each subtype is determined, and the cardiovascular event risk stratification result corresponding to each subtype is calculated based on the posterior probability vector.

2. The method for predicting cardiovascular event risk in COPD patients based on multimodal data as described in claim 1, characterized in that, The process of acquiring demographic data, blood biomarker data, and echocardiographic data of ECOPD patients, and constructing a multimodal binary feature vector, includes: Demographic data, blood biomarker data, and echocardiogram data of ECOPD patients were obtained from the electronic medical record system. The demographic data, blood biomarker data, and echocardiogram data are divided into a set of continuous variables and a set of qualitative variables. The set of continuous variables is subjected to threshold binary encoding to obtain encoded continuous variables, and a multimodal binary feature vector is constructed based on the encoded continuous variables and the set of qualitative variables.

3. The method for predicting cardiovascular event risk in COPD patients based on multimodal data as described in claim 2, characterized in that, The process of performing univariate difference significance tests, lasso regression compression screening, and multivariate logistic regression independent effect confirmation on the multimodal binary feature vectors yields a core feature set, including: The ECOPD patients were divided into a positive group and a negative group, and a one-way significance test was performed on the multimodal binary feature vectors between the positive group and the negative group to obtain a set of significantly different variables. The set of significantly different variables is input into the lasso regression model for coefficient compression to obtain the initial screening feature set. The initial screening feature set is input into a multivariate logistic regression model for independent prediction effect confirmation to obtain the core feature set, which includes age, blood glucose, eosinophil percentage, serum creatinine, left atrial diameter, pulmonary artery diameter, fractional shortening rate, left atrial enlargement, and aortic valve degeneration.

4. The method for predicting cardiovascular event risk in COPD patients based on multimodal data as described in claim 1, characterized in that, The step of determining the posterior probability vector of each patient's subtype based on the core feature set, and calculating the cardiovascular event risk stratification result corresponding to each subtype based on the posterior probability vector, includes: The core feature set is used as an indicator variable and input into latent class analysis models with different numbers of categories. The Bayesian information criterion, the consistent Akaike information criterion, the Bayesian information criterion corrected by the sample size, and the entropy value corresponding to each latent class analysis model are calculated to obtain the set of fitting indices for each latent class analysis model. Based on the set of fitting indices, perform bootstrap likelihood ratio test and Roe-Mendel-Rubin likelihood ratio test on the latent class analysis models with adjacent number of classes to determine the latent class analysis model corresponding to the optimal number of classes. The core feature set is input into the latent class analysis model corresponding to the optimal number of categories to calculate the posterior probability vector of each patient belonging to each subtype; Based on the comparison and verification of the posterior probability vector, the cardiovascular event risk stratification results corresponding to each subtype are obtained.

5. The method for predicting cardiovascular event risk in COPD patients based on multimodal data as described in claim 4, characterized in that, The step of inputting the core feature set into the latent class analysis model corresponding to the optimal number of categories, and calculating the posterior probability vector of each patient belonging to each subtype, includes: The conditional response probability distributions of each subtype in the latent category analysis model corresponding to the optimal number of categories on each indicator variable of the core feature set are compared to obtain the feature patterns of each subtype. Based on the aforementioned characteristic patterns, the subtype with the highest probability of responding to abnormal levels of left atrial diameter, pulmonary artery diameter, and left atrial enlargement is identified as the cardiopulmonary structural remodeling with metabolic abnormalities subtype; the subtype with the lowest probability of responding to advanced age is identified as the relatively young subtype with normal cardiac and renal function; and the subtype with the highest probability of responding to aortic valve degeneration is identified as the cardiac aging subtype characterized by valvular degeneration. The core feature set of each patient is substituted into the latent class analysis model corresponding to the optimal number of categories to calculate the posterior probability, thereby obtaining the posterior probability vector of each patient belonging to each subtype.

6. The method for predicting cardiovascular event risk in COPD patients based on multimodal data as described in claim 5, characterized in that, The comparison and verification based on the posterior probability vector yields the cardiovascular event risk stratification results for each subtype, including: The subtype corresponding to the maximum posterior probability in the posterior probability vector is determined as the modality classification result for each patient. The modality classification results are input into a logistic regression model to calculate the first association strength between each subtype and the target adverse cardiovascular events and cardiovascular diseases. Based on the posterior probability vector, classification error correction weights are constructed, and based on the classification error correction weights, remote outcome analysis is performed on the modality classification results, the target adverse cardiovascular events, and the cardiovascular diseases to obtain the second association strength result. The first correlation strength result and the second correlation strength result are compared and verified to obtain the cardiovascular event risk stratification results corresponding to each subtype.

7. A cardiovascular event risk prediction system for COPD patients based on multimodal data, characterized in that, The steps for implementing the method for predicting cardiovascular event risk in COPD patients based on multimodal data as described in any one of claims 1 to 6 include: The data acquisition module is used to acquire demographic data, blood biomarker data, and echocardiographic data of ECOPD patients, and to construct a multimodal binary feature vector. The feature analysis module is used to perform single-factor difference significance test, lasso regression compression screening, and multi-factor logistic regression independent effect confirmation on the multimodal binary feature vector to obtain the core feature set; The risk stratification module is used to determine the posterior probability vector of each patient belonging to each subtype based on the core feature set, and to calculate the cardiovascular event risk stratification result corresponding to each subtype based on the posterior probability vector.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method for predicting cardiovascular event risk in COPD patients based on multimodal data as described in any one of claims 1 to 6.

9. A readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for predicting cardiovascular event risk in COPD patients based on multimodal data as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the method for predicting cardiovascular event risk in COPD patients based on multimodal data as described in any one of claims 1 to 6.