Multi-mode cardiovascular disease detection method based on comprehensive view analysis
By using a multimodal detection method with comprehensive view analysis in the diagnosis of cardiovascular disease, multimodal feature fusion is solved by using retinal fundus images and clinical indicators, the problem of difficulty in comprehensively predicting CVD progress in the prior art is solved, and higher detection accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510069917.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to comprehensively predict disease progression in the diagnosis of cardiovascular disease (CVD), and the diagnostic methods are limited to single-modal data analysis, and cannot effectively utilize multiple modal diagnostic information.
A multimodal cardiovascular disease detection method is proposed for comprehensive view analysis. By segmenting fundus blood vessels from retinal fundus images, combining retinal fundus images and non-invasive clinical indicators, a detection model including a comprehensive fundus image extraction module, a clinical index feature extraction module and a multi-order belief interactive feature fusion module are constructed, and multi-modal feature fusion and cardiovascular disease classification are carried out.
Through the comprehensive analysis of multimodal data, this method enhances a deep understanding of CVD events, improves the accuracy and reliability of cardiovascular disease detection, and can predict disease progress more comprehensively.
Smart Images

Figure CN120015284A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer-aided diagnosis, and in particular to a multimodal cardiovascular disease detection method based on comprehensive view analysis. Background Art
[0002] Cardiovascular disease (CVD) is one of the leading causes of illness and death worldwide. According to the World Health Organization (WHO), CVD causes approximately 19.8 million deaths each year, and the number is increasing year by year. The increasing prevalence of CVD is a major public health issue and has put considerable pressure on healthcare systems and financial resources.
[0003] Cardiovascular diseases mainly include coronary heart disease, heart failure, stroke and other pathological conditions. According to research, the occurrence of CVD is usually closely related to factors such as hypertension, high cholesterol, BMI, etc., and presents the characteristics of long-term concealment. Early diagnosis and treatment have been shown to significantly reduce the health risks brought by the disease and improve the quality of life of patients. Therefore, it is particularly important to carry out CVD screening. However, due to the uneven distribution of medical resources, especially in rural and remote areas, CVD patients may have difficulty in obtaining screening and diagnostic services in a timely manner. In recent years, researchers have found that the diagnosis of CVD is closely related to retinal fundus images and specific clinical indicators. With the rapid development of deep learning technology, CVD automatic detection methods based on deep learning have gradually attracted attention. Such methods have the advantages of high efficiency and fast speed, which to a certain extent alleviate the diagnostic difficulties caused by the imbalance of medical resources. However, it should be noted that the current diagnostic methods are still limited to the analysis of single-modal data, lack the analysis of multi-modal diagnostic information, and cannot comprehensively predict the progression of CVD. Therefore, this paper proposes a multi-modal cardiovascular disease detection method based on comprehensive view analysis. This method first segments the fundus blood vessels from the retinal fundus image to obtain a blood vessel segmentation map, and then uses the retinal fundus image and related non-invasive clinical indicators as the input of the multimodal CVD detection model of comprehensive view analysis. The model includes a fundus image comprehensive feature extraction module, a clinical indicator feature extraction module and a multi-order belief interactive feature fusion module. The fundus image comprehensive feature extraction module is used to comprehensively analyze the retinal fundus image and the blood vessel segmentation map to obtain image features, and the clinical indicator feature extraction module is used to extract features of clinical indicators. The multi-order belief interactive feature fusion module is used to fuse the two extracted features, and the fused features are input into the classification network to obtain the CVD detection results. Finally, the multimodal CVD detection model of comprehensive view analysis is trained to convergence and applied to CVD classification and lesion segmentation. This method has been widely used in the field of CVD diagnosis. Summary of the invention
[0004] The purpose of the present invention is to solve the problem that CVD progression analysis is difficult, and it is difficult to diagnose and prevent it as early as possible, and to propose a multimodal cardiovascular disease detection method with comprehensive view analysis.
[0005] The above invention objectives are mainly achieved through the following technical solutions:
[0006] S1. Screen the retinal fundus images that meet the quality standards and include them in the data set. Collect the patients' non-invasive clinical index information and form a corresponding relationship with the retinal fundus images to obtain the CVD detection data set.
[0007] S2. Take the retinal fundus image as input and the corresponding vascular map as label to train the fundus vascular segmentation model. Use the fundus vascular segmentation model to extract the vascular area of the fundus image of the CVD detection dataset to obtain the vascular segmentation map. The steps are as follows:
[0008] (1) Based on the retinal fundus image segmentation dataset, a fundus vascular segmentation model is constructed, and the model is trained using the Dice coefficient loss as the loss function. The Dice coefficient loss function formula is as follows:
[0009]
[0010] Where L Dice is the mean square error, y i With f(x i ) represent the true value and predicted value respectively, and N represents the number of pixels;
[0011] (2) After the fundus vascular segmentation model converges, the retinal fundus image in the CVD prediction dataset is input to obtain the corresponding vascular segmentation map.
[0012] S3, respectively taking the retinal fundus image and the corresponding vascular segmentation map and clinical index as input, and taking the corresponding CVD diagnosis result as label to train the single modality CVD detection model, and obtaining the fundus image comprehensive feature extraction module including the fundus image feature extraction module and the vascular segmentation map, and the clinical index feature extraction module, wherein the steps are as follows:
[0013] (1) The fundus image comprehensive feature extraction module takes the retinal fundus image and the vascular segmentation map as input, constructs a fundus image feature extraction module and a vascular segmentation map feature extraction module, obtains the overall features of the fundus image and the vascular morphology features, and fuses the two features to obtain the comprehensive features of the fundus image;
[0014] (2) The clinical indicator feature extraction module takes the clinical indicator as input to obtain the clinical indicator features;
[0015] (3) Construct a single-modal CVD detection model and train the model using cross entropy loss as the loss function;
[0016] (4) After the single-modality CVD prediction model converges, the model parameters of the fundus image comprehensive feature extraction module and the clinical indicator feature extraction module of the model are obtained respectively.
[0017] S4. Construct a multimodal cardiovascular disease detection model for comprehensive view analysis, which includes a fundus image comprehensive feature extraction module, a clinical indicator feature extraction module, a multi-order belief interaction feature fusion module, and a cardiovascular disease classification task head. The multimodal features obtained from different feature extraction modules are fused using the multi-order belief interaction feature fusion module and input into the classification task head to obtain the cardiovascular disease detection results. The steps are as follows:
[0018] (1) extracting comprehensive fundus image features and clinical index features using a fundus image comprehensive feature extraction module and a clinical index feature extraction module;
[0019] (2) Construct a multi-order belief interaction feature fusion module, take the features obtained from different feature extraction modules as input, and use the basic belief allocation method and Dempster-Shafer theory to perform two-stage fusion of the features extracted from the two modalities to obtain multimodal fusion features to achieve complementarity between multimodal information. The steps are as follows:
[0020] ① Two basic belief assignment (BBA) methods are used to assign an evidence or confidence level e to the comprehensive features of fundus images and clinical index features respectively:
[0021]
[0022] Where z is the comprehensive characteristics of fundus images and clinical index characteristics, represents the total amount of evidence, and N represents the number of categories in the classification task;
[0023] ② Calculate the Dirichlet distribution parameter α corresponding to the two modalities of fundus image and clinical index respectively. The calculation formula is as follows:
[0024] α=e n +1 (4)
[0025] where e n The evidence or confidence level for the comprehensive features of fundus images and clinical index features;
[0026] ③ Calculate the belief b and uncertainty u of each modal feature obtained by the two basic belief allocation methods respectively to provide conditions for subsequent fusion:
[0027]
[0028] in Corresponding to the Dirichlet distribution, α is the Dirichlet distribution parameter;
[0029] ④ According to the belief b and uncertainty u obtained by the two BBA methods, first-order fusion is performed in each mode:
[0030]
[0031] in To combine beliefs, is the combined uncertainty, ° is the Hadamard product, and γ is the product based on the equation The normalization coefficient of
[0032] ⑤ According to the combined belief and combined uncertainty of different modes obtained by the first-order fusion, the second-order fusion is performed between different modes to obtain the final multimodal fusion features:
[0033]
[0034] in and are the uncertainty and belief obtained after the first-order fusion of non-invasive image modalities, and are the uncertainty and belief obtained after the first-order fusion of clinical index modalities, and f fusion is the obtained multimodal fusion feature.
[0035] S5. The multimodal cardiovascular disease detection model based on comprehensive view analysis is trained to convergence and used for cardiovascular disease classification.
[0036] Effects of the Invention
[0037] The present invention provides a multimodal cardiovascular disease detection method for comprehensive view analysis. The method first segments fundus blood vessels from a retinal fundus image to obtain a blood vessel segmentation map, and then uses the retinal fundus image and related non-invasive clinical indicators as inputs of a multimodal CVD detection model for comprehensive view analysis. The model includes a fundus image comprehensive feature extraction module, a clinical indicator feature extraction module, and a multi-order belief interactive feature fusion module. The fundus image comprehensive feature extraction module is used to perform a comprehensive analysis of the retinal fundus image and the blood vessel segmentation map to obtain image features, the clinical indicator feature extraction module is used to extract features of the clinical indicators, and the multi-order belief interactive feature fusion module is used to fuse the two extracted features, and the fused features are input into a classification network to obtain a CVD detection result. Finally, the multimodal CVD detection model for comprehensive view analysis is trained to convergence and applied to CVD classification and lesion segmentation. Experiments show that the advantages of this method are as follows: (1) The vascular segmentation map is used to assist the cross-view learning of the retinal fundus image, and the model's feature perception of the retinal blood vessels is enhanced. (2) The two multimodal data of retinal fundus images and non-invasive clinical indicators are used as input to enhance the model's deep understanding of CVD events. (3) The multimodal features are integrated and complemented by the multi-order belief interaction feature fusion method to improve the accuracy and reliability of CVD detection. The present invention is applied to CVD detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 A flow chart of a multimodal cardiovascular disease detection method for comprehensive view analysis in an example of the present invention;
[0039] Figure 2 It is a flowchart of extracting fundus blood vessel segmentation map in an example of the present invention;
[0040] Figure 3 It is a training flow chart of the fundus image comprehensive feature extraction model and the clinical indicator feature extraction model in the example of the present invention;
[0041] Figure 4 A structural diagram of a multimodal cardiovascular disease detection model for comprehensive view analysis in an example of the present invention;
[0042] Figure 5 This is a structural diagram of the multi-order belief interaction feature fusion module in an example of the present invention. Specific implementation methods Specific implementation method one:
[0044] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] like Figure 1 As shown in FIG. 1 , the multimodal cardiovascular disease detection method based on comprehensive view analysis includes the following steps:
[0046] S1. Low-quality retinal fundus images with blur, low contrast, low resolution and artifacts were screened out from the UK Biobank database, and retinal fundus images that met the quality standards were included in the data set. Five non-invasive clinical index information of patients, including age, gender, systolic blood pressure, body mass index and smoking status, were collected and correspondingly related to the retinal fundus images to obtain the CVD detection data set.
[0047] S2. Use the fundus vascular segmentation dataset Drive to train the fundus vascular segmentation model FR-UNet, and use the trained FR-UNet to extract fundus vascular from the retinal fundus images in the CVD detection dataset to obtain the corresponding vascular segmentation map;
[0048] S3, respectively using the retinal fundus image and the corresponding vascular segmentation map and clinical indicators in the CVD detection data set as input, and using the corresponding CVD diagnosis results as labels to train the single-modality CVD detection model, and obtaining a fundus image comprehensive feature extraction module including a fundus image feature extraction module and a vascular segmentation map, and a clinical indicator feature extraction module;
[0049] S4. Construct a multimodal cardiovascular disease detection model for comprehensive view analysis, which includes a comprehensive image feature extraction module, a clinical indicator feature extraction module, a multi-order belief interaction feature fusion module, and a cardiovascular disease classification task head. They use DenseNet-201 as the backbone network and use the multi-order belief interaction feature fusion module to fuse the multimodal features obtained from different feature extraction modules and input them into the classification task head to obtain the cardiovascular disease detection results.
[0050] S5. The multimodal cardiovascular disease detection model of comprehensive view analysis is trained to convergence using the CVD detection dataset and used for cardiovascular disease classification.
[0051] The embodiments of the present invention are described in detail below:
[0052] The embodiment of the present invention is specifically implemented as follows.
[0053] S1. Low-quality retinal fundus images with blur, low contrast, low resolution and artifacts were screened out from the UK Biobank database, and retinal fundus images that met the quality standards were included in the dataset. Five non-invasive clinical indicator information including age, gender, systolic blood pressure, body mass index and smoking status of patients were collected and correspondingly related to the retinal fundus images to obtain the CVD detection dataset.
[0054] S2, such as Figure 2 As shown in the figure, the fundus vascular segmentation dataset Drive is used to train the fundus vascular segmentation model FR-UNet, and the trained FR-UNet is used to extract fundus vascular from the retinal fundus images in the CVD detection dataset to obtain the corresponding vascular segmentation map. The steps are as follows:
[0055] (1) The fundus images in the retinal fundus image vascular segmentation dataset Drive are used as input, and the vascular segmentation map is used as the label to train the model;
[0056] (2) Construct the fundus vascular segmentation model FR-UNet and train the model using the Dice coefficient loss as the loss function. The Dice coefficient loss function formula is as follows:
[0057]
[0058] Where L Dice is the mean square error, y i With f(x i ) represent the true value and predicted value respectively, and N represents the number of pixels;
[0059] (3) After the fundus vascular segmentation model converges, the retinal fundus image in the CVD prediction dataset is input to obtain the corresponding vascular segmentation map.
[0060] S3, such as Figure 3 As shown in the figure, the retinal fundus images and the corresponding vascular segmentation maps and clinical indicators in the CVD detection dataset are used as inputs, and the corresponding CVD diagnosis results are used as labels to train the single-modality CVD detection model, and the fundus image comprehensive feature extraction module including the fundus image feature extraction module and the vascular segmentation map and the clinical indicator feature extraction module are obtained. The steps are as follows:
[0061] (1) The retinal fundus images and corresponding vascular segmentation maps and clinical indicators in the CVD detection dataset are used as input, and the CVD detection results are used as labels to construct a dataset, and the dataset is divided into a training set, a validation set, and a test set;
[0062] (2) The fundus image comprehensive feature extraction module takes the retinal fundus image and the vascular segmentation map as input, respectively, and uses DenseNet-201 as the backbone to construct the fundus image feature extraction module and the vascular segmentation map feature extraction module to obtain the overall features of the fundus image and the vascular morphology features, and then fuses the two features to obtain the comprehensive features of the fundus image;
[0063] (3) The clinical indicator feature extraction module takes the clinical indicator as input, constructs a clinical indicator feature extraction module CNN, and obtains the clinical indicator features;
[0064] (4) Use the training set to train the model with cross entropy loss as the loss function. The formula of the cross entropy loss function is as follows:
[0065] L BCE =-y i log(f(x i ))+(1-y i )log(1-f(x i )) (2)
[0066] Where L BCE is the cross entropy loss, y i With f(x i ) represent the true value and predicted value respectively;
[0067] (5) After the model converges, the fundus image comprehensive feature extraction module and clinical indicator feature extraction module of the model are obtained respectively.
[0068] S4, such as Figure 4 As shown in the figure, a multimodal cardiovascular disease detection model based on comprehensive view analysis is constructed. The model includes a comprehensive image feature extraction module, a clinical indicator feature extraction module, a multi-order belief interaction feature fusion module, and a cardiovascular disease classification task head. They use DenseNet-201 and CNN as the backbone network, and use the multi-order belief interaction feature fusion module to fuse the multimodal features obtained from different feature extraction modules and input them into the classification task head to obtain the cardiovascular disease detection results. The steps are as follows:
[0069] (1) Using the fundus image comprehensive feature extraction module DenseNet201 and the clinical index feature extraction module CNN to extract the fundus image comprehensive features and clinical index features;
[0070] (2) Construct a multi-order belief interaction feature fusion module. Take the features obtained from different feature extraction modules as input, and use the basic belief assignment (BBA) and Dempster-Shafer theory to perform two-stage fusion of the features extracted from the two modalities to obtain multimodal fusion features to achieve complementarity between multimodal information. The steps are as follows:
[0071] ① Two basic belief assignment (BBA) methods are used to assign an evidence or confidence level e to the comprehensive features of fundus images and clinical index features respectively:
[0072]
[0073] Where z is the comprehensive characteristics of fundus images and clinical index characteristics, represents the total amount of evidence, and N represents the number of categories in the classification task;
[0074] ② Calculate the Dirichlet distribution parameter α corresponding to the two modalities of fundus image and clinical index respectively. The calculation formula is as follows:
[0075] α=e n +1 (5)
[0076] where e n The evidence or confidence level for the comprehensive features of fundus images and clinical index features;
[0077] ③ Calculate the belief b and uncertainty u of each modal feature obtained by the two basic belief allocation methods respectively to provide conditions for subsequent fusion:
[0078]
[0079] in Corresponding to the Dirichlet distribution, α is the Dirichlet distribution parameter;
[0080] ④ According to the belief b and uncertainty u obtained by the two BBA methods, first-order fusion is performed in each mode:
[0081]
[0082] in To combine beliefs, is the combined uncertainty, is the Hadamard product, and γ is based on the equation The normalization coefficient of
[0083] ⑤ According to the combined belief and combined uncertainty of different modes obtained by the first-order fusion, the second-order fusion is performed between different modes to obtain the final multimodal fusion features:
[0084]
[0085] in and are the uncertainty and belief obtained after the first-order fusion of non-invasive image modalities, and are the uncertainty and belief obtained after the first-order fusion of clinical index modalities, and f fusion is the obtained multimodal fusion feature.
[0086] S5. The multimodal cardiovascular disease detection model of comprehensive view analysis is trained to convergence using the CVD detection dataset and used for cardiovascular disease classification.
Claims
1. A multimodal cardiovascular disease detection method based on comprehensive view analysis, characterized in that: The steps include: S1. Screen retinal fundus images that meet the quality standards and include them in the data set. Collect the non-invasive clinical index information of the patients and form a corresponding relationship with the retinal fundus images to obtain the CVD detection data set. S2, using the retinal fundus image as input and the corresponding vascular map as a label to train the fundus vascular segmentation model, and using the fundus vascular segmentation model to extract the vascular area of the fundus image of the CVD detection dataset to obtain a vascular segmentation map; S3, respectively taking the retinal fundus image and the corresponding vascular segmentation map and clinical index as input, and taking the corresponding CVD diagnosis result as label to train the single modality CVD detection model, and after convergence, obtaining the feature backbone of the model, i.e., the fundus image comprehensive feature extraction module including the fundus image feature extraction module and the vascular segmentation map, and the clinical index feature extraction module; S4. Construct a multimodal cardiovascular disease detection model for comprehensive view analysis, which includes a comprehensive image feature extraction module, a clinical indicator feature extraction module, a multi-order belief interaction feature fusion module, and a cardiovascular disease classification task head. The multimodal features obtained from different feature extraction modules are fused using the multi-order belief interaction feature fusion module and input into the classification task head to obtain the cardiovascular disease detection results. S5. The multimodal cardiovascular disease detection model based on comprehensive view analysis is trained to convergence and used for cardiovascular disease classification.
2. The multimodal cardiovascular disease detection method of comprehensive view analysis according to claim 1, characterized in that: In step S2, the retinal fundus image is used as input, and the corresponding vascular map is used as a label to train the fundus vascular segmentation model, obtain the fundus vascular segmentation model and extract the fundus vascular of the CVD detection data set, and obtain the vascular segmentation map, which includes the following steps: S21. Based on the retinal fundus image segmentation dataset, a fundus vascular segmentation model is constructed, and the model is trained using the Dice coefficient loss as the loss function. The Dice coefficient loss function formula is as follows: Where L Dice is the mean square error, y i With f(x i ) represent the true value and predicted value respectively, and N represents the number of pixels; S22. After the fundus vascular segmentation model converges, the retinal fundus image in the CVD prediction data set is input to obtain the corresponding vascular segmentation map.
3. The multimodal cardiovascular disease detection method of comprehensive view analysis according to claim 1, characterized in that: In step S3, the retinal fundus image and the corresponding vascular segmentation map and clinical indicators are respectively used as input, and the corresponding CVD diagnosis results are used as labels to train the single-modality CVD detection model, and obtain the fundus image comprehensive feature extraction module including the fundus image feature extraction module and the vascular segmentation map, and the clinical indicator feature extraction module, which includes the following steps: S31, the fundus image comprehensive feature extraction module takes the retinal fundus image and the blood vessel segmentation map as input, constructs a fundus image feature extraction module and a blood vessel segmentation map feature extraction module, obtains the overall feature of the fundus image and the blood vessel morphology feature, and fuses the two features to obtain the comprehensive feature of the fundus image; S32, a clinical indicator feature extraction module takes the clinical indicator as input to obtain clinical indicator features; S34, constructing a single-modal CVD detection model, and training the model using cross entropy loss as a loss function; S35. After the single-modality CVD prediction model converges, the model parameters of the fundus image comprehensive feature extraction module and the clinical indicator feature extraction module of the model are respectively obtained.
4. The multimodal cardiovascular disease detection method of comprehensive view analysis according to claim 1, characterized in that: In step S4, a multimodal cardiovascular disease detection model for comprehensive view analysis is constructed, which includes a fundus image comprehensive feature extraction module, a clinical indicator feature extraction module, a multi-order belief interaction feature fusion module, and a cardiovascular disease classification task head. The multimodal features obtained by different feature extraction modules are fused by using the multi-order belief interaction feature fusion module and input into the classification task head to obtain the cardiovascular disease detection result, which includes the following steps: S41, extracting comprehensive fundus image features and clinical indicator features using a fundus image comprehensive feature extraction module and a clinical indicator feature extraction module; S42. Construct a multi-order belief interaction feature fusion module, take the features obtained from different feature extraction modules as input, and use the basic belief allocation method and Dempster-Shafer theory to perform two-stage fusion of the features extracted from the two modalities to obtain multimodal fusion features to achieve complementarity between multimodal information.
5. The multimodal cardiovascular disease detection model for building comprehensive view analysis according to claim 4, characterized in that: In step S42, a multi-order belief interaction feature fusion module is constructed, and the features obtained by different feature extraction modules are used as input. The features of the two modalities extracted are fused in two stages through basic belief assignment (BBA) and Dempster-Shafer theory to obtain multimodal fusion features to achieve complementarity between multimodal information. The steps include: S421. Use two basic belief assignment (BBA) methods to assign an evidence or confidence level e to the comprehensive features of fundus images and clinical indicator features respectively: Where z is the comprehensive characteristics of fundus images and clinical index characteristics, represents the total amount of evidence, and N represents the number of categories in the classification task; S422, respectively calculate the Dirichlet distribution parameter α corresponding to the two modalities of fundus image and clinical index, and the calculation formula is as follows: α=e n +1 (4) where e n The evidence or confidence level for the comprehensive features of fundus images and clinical index features; S423, respectively calculate the belief b and uncertainty u of each modal feature obtained by the two basic belief allocation methods to provide conditions for subsequent fusion: in Corresponding to the Dirichlet distribution, α is the Dirichlet distribution parameter; S424. The belief b and uncertainty u obtained by the two BBA methods are first-order fused in each mode: in To combine beliefs, is the combined uncertainty, is the Hadamard product, and γ is based on the equation The normalization coefficient of S425. Based on the combined beliefs and combined uncertainties of different modalities obtained by the first-order fusion, a second-order fusion is performed between different modalities to obtain the final multimodal fusion features: in and are the uncertainty and belief obtained after the first-order fusion of non-invasive image modalities, and are the uncertainty and belief obtained after the first-order fusion of clinical index modalities, and f fusion is the obtained multimodal fusion feature.
Citation Information
Cited By
Ophthalmic disease image recognition method based on double-flow OfficientNet and decision tree
CN120236317A
Retinal vessel marker concept perception and multi-mode fusion coronary artery lesion detection method
CN121074010A