Heart assessment rapid diagnosis method based on multi-modal data fusion and related device
Through the combination of multimodal data fusion of Unet3+, Transformer and BERT models, the noise processing, real-time and personalized adaptation problems of echocardiography segmentation are solved, precise segmentation of heart structure and automatic calculation of ejaculation fractions are realized, and personalized diagnostic support is provided.
Patent Information
- Application Number
- CN202510531480.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
AI Technical Summary
The existing echocardiography segmentation method has blurred and inaccurate segmentation boundaries when processing complex noise images, which is difficult to meet the needs of real-time clinical diagnosis, ignores individual differences in patients, and lacks ability to capture cardiac dynamic changes, so it cannot provide personalized diagnostic support.
Unet3+, Transformer and BERT models combined with multimodal data fusion, echocardiography features are extracted through wavelet convolution, and patient attribute information is fused to achieve accurate segmentation of heart structure and automatic calculation of ejaculation fraction.
It improves the accuracy and real-time nature of echocardiography segmentation, enhances the adaptability to individual differences, and can quickly and accurately provide personalized diagnostic support, reflecting the functional changes of the heart under different physiological states.
Smart Images

Figure CN120452738A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image segmentation, and in particular to a rapid cardiac assessment and diagnosis method based on multimodal data fusion and related devices. Background Art
[0002] Cardiovascular disease (CVD) has become a major threat to global health. Its morbidity and mortality continue to rise with changing lifestyles and an aging population, placing a heavy burden on countless patients and their families. Early diagnosis and prevention are crucial to reducing CVD-related morbidity and mortality, relying on continuous advances in medical imaging to improve diagnostic and treatment outcomes. Advances in cardiovascular research and medical imaging have significantly improved diagnostic accuracy, enabling noninvasive assessment of cardiac structure and function through modalities such as magnetic resonance imaging (MRI), computed tomography (CT), and ultrasound. As an advanced ultrasound technique, echocardiography enables comprehensive assessment of cardiac anatomy and function, which is crucial for early cardiovascular detection and diagnosis. Its real-time imaging, lack of radiation, cost-effectiveness, and widespread clinical application further enhance its diagnostic value. However, the diagnostic value of echocardiography depends heavily on accurate segmentation of cardiac structures.
[0003] Deep learning (DL) provides an end-to-end learning paradigm that can automatically extract complex features for image segmentation. In 2015, the emergence of the U-Net network revolutionized medical image segmentation and achieved remarkable results. For example, Zyuzin et al. developed an automatic segmentation method based on U-Net for depicting the left ventricular endocardial border in two-dimensional echocardiography, and improved the model performance through technical optimization. Wan et al. proposed a semi-supervised segmentation algorithm for four-chamber echocardiography videos based on multi-level edge perception and calibration fusion modules, which improved the segmentation accuracy. For echocardiography sequence segmentation, Li et al. proposed an M4S-Net for echocardiography sequence segmentation, which used multi-level shape priors and optical flow optimization modules to improve temporal stability during dynamic cardiac motion, achieving real-time segmentation of echocardiography.
[0004] Despite significant progress in medical image segmentation using deep learning (DL), most existing methods still have significant limitations in adapting to inter-patient anatomical variability. Key personal attributes such as age and gender are often overlooked, resulting in segmentation models that lack personalization and fail to meet the diverse needs of clinical populations. Multimodal fusion has emerged as a promising strategy in medical imaging, integrating complementary information from patient clinical records. Zillner et al. leveraged image metadata and ontology reasoning to enhance automated lymphoma staging. Wu et al. introduced the LV-SAM model based on SAM-Med2D, which combines a multiscale adapter, a multimodal cue encoder, and a multiscale decoder to improve the accuracy of LV segmentation and LVEF estimation. Similarly, Gu et al. demonstrated that incorporating demographic features significantly improved tumor segmentation compared to attribute-independent models. Similarly, demographic features such as age, gender, and BMI are crucial in the task of segmenting cardiac anatomy. Cardiac anatomy varies with age and gender, impacting echocardiographic segmentation. Therefore, incorporating these demographic features is crucial for achieving more accurate echocardiographic segmentation. At the same time, there is currently no real-time personalized integrated algorithm that combines patient attribute information to solve the problem of rapid diagnostic algorithm for cardiac assessment by echocardiography to solve the problem of real-time difficulties in personalized patient diagnosis and treatment.
[0005] The fields of echocardiographic image segmentation and ejection fraction calculation currently face the following major challenges: ① Traditional segmentation methods struggle to accurately identify cardiac structures, especially when processing images containing complex noise. Segmentation boundaries are blurred and inaccurate, severely impacting the reliability of subsequent ejection fraction calculations. For example, in the presence of speckle noise and low contrast, traditional thresholding and edge detection algorithms often miss critical cardiac contour information, resulting in segmentation results that fail to accurately reflect the heart's actual morphology. ② Existing algorithms struggle to meet the operational efficiency requirements of real-time clinical diagnosis. Complex computational processes result in lengthy image analysis times, hindering physicians' ability to provide immediate diagnostic support. Particularly when processing large datasets or performing high-resolution image analysis, some deep learning-based models require long inference times due to their large number of parameters, making them impractical in emergency situations. ③ Most methods ignore the impact of individual patient differences on cardiac structure, resulting in poor model adaptability across different patient populations and limited ability to provide personalized diagnostic support. For example, heart size and morphology vary significantly between patients of different age groups. Existing models often use uniform segmentation parameters and lack the flexibility to adapt to these variations, limiting their applicability to diverse patient populations. ④ Existing models are inadequate in capturing dynamic changes in the heart, making it difficult to accurately reflect changes in cardiac function under different physiological conditions. Current models focus primarily on analyzing static images and lack effective means of capturing the complex movements and morphological changes of the heart during the cardiac cycle. This makes it difficult for doctors to fully understand the heart's functional state, which in turn affects the accurate diagnosis of certain cardiomyopathies or valvular diseases. Summary of the Invention
[0006] In order to solve the above technical problems, the present invention proposes a rapid diagnostic method for cardiac assessment based on multimodal data fusion and related devices, which overcome the shortcomings of traditional methods in segmentation accuracy, efficiency, personalized adaptability and dynamic capture capabilities, and innovatively combine advanced network architecture and feature extraction technology to achieve efficient, accurate and personalized echocardiographic analysis.
[0007] In order to achieve the above object, the technical solution of the present invention is as follows:
[0008] A rapid diagnostic method for cardiac assessment based on multimodal data fusion, comprising the following steps:
[0009] Acquiring and preprocessing medical data, wherein the medical data includes echocardiographic image sequence data and patient attribute information medical record data;
[0010] Extracting image features of the echocardiographic image sequence data; extracting encoding features of patient attribute information medical record data using a pre-trained BERT model; inputting the image features into a Unet3+ model encoder, and utilizing multi-scale feature extraction and fusion capabilities to obtain multi-scale features; performing bidirectional conversion between image and text based on the PICA module of the Transformer encoder using the multi-scale features and encoding features to obtain deeply fused features;
[0011] The fused features are input into the Unet3+ model decoder to output the cardiac structure segmentation results, and the ejection fraction is automatically calculated based on the cardiac structure segmentation results.
[0012] Preferably, the attribute information includes age, gender, BMI and medical history.
[0013] Preferably, the image features of the echocardiographic image sequence data are extracted using the following formula:
[0014] F wf =αF wavelet +(1-α)F conv
[0015] Among them, F wavelet is the wavelet convolution feature, F conv is the common convolution feature, and α is the fusion weight coefficient.
[0016] Preferably, the multi-scale features and the coding features are bidirectionally converted between the image and the text in the PICA module of the Transformer encoder to obtain the deeply fused fusion features, which specifically includes the following steps:
[0017] Use the attention mechanism to calculate the correlation weights between modal features;
[0018]
[0019] Among them, S ij is the similarity between the i-th modal feature and the j-th modal feature, α ij is the corresponding attention weight, S ik represents the similarity between the i-th modal feature and the k-th modal feature, and N is the number of modalities;
[0020] The multimodal features are weighted and fused according to the correlation weights to obtain the fusion feature F multi :
[0021]
[0022]
[0023] Among them, α i is the comprehensive attention weight of the i-th modality, is the eigenvector of the i-th mode.
[0024] Preferably, the similarity S ij It is obtained by calculating the dot product similarity or cosine similarity between the features of two modalities.
[0025] Preferably, the ejection fraction is automatically calculated based on the cardiac structure segmentation result, and the specific steps are as follows:
[0026] Based on the cardiac structure segmentation results, the volume of the heart at the end of systole and end of diastole is extracted as follows:
[0027]
[0028] Among them, V ES and V ED are the cardiac volumes at end-systole and end-diastole, respectively. and are the heart areas of the i-th slice at the end of systole and end of diastole, d is the slice spacing, N ES and N ED are the number of slices at end-systole and end-diastole, respectively;
[0029] To calculate the ejection fraction value, the formula is as follows:
[0030]
[0031] Where EF is the ejection fraction.
[0032] Preferably, the method further comprises the following steps:
[0033] The cardiac structure segmentation results were evaluated using the Dyss similarity and Hausdorff distance as evaluation indicators;
[0034] Evaluate the accuracy and AUC of extracting the heart at end-diastole and end-systole based on the cardiac structure segmentation results;
[0035] The accuracy of ejection fraction calculation was evaluated using the mean absolute error.
[0036] A rapid diagnostic device for cardiac assessment based on multimodal data fusion, comprising: an acquisition module, an extraction and fusion module, and a segmentation and calculation module, wherein:
[0037] An acquisition module, configured to acquire and pre-process medical data, wherein the medical data includes echocardiographic image sequence data and patient attribute information and medical record data;
[0038] An extraction and fusion module is configured to extract image features from the echocardiographic image sequence data; extract encoding features from the patient attribute information medical record data using a pre-trained BERT model; input the image features into a Unet3+ model encoder, and utilize multi-scale feature extraction and fusion capabilities to obtain multi-scale features; and perform bidirectional conversion between image and text using the Transformer encoder's PICA module to obtain deeply fused features.
[0039] The segmentation and calculation module is used to input the fusion features into the Unet3+ model decoder, output the cardiac structure segmentation results, and automatically calculate the ejection fraction based on the cardiac structure segmentation results.
[0040] Based on the above content, the present invention further discloses a computer device, comprising: a memory for storing a computer program; and a processor for implementing any of the above methods when executing the computer program.
[0041] Based on the above content, the present invention further discloses a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, any of the above methods is implemented.
[0042] Based on the above technical solution, the present invention has the following beneficial effects: Using Unet3+, Transformer, and Bert models as its core technologies, the present invention extracts echocardiographic features through wavelet convolution, effectively overcoming the limitations of traditional methods in processing echocardiographic noise and meeting clinical real-time requirements. Simultaneously, the introduction of patient attribute information significantly enhances the segmentation capability of echocardiographic sequences, reduces noise interference, and enables automatic recognition of end-systolic and end-diastolic images of two-chamber and four-chamber hearts in the sequences. Based on this, the ejection fraction is automatically calculated, providing clear, accurate, and personalized images and cardiac ejection fraction data for clinical diagnosis. The multimodal information fusion submodule, the PICA module, integrates echocardiographic images with multi-source patient data, such as gender, age, BMI, and medical history. By constructing an adaptation network, image features are personalized, enabling the model to better adapt to individual patient differences and improving the accuracy of segmentation and ejection fraction calculation. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a flowchart of a rapid diagnostic method for cardiac assessment based on multimodal data fusion in an embodiment;
[0044] Figure 2 is a schematic diagram of the extraction and fusion process in a rapid diagnosis method for cardiac assessment based on multimodal data fusion in an embodiment;
[0045] Figure 3 is a schematic diagram of a fusion process in a rapid diagnosis method for cardiac assessment based on multimodal data fusion in an embodiment;
[0046] Figure 4 The present invention is a schematic diagram of the process of segmenting and calculating ejection fraction in a rapid diagnostic method for cardiac assessment based on multimodal data fusion in an embodiment. DETAILED DESCRIPTION
[0047] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention.
[0048] See also Figures 1 to 4 This embodiment provides a rapid cardiac assessment and diagnosis method based on multimodal data fusion, comprising the following steps:
[0049] Step 1: Acquire and pre-process medical data, wherein the medical data includes echocardiographic image sequence data and patient attribute information medical record data.
[0050] In this embodiment, echocardiographic image sequence data is acquired and filtered to remove noise to improve image quality. At the same time, patient attribute information (such as age, gender, BMI, etc.) and medical history data are collected and embedded in the data for fusion with image features.
[0051] In the data preprocessing stage, the echocardiographic image sequence data is converted into a format suitable for model processing, and normalization and other operations are performed to ensure the consistency and validity of the data; the echocardiographic image sequence data is further denoised using filters to enhance the quality and clarity of the image, laying the foundation for subsequent feature extraction and segmentation tasks.
[0052] Step 2: Extract image features of the echocardiographic image sequence data; Use the pre-trained BERT model to extract the encoding features of the patient attribute information medical record data; Input the image features into the Unet3+ model encoder, and use the multi-scale feature extraction and fusion capabilities to obtain multi-scale features; Perform bidirectional conversion between image and text based on the multi-scale features and encoding features in the PICA module of the Transformer encoder to obtain deeply fused features.
[0053] In this embodiment, the computer extracts image features from echocardiographic images and fuses the encoded features of patient attribute information and medical record data with the image features. This multimodal, multi-level feature fusion enriches the model's understanding of cardiac status and improves the accuracy of segmentation and ejection fraction calculation.
[0054] Wavelet convolution is used to decompose echocardiographic images at different scales, highlighting the key features of cardiac structure and suppressing noise interference; ordinary convolution is used to extract local features of echocardiographic images, reducing the dependence on multi-layer ordinary convolution. Specifically, see Figure 4 , the dotted box 1 is a multi-layer denoising filter, which contains 3 traditional denoising filters, and performs preliminary denoising on the input echocardiographic image. The dotted box 2 is a multi-layer convolution layer, which is composed of multiple 3x3 convolution layers, which is used to extract the basic features of the image and further denoise, providing a clear feature map for subsequent processing. The wavelet convolution layer in the dotted box 3 performs decomposition at different scales on the denoised feature map, highlighting the key features of the cardiac structure and suppressing noise. This is followed by the Dropout layer in the dotted box 4, which prevents the model from overfitting, randomly discards some neuron outputs, and enhances generalization and robustness. Finally, the maximum pooling layer in the dotted box 5 downsamples the feature map output by the wavelet convolution layer, reducing the size and computational complexity, while retaining and highlighting important features, and enhancing the model's ability to abstractly represent features. The entire process achieves efficient extraction and preliminary processing of echocardiographic image features. The specific steps are:
[0055] The echocardiographic image is input into the wavelet convolution and ordinary convolution network for multi-scale decomposition and local feature extraction to obtain feature maps and local features at different scales;
[0056] I dec =W dec (I)
[0057] Where I is the input echocardiographic image, W dec is the wavelet decomposition operation, I dec is the feature map after decomposition;
[0058] The decomposed feature map is thresholded to remove the noise part and enhance the cardiac structure features to obtain the wavelet convolution feature F. wavelet ;
[0059] The processed feature map is fused with the local features to obtain a feature representation rich in cardiac structural information and low in noise;
[0060] F wf =αF wavelet +(1-α)F conv
[0061] Among them, F wavelet is the wavelet convolution feature, F conv is the common convolution feature, α is the fusion weight coefficient, F wf is the fused image feature.
[0062] Specifically, we integrate echocardiographic images with other multimodal information (patient attribute information and medical records), such as gender, age, BMI, and medical history, to enrich the model's understanding of cardiac status. The implementation steps are as follows:
[0063] The echocardiographic images and patient attribute information medical record data are preprocessed and converted into a unified feature dimension. The formula is as follows:
[0064] F modality =φ i (M i )
[0065] Among them, M i is the data of mode i, φ i is the corresponding preprocessing function, F modality is the preprocessed feature. The purpose of this step is to eliminate the heterogeneity between different modal data so that they can be effectively fused in the same feature space. The echocardiographic images and patient attribute information medical record data are preprocessed and converted into a unified feature dimension. Specifically, for echocardiographic images, the preprocessing function φ i Specifically, the image is first filtered and denoised, including Gaussian filtering, median filtering, and wavelet filtering, to remove noise interference and enhance cardiac structural features. The image pixel values are then normalized to the range of [0,1] to ensure that the pixel values between different images are comparable. Finally, data enhancement is performed by randomly rotating the image to improve the generalization ability of the model. Finally, the image features are reshaped to ensure that the image features and the patient attribute information medical record data features have the same feature dimension, so that they can be effectively fused in the same feature space, providing high-quality input for subsequent feature extraction and fusion modules. For patient attribute information medical record data, the preprocessing function φ i Specifically, the BERT encoding method was adopted to encode other multimodal information (patient attribute information and medical records data), such as gender, age, BMI, and medical records, into 128-dimensional feature vectors to ensure that the converted features are comparable and fusible.
[0066] The attention mechanism is used to calculate the correlation weight between the two modal features of echocardiographic images and patient attribute information medical records data. The formula is as follows:
[0067]
[0068] Among them, S ij is the similarity between the i-th modal feature and the j-th modal feature, S ikRepresents the similarity between the i-th modal feature and the k-th modal feature. Here k is an index variable used to traverse all modal features (from 1 to N) in order to calculate the similarity between the i-th modal feature and all other modal features, α ij is the corresponding attention weight. Similarity S ij The attention weight is obtained by calculating the dot product similarity metric between the features of two modalities. The attention weight reflects the relative importance and correlation between the features of different modalities, enabling the model to automatically focus on the modal information most relevant to the current task.
[0069] According to the weight, the echocardiogram features and the patient attribute information medical record data features are weighted fused to obtain the fusion feature F multi , the formula is as follows:
[0070]
[0071]
[0072] in, is the eigenvector of the i-th mode, α i is the comprehensive attention weight of the i-th modality and the comprehensive attention weight α i It is obtained by normalizing the attention weights between all modal features, ensuring that the weight distribution in the fusion process is reasonable and the sum is 1, β i is the scaling factor for the i-th modal feature, which is used to adjust the influence of this modal feature in the fusion process. This weighted fusion approach not only highlights the modal features that are more important to the current task, but also effectively suppresses irrelevant or redundant modal information, thereby improving the model's understanding and analysis of cardiac status.
[0073] In step 3, the fused features are input into the Unet3+ model decoder, the cardiac structure segmentation results are output, and the ejection fraction is automatically calculated based on the cardiac structure segmentation results.
[0074] In this embodiment, specifically, with Unet3+, Transformer and Bert models as the core, a pre-trained BERT model is used to integrate patient attribute information medical record data, and natural language descriptions are encoded into semantically rich embedding vectors. These text embeddings are jointly processed by the Transformer encoder with the multi-scale features from the Unet3+ model encoder. The PICA module achieves a deep fusion of patient attributes and echocardiographic image features. This integration enables the model to take into account age-related cardiac changes and gender-specific anatomical differences, thereby significantly improving segmentation accuracy and clinical adaptability, and avoiding the use of a large number of fully connected layers to integrate information from different modalities, thereby reducing the number of fully connected layers. The fused feature Fmulti The Unet3+ model decoder is fed into the model, leveraging its multi-scale feature fusion capabilities to obtain segmentation features. Finally, the designed segmentation head converts the resulting features into specific pixel-level segmentation masks, enabling automatic recognition of end-systolic and end-diastolic images of two- and four-chamber hearts in echocardiography.
[0075] In this embodiment, the ejection fraction calculation module calculates the ejection fraction based on the output cardiac structure segmentation result. The ejection fraction calculation formula is used to calculate directly based on the output cardiac structure segmentation result, without the need for complex feature conversion or data mapping through multiple layers of fully connected layers, further reducing the use of fully connected layers. The volume information of the heart at different cardiac cycle stages is extracted from the segmentation results, and the specific positions of the end-systole and end-diastole of the two-chamber and four-chamber hearts are determined by analyzing the heart volume change curve within the cardiac cycle. At the determined end-systole and end-diastole, the corresponding heart volume data are respectively extracted, and the ejection fraction calculation formula is used to calculate the ejection fraction value based on the extracted end-systole and end-diastole volume data. The specific steps are as follows:
[0076] The volume information of the heart at different cardiac cycle stages is extracted from the segmentation results.
[0077]
[0078] Among them, V ES and V ED are the cardiac volumes at end-systole and end-diastole, respectively. and are the heart areas of the nth slice at the end of systole and end of diastole, d is the slice spacing, N ES and N ED are the number of slices at end-systole and end-diastole, respectively.
[0079] By analyzing the cardiac volume change curve during the cardiac cycle, the specific positions of the end-systole and end-diastole of the two-chamber and four-chamber hearts can be determined.
[0080] At the determined end-systole and end-diastole, the corresponding heart volume data are extracted respectively.
[0081] The ejection fraction value was calculated based on the extracted end-systolic and end-diastolic volume data using the ejection fraction calculation formula.
[0082]
[0083] Where EF is the ejection fraction.
[0084] For model training, a large dataset of annotated echocardiographic image sequences was used, along with a corresponding loss function, to learn and update model parameters using an optimization algorithm to improve segmentation and ejection fraction calculation performance. The backbone network was a U-Net3+, consisting of two upsampling and downsampling layers, an input layer, an output layer, and an intermediate layer. The input size was 256×256×32. Training was performed using the Adam optimizer with a learning rate of 2e-5 and 1000 training iterations. The overall network framework was implemented using TensorFlow and trained on three NVIDIA RTX 3090 GPUs (24GB) with a batch size of 5.
[0085] The prediction and segmentation effect was evaluated using Dice similarity (DICE) and Hausdorff distance (HD) of echocardiographic images; the accuracy and AUC of the predicted end-diastole and end-systole of the two-chamber and four-chamber heart were evaluated; the accuracy of ejection fraction calculation was evaluated using mean absolute error (MAE).
[0086] Table 1 Test experiment evaluation index results
[0087]
[0088] As shown in Table 1, the experimental results demonstrate that this method outperforms traditional methods in terms of segmentation accuracy, real-time performance, personalized adaptability, and dynamic capture capabilities. In terms of segmentation accuracy, this method fuses multimodal and multi-level features, combined with wavelet convolution to extract echocardiographic features, resulting in clearer and more accurate cardiac structural boundaries. Furthermore, in terms of real-time performance, the optimization of the algorithm process and model architecture enables this method to quickly respond to clinical needs and provide timely diagnostic support to physicians. Furthermore, the introduction of patient attribute information and medical record data significantly enhances the adaptability of this method to different patients, enabling it to automatically adjust segmentation parameters based on individual patient differences, thereby improving diagnostic accuracy. Finally, this method also improves its ability to capture dynamic changes in the heart, more accurately reflecting changes in cardiac function under different physiological states and providing more comprehensive information for clinical diagnosis.
[0089] It should be understood that, although the various steps in the above flow chart are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the above flow chart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0090] In one embodiment, a rapid cardiac assessment and diagnosis device based on multimodal data fusion is provided, comprising: an acquisition module, an extraction and fusion module, and a segmentation and calculation module, wherein:
[0091] An acquisition module, configured to acquire and pre-process medical data, wherein the medical data includes echocardiographic image sequence data and patient attribute information and medical record data;
[0092] An extraction and fusion module is configured to extract image features from the echocardiographic image sequence data; extract encoding features from the patient attribute information medical record data using a pre-trained BERT model; input the image features into a Unet3+ model encoder, and utilize multi-scale feature extraction and fusion capabilities to obtain multi-scale features; and perform bidirectional conversion between image and text using the Transformer encoder's PICA module to obtain deeply fused features.
[0093] The segmentation and calculation module is used to input the fusion features into the Unet3+ model decoder, output the cardiac structure segmentation results, and automatically calculate the ejection fraction based on the cardiac structure segmentation results.
[0094] The apparatus described in the above embodiments may be implemented by a computer chip or entity, or by a product having a certain function. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0095] In one embodiment, a computer device is provided, comprising: a memory for storing a computer program; and a processor for implementing the steps of the above-mentioned rapid diagnosis method for cardiac assessment based on multimodal data fusion when executing the computer program.
[0096] In one embodiment, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned rapid diagnosis method for cardiac assessment based on multimodal data fusion are implemented.
[0097] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0098] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0099] Each embodiment in this specification is described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For relevant parts, refer to the description of the method embodiment.
[0100] The above is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A rapid diagnostic method for cardiac assessment based on multimodal data fusion, characterized in that: The steps include: Acquiring and preprocessing medical data, wherein the medical data includes echocardiographic image sequence data and patient attribute information medical record data; Extracting image features of the echocardiographic image sequence data; extracting encoding features of patient attribute information medical record data using a pre-trained BERT model; inputting the image features into a Unet3+ model encoder, and utilizing multi-scale feature extraction and fusion capabilities to obtain multi-scale features; performing bidirectional conversion between image and text based on the PICA module of the Transformer encoder using the multi-scale features and encoding features to obtain deeply fused features; The fused features are input into the Unet3+ model decoder to output the cardiac structure segmentation results, and the ejection fraction is automatically calculated based on the cardiac structure segmentation results.
2. The rapid diagnostic method for cardiac assessment based on multimodal data fusion according to claim 1, characterized in that: The attribute information includes age, gender, BMI and medical history.
3. The rapid cardiac assessment and diagnosis method based on multimodal data fusion according to claim 1, characterized in that: The image features of the echocardiographic image sequence data are extracted using the following formula: F wf =αF wavelet +(1-α)F conv Among them, F wavelet is the wavelet convolution feature, F conv is the common convolution feature, and α is the fusion weight coefficient.
4. The rapid cardiac assessment and diagnosis method based on multimodal data fusion according to claim 1, characterized in that: The multi-scale features and the coding features are converted bidirectionally between the image and the text based on the PICA module of the Transformer encoder to obtain the deeply fused fusion features, which specifically includes the following steps: Use the attention mechanism to calculate the correlation weights between modal features; Among them, S ij is the similarity between the i-th modal feature and the j-th modal feature, α ij is the corresponding attention weight, S ik represents the similarity between the i-th modal feature and the k-th modal feature, and N is the number of modalities; The multimodal features are weighted and fused according to the correlation weights to obtain the fusion feature F multi : Among them, α i is the comprehensive attention weight of the i-th modality, is the eigenvector of the i-th mode.
5. The rapid cardiac assessment and diagnosis method based on multimodal data fusion according to claim 1, characterized in that: The similarity S ij It is obtained by calculating the dot product similarity or cosine similarity between the features of two modalities.
6. The rapid cardiac assessment and diagnosis method based on multimodal data fusion according to claim 1, characterized in that: The ejection fraction is automatically calculated based on the cardiac structure segmentation results. The specific steps are as follows: Based on the cardiac structure segmentation results, the volume of the heart at the end of systole and end of diastole is extracted as follows: Among them, V ES and V ED are the cardiac volumes at end-systole and end-diastole, respectively. and are the heart areas of the i-th slice at the end of systole and end of diastole, d is the slice spacing, N ES and N ED are the number of slices at end-systole and end-diastole, respectively; To calculate the ejection fraction value, the formula is as follows: Where EF is the ejection fraction.
7. The rapid cardiac assessment and diagnosis method based on multimodal data fusion according to claim 6, characterized in that: The following steps are also included: The cardiac structure segmentation results were evaluated using the Dyss similarity and Hausdorff distance as evaluation indicators; The accuracy and AUC of extracting the heart at end-diastole and end-systole based on the cardiac structure segmentation results were evaluated; The accuracy of ejection fraction calculation was evaluated using the mean absolute error.
8. A rapid diagnostic device for cardiac assessment based on multimodal data fusion, characterized in that: include: Acquisition module, extraction and fusion module and segmentation and calculation module, among which, An acquisition module, configured to acquire and pre-process medical data, wherein the medical data includes echocardiographic image sequence data and patient attribute information and medical record data; An extraction and fusion module is configured to extract image features from the echocardiographic image sequence data; extract encoding features from the patient attribute information medical record data using a pre-trained BERT model; input the image features into a Unet3+ model encoder, and utilize multi-scale feature extraction and fusion capabilities to obtain multi-scale features; and perform bidirectional conversion between image and text using the Transformer encoder's PICA module to obtain deeply fused features. The segmentation and calculation module is used to input the fusion features into the Unet3+ model decoder, output the cardiac structure segmentation results, and automatically calculate the ejection fraction based on the cardiac structure segmentation results.
9. A computer device, characterized in that: The device comprises: a memory for storing a computer program; A processor, configured to implement the method according to any one of claims 1 to 7 when executing the computer program.
10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.