Transform and double-branch feature decoupling-based multi-modal medical image fusion method, system and equipment and medium
Through the method based on Transformer and dual-branch feature decoupling, the problem of insufficient global information modeling in multimodal medical image fusion is solved, high-quality image fusion is achieved, and the accuracy and comprehensiveness of medical image analysis is improved.
Patent Information
- Application Number
- CN202510703131.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-05-29
AI Technical Summary
The existing multimodal medical image fusion method is difficult to effectively capture long-distance dependencies, resulting in insufficient global information modeling capabilities and insufficient information complementary relationships between modals, affecting the quality of the fusion image.
Using a method based on Transformer and dual-branch feature decoupling, the global and local features are extracted and fused respectively through edge enhancement, dual-modal cross attention and feature decoupling modules to generate high-quality fusion images.
It significantly improves the image fusion effect, can better capture and combine the details and structural information between different modes, and provides more accurate and comprehensive support for medical image analysis.
Smart Images

Figure CN120543394A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image processing, and in particular relates to a multimodal medical image fusion method, system, device and medium based on Transformer and dual-branch feature decoupling. Background Art
[0002] The development of medical imaging technology has provided doctors with a variety of means to observe the internal structure and function of the human body. From the earliest X-ray imaging to modern computed tomography (CT), magnetic resonance imaging (MRI), positron emission tomography (PET), and single-photon emission computed tomography (SPECT), medical images acquired by different imaging devices present different modalities. However, single-modality images contain limited information and often cannot fully describe the imaging target, thus failing to meet the extensive information requirements for clinical diagnosis. Therefore, to observe all features of a site in a single image, it is necessary to extract effective information from multimodal medical images and fuse the complementary information from multiple original medical images. The fused image can provide a more comprehensive and reliable description of the lesion. By fusing source images of the same target acquired by different devices, a fused image with richer information and a more comprehensive description of the target can be generated. This fused image can obtain information about human organs and tissues that cannot be observed from single-modality images alone. This information can be used to detect, classify, and accurately localize abnormal lesions in human tissues, and has important scientific significance and practical application value in medical research, clinical diagnosis, surgical navigation, and prognosis evaluation.
[0003] Traditional multimodal medical image fusion methods typically rely on convolutional neural networks (CNNs) or transform domain methods for feature extraction and fusion. However, due to the inherent limitations of their local receptive field, CNNs struggle to effectively capture long-range dependencies in medical images, which results in insufficient global information modeling capabilities for fusion models. In recent years, significant progress has been made in the application of the Transformer architecture to computer vision tasks. Its self-attention mechanism can effectively model global feature relationships and achieve deeper modality complementarity during feature fusion. Existing multimodal image fusion methods often directly mix image features from different modalities, resulting in information redundancy and optimization difficulties. In the multimodal image fusion process, the flow of information between modalities is crucial. However, existing methods typically rely on simple channel splicing or attention weighting mechanisms, failing to fully model the information complementarity between modalities. This results in some modal information being overemphasized or ignored, thus affecting the quality of the fused image. Summary of the Invention
[0004] The purpose of the present invention is to provide a multimodal medical image fusion method, system, device and medium based on Transformer and dual-branch feature decoupling to solve the problems existing in the above-mentioned prior art.
[0005] To achieve the above objectives, the present invention provides a multimodal medical image fusion method based on Transformer and dual-branch feature decoupling, comprising:
[0006] Acquiring a multimodal medical image, and performing normalization processing on the multimodal medical image to obtain a preprocessed multimodal medical image;
[0007] The preprocessed multimodal medical image is input into a medical image fusion model for image fusion to obtain a fused image; wherein the medical image fusion model includes a shallow feature extraction module, a bimodal cross attention module, a dual-branch feature decoupling module and a splicing fusion module connected in sequence.
[0008] Optionally, the normalizing the multimodal medical image specifically includes:
[0009] The multimodal medical image is format-converted and the format-converted multimodal medical image is normalized to obtain a preprocessed multimodal medical image.
[0010] Optionally, the training process of the medical image fusion model specifically includes:
[0011] Acquiring training data, wherein the training data includes multimodal medical training images and corresponding fusion results;
[0012] An initial medical image fusion model is constructed, the training data is input into the initial medical image fusion model for image fusion, and training is performed with the goal of minimizing the loss between the initial training result after image fusion and the fusion result corresponding to the multimodal medical training image to obtain a trained medical image fusion model.
[0013] Optionally, the processing of the medical image fusion model specifically includes:
[0014] The preprocessed multimodal medical image is input into the shallow feature extraction module to perform edge feature enhancement and extract shallow features of the image through a general knowledge encoder to obtain a multi-scale initial feature representation;
[0015] Inputting the multi-scale initial feature representation into the bimodal cross-attention module, enhancing cross-modal feature interaction through the gradient extraction block and the cross-attention block to obtain gradient enhanced features and cross-attention features, and fusing the obtained gradient enhanced features and cross-attention features by element-by-element addition to obtain weighted optimized fused features;
[0016] The weighted optimized fusion features are input into the dual-branch feature decoupling module to extract global features and local features;
[0017] The global features and the local features are input into a splicing and fusion module for feature splicing and reconstruction to generate a fused image.
[0018] Optionally, the processing of the shallow feature extraction module specifically includes:
[0019] The initial feature of the input multimodal medical image X is X0. Under the dense connection mechanism, the feature X is generated. l :
[0020] X l =H d ([X0,X1,...,X l-1 ])
[0021] Where H d (·) represents the combination of 3×3 convolution and Leaky ReLU activation;
[0022] The gradient information is extracted by applying the anisotropic Sobel operator through the gradient enhancement unit to enhance the structural contour and edge details of the image; let the Sobel filter K x , K y Gradient calculations acting on the X and Y directions respectively:
[0023] G x =W(X)*K x ,G y =W(X)*K y
[0024] Where W(X) represents the discrete wavelet transform of the input image, G X is the gradient calculation result in the X direction, G Y is the gradient calculation result in the Y direction;
[0025] The residual fusion strategy is used to combine the original features with the enhanced gradient information to obtain the output features Then, through the general knowledge encoder, generate the multi-scale initial feature representation
[0026] Optionally, the processing of the bimodal cross attention module specifically includes:
[0027] Constructing a cross-modal complementary network, the cross-modal complementary network comprising a cross-attention block and a gradient extraction block, the cross-attention block consisting of a first cross-attention block and a second cross-attention block interacting with each other;
[0028] The initial feature representation is subjected to 3×3 convolution and ReLU activation function, and the input features of modality A and modality B are used as the components of the first gradient extraction block, the first cross attention block, the second cross attention block, and element-by-element addition respectively. The final result of feature addition is the weighted optimized fusion feature.
[0029] Optionally, the processing process of the dual-branch feature decoupling module specifically includes:
[0030] The weighted optimized fusion features are input into the dual-branch feature decoupling module, and the low-frequency common features of the input data are extracted through the global representation encoder to obtain the global features; the high-frequency unique features of the input data are extracted through the local texture encoder to obtain the local features.
[0031] Optionally, the processing of the splicing and fusion module specifically includes:
[0032] The global features and local features of different modalities are spliced separately in the channel dimension to obtain the fused global features and the fused local features;
[0033] Combine the fused global features and the fused local features to obtain fused features;
[0034] The fused features are input into the decoder for feature reconstruction to generate a fused image.
[0035] A multimodal medical image fusion system based on Transformer and dual-branch feature decoupling, including:
[0036] a data acquisition module, configured to acquire a multimodal medical image and perform normalization processing on the multimodal medical image to obtain a preprocessed multimodal medical image;
[0037] An image fusion module is used to input the preprocessed multimodal medical images into a medical image fusion model for image fusion to obtain a fused image; wherein the medical image fusion model includes a shallow feature extraction module, a bimodal cross attention module, a dual-branch feature decoupling module and a splicing fusion module connected in sequence.
[0038] An electronic device includes a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform a multimodal medical image fusion method based on Transformer and dual-branch feature decoupling.
[0039] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements a multimodal medical image fusion method based on Transformer and dual-branch feature decoupling.
[0040] The technical effects of the present invention are:
[0041] In order to solve the problems existing in the prior art, the present invention proposes a multimodal medical image fusion method based on Transformer and dual-branch feature decoupling. First, the edge enhancement strategy is used to process the multimodal features of the initial input to enhance the detail information of the image. Then, the processed features are input into a general knowledge encoder to extract shallow features. Then, these shallow features are sent to a dual-modal cross-attention network containing a gradient block and an attention block to further extract and fuse the correlation information between the modalities. On this basis, the present invention designs a dual-branch feature decoupling architecture to decouple the global and local features of the dual modalities respectively. Finally, the decoupled global and local features are respectively sent to the corresponding fusion layers, and the final fused image is generated through feature reconstruction. This method significantly improves the image fusion effect by comprehensively fusing the features of dual-modal medical images, and can better capture and combine the details and structural information between different modalities, thereby providing more accurate and comprehensive support for medical image analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0043] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0044] Figure 1 This is a flow chart of a multimodal medical image fusion method based on Transformer and dual-branch feature decoupling in an embodiment of the present invention;
[0045] Figure 2 Schematic diagram of a multimodal medical image fusion method based on Transformer and dual-branch feature decoupling in an embodiment of the present invention;
[0046] Figure 3 Detailed schematic diagram of the edge enhancement module in the overall schematic diagram of the multimodal medical image fusion method based on Transformer and dual-branch feature decoupling in an embodiment of the present invention;
[0047] Figure 4Detailed schematic diagram of a bimodal cross attention module in the overall schematic diagram of a multimodal medical image fusion method based on Transformer and dual-branch feature decoupling in an embodiment of the present invention;
[0048] Figure 5 Comparison of visualization results between this embodiment and other algorithms on the Havard multimodal medical image fusion dataset MRI-PET in an embodiment of the present invention;
[0049] Figure 6 Comparison of visualization results between this embodiment and other algorithms on the Havard multimodal medical image fusion dataset MRI-CT in an embodiment of the present invention;
[0050] Figure 7 This is a comparison of visualization results between this embodiment and other algorithms in the Havard multimodal medical image fusion dataset MRI-SPECT in an embodiment of the present invention. DETAILED DESCRIPTION
[0051] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as limiting the present invention, but rather as a more detailed description of certain aspects, features, and embodiments of the present invention.
[0052] It should be understood that the terms described herein are intended only to describe particular embodiments and are not intended to limit the present invention. In addition, for numerical ranges herein, it should be understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. Each smaller range between any intermediate value within a stated value or stated range and any other stated value or intermediate value within the stated range is also encompassed by the present invention. The upper and lower limits of these smaller ranges may be independently included or excluded within the scope.
[0053] It will be apparent to those skilled in the art that various modifications and variations may be made to the specific embodiments of the present invention without departing from the scope or spirit of the invention. Other embodiments will be apparent to those skilled in the art from the present invention. The present description and examples are intended to be illustrative only.
[0054] The words “include,” “including,” “have,” “contain,” etc. used in this article are open-ended terms, meaning including but not limited to.
[0055] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0056] like Figure 1 - Figure 7As shown, this embodiment provides a multimodal medical image fusion method based on Transformer and dual-branch feature decoupling, including: acquiring a multimodal medical image, normalizing the multimodal medical image to obtain a preprocessed multimodal medical image; inputting the preprocessed multimodal medical image into a medical image fusion model for image fusion to obtain a fused image; wherein the medical image fusion model includes a shallow feature extraction module, a dual-modal cross-attention module, a dual-branch feature decoupling module and a splicing fusion module connected in sequence.
[0057] To solve the problems existing in existing technologies, a fusion method that combines feature decoupling, CNN-Transformer interaction, and bimodal cross-attention modules is needed to ensure that the model can fully mine and fuse the complementary information of multimodal medical images, thereby improving the global consistency and local detail expression capabilities of the fused image and achieving higher quality medical image fusion.
[0058] To achieve the above objectives, this embodiment proposes a multimodal medical image fusion method based on Transformer and dual-branch feature decoupling. First, the edge enhancement strategy is used to process the multimodal features of the initial input to enhance the detail information of the image. Then, the processed features are input into the general knowledge encoder to extract shallow features. Then, these shallow features are sent to the dual-modal cross-attention network containing gradient blocks and attention blocks to further extract and fuse the correlation information between modalities. On this basis, this embodiment designs a dual-branch feature decoupling architecture to decouple the global and local features of the dual modalities respectively. Finally, the decoupled global and local features are respectively sent to the corresponding fusion layers, and the final fused image is generated through feature reconstruction. This method significantly improves the image fusion effect by fully fusing the features of dual-modal medical images, and can better capture and combine the details and structural information between different modalities, thereby providing more accurate and comprehensive support for medical image analysis.
[0059] This embodiment proposes a high- and low-frequency feature decoupling strategy, which realizes independent modeling of global and local features through a dual-branch architecture, improves the feature extraction capability of the fusion model, and improves the flexibility and scalability of the model. This embodiment combines CNN and Transformer to design an efficient multimodal medical image fusion architecture to give full play to the local feature extraction capability of CNN and the global modeling capability of Transformer, effectively improving the clarity and information integrity of the fused image. This embodiment proposes a dual-modality cross-attention module (Dual-Modality Cross-Attention, DMCA) to enhance the complementary information interaction between modalities, improve the quality of fusion features, maintain the integrity of key information, and thus improve the generalization ability of the fusion model.
[0060] The main steps of this embodiment are as follows:
[0061] Step 1: Preprocess the input multimodal medical images, convert the medical image data in JPG format into h5 format, and normalize the data to improve the stability and effectiveness of subsequent feature extraction.
[0062] Step 2: Edge features are enhanced on the normalized multimodal images, which are then fed into the Transformer-based Universal Knowledge Encoder (UKE) to extract shallow features. The encoder uses convolutional mapping layers to transform the input features and generate a multi-scale initial feature representation.
[0063] Construct a feature enhancement network including edge enhancement module (EEM). The edge enhancement module consists of multi-scale feature extraction unit and gradient enhancement unit to fully preserve the edge features of the image and optimize the local contrast. Specifically, the multi-scale feature extraction unit consists of dense blocks. The initial feature of the input image X is X0. Under the dense connection mechanism, the feature X is generated. l , the generation process is as follows:
[0064] X l =H d ([X0,X1,...,X l-1 ])
[0065] Among them, H d (·) represents the combination of 3×3 convolution and Leaky ReLU activation to ensure the stability and nonlinear expression of feature extraction. The gradient enhancement unit applies anisotropic Sobel operator to extract gradient information to enhance the structural contour and edge details of the image. Let the Sobel filter K x , K y Gradient calculations acting on the X and Y directions respectively:
[0066] G x =W(X)*K x ,G y =W(X)*K y
[0067] Among them, W(X) represents the discrete wavelet transform (DWT) of the input image to improve the robustness of the gradient feature. Finally, the feature enhancement module uses the residual fusion strategy to combine the original features with the enhanced gradient information to obtain the output feature Then, through UKE, a multi-scale initial feature representation is generated
[0068] Step 3: The initial feature representation obtained in step 2 is The input is fed into the Dual-Modality CrossAttention (DMCA) module, which contains the Gradient Extraction Block (GEB) and the Cross Attention Block (CAB) to enhance cross-modal feature interactions.
[0069] Construct a cross-modal complementary network including multiple cross-attention blocks, and pass the features output in step 2 through 3×3 convolution and ReLU activation function respectively. The input features of modalities A and B will be used as components of GEB, CAB1, CAB2 and element-by-element addition respectively, and the final result of feature addition is the output. Specifically, in the cross-attention block, the key-query-value (KQV) mechanism is adopted to improve the feature fusion effect through cross-modal information interaction. CAB consists of two interacting sub-modules (CAB1 and CAB2), which process features of different modalities respectively, and realize information sharing and alignment through cross-modal attention calculation. The calculation method is as follows:
[0070]
[0071] Among them, the Attention mechanism is defined as:
[0072]
[0073] After calculation by CAB1 and CAB2, the information of modal B is incorporated into the features of modal A, and the information of modal A is incorporated into the features of modal B, achieving deep cross-modal interaction.
[0074] Finally, the gradient enhancement features and the cross-attention features are fused through element-wise addition to obtain the optimized cross-modal feature representation:
[0075]
[0076] Step 4: The weighted optimized fusion features obtained in step 3 are input into the dual-branch feature disentanglement module. The global representation encoder (GRE) is used to extract low-frequency common features, and the local texture encoder (LTE) is used to extract high-frequency unique features.
[0077] A dual-branch feature decoupling and adaptive representation modeling framework is constructed. The Transformer-based Global Representation Encoder (GRE) is used to extract common features. At the same time, a reversible transform is introduced into the Local Texture Encoder (LTE) to accurately retain high-frequency information. The adversarial feature decoupling loss is used to ensure feature independence. Specifically, the features output in step 3 are The high-frequency features and low-frequency features of the corresponding modality are extracted through two branches respectively. The global representation encoder adopts the self-attention mechanism to enhance the cross-region information correlation through global context modeling, thereby extracting low-frequency structural information. This embodiment introduces a Transformer-based feature encoder to perform global semantic interaction in multi-layer self-attention calculations to preserve the overall scene structure while avoiding local redundant information interference. Specifically, the input feature and After linear projection, it is transformed into keys, queries, and values, and the global feature representation is calculated through the Multi-Head Self-Attention mechanism (MHSA):
[0078]
[0079] Where d is the feature dimension. After stacking multiple layers of Transformer, the global feature representation GRE captures long-range dependencies in the feature space and enhances the ability to model low-frequency structural information.
[0080] However, the global representation method based on self-attention is prone to loss of local details during the feature fusion process. To this end, this embodiment introduces a local texture encoder (LTE), which is built based on a neural network to decouple and reconstruct high-frequency detail information in a high-fidelity manner. This embodiment uses conditional reversible transformation (CRT) to transform the input features without introducing information loss. and Since reversibility ensures lossless information transmission, this embodiment can maintain the original high-frequency texture after decoupling and accurately restore details during reconstruction.
[0081] This embodiment introduces an adaptive feature disentanglement mechanism between GRE and LTE, ensuring the mutual complementation of local and global features through a dynamic feature transformation strategy. Specifically, this embodiment constructs a feature disentanglement adversarial loss (FDA-Loss) to encourage the minimization of the mutual information between the two feature representations:
[0082]
[0083] Among them, D KL (·||·) represents the Kullback-Leibler divergence, which is used to measure the statistical distance between two feature distributions. By minimizing D KL (P(Φ GRE )||P(Φ LTE )), this embodiment makes the extracted features more independent in a statistical sense, avoids information omissions between modalities, thereby improving the discriminability and reconstructibility of the final fusion features, and ultimately achieves efficient decoupling of multimodal dual-branch features.
[0084] Step 5: Input the global and local features obtained through dual-branch decoupling in step 4 into the fusion layer for feature splicing and reconstruction, and finally generate a fused image.
[0085] For the two pairs of high and low frequency features obtained by decoupling the two modes obtained in step 4, in order to fully explore the complementarity of these four features, this embodiment adopts a step-by-step fusion strategy, that is, the high frequency features of modes A and B are combined into and Sent to the local fusion layer, low-frequency features and Sent to the global fusion layer.
[0086] In the local fusion layer, specifically, this embodiment concatenates the high-frequency features of modal A and modal B in the channel dimension:
[0087]
[0088] Since high-frequency features are often affected by noise and local variations, this embodiment introduces an Adaptive Feature Weighting (AFW) mechanism to calculate the weights of the high-frequency features of each modality:
[0089]
[0090] Among them, σ(·) is the Sigmoid activation function, W LTE and b LTEis a learnable parameter. To further enhance the information expression capability, this embodiment uses a gated feature fusion mechanism to perform a weighted combination of high-frequency features:
[0091]
[0092] Similar to high-frequency feature fusion, this embodiment first concatenates low-frequency features in the channel dimension:
[0093]
[0094] Then, the adaptive weights of the low-frequency features are calculated:
[0095]
[0096] The fused low-frequency features are expressed as:
[0097]
[0098] After completing the fusion of high-frequency and low-frequency features, this embodiment needs to combine them together to form the final fusion feature:
[0099]
[0100] The final fusion features are sent to the decoder for feature reconstruction to generate a fused image.
[0101] This example addresses the shortcomings of existing multimodal medical image fusion methods in terms of global and local information modeling, intermodal feature interaction, and edge detail preservation. A multimodal medical image fusion method based on the Transformer and dual-branch feature decoupling is proposed. Ultimately, an adaptive fusion strategy enables the model to dynamically adjust the contribution of information from different modalities, thereby generating more accurate, clear, and stable fused images in medical image fusion tasks.
[0102] like Figure 1 The flowchart of the multimodal medical image fusion method based on Transformer and dual-branch feature decoupling in this embodiment is shown as follows: Figure 2 、 Figure 3 、 Figure 4 FIG. 1 is a schematic diagram of a multimodal medical image fusion method based on Transformer and dual-branch feature decoupling according to this embodiment, and a detailed schematic diagram of the edge enhancement module EEM and the dual-modal cross attention module DMCA included therein, which are described in detail as follows:
[0103] (1) Preprocess the input multimodal medical images, convert the medical image data in JPG format into h5 format, and normalize the data to improve the stability and effectiveness of subsequent feature extraction.
[0104] (2) The multimodal image is input into the edge feature enhancement module in two ways. One branch passes the image feature through the dense block, and the gradient enhancement unit of the other branch applies the anisotropic Sobel operator to extract the gradient information to enhance the structural contour and edge details of the image. Finally, the feature enhancement module uses the residual fusion strategy to combine the original features with the enhanced gradient information to obtain the output feature Then, they are input into the Transformer-based general knowledge encoder to extract shallow features. The encoder uses a convolutional mapping layer to transform the input features and generate multi-scale initial feature representations.
[0105] (3) The initial features are input into a cross-modal complementary network consisting of multiple cross-attention blocks, and the bimodal features are respectively subjected to 3×3 convolution and ReLU activation functions. The input features of modalities A and B will be used as components of GEB, CAB1, CAB2, and element-wise addition, and the final result of feature addition is the output. Specifically, in the cross-attention block, the key-query-value (KQV) mechanism is adopted to improve the feature fusion effect through cross-modal information interaction. CAB consists of two interacting sub-modules (CAB1 and CAB2), which process features of different modalities respectively and realize information sharing and alignment through cross-modal attention calculation. After calculation by CAB1 and CAB2, the information of modal B is integrated into the features of modal A, and the information of modal A is integrated into the features of modal B, realizing deep cross-modal interaction. Finally, the gradient enhancement features and the cross-attention features are fused through element-wise addition to obtain the optimized cross-modal feature representation.
[0106] (4) The optimized features are input into the dual-branch feature decoupling module. The global representation encoder is used to extract low-frequency common features, and the local texture encoder is used to extract high-frequency unique features. The global representation encoder is based on the self-attention mechanism of Transformer, which can effectively extract global features. The local texture encoder uses conditional reversible transformation to transform the input features. and Bidirectional mapping is performed. Because reversibility ensures lossless information transmission, this embodiment can preserve the original high-frequency texture after decoupling and accurately restore details during reconstruction. This embodiment introduces an adaptive feature decoupling mechanism between GRE and LTE, ensuring the mutual complementation of local and global features through a dynamic feature transformation strategy.
[0107] The global and local features obtained by dual-branch decoupling are input into the fusion layer for feature splicing and reconstruction, and finally a fused image is generated.
[0108] The present embodiment is further described below through specific examples:
[0109] This example uses the MSRS, TNO, and RoadScene datasets for training and testing, and tests and generalizes the method on the Harvard WholeBrainAtlas medical dataset. The specific dataset processing flow is as follows:
[0110] In the MSRS dataset, this example uses 1083 pairs of images for training and 361 pairs for testing. For the RoadScene dataset, this example uses 30 of the 50 pairs of images for training and the remaining 20 for testing. For the TNO dataset, this example uses 300 of the 361 pairs of images for training and 61 for testing. This example preprocesses the images in these datasets, resizing them to 256×256 pixels and normalizing pixel intensities to the range [0, 1].
[0111] The experimental results of this example were qualitatively compared with other methods. A multimodal medical image fusion method based on Transformer and dual-branch feature decoupling achieved high-quality fusion performance on the Harvard Whole Brain Atlas medical dataset. Specifically, on the MRI-PET dataset, the proposed method is compared with the unified unsupervised image fusion network (U2Fusion), the cross-domain long-range learning for general image fusion via swin transformer (SwinFusion), the compressed decomposition network (SDNet), the general semantic-guided network with coupled mask ensemble for medical image fusion (GeSeNet), a generative infrared and visible light image fusion adversarial network (FusionGAN), the dual-discriminator conditional generative adversarial network for multi-resolution image fusion (DDcGAN), and the correlation-driven dual-branch feature decomposition model for multi-modality image fusion. The fusion quality is significantly improved compared to Fusion (CDDFuse) (see Table 1, Table 2, Table 3). This embodiment's multimodal medical image fusion method based on Transformer and dual-branch feature decoupling effectively helps the model achieve good fusion performance.
[0112] Table 1 Comparison results with other methods using Harvard Whole Brain Atlas medical dataset MRI-PET
[0113]
[0114]
[0115] Table 2 Comparison results with other methods using Harvard Whole Brain Atlas medical dataset MRI-CT
[0116]
[0117] Table 3 Comparison results with other methods using Harvard Whole Brain Atlas medical dataset MRI-SPECT
[0118]
[0119]
[0120] The specific visualization results are as follows Figure 5 、 Figure 6 、 Figure 7 As shown in the results, it can be seen that the fusion results of this embodiment have better inter-modal information complementarity than other algorithms. Therefore, this embodiment adopts the above-mentioned multimodal medical image fusion method based on Transformer and dual-branch feature decoupling to achieve the combination of cross-modal image feature enhancement and cross-modal consistency learning to improve fusion quality.
[0121] This embodiment may also provide a multimodal medical image fusion system based on Transformer and dual-branch feature decoupling, including:
[0122] a data acquisition module, configured to acquire a multimodal medical image and perform normalization processing on the multimodal medical image to obtain a preprocessed multimodal medical image;
[0123] An image fusion module is used to input the preprocessed multimodal medical images into a medical image fusion model for image fusion to obtain a fused image; wherein the medical image fusion model includes a shallow feature extraction module, a bimodal cross attention module, a dual-branch feature decoupling module and a splicing fusion module connected in sequence.
[0124] This embodiment can also provide an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute a multimodal medical image fusion method based on Transformer and dual-branch feature decoupling.
[0125] Optionally, this embodiment further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the multimodal medical image fusion method based on Transformer and dual-branch feature decoupling.
[0126] The above description is merely a preferred embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A multimodal medical image fusion method based on Transformer and dual-branch feature decoupling, characterized by: include: Acquiring a multimodal medical image, and performing normalization processing on the multimodal medical image to obtain a preprocessed multimodal medical image; The preprocessed multimodal medical image is input into a medical image fusion model for image fusion to obtain a fused image; wherein the medical image fusion model includes a shallow feature extraction module, a bimodal cross attention module, a dual-branch feature decoupling module and a splicing fusion module connected in sequence.
2. The method according to claim 1, characterized in that The training process of the medical image fusion model specifically includes: Acquiring training data, wherein the training data includes multimodal medical training images and corresponding fusion results; An initial medical image fusion model is constructed, the training data is input into the initial medical image fusion model for image fusion, and training is performed with the goal of minimizing the loss between the initial training result after image fusion and the fusion result corresponding to the multimodal medical training image to obtain a trained medical image fusion model.
3. The method according to claim 1, characterized in that The processing process of the medical image fusion model specifically includes: The preprocessed multimodal medical image is input into the shallow feature extraction module to perform edge feature enhancement and extract shallow features of the image through a general knowledge encoder to obtain a multi-scale initial feature representation; Inputting the multi-scale initial feature representation into the bimodal cross-attention module, enhancing cross-modal feature interaction through the gradient extraction block and the cross-attention block to obtain gradient enhanced features and cross-attention features, and fusing the obtained gradient enhanced features and cross-attention features by element-by-element addition to obtain weighted optimized fused features; The weighted optimized fusion features are input into the dual-branch feature decoupling module to extract global features and local features; The global features and the local features are input into a splicing and fusion module for feature splicing and reconstruction to generate a fused image.
4. The method according to claim 3, characterized in that The processing process of the shallow feature extraction module specifically includes: The initial feature of the input multimodal medical image X is X0. Under the dense connection mechanism, the feature X is generated. l : X l =H d ([X0, X1,…, X l-1 ]) Where H d (·) represents the combination of 3×3 convolution and Leaky ReLU activation; The gradient information is extracted by applying the anisotropic Sobel operator through the gradient enhancement unit to enhance the structural contour and edge details of the image; let the Sobel filter K x , K y Gradient calculations acting on the X and Y directions respectively: G x =W(X)*K x ,G y =W(X)*K y Where W(X) represents the discrete wavelet transform of the input image, G X is the gradient calculation result in the X direction, G Y is the gradient calculation result in the Y direction; The residual fusion strategy is used to combine the original features with the enhanced gradient information to obtain the output features Then, through the general knowledge encoder, generate the multi-scale initial feature representation 5. The method according to claim 3, characterized in that The processing process of the bimodal cross attention module specifically includes: Constructing a cross-modal complementary network, the cross-modal complementary network comprising a cross-attention block and a gradient extraction block, the cross-attention block consisting of a first cross-attention block and a second cross-attention block interacting with each other; The initial feature representation is subjected to 3×3 convolution and ReLU activation function, and the input features of modality A and modality B are used as the components of the first gradient extraction block, the first cross attention block, the second cross attention block, and element-by-element addition respectively. The final result of feature addition is the weighted optimized fusion feature.
6. The method according to claim 3, characterized in that The processing process of the dual-branch feature decoupling module specifically includes: The weighted optimized fusion features are input into the dual-branch feature decoupling module, and the low-frequency common features of the input data are extracted through the global representation encoder to obtain the global features; the high-frequency unique features of the input data are extracted through the local texture encoder to obtain the local features.
7. The method according to claim 3, characterized in that The processing process of the splicing and fusion module specifically includes: The global features and local features of different modalities are spliced separately in the channel dimension to obtain the fused global features and the fused local features; Combine the fused global features and the fused local features to obtain fused features; The fused features are input into the decoder for feature reconstruction to generate a fused image.
8. A multimodal medical image fusion system based on Transformer and dual-branch feature decoupling, characterized by: include: a data acquisition module, configured to acquire a multimodal medical image and perform normalization processing on the multimodal medical image to obtain a preprocessed multimodal medical image; An image fusion module is used to input the preprocessed multimodal medical images into a medical image fusion model for image fusion to obtain a fused image; wherein the medical image fusion model includes a shallow feature extraction module, a bimodal cross attention module, a dual-branch feature decoupling module and a splicing fusion module connected in sequence.
9. An electronic device, characterized in that: It includes a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform a multimodal medical image fusion method based on Transformer and dual-branch feature decoupling according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that It stores a computer program, which, when executed by a processor, implements a multimodal medical image fusion method based on Transformer and dual-branch feature decoupling as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Transform method of tracking structure for image restoration
CN115619685A
Global and local feature interactive parallel multi-modal medical image fusion method
CN117974468A
Multi-modal medical image fusion method, system, device, medium and program product
CN119624793A
Medical image segmentation method
WO2024098318A1
Cited By
Medical image fusion method, system and equipment based on space-frequency domain feature interaction analysis and medium
CN121121374A
Image data compression method and device and storage medium
CN121357335A