Medical Image Fusion Method and System Based on Large Kernel Attention Mechanism

By introducing large-nuclear attention mechanism and frequency domain dynamic aggregation processing in medical image fusion, the problem that existing methods are difficult to capture global context and local details is solved, and high-quality medical image fusion is achieved.

CN119323523BActive Publication Date: 2025-06-24TAISHAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411854221.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-06-24
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

Existing medical image fusion methods are difficult to effectively capture global context information and local details in multimodal images, resulting in low quality of the fusion image and prone to artifacts and blur problems.

Method used

The medical image fusion method based on the large-nuclear attention mechanism is adopted, combined with multi-scale convolution operation, channel attention and spatial attention mechanism, and the precise fusion of image details and structure is achieved through dynamic aggregation of frequency domains.

Benefits of technology

It improves the retention ability of the fused image in global information and local details, reduces artifacts and blur problems, and improves image reconstruction quality and diagnostic accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119323523B_ABST
    Figure CN119323523B_ABST
Patent Text Reader

Abstract

The present invention proposes a medical image fusion method and system based on a large kernel attention mechanism, belonging to the technical field of medical image fusion. By introducing the large kernel attention mechanism and frequency domain transformation processing, and combining multi-scale convolution operations, channel attention, and spatial attention mechanisms, a higher balance is achieved in the retention of global information and local details in the fused image. At the same time, through the separation processing of amplitude and phase information in the frequency domain, the precise fusion of image details and structures is realized, avoiding the common artifacts and blurring problems in traditional methods. In addition, the present invention designs a densely connected decoder module, enhancing the ability to retain details during image reconstruction, and through an adaptive feature fusion mechanism, it is more intelligent and accurate in processing the features of different modality images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image fusion, and particularly relates to a medical image fusion method and system based on a large kernel attention mechanism. Background Technique

[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] With the development of modern medical imaging technology, multi-modal medical imaging technologies such as MRI (Magnetic Resonance Imaging) and PET (Positron Emission Tomography) have been widely used in clinical diagnosis. MRI can provide high-resolution soft tissue anatomical information, while PET can reflect cell metabolic activities. Through the fusion of these two types of images, doctors can observe the pathological conditions of patients more comprehensively, thereby improving the accuracy of diagnosis and the treatment effect. However, how to efficiently and accurately fuse multi-modal images has become an important research direction in the field of medical image processing.

[0004] In the prior art, common image fusion methods mainly include pixel-level, feature-level, and decision-level fusion methods. Pixel-level fusion methods directly process and synthesize the pixel values of the source images. Typical algorithms include wavelet transform, pyramid transform, etc. However, these methods are prone to losing important structural information or introducing artifacts, resulting in low-quality fused images. Feature-level fusion methods fuse the images by extracting the features of the source images. Although the quality of the fused images is improved, the feature selection and processing processes are complex and often require a large amount of computing resources. Decision-level fusion methods use the classification or recognition results of the images for high-level decision-making, but this method usually requires prior knowledge and is difficult to meet the requirements of multi-modal image fusion in complex scenarios.

[0005] In recent years, deep learning technology has been introduced into the field of image fusion, and significant progress has been made. Through deep learning models such as convolutional neural networks (CNNs), image fusion methods can more effectively learn the deep features of images, thereby realizing more intelligent image fusion processing. Some studies have proposed using self-attention mechanisms to enhance the feature selection and representation capabilities during the fusion process, but existing methods often ignore the full utilization of multi-scale information in the feature extraction process. Multi-scale features are of great significance in medical image fusion because different modal images may contain key information at different scales.

[0006] In addition, the attention mechanisms in the prior art usually use a relatively small kernel size for convolution operations. Although they can improve the local perception ability of the model, they cannot fully capture the global context information in medical images. Therefore, how to design a more effective attention mechanism, especially for capturing large-scale context, has become the key to improving the image fusion effect. Summary of the Invention

[0007] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a medical image fusion method and system based on a large kernel attention mechanism, which introduces a large kernel attention mechanism and frequency domain transformation processing, and combines multi-scale convolution operations, channel attention and spatial attention mechanisms, so that the fused image achieves a higher balance in retaining global information and local details. At the same time, by separating the amplitude and phase information in the frequency domain, the precise fusion of image details and structures is realized, avoiding the common artifacts and blurring problems in traditional methods. In addition, the present invention also designs a densely connected decoder module, which enhances the ability to retain details during image reconstruction, and through an adaptive feature fusion mechanism, is more intelligent and accurate in processing the features of different modality images.

[0008] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:

[0009] The first aspect of the present invention provides a medical image fusion method based on a large kernel attention mechanism;

[0010] A medical image fusion method based on a large kernel attention mechanism includes:

[0011] Obtain multi-modal medical images and preprocess the obtained multi-modal medical images;

[0012] Input the preprocessed multi-modal medical images into a two-branch encoder, and perform preliminary feature extraction through a multi-scale convolutional layer;

[0013] Use the large kernel attention mechanism to perform group normalization, projection and segmentation processing on the preliminarily extracted features, and obtain multiple segmented features;

[0014] Input the multiple segmented features into large kernel attention sub-modules and corresponding spatial feature extraction convolutional layers of different scales respectively to obtain the global and local features of the image: splice the global and local features and project them to the original dimension;

[0015] Use spectral dynamic aggregation to perform Fourier transform on the multi-scale features projected to the original dimension, and process the amplitude and phase information of the multi-scale features to achieve dynamic aggregation in the frequency domain;

[0016] Use multi-scale convolution, channel attention and spatial attention mechanisms to process the features after Fourier transform, and realize the adaptive fusion of features of different modalities and different regions and the adaptive weight adjustment;

[0017] Use a densely connected decoder to process the fused features to obtain the final fused image representation.

[0018] As a further technical solution, the multimodal medical images include MRI magnetic resonance imaging and PET positron emission tomography images.

[0019] As a further technical solution, the process of preprocessing the acquired multimodal medical images is as follows: adjusting the input images to a standard size of 128×128, and ensuring that the pixel value ranges of the multimodal medical images are consistent through normalization operations.

[0020] As a further technical solution, the process of inputting the preprocessed multimodal medical images into the dual-branch encoder and performing preliminary feature extraction through multi-scale convolutional layers is as follows:

[0021] The preprocessed MRI and PET images are respectively input into their respective branch encoders for feature extraction through multi-scale convolutional layers; the multi-scale convolution uses convolutional kernels of different sizes to capture feature information of different scales in the medical images;

[0022] ;

[0023] ;

[0024] where, and represent the multi-scale convolution extractors of the MRI and PET branches; the output and are the preliminary feature representations of different modalities; and are the preprocessed MRI and PET images respectively.

[0025] As a further technical solution, the process of respectively inputting the segmented multiple medical image features into large kernel attention sub-modules of different scales and corresponding spatial feature extraction convolutional layers to obtain the global and local features of the images is as follows:

[0026]

[0027] In the formula, , , are the global feature maps obtained after being processed by large kernel attention sub-modules of different scales; , , are the local feature maps obtained after being processed by spatial feature extraction convolutional layers of different scales; LKA3, LKA5, LKA7 are large kernel attention sub-modules of different scales; X3, X5, X7 are the corresponding 3x3, 5x5, 7x7 spatial feature extraction convolutional layers; , , Multiple medical image features after segmentation;

[0028] The process of splicing the global and local features and then projecting them to the original dimension is as follows: Multiply the large kernel convolution result by the spatial feature extraction result to enhance the representation ability of multi-scale features:

[0029]

[0030] In the formula and are the feature representations after multi-scale feature fusion, containing the combined information of global and local features;

[0031] The spliced multi-scale features are projected back to the original dimension through a projection convolution layer to ensure that the feature size is consistent with the input:

[0032] ;

[0033] ;

[0034] In the formula, and are the final feature representations obtained after processing by the last projection layer, and their dimensions are consistent with the input features; is the projection convolution layer.

[0035] As a further technical solution, the process of using spectral dynamic aggregation to perform Fourier transform on the multi-scale features projected to the original dimension and processing the amplitude and phase information of the multi-scale features to achieve dynamic aggregation in the frequency domain is as follows:

[0036] The spectral dynamic aggregation module performs Fourier transform on the multi-scale features projected to the original dimension, converting them from the spatial domain to the frequency domain , and respectively obtaining the amplitude and phase information of the features;

[0037] ;

[0038] ;

[0039] Among them, represents the Fourier transform operation, and respectively represent the representations of MRI and PET image features in the frequency domain; after Fourier transform, the features are decomposed into amplitude information and phase information ;

[0040] ;

[0041] ;

[0042] Among them, and respectively represent the amplitude information of MRI and PET images, and are phase information, is the Euler representation on the complex plane, where i is the imaginary unit, is the phase angle.

[0043] As a further technical solution, a medical image fusion method based on a large kernel attention mechanism further includes loss function design.

[0044] The second aspect of the present invention provides a medical image fusion system based on a large kernel attention mechanism.

[0045] A medical image fusion system based on a large kernel attention mechanism includes:

[0046] A multi-modal medical image acquisition module, configured to: acquire multi-modal medical images and preprocess the acquired multi-modal medical images;

[0047] A large kernel attention mechanism processing module, configured to: input the preprocessed multi-modal medical images into a dual-branch encoder, and perform preliminary feature extraction through multi-scale convolutional layers; use the large kernel attention mechanism to perform group normalization, projection, and segmentation processing on the preliminarily extracted features to obtain multiple segmented features;

[0048] A multi-scale large kernel convolution and spatial feature extraction module, configured to: input the multiple segmented features into different-scale large kernel attention sub-modules and corresponding spatial feature extraction convolutional layers respectively to obtain the global and local features of the image: perform splicing processing on the global and local features and then project them to the original dimension;

[0049] A spectral dynamic aggregation module, configured to: perform Fourier transform on the multi-scale features projected to the original dimension by using the spectral dynamic aggregation module, and process the amplitude and phase information of the multi-scale features to achieve dynamic aggregation in the frequency domain;

[0050] A feature fusion and adaptive weight adjustment module, configured to: process the features after Fourier transform by using multi-scale convolution, channel attention, and spatial attention mechanisms to achieve adaptive fusion of features of different modalities and different regions and adaptive weight adjustment;

[0051] A medical image representation module, configured to: process the fused features using a densely connected decoder to obtain a finally fused medical image representation.

[0052] A third aspect of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in a medical image fusion method based on a large kernel attention mechanism as described in the first aspect of the present invention.

[0053] A fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a medical image fusion method based on a large kernel attention mechanism as described in the first aspect of the present invention.

[0054] The above one or more technical solutions have the following beneficial effects:

[0055] By combining multi-scale feature extraction, large kernel attention mechanism, and spectral dynamic aggregation, the present invention efficiently processes the features of MRI and PET images in the frequency domain. Compared with traditional methods, the present invention can not only capture the local details and global structure of images, but also effectively improve the modeling ability of long-range dependent features through large kernel convolution. At the same time, the proposed global-local adaptive fusion strategy can dynamically adjust the feature weights according to the importance of different modalities and regions, ensuring the accurate fusion of information.

[0056] In addition, the present invention uses a dual-branch encoder to extract the preprocessed MRI and PET images, which is beneficial to extracting and fusing the multi-modal features of the images; the decoder realizes the full utilization of multi-scale features through a densely connected structure, further improving the quality of image reconstruction, reducing information loss and artifacts. Different from the traditional symmetric encoder-decoder structure, the asymmetric encoder-decoder architecture provided by the present invention is more suitable for the characteristics of multi-modal medical image fusion, and finally outputs a fusion image with better quality. Further, by introducing phase and amplitude losses, the phase and amplitude consistency in the frequency domain is enhanced, ensuring the high quality of the fusion image. In practical applications, the present invention can effectively improve the accuracy of medical image fusion, ensure richer detail presentation and more accurate diagnostic information, and provide strong technical support for medical image analysis.

[0057] The advantages of the additional aspects of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0059] Figure 1 It is the flowchart of the method for the first embodiment.

[0060] Figure 2 It is the schematic diagram of the large kernel attention mechanism module structure in the first embodiment.

[0061] Figure 3 It is the schematic diagram of the spectrum dynamic aggregation process in the first embodiment.

[0062] Figure 4a It is the MRI image in the first embodiment; Figure 4b It is the PET image in the first embodiment; Figure 4c It is the fused image.

[0063] Figure 5 It is the system structure diagram of the second embodiment. Detailed implementation manners

[0064] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0065] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention.

[0066] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0067] MRI and PET images respectively provide anatomical structure and functional metabolism information. There are significant differences in feature distribution, spatial resolution, etc. between the images of these two modalities. Existing image fusion methods, such as pixel-level and feature-level fusion, usually cannot effectively capture the complementary information between these two modalities; at the same time, most medical image fusion methods are limited to processing in the spatial domain and do not fully utilize frequency domain information.

[0068] In medical images, it is necessary to retain the global anatomical structure information (such as the outline and morphology of organs) and also to depict the details of local lesion areas in detail (such as the morphological features of lesions). Existing small kernel convolutions and local attention mechanisms perform well in processing local features, but have limited ability to capture global context information and are prone to losing some key global structure information.

[0069] Based on the above problems, the present invention proposes a medical image fusion method and system based on a large kernel attention mechanism, which will be described in detail below with specific embodiments.

[0070] Example 1

[0071] This embodiment discloses a medical image fusion method based on a large kernel attention mechanism;

[0072] As Figure 1 shown, a medical image fusion method based on a large kernel attention mechanism includes:

[0073] Step S1, obtain multi-modal medical images and preprocess the obtained multi-modal medical images; wherein, in this embodiment, the multi-modal medical images are MRI images and PET images, and the process of preprocessing the MRI images and PET images includes: standardization of the image size, that is, adjusting the input images to a standard size of 128×128, and ensuring that the pixel value ranges of the medical images of each modality are consistent through a normalization operation. This step helps to eliminate the scale differences between different image modalities and lays a foundation for subsequent feature extraction.

[0074] In the data flow process, the first operations completed are as follows:

[0075] ;

[0076] ;

[0077] Among them, and are the preprocessed MRI image and PET image respectively.

[0078] Step S2, input the preprocessed multi-modal medical images into a dual-branch encoder, and perform preliminary feature extraction through multi-scale convolutional layers; use the large kernel attention mechanism to perform group normalization, projection, and segmentation processing on the preliminarily extracted features to obtain multiple segmented features;

[0079] Specifically, in combination with Figure 2 , in step S2, the preprocessed MRI and PET images are respectively input into their respective encoders, and feature extraction is first performed through multi-scale convolutional layers. In this embodiment, the encoder adopts a dual-branch structure, each branch contains multiple convolutional blocks, each convolutional block consists of a standard convolution, batch normalization, activation function, and large kernel attention module, and the dual-branch encoder uses skip connections to connect features layer by layer to enhance feature expression, and finally the dual-branch features are adaptively fused through a fusion module.

[0080] In the dual-branch encoder, the multi-scale convolutional layers use convolutional kernels of different sizes (such as 3×3, 5×5, 7×7) to capture feature information of different scales in the images.

[0081] ;

[0082] ;

[0083] Among them, and represent multi-scale convolutional extractors for the MRI and PET branches, and the output and are preliminary feature representations of different modalities.

[0084] For the preliminarily extracted features and perform group normalization to reduce the difference in feature distribution between channels and improve training stability.

[0085] ;

[0086] ;

[0087] In the formula, is group normalization.

[0088] The normalized features are projected into a higher-dimensional space through the projection convolutional layer project. is a 1×1 convolutional layer used to expand the dimension of the features so that subsequent multi-scale convolutional operations can run in a richer feature space. The specific projection operation is:

[0089] ;

[0090] ;

[0091] , are the projected feature maps; the number of channels of the projected features is expanded from C to 2C.

[0092] The projected features are divided into two parts, denoted as and , among which, represents the part used to maintain the residual connection after feature segmentation; will be further divided into three parts , and ;

[0093]

[0094] Step S3, input the multiple segmented features into large kernel attention sub-modules of different scales and corresponding spatial feature extraction convolutional layers respectively to obtain the global and local features of the image: splice the global and local features and then project them to the original dimension;

[0095] The segmented , and are respectively input into the large kernel attention sub-modules (LKA3, LKA5, LKA7) of different scales and the corresponding spatial feature extraction convolutional layers (X3, X5, X7) to process global and local features respectively:

[0096]

[0097] In the formula, , , are the global feature maps obtained after being processed by the large kernel attention sub-modules of different scales; , , are the local feature maps obtained after being processed by the spatial feature extraction convolutional layers of different scales; LKA3, LKA5, LKA7 are the large kernel attention sub-modules of different scales, and each LKA module contains a combination of depth convolution and dilated convolution to expand the receptive field; X3, X5, X7 are the corresponding 3x3, 5x5, 7x7 spatial feature extraction convolutional layers; , , are multiple medical image features after segmentation.

[0098] Multiply the large kernel convolution result and the spatial feature extraction result to enhance the representation ability of multi-scale features:

[0099]

[0100] In the formula , is the feature representation after multi-scale feature fusion, which contains the combined information of global and local features;

[0101] The concatenated multi-scale features pass through the projection convolutional layer and are projected back to the original dimension to ensure that the feature size is consistent with the input:

[0102] and ;

[0103] In the formula, , respectively represent the final feature representations of the MRI and PET images after being processed by the projection convolutional layer. These feature dimensions are consistent with the input and are prepared for subsequent frequency domain processing.

[0104] Step S4, perform Fourier transform on the multi-scale features projected onto the original dimension using spectral dynamic aggregation, and process the amplitude and phase information of the multi-scale features to achieve dynamic aggregation in the frequency domain;

[0105] Combined with Figure 3 , first, the features extracted by the dual-branch encoder (such as MRI and PET encoders) and After being processed by multi-scale large kernel convolution, they are input into the spectral dynamic aggregation module. First, the spectral dynamic aggregation module performs Fourier transform on the input features, converting them from the spatial domain to the frequency domain , and respectively obtaining the amplitude and phase information of the features.

[0106] ;

[0107] Among them, represents the Fourier transform operation, and respectively represent the representations of MRI and PET image features in the frequency domain. After Fourier transform, the features are decomposed into amplitude information and phase information :

[0108] and ;

[0109] Among them, and respectively represent the amplitude information of MRI and PET images, and are the phase information, is the Euler representation on the complex plane, i is the imaginary unit, is the phase angle, and this complex form of representation can simultaneously contain the amplitude information and the phase information , which helps to complete the result and detail features of the image in the frequency domain.

[0110] The amplitude information reflects the global structure and brightness distribution in the image. Therefore, the amplitude processing part in the spectral dynamic aggregation module dynamically adjusts the amplitude information of MRI and PET features. Specifically, convolutional layers and non-linear activation functions (such as Leaky ReLU) are used to enhance the amplitude information to ensure that the global structure information is retained and strengthened. The processing of the amplitude information can be expressed as:

[0111] ;

[0112] ;

[0113] and extracts the amplitude information from the Fourier-transformed MRI and PET features, represents a convolutional layer for enhancing the amplitude information, is a non-linear activation function. The processed amplitude information passes through the second convolutional layer for further adjustment to ensure its adaptability to subsequent feature fusion tasks.

[0114] Finally, the processed amplitude information and are recombined back into the frequency-domain features.

[0115] Secondly, in the process of processing the amplitude and phase information of multi-scale features to achieve dynamic aggregation in the frequency domain, the phase information contains edge, texture, and detail information in the image, which is crucial for the visual effect of the image. The FSDA module processes the phase information to enhance the edges and details of the image. In the specific implementation, convolutional operations and attention mechanisms are used to dynamically adjust the phase information. The processing process of the phase information is as follows: The phase information mainly reflects the detail and texture features of the image, such as edges, contours, and local details. Therefore, the phase processing part in the spectral dynamic aggregation module dynamically adjusts the phase information of the MRI and PET features to ensure that the detail information is enhanced and optimized. Specifically, the phase processing module enhances the phase information through convolutional layers, non-linear activation functions, and phase importance weighting to ensure that details and textures are preserved during the fusion process.

[0116] Specifically, extract the phase information from the Fourier-transformed MRI and PET features and :

[0117] ;

[0118] ;

[0119] The input phase information and is processed through convolutional layers to enhance important features such as edges and details contained in the phase information. This process is represented by the following formula:

[0120] ;

[0121] ;

[0122] where, represents a convolutional layer for enhancing the phase information, It is a non - linear activation function (such as ReLU) used to further enhance details.

[0123] To dynamically adjust the contribution of phase information in different regions to details, the phase processing module introduces a phase importance weighting mechanism. This mechanism calculates the weight of the phase information , and weights the processed phase information according to the weight to enhance the attention to key regions. The calculation formula for the phase importance weight is:

[0124] ;

[0125] ;

[0126] Weight the phase information using the phase importance weight:

[0127] ;

[0128] ;

[0129] In the formula, is the phase weight generated by the convolutional layer, and the activation function is used to control the weight within the range of . After the phase importance weighting process, the obtained and are the phase information that has enhanced details and textures. Finally, these weighted phase information will be used to combine with the processed amplitude information to generate a new complex - valued feature representation:

[0130] ;

[0131] ;

[0132] Inverse Fourier transform the processed frequency - domain features to return to the spatial domain and generate new spatial features. The formula for the inverse Fourier transform is:

[0133] ;

[0134] ;

[0135] After returning from the frequency domain to the spatial domain, to further adaptively control the intensity of these features, the large - kernel attention module introduces a learnable scaling parameter . This scaling parameter is a learnable scalar matrix that acts on each channel to dynamically adjust the amplitude of the output features:

[0136] ;

[0137] ;

[0138] Among them, is a learnable parameter corresponding to each channel, and the optimal scaling ratio of each channel is gradually learned through training. This scaling mechanism ensures that the model can flexibly adjust the output intensity in different feature situations, preventing features from being too strong or too weak. At the same time, in order to retain the key information in the input features, the large kernel attention module finally adds the original input features and to the features and after scaling adjustment:

[0139] ;

[0140] ;

[0141] Thus, the final Fourier-transformed features and are obtained. This operation enables the original input features and the features processed in the frequency domain and adjusted by scaling to act together, preventing information loss while enhancing the feature representation ability.

[0142] Step S5: Use multi-scale convolution, channel attention, and spatial attention mechanisms to process the Fourier-transformed features to achieve adaptive fusion of features in different modalities and different regions and adaptive weight adjustment;

[0143] Specifically, the Fourier-transformed features and are first processed by convolutions of different scales to capture local and global information respectively. The convolution kernel extracts detailed information, the convolution kernel is used for medium-scale features, while the convolution kernel captures global context information. The specific formula is as follows:

[0144]

[0145] Add the features of different scales to form feature representations of each scale:

[0146] ;

[0147] During the calculation using the channel attention mechanism, for the features after multi-scale convolution processing, an adaptive average pooling (avg_pool) operation is first performed to extract the global information of each channel:

[0148] ;

[0149] ;

[0150] Next, the pooled result is used to calculate the channel attention weights through two layers of convolution and , and the Sigmoid activation function is used:

[0151] ;

[0152] Finally, the frequency-domain features are weighted:

[0153] ;

[0154] ;

[0155] Subsequently, the spatial attention mechanism is used to perform spatial weighting on the weighted frequency-domain features. First, convolution operations are used to extract the spatial attention weights and :

[0156] ;

[0157] ;

[0158] The frequency-domain features are weighted using the spatial attention weights:

[0159] ;

[0160] ;

[0161] After channel and spatial weighting processing, a convolution is used to process the fused frequency-domain features to ensure the consistency of the channel dimension:

[0162] ;

[0163] Finally, the original frequency-domain features and are added to the weighted fused features using a residual connection to generate the final frequency-domain fused features:

[0164] ;

[0165] Through the residual connection, the original frequency-domain information is retained, and the robustness of the processed fused features is enhanced.

[0166] Step S6, use the densely connected decoder to process the fused features to obtain the final fused image representation.

[0167] In this embodiment, the encoder adopts a dense connection structure, and the dense connection structure of the decoder focuses on optimizing the quality of the fusion result. It is used to first expand the fusion features to a high-dimensional space and then optimize the features through multiple dense connection blocks. Inside each dense connection block, consecutive convolution operations and feature concatenation are employed; finally, an optimized fusion image is output.

[0168] The fusion features output by the encoder First, through a convolutional layer for expansion to adapt to the processing dimension of the decoder:

[0169] ;

[0170] In the dense connection decoding block, features are concatenated layer by layer to ensure smooth and effective information transmission during the decoding process. Let the feature of the th layer of the dense connection block be , then the output of each layer can be expressed as:

[0171] ;

[0172] Through the dense connection method, each layer will concatenate all the previous features with the current output, thus forming a feature representation that accumulates layer by layer.

[0173] The final output of the dense connection decoding block will pass through a Figure 4c convolutional layer to compress the feature dimension to the original number of image channels, outputting the final image representation of the decoder. The fusion result is as Figure 4a shown, where Figure 4b is the MRI image and

[0174] ;

[0175] The dense connection decoder realizes further optimization and refinement of the fusion features through multiple layers of dense connections, can effectively enhance image details, and improve the visual quality of the fusion result. Specifically, the decoder first expands the low-dimensional fusion features to a higher-dimensional space, then gradually optimizes the feature representation through multiple dense connection blocks, and finally outputs a fusion image with better quality. It not only ensures the effective retention of detailed information during the fusion process but also can improve the overall quality of the fusion image through progressive optimization.

[0176] Furthermore, in this embodiment, a phase and amplitude loss function is also introduced. The phase and amplitude loss functions respectively calculate the differences in the phase information and amplitude information between the fusion image and the MRI and PET images, and measure the differences in phase and amplitude through the mean square error (MSE) respectively.

[0177] (a) Phase loss

[0178] The phase loss measures the phase difference by calculating the mean square error of the phases of the fused image and the MRI and PET images in the frequency domain. The formula for the phase loss is as follows:

[0179] ;

[0180] where represents the phase information of the fused image, and represent the phase information of the MRI and PET images, respectively.

[0181] (b) Amplitude loss

[0182] Similarly, the amplitude loss measures the amplitude difference by calculating the mean square error of the amplitudes of the fused image and the MRI and PET images in the frequency domain. The formula for the amplitude loss is as follows:

[0183] ;

[0184] where: represents the amplitude information of the fused image, and represent the amplitude information of the MRI and PET images, respectively.

[0185] It also includes an intensity loss function , a gradient loss function and a complete loss function. Among them, the intensity loss function is used to measure the pixel difference between the fused image and the MRI and PET images in the spatial domain. This loss function calculates the mean square error between the fused image and the MRI image, and the mean square error between the fused image and the PET image, and gives a weight of 0.5 to the loss of the PET image.

[0186] The formula is as follows:

[0187] ;

[0188] where, represents the pixel value of the fused image, and represent the pixel values of the MRI and PET images, respectively.

[0189] The gradient loss function is used to measure the difference in gradient information between the fused image and the MRI and PET images. This helps to preserve the edge and detail information in the image during the fusion process.

[0190] Let represent the gradient operator of the image (e.g., Sobel operator), and the formula of the gradient loss function is as follows:

[0191] ;

[0192] where: represents the gradient information of the fused image, and represent the gradient information of the MRI and PET images respectively.

[0193] Finally, the total loss function can be optimized by combining the phase and amplitude loss, intensity loss, and gradient loss, and the formula is as follows:

[0194] ;

[0195] where, and are the weight parameters of each loss term, used to control the influence of different loss terms in the total loss. The above weight parameters are set to 1, 1, 1, and 5 in sequence. Through the combination of the phase and amplitude loss, intensity loss, and gradient loss, the model can simultaneously focus on the phase information, amplitude information, pixel intensity, and gradient information of the image during the fusion process, so as to more comprehensively retain the key information of the MRI and PET images in the spatial domain and frequency domain.

[0196] Example Two

[0197] This example discloses a medical image fusion system based on a large kernel attention mechanism;

[0198] As Figure 5 shown, a medical image fusion system based on a large kernel attention mechanism includes:

[0199] A multi-modal medical image acquisition module, configured to: acquire multi-modal medical images and preprocess the acquired multi-modal medical images;

[0200] A large kernel attention mechanism processing module, configured to: input the preprocessed multi-modal medical images into a dual-branch encoder, and perform preliminary feature extraction through a multi-scale convolutional layer; use the large kernel attention mechanism to perform group normalization, projection, and segmentation processing on the preliminarily extracted features to obtain multiple segmented features;

[0201] A multi-scale large kernel convolution and spatial feature extraction module, configured to: input the multiple segmented features into different-scale large kernel attention sub-modules and corresponding spatial feature extraction convolutional layers respectively to obtain the global and local features of the image: perform splicing processing on the global and local features and then project them to the original dimension;

[0202] A spectrum dynamic aggregation module, configured to: perform Fourier transform on multi-scale features projected onto the original dimension by using the spectrum dynamic aggregation module, and process the amplitude and phase information of the multi-scale features to achieve dynamic aggregation in the frequency domain;

[0203] A feature fusion and adaptive weight adjustment module, configured to: process the features after Fourier transform by using multi-scale convolution, channel attention, and spatial attention mechanisms to achieve adaptive fusion of features of different modalities and different regions and adaptive weight adjustment;

[0204] A medical image representation module, configured to: process the fused features by using a densely connected decoder to obtain the finally fused medical image representation.

[0205] Embodiment III

[0206] The purpose of this embodiment is to provide a computer-readable storage medium.

[0207] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in a medical image fusion method based on a large kernel attention mechanism as described in Embodiment 1 of the present disclosure.

[0208] Embodiment IV

[0209] The purpose of this embodiment is to provide an electronic device.

[0210] An electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a medical image fusion method based on a large kernel attention mechanism as described in Embodiment 1 of the present disclosure.

[0211] The steps involved in the devices in the above Embodiments II, III, and IV correspond to those in Method Embodiment 1. For specific implementation manners, reference may be made to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0212] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. The present invention is not limited to any specific combination of hardware and software.

[0213] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts on the basis of the technical solutions of the present invention are still within the protection scope of the present invention.

Claims

1. A medical image fusion method based on large kernel attention mechanism, characterized in that: include: Acquire multimodal medical images, and preprocess the acquired multimodal medical images; The preprocessed multimodal medical image is input into the dual-branch encoder, and preliminary feature extraction is performed through the multi-scale convolutional layer; the preliminarily extracted features are group normalized, projected, and segmented using the large-core attention mechanism to obtain multiple features after segmentation. , , ; , , is the medical image feature segmented by a, and a is segmented by the projected feature, as shown below: ; ; The segmented features are respectively input into the large-core attention submodules of different scales and the corresponding spatial feature extraction convolutional layers to obtain the global and local features of the image, and the global and local features are spliced ​​and projected to the original dimension; wherein the global and local features of the image are obtained as follows: In the formula, , , It is the global feature map obtained after being processed by large-core attention submodules of different scales; , , It is the local feature map obtained after processing by convolutional layers of spatial feature extraction of different scales; LKA3, LKA5, and LKA7 are large-core attention submodules of different scales; X3, X5, and X7 are the corresponding 3x3, 5x5, and 7x7 spatial feature extraction convolutional layers; Spectrum dynamic aggregation is used to perform Fourier transform on the multi-scale features projected to the original dimension, and the amplitude and phase information of the multi-scale features are processed to achieve dynamic aggregation in the frequency domain; specifically: The spectral dynamic aggregation module performs Fourier transform on the multi-scale features projected to the original dimension, transforming them from the spatial domain Convert to frequency domain , respectively obtain the amplitude and phase information of the feature; ; ; in, represents the Fourier transform operation, and Respectively represent the representation of MRI and PET image features in the frequency domain; after Fourier transform, the features are decomposed into amplitude information and phase information ; ; ; in, and are the amplitude information extracted from the MRI and PET features after Fourier transformation, and is the phase information, is the Euler representation on the complex plane, i is the imaginary unit, is the phase angle; The convolutional layer and nonlinear activation function are used to enhance the amplitude information. The processing of amplitude information is expressed as: ; ; and is the amplitude information extracted from the Fourier transformed MRI and PET features, represents the convolutional layer used to enhance the amplitude information, is a nonlinear activation function; the processed amplitude information is passed through the second convolutional layer Make adjustments; The processed amplitude information and Reassemble back to frequency domain features; Convolution operation and attention mechanism are used to dynamically adjust the phase information; Extracting phase information from Fourier transformed MRI and PET features and : ; ; Input phase information and Through the convolution layer, the important features of edges and details contained in the phase information are enhanced; this process is expressed by the following formula: ; ; By calculating the weight of the phase information , the processed phase information is weighted according to the weight to enhance the focus on the key area; the calculation formula of the phase importance weight is: ; ; The phase information is weighted using the phase importance weight: ; ; In the formula, is the phase weight generated by the convolutional layer, using The activation function controls the weights In the range of; after the phase importance weighting process is completed, the obtained and It is the phase information that has enhanced details and textures; The weighted phase information is combined with the processed amplitude information to generate a new complex feature representation: ; ; The processed frequency domain features are transformed by inverse Fourier transform Return to the spatial domain and generate new spatial features; the formula for inverse Fourier transform is: ; ; Introduce a learnable scaling parameter to dynamically adjust the magnitude of the output features: ; ; in, It is a learnable parameter corresponding to each channel, and the optimal scaling ratio of each channel is gradually learned through training; The original input features are transformed into and With scaled features and Add: ; ; Get the final Fourier transformed features and ; The features after Fourier transformation are processed by using multi-scale convolution, channel attention and spatial attention mechanisms to achieve adaptive fusion of features of different modalities and regions and adaptive weight adjustment. Specifically, Perform multi-scale convolution processing on the final Fourier transformed features to obtain feature representations at each scale; For the obtained feature representations of each scale, an adaptive average pooling operation is performed to extract the global information of each channel: The pooled result is passed through two layers of convolution to calculate the channel attention weight and ,in ; ; Weight the frequency domain features: ; ; The spatial attention mechanism is used to spatially weight the weighted frequency domain features; the convolution operation is used to extract the spatial attention weights. and : ; ; Use spatial attention weights to weight frequency domain features: ; ; After channel and spatial weighting, use Convolution processes the fused frequency domain features to ensure the consistency of channel dimensions: ; Use residual connection to transform the original frequency domain and Add it to the weighted fusion feature to generate the final frequency domain fusion feature: ; in, Waiting for the final frequency domain fusion; The fused features are processed using a densely connected decoder to obtain a final fused image representation. Specifically, the densely connected decoder first expands the fused features to a high-dimensional space, and then optimizes the features through multiple densely connected blocks, and continuous convolution operations and feature splicing are used inside each densely connected block.

2. The medical image fusion method based on large-core attention mechanism according to claim 1, characterized in that: The multimodal medical images include MRI magnetic resonance imaging and PET positron emission tomography images.

3. The medical image fusion method based on large kernel attention mechanism as claimed in claim 1, characterized in that: The process of preprocessing the acquired multimodal medical image is: adjusting the input image to a standard size of 128×128, and ensuring that the pixel value range of the multimodal medical image is consistent through a normalization operation.

4. The medical image fusion method based on large kernel attention mechanism according to claim 1, characterized in that: The process of inputting the preprocessed multimodal medical image into the dual-branch encoder and performing preliminary feature extraction through the multi-scale convolutional layer is as follows: The preprocessed MRI and PET images are respectively input into their respective branch encoders, and feature extraction is performed through a multi-scale convolution layer; the multi-scale convolution uses convolution kernels of different sizes to capture feature information of different scales in the medical image; ; ; in, and Represents the multi-scale convolutional extractor for the MRI and PET branches; the output and It is the preliminary feature representation of different modalities; and They are the preprocessed MRI and PET images respectively.

5. The medical image fusion method based on large kernel attention mechanism as claimed in claim 1, characterized in that: The process of concatenating the global and local features and projecting them to the original dimension is: multiplying the large kernel convolution result and the spatial feature extraction result to enhance the representation capability of multi-scale features: In the formula , It is the feature representation after multi-scale feature fusion, which contains the combined information of global and local features; The concatenated multi-scale features are projected back to the original dimension through the projection convolution layer to ensure that the feature size is consistent with the input: ; ; In the formula, , It is the final feature representation obtained after the last projection layer, and its dimension is consistent with the input feature; It is the projection convolution layer.

6. The medical image fusion method based on large kernel attention mechanism according to claim 1, characterized in that: A medical image fusion method based on large kernel attention mechanism, also including loss function design.

7. A medical image fusion system based on a large core attention mechanism, using the medical image fusion method based on a large core attention mechanism as described in any one of claims 1 to 6, characterized in that: include: The multimodal medical image acquisition module is configured to: acquire the multimodal medical image and preprocess the acquired multimodal medical image; The large core attention mechanism processing module is configured to: input the preprocessed multimodal medical image into the dual-branch encoder, perform preliminary feature extraction through the multi-scale convolution layer; use the large core attention mechanism to perform group normalization, projection and segmentation processing on the preliminary extracted features to obtain multiple features after segmentation; The multi-scale large-kernel convolution and spatial feature extraction module is configured to: input the segmented multiple features into large-kernel attention sub-modules of different scales and corresponding spatial feature extraction convolution layers respectively to obtain global and local features of the image; and concatenate the global and local features and then project them to the original dimension; The spectrum dynamic aggregation module is configured to: perform Fourier transform on the multi-scale features projected to the original dimension using the spectrum dynamic aggregation module, and process the amplitude and phase information of the multi-scale features to achieve dynamic aggregation in the frequency domain; The feature fusion and adaptive weight adjustment module is configured to: process the features after Fourier transformation by using multi-scale convolution, channel attention and spatial attention mechanisms, and realize adaptive fusion of features of different modalities and different regions and adaptive weight adjustment; The medical image representation module is configured to: process the fused features using a densely connected decoder to obtain a final fused medical image representation.

8. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps in a medical image fusion method based on a large-core attention mechanism as described in any one of claims 1 to 6 are implemented.

9. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the medical image fusion method based on the large core attention mechanism as described in any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Small-parameter medical image segmentation system and method based on LKA-Unet model

    CN116385453A

  • Multi-modal medical image fusion based on expansion convolution and attention GCN

    CN117392494A

  • Medical image segmentation method, device and equipment and computer readable storage medium

    CN119131383A