Medical image fusion method based on wavelet domain and hierarchical state space fusion
Through the fusion method of wavelet domain and hierarchical state space, combined with wavelet transformation and convolutional neural network, the problem of high complexity of global feature extraction and computing in medical image fusion is solved, and efficient and accurate cross-modal information fusion is achieved.
Patent Information
- Application Number
- CN202510517365.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-18
AI Technical Summary
In the process of feature extraction and information integration, the existing medical image fusion methods have problems such as difficulty in effectively modeling global features, high computational complexity of deep learning methods, and poor balance between high-frequency details and low-frequency structures in the process of feature extraction and information integration.
Using a method based on wavelet domain and hierarchical state space fusion, a multi-modal image fusion network of a wavelet transform module and an encoder, a fusion module and a decoder, combined with a wavelet transform and a convolutional neural network, the fusion of global spatial features and fine-grained local features are achieved, and the computing efficiency is optimized.
It improves the efficiency and accuracy of cross-modal information fusion, reduces the computational complexity, is suitable for high-resolution medical image processing, and provides more efficient and accurate fusion solutions.
Smart Images

Figure CN120339775A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image processing, and particularly relates to a medical image fusion method based on the fusion of wavelet domain and hierarchical state space. Background Art
[0002] Medical image fusion, as a frontier technology in the fields of information fusion and digital image processing, aims to fully integrate image data from multiple medical imaging modalities, so as to simultaneously present key anatomical structures and functional information in a single image, thereby providing a more accurate basis for clinical diagnosis and treatment.
[0003] In the prior art, different imaging modalities have their own unique advantages and limitations, and they usually show obvious complementary characteristics in obtaining anatomical and functional information. For example, magnetic resonance imaging (MRI) can accurately present anatomical details due to its high resolution and excellent soft tissue contrast, but it is insufficient in reflecting functional characteristics such as cell metabolic activities and blood flow dynamics; while single photon emission computed tomography (SPECT) and positron emission tomography (PET) are good at revealing the metabolic and blood flow information of lesion areas, but their spatial resolution is relatively low, making it difficult to provide fine anatomical structure details. Therefore, by fusing the image data collected by different imaging modalities, not only can the deficiencies in information expression of a single imaging technology be made up for, but also high-precision anatomical information and rich functional characteristics can be retained simultaneously in a composite image. This fusion image technology is of great significance in clinical applications such as lesion localization and pathological analysis, providing strong support for clinical practice.
[0004] Currently, medical image fusion technologies can be divided into two categories: traditional methods and deep learning methods. Traditional fusion methods are usually divided into several categories: in traditional multi-modal image fusion methods, feature extraction and fusion are usually carried out in the spatial domain or the transform domain. The transform domain method decomposes the source image into different frequency sub-bands through mathematical transformation, and then reconstructs the image by inverse transformation after fusing the coefficients. The spatial domain method directly performs fusion at the pixel or local region level, selects key pixels or regions depending on saliency measurement, and completes the fusion process by combining linear or non-linear operations.
[0005] In recent years, medical image fusion methods based on deep learning have shown significant advantages. Compared with traditional handcrafted feature engineering, deep neural networks have powerful feature representation capabilities and adaptive learning mechanisms, which can effectively capture the complex spatial distributions and semantic associations in medical images, significantly improving the performance of cross-modal information fusion. Currently, deep learning methods for medical image fusion mainly use CNN and Transformer architectures for feature extraction and reconstruction. Although these methods have achieved good results in detail enhancement and global information extraction, they still face some challenges. Convolutional neural networks are restricted by their local receptive fields, which limits their ability to capture long-term dependencies, especially when dealing with medical images with complex structures and details. In addition, due to a large number of matrix calculations, the computational complexity of the attention mechanism is quadratically proportional to the input image size, resulting in a high computational cost. This is particularly problematic for high-resolution images, leading to a significant increase in computational load and a decrease in training and inference efficiency. Although recent fusion models attempt to utilize the advantages of convolutional networks and transformers through hybrid methods, the computational overhead remains a major challenge and requires further optimization to improve fusion efficiency and performance. Summary of the Invention
[0006] Aiming at the above deficiencies in the prior art, the medical image fusion method based on wavelet domain and hierarchical state space fusion provided by the present invention solves the deficiencies existing in the prior medical image fusion methods in the process of feature extraction and information integration, including problems such as traditional methods being difficult to effectively model global features, high computational complexity of deep learning methods, and poor balance between high-frequency details and low-frequency structures in the fusion results.
[0007] In order to achieve the above invention purpose, the technical solution adopted by the present invention is: a medical image fusion method based on wavelet domain and hierarchical state space fusion, comprising the following steps:
[0008] Perform color space conversion on the medical source image to obtain the medical image to be fused;
[0009] Input the medical image to be fused into the constructed multi-modal image fusion network, and reconstruct to obtain the fused medical image;
[0010] The multi-modal image fusion network includes an encoder, a fusion module, and a decoder;
[0011] The encoder includes a wavelet transform module, which decomposes and reconstructs the input image through wavelet transform and convolutional operations, and extracts image features combining shallow and multi-layer frequency domain information; the fusion module realizes the fusion of global spatial features and fine-grained local features through a state space model and a convolutional neural network, and obtains a fused feature map.
[0012] Further, the medical source image in the RGB color space is converted to the YUV color space, and the obtained Y-channel image and the single-channel medical source image are used as the medical images to be fused.
[0013] Further, the encoder includes first to fourth feature extraction modules connected in sequence. Each layer of the feature extraction module includes an E-Conv convolutional layer and a wavelet transform module connected in sequence;
[0014] The process of the feature extraction module processing the input image is expressed as:
[0015]
[0016] where represents the image feature output after the nth input image is processed by the feature extraction module corresponding to the jth layer scale. WT-Block(·) represents the feature extraction module, E-Conv represents the convolutional operation, MaxPool(·) represents the max pooling operation. The subscript j represents the scale level index, and the subscript n represents the input image index, n = 1, 2, j = 1, 2, 3, 4.
[0017] Further, the process of the wavelet transform module processing the input image is as follows:
[0018] Perform the first-layer discrete wavelet transform on the input image to obtain four corresponding sub-band components, including the low-frequency component, the high-frequency component in the horizontal direction, the high-frequency component in the vertical direction, and the high-frequency component in the diagonal direction;
[0019] Perform the second-layer discrete wavelet transform on the low-frequency component to obtain four corresponding sub-band components;
[0020] Perform depth convolution on the four sub-band components obtained from the first-layer discrete wavelet transform and the second-layer discrete wavelet transform respectively to obtain multi-scale image features at different frequencies;
[0021] Perform the first-layer inverse discrete wavelet transform on the multi-scale image features after the second-layer discrete wavelet transform and depth convolution, and merge the low-frequency component and the high-frequency component at each layer to obtain the reconstructed features combining high-frequency features and low-frequency features;
[0022] Fuse the reconstructed features with the multi-scale image features after the first-layer discrete wavelet transform and depth convolution, and perform the second-layer inverse discrete wavelet transform on the fused features to obtain the preliminary reconstructed image features;
[0023] Fuse the preliminary reconstructed image with the input image after the 3×3 convolution operation, and process it through batch normalization and the GELU activation function to obtain the image features combining the fused shallow features and the reconstructed features.
[0024] Further, the process of the wavelet transform module decomposing the input image is expressed as:
[0025]
[0026] The reconstruction process is expressed as:
[0027]
[0028] In the formula, represents the four sub-band components obtained by the i-th layer of wavelet decomposition, which are the low-frequency component, the high-frequency component in the horizontal direction, the high-frequency component in the vertical direction, and the high-frequency component in the diagonal direction in sequence. WT(·) represents the discrete wavelet transform, Conv(·) represents the convolution operation, and Z (1) represents the reconstructed feature obtained by the first-layer inverse discrete wavelet transform, and Z (2) represents the reconstructed feature obtained by the second-layer inverse discrete wavelet transform. IWT(·) represents the inverse discrete wavelet transform, represents the input image, BN(·) represents batch normalization, GELU(·) represents the GELU activation function, Output represents the image feature obtained by reconstruction, and the subscript i represents the decomposition level index, where i = 1, 2.
[0029] Further, the image features of different scales output by the first to fourth feature extraction modules respectively pass through a fusion module for feature fusion, and a fused feature map is output;
[0030] Each of the fusion modules includes a Split operation layer, a first branch, a second branch, a fusion layer, and a 1×1 convolution layer;
[0031] The output end of the Split operation layer is connected to the parallel first branch and second branch, and the output ends of the first branch and the second branch are sequentially connected to the fusion layer and the 1×1 convolution layer.
[0032] Further, the first branch includes a first LN layer, a third branch, a fourth branch, and a first linear layer;
[0033] The output end of the LN is connected to the parallel third branch and fourth branch. The third branch includes a second linear layer, a depthwise separable convolution layer, a 2D selective scanning layer, a second LN layer, and a channel fusion layer connected in sequence;
[0034] The fourth branch includes a third linear layer and a SiLU activation function connected in sequence. The output end of the SiLU activation function is connected to the input end of the channel fusion layer, and the output end of the channel fusion layer is connected to the input end of the first linear layer;
[0035] The second branch includes a pointwise convolution layer and a depthwise convolution layer connected in sequence.
[0036] Furthermore, the decoder starts from the fusion feature maps corresponding to the fourth to the first feature extraction modules, performs upsampling and D-Conv operations layer by layer, and after each D-Conv operation, splices the feature map corresponding to the current layer with the feature map of the previous layer, and then reconstructs a medical image that fuses multi-scale features.
[0037] Furthermore, the decoder outputs the fused medical image I f which is expressed as:
[0038]
[0039] where D-Conv represents a series of convolution operations, including convolution, batch normalization, and activation functions, ↑ represents the upsampling operation, represents the fusion feature map output by the fusion module corresponding to the j-th layer scale, and [] represents the splicing operation of feature maps, which is used to connect feature maps of multiple different scales.
[0040] The beneficial effects of the present invention are as follows:
[0041] (1) The multi-modal image fusion network constructed by the present invention uses a U-shaped encoder-decoder structure to extract multi-scale disease-related information from complex backgrounds. In the encoder, a multi-scale receptive field expansion mechanism is designed through a wavelet convolutional network, which expands the receptive field of the convolutional operation and improves the ability of the encoding network to extract coarse and fine-grained features. In the fusion module, the module combines the dynamic modeling ability of the state space model and the local feature extraction ability of the convolutional neural network. The former effectively captures long-range correlations through multi-directional scanning, facilitating the extraction of global spatial features of pathological structures, while the latter focuses on retaining fine-grained local details in medical images and efficiently captures global spatial features and local details with linear computational complexity.
[0042] (2) The method of the present invention combines wavelet transform and selective state space model in the medical image fusion task to achieve joint feature representation in the frequency domain and spatial domain, which can take into account both detail and global information, as well as significant information in different feature extraction stages, effectively improving the cross-modal information interaction ability, thereby constructing a clearer and more complete fused image and improving the fusion effect.
[0043] (3) The present invention optimizes the computational efficiency, uses state space modeling to replace the traditional attention structure, reduces the computational complexity on the premise of ensuring the fusion quality, making it more suitable for high-resolution medical image processing and clinical applications, and providing a more efficient and accurate solution for medical image fusion. Description of the Drawings
[0044] Figure 1Flowchart of the medical image fusion method based on wavelet domain and hierarchical state space fusion provided by the present invention.
[0045] Figure 2 Structural diagram of the multi-modal image fusion network provided by the present invention.
[0046] Figure 3 Schematic diagram of the structure of the wavelet transform module provided by the present invention.
[0047] Figure 4 Schematic diagram of the structure of the fusion module provided by the present invention. Detailed implementation manners
[0048] The following describes the detailed implementation manners of the present invention to facilitate those skilled in the art to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed implementation manners. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
[0049] Please refer to Figures 1 to 3 , the present invention provides a medical image fusion method based on wavelet domain and hierarchical state space fusion, which realizes medical image fusion based on the constructed multi-modal image fusion network, not only improves the efficiency and accuracy of cross-modal information fusion, but also optimizes the consumption of computing resources, providing a more efficient and accurate fusion method for medical image analysis.
[0050] As Figure 1 shown, the method of the present invention includes the following steps:
[0051] Perform color space conversion on the medical source image to obtain the medical image to be fused;
[0052] Input the medical image to be fused into the constructed multi-modal image fusion network, and reconstruct the fused medical image;
[0053] As Figure 2 shown, the multi-modal image fusion network in the present invention includes an encoder, a fusion module, and a decoder;
[0054] Among them, the encoder includes a wavelet transform module, and the wavelet transform module decomposes and reconstructs the input image through wavelet transform and convolution operations, and extracts image features combining shallow and multi-layer frequency domain information; the fusion module realizes the fusion of global spatial features and fine-grained local features through a state space model and a convolutional neural network, and obtains a fused feature map.
[0055] In the implementation process of the method of the present invention, the discrete wavelet transform is combined with the state space model, and the multi-frequency information of the image is captured by the wavelet-guided feature extraction module; in the fusion process, a multi-scale fusion module is used to extract local details and global spatial features, so that the fused medical image is superior to the prior art in terms of structural integrity, retention of lesion area details, and enhancement of global contrast.
[0056] The following further describes the implementation of the present invention in conjunction with Embodiments 1 to 2.
[0057] Embodiment 1:
[0058] This embodiment focuses on preprocessing the medical source images to be fused to obtain the medical images input into the multi-modal image fusion network.
[0059] In this embodiment, the medical source image in the RGB color space is converted to the YUV color space, and the obtained Y-channel image and the single-channel medical source image are used as the medical images to be fused.
[0060] In a specific example of this embodiment, medical source images such as PET and SPECT images are color images, while the MPI image is a single-channel gray image. To solve the biomedical fusion image of the three-channel RGB function and the single-channel gray structure, the PET and SPECT images are therefore converted to the YUV color space to obtain Y, U, and V three-channel images, and then the Y channels of the PET and SPCT images and the MRI image are input into the subsequent multi-modal image fusion network.
[0061] Among them, the conversion formula is as follows:
[0062] Y = 0.299R + 0.587G + 0.114B
[0063] U = -0.14713R - 0.28886G + 0.436B
[0064] V = 0.615R - 0.51499G - 0.10001B
[0065] In the formula, the Y channel represents the luminance information of the image, and the U and V channels represent the color (chromaticity) information of the image.
[0066] Embodiment 2:
[0067] This embodiment expands on the structure of the multi-modal image fusion network and the process of realizing the fusion of the input images.
[0068] In this embodiment, in Figure 2Among them, the encoder includes first to fourth feature extraction modules connected in sequence, and each layer of the feature extraction module includes an E-Conv convolutional layer and a wavelet transform module WT-Block connected in sequence.
[0069] Among them, the process of the feature extraction module processing the input image is expressed as:
[0070]
[0071] In the formula, represents the image feature output after the processing of the nth input image by the feature extraction module corresponding to the jth layer scale. WT-Block(·) represents the feature extraction module, E-Conv represents the convolutional operation, MaxPool(·) represents the max pooling operation, I0 represents the input image of the first feature extraction module, the subscript j represents the scale level index, the subscript n represents the input image index, n = 1, 2 (two source images for fusion), and j = 1, 2, 3, 4.
[0072] As Figure 3 shown is the specific structure of the above-mentioned wavelet transform module WT-Block, and its process of processing the input image is:
[0073] Perform the first-layer discrete wavelet transform (DWT) on the input image to obtain four corresponding sub-band components (LL), including the low-frequency component (LL), the high-frequency component in the horizontal direction (LH), the high-frequency component in the vertical direction (HL), and the high-frequency component in the diagonal direction (HH);
[0074] Perform the second-layer discrete wavelet transform (DWT) on the low-frequency component to obtain four corresponding sub-band components;
[0075] Perform depth convolution (Depth Conv) on the four sub-band components obtained from the first-layer discrete wavelet transform and the second-layer discrete wavelet transform respectively to obtain multi-scale image features at different frequencies;
[0076] Perform the first-layer inverse discrete wavelet transform (IDWT) on the multi-scale image features after the second-layer discrete wavelet transform and depth convolution, and combine the low-frequency component and the high-frequency component at each layer to obtain a reconstructed feature combining high-frequency features and low-frequency features;
[0077] Fuse the reconstructed feature with the multi-scale image features after the first-layer discrete wavelet transform and depth convolution, and perform the second-layer inverse discrete wavelet transform (IDWT) on the fused feature to obtain the preliminary reconstructed image feature;
[0078] The preliminary reconstructed image is fused with the input image that has undergone a 3×3 convolution operation, and is processed by batch normalization (BN) and the GELU activation function to obtain image features that fuse shallow features and reconstructed features.
[0079] In this embodiment, the wavelet transform performs two-layer discrete wavelet transform on the input image. On the basis of the two-layer discrete wavelet transform, in the process of two-layer inverse discrete wavelet transform, Z includes the reconstructed features obtained from the first and second layers of inverse discrete wavelet transform, and the Z of the second layer (2) is obtained by performing inverse discrete wavelet transform on the low- and high-frequency components of the second layer. Then, the Z of the second layer (2) is added to the low-frequency component in the first-layer discrete wavelet transform to obtain a new first-layer which is then combined with the high-frequency component of the first layer to perform inverse discrete wavelet transform to obtain the reconstructed feature Z of the first layer (1) , and finally, the reconstructed feature Z of the first layer (1) is added to the input image after convolution to obtain the image features output by the wavelet transform module.
[0080] In the above process, the wavelet transform module applies the same conversion process to the input medical source image PET / SPECT. This module combines wavelet transform and convolution operations to decompose and reconstruct the input image, improving the computational efficiency while expanding the original receptive field and feature integrity of the convolution.
[0081] Specifically, the process of the wavelet transform module decomposing the input image is expressed as:
[0082]
[0083] The reconstruction process is expressed as:
[0084]
[0085] In the formula, represents the four sub-band components obtained from the i-th layer of wavelet decomposition, which are the low-frequency component, the high-frequency component in the horizontal direction, the high-frequency component in the vertical direction, and the high-frequency component in the diagonal direction in sequence. WT(·) represents the discrete wavelet transform, Conv(·) represents the convolution operation, and Z (1) represents the reconstructed feature obtained from the first layer of inverse discrete wavelet transform, Z (2) represents the reconstructed feature obtained from the second layer of inverse discrete wavelet transform, IWT(·) represents the inverse discrete wavelet transform, represents the input image, BN(·) represents batch normalization, GELU(·) represents the GELU activation function, Output represents the reconstructed image features, and the subscript i represents the decomposition level index, i = 1, 2.
[0086] After multi-level wavelet decomposition, progressive fusion is performed at different levels. The adjusted high-frequency features are combined with the low-frequency features and participate in subsequent feature reconstruction. In the inverse discrete wavelet transform, the high-frequency components and low-frequency components are merged at each layer to restore the frequency of the original input image.
[0087] The above decomposition and reconstruction process combines shallow and multi-level frequency domain information, enhancing the model's performance in detail retention and feature restoration.
[0088] In this embodiment, as Figure 2 shown, the different-scale image features output by the first to fourth feature extraction modules respectively pass through a fusion module for feature fusion, and a fused feature map is output.
[0089] As Figure 4 shown is the specific structure of each fusion module. Each fusion module includes a Split operation layer, a first branch, a second branch, a fusion layer, and a 1×1 convolutional layer;
[0090] The output end of the Split operation layer is connected to the parallel first branch and second branch, and the output ends of the first branch and second branch are sequentially connected to the fusion layer and the 1×1 convolutional layer.
[0091] Among them, the first branch includes a first LN layer, a third branch, a fourth branch, and a first linear layer;
[0092] The output end of the LN is connected to the parallel third branch and fourth branch. The third branch includes a second linear layer, a depthwise separable convolutional layer, a 2D selective scan layer, a second LN layer, and a channel fusion layer connected in sequence;
[0093] The fourth branch includes a third linear layer and a SiLU activation function connected in sequence. The output end of the SiLU activation function is connected to the input end of the channel fusion layer, and the output end of the channel fusion layer is connected to the input end of the first linear layer;
[0094] The second branch includes a pointwise convolution and a depthwise convolution layer connected in sequence.
[0095] In this embodiment, the above-mentioned fusion module adopts a dual-branch structure, and the output of the encoder is fed into the first branch and the second branch respectively. In this embodiment, in order to reduce the loss of local features, a first branch consisting of a deep convolution layer and a point-by-point convolution layer is designed to capture local information in medical images by taking advantage of the continuous convolution layer. The other sub-output is input to the second branch, where the input is first normalized using layer normalization to ensure the balance and stability between different inputs. It is then split into two branches. In the first sub-branch (third branch) of the second branch, the input undergoes a linear layer, a depth-separable convolution, and an activation function, after which it is sent to the 2D selective scanning module for further feature extraction, and then layer normalization is performed to normalize the features. In the second sub-branch (fourth branch) in the second branch, the input passes through a linear layer and an activation function, and is element-wise multiplied with the output of the third branch to merge the two branches. Finally, the outputs of the two branches are fused along the channel dimension of the feature map, and the feature map is fused and output using a convolution layer with a convolution kernel size of 1.
[0096] In this embodiment, Figure 1 In the decoder, starting from the fused feature maps corresponding to the fourth to first feature extraction modules, upsampling and D-Conv operations are performed layer by layer. After each D-Conv operation, the feature map corresponding to the current layer is concatenated with the feature map of the previous layer to reconstruct a medical image that fuses multi-scale features.
[0097] Specifically, the decoder outputs the fused medical image I f It is expressed as:
[0098]
[0099] In the formula, D-Conv represents a series of convolution operations, including convolution, batch normalization, and activation function, ↑ represents upsampling operation, It represents the fused feature map output by the fusion module corresponding to the j-th scale. [] represents the concatenation operation of the feature map, which is used to connect feature maps of multiple different scales.
[0100] The method of the present invention designs a CNN-SSM collaborative medical image fusion method, aiming to address the limitations of existing methods. The proposed multimodal image fusion network adopts a wavelet convolution receptive field expansion mechanism to enhance the ability to extract complementary information at different scales; subsequently, the fusion module combines the dynamic modeling capability of the state-space model with the local feature extraction advantage of the convolutional neural network. Image fusion improves computational efficiency while taking into account both details and global features.
[0101] In the present invention, specific embodiments are used to illustrate the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.
[0102] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on these technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A medical image fusion method based on the fusion of wavelet domain and hierarchical state space, characterized in that Including the following steps: Perform color space conversion on the medical source image to obtain the medical image to be fused; Input the medical image to be fused into the constructed multi-modal image fusion network to reconstruct the fused medical image; The multi-modal image fusion network includes an encoder, a fusion module, and a decoder; The encoder includes a wavelet transform module. The wavelet transform module decomposes and reconstructs the input image through wavelet transform and convolution operations, and extracts image features that combine shallow and multi-layer frequency domain information. The fusion module realizes the fusion of global spatial features and fine-grained local features through a state space model and a convolutional neural network to obtain a fused feature map.
2. The method according to claim 1, wherein Convert the medical source image in the RGB color space to the YUV color space, and use the obtained Y-channel image and the single-channel medical source image as the medical image to be fused.
3. The method according to claim 1, wherein The encoder includes first to fourth feature extraction modules connected in sequence. Each layer of the feature extraction module includes an E-Conv convolutional layer and a wavelet transform module connected in sequence; The process of the feature extraction module processing the input image is expressed as: Wherein, represents the image feature output after the processing of the feature extraction module corresponding to the nth input image at the jth layer scale, WT-Block(·) represents the feature extraction module, E-Conv represents the convolution operation, MaxPool(·) represents the max pooling operation, the subscript j represents the scale level index, the subscript n represents the input image index, n = 1, 2, j = 1, 2, 3, 4.
4. The method according to claim 3, characterized in that The process of the wavelet transform module processing the input image is: Perform the first-layer discrete wavelet transform on the input image to obtain four corresponding sub-band components, including a low-frequency component, a high-frequency component in the horizontal direction, a high-frequency component in the vertical direction, and a high-frequency component in the diagonal direction; Perform the second-layer discrete wavelet transform on the low-frequency component to obtain four corresponding sub-band components; Perform depth convolution on the four sub-band components obtained by the first-layer discrete wavelet transform and the second-layer discrete wavelet transform respectively to obtain multi-scale image features at different frequencies; Perform the first-layer inverse discrete wavelet transform on the multi-scale image features after the second-layer discrete wavelet transform and depth convolution, and combine the low-frequency component and the high-frequency component at each layer to obtain a reconstructed feature that combines high-frequency features and low-frequency features; Fuse the reconstructed feature with the multi-scale image features after the first-layer discrete wavelet transform and depth convolution, and perform the second-layer inverse discrete wavelet transform on the fused features to obtain preliminary reconstructed image features; Fuse the preliminary reconstructed image with the input image after 3×3 convolution operation, and process it through batch normalization and the GELU activation function to obtain image features that fuse shallow features and reconstructed features.
5. The method according to claim 1, wherein The process of the wavelet transform module decomposing the input image is expressed as: The reconstruction process is expressed as: Wherein, represents the four sub-band components obtained by the i-th layer of wavelet decomposition, which are the low-frequency component, the high-frequency component in the horizontal direction, the high-frequency component in the vertical direction, and the high-frequency component in the diagonal direction in sequence. WT(·) represents the discrete wavelet transform, Conv(·) represents the convolution operation, and Z (1) represents the reconstructed feature obtained by the first layer of inverse discrete wavelet transform, and Z (2) represents the reconstructed feature obtained by the second layer of inverse discrete wavelet transform, and IWT(·) represents the inverse discrete wavelet transform. represents the input image, BN(·) represents batch normalization, GELU(·) represents the GELU activation function, Output represents the image feature obtained by reconstruction, and the subscript i represents the decomposition level index, where i = 1, 2.
6. The method according to claim 3, wherein The different-scale image features output by the first to fourth feature extraction modules are respectively subjected to feature fusion through a fusion module to output a fused feature map; Each fusion module includes a Split operation layer, a first branch, a second branch, a fusion layer, and a 1×1 convolutional layer; The output end of the Split operation layer is connected to the parallel first branch and second branch, and the output ends of the first branch and the second branch are sequentially connected to the fusion layer and the 1×1 convolutional layer.
7. The method according to claim 6, wherein The first branch includes a first LN layer, a third branch, a fourth branch, and a first linear layer; The output end of the LN is connected to the third branch and the fourth branch arranged in parallel. The third branch includes a second linear layer, a depthwise separable convolutional layer, a 2D selective scanning layer, a second LN layer, and a channel fusion layer connected in sequence; The fourth branch includes a third linear layer and a SiLU activation function connected in sequence. The output end of the SiLU activation function is connected to the input end of the channel fusion layer, and the output end of the channel fusion layer is connected to the input end of the first linear layer; The second branch includes a depthwise convolution and a pointwise convolution layer connected in sequence.
8. The method according to claim 6, wherein Starting from the fused feature maps corresponding to the fourth to first feature extraction modules, the decoder performs upsampling and D-Conv operations layer by layer. After each D-Conv operation, the feature map corresponding to the current layer is concatenated with the feature map of the previous layer, and then a medical image with multi-scale features fused is reconstructed.
9. The method according to claim 8, wherein The decoder outputs the fused medical image I f which is expressed as: Wherein, D-Conv represents a series of convolution operations, including convolution, batch normalization, and activation functions, ↑ represents an upsampling operation, represents the fused feature map output by the fusion module corresponding to the j-th layer scale, and [] represents the splicing operation of the feature map, which is used to connect feature maps of multiple different scales.