Medical image segmentation method based on spatial domain and frequency domain feature attention

Through the medical image segmentation method based on spatial and frequency domain feature attention, the problem of insufficient model generalization ability in multimodal medical image segmentation is solved, and a higher accuracy and robust image segmentation effect is achieved, especially in multi-scale and detailed segmentation of lesion areas.

CN120298432APending Publication Date: 2025-07-11CHANGCHUN UNIV OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510399418.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing medical image segmentation model has limited generalization ability when processing multimodal data sets, making it difficult to effectively extract multi-scale feature information and perform accurate edge texture segmentation. Especially in different mode data, the lesion size difference is significant, resulting in unsatisfactory segmentation effect.

Method used

The medical image segmentation method based on spatial domain and frequency domain feature attention is adopted, and a model is created through encoder, decoder, jump connection block and weighted loss function. A multi-scale hybrid convolution module, a convolution multi-head self-attention module, a frequency enhancement channel attention module and a dynamic grouping convolution spatial attention module are used to train and test with medical image data sets of different modes to improve the multi-scale understanding and noise suppression ability of the model to lesion areas.

Benefits of technology

The segmentation accuracy and robustness of the model for multimodal medical images is improved, the multi-scale and diverse understanding of the lesion area is enhanced, the segmentation ability of boundary and detailed areas is improved, noise interference is reduced, and the reliability of segmentation results is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298432A_ABST
    Figure CN120298432A_ABST
Patent Text Reader

Abstract

The invention provides a medical image segmentation method based on spatial domain and frequency domain feature attention, and the method comprises the steps: 1, selecting public data sets of different modalities, dividing the data sets into a training set, a test set and a verification set, and carrying out the preprocessing of the data set images; secondly, creating a medical image segmentation model based on an encoder, a decoder, a jump connection block and a weighted loss function, wherein the decoder is composed of a multi-scale hybrid convolution module, a convolution multi-head self-attention module, a frequency enhancement channel attention module and a dynamic grouping space attention module; thirdly, training, verifying and testing a medical image segmentation model through the data set; and 4, deploying the medical image segmentation model which is tested and has a good effect on the server, and carrying out a medical image segmentation task by using the medical image segmentation model. The method has the advantages that the robustness and generalization of the medical image segmentation model and the segmentation precision on different modal data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image segmentation. Specifically, it designs a medical image segmentation method based on spatial domain and frequency domain feature attention, and achieves better segmentation effects on multiple medical image datasets of different modalities. Background Art

[0002] With the rapid development of deep learning and artificial intelligence technologies, the application of computer vision in the field of medical image processing has become increasingly widespread. Especially in computer-aided diagnosis and surgical navigation, image segmentation technology plays a core role. Medical imaging technologies (such as MRI, CT, ultrasound, and PET, etc.) can provide detailed information about the internal structure of the human body, helping doctors analyze disease conditions more accurately, thereby improving the accuracy of diagnosis.

[0003] The core task of medical image segmentation is to accurately identify and separate different organs, tissues, and lesion areas, enabling doctors to clearly observe the morphology and scope of the lesions, and providing a scientific basis for clinical decision-making. This technology not only improves the reliability of diagnosis but also optimizes treatment plans, ensuring the efficiency and consistency of the medical process. In practical applications, efficient and accurate image segmentation can help doctors quickly lock in the lesions, evaluate the degree of lesions, and provide strong support for personalized treatment. There are significant differences between medical images and natural RGB images. Medical images are generally affected by noise, artifacts, etc., which significantly increase the difficulty of feature extraction. In practical applications, accurately annotating lesion tissues on medical images often relies on a large number of manual operations by experienced doctors. While consuming a lot of time and energy, there may also be problems such as annotation errors, further affecting the accuracy and reliability of the segmentation results.

[0004] With the progress of computer technology, deep learning has begun to be applied to the field of medical image processing, especially showing good results in medical image segmentation tasks. In recent years, convolutional neural networks (CNNs) and Transformer models have achieved remarkable results in the field of medical image segmentation due to their excellent feature extraction capabilities. These models can accurately capture the key features in the images, providing strong technical support for disease diagnosis, treatment plan formulation, and prognosis assessment.

[0005] However, in the field of medical image segmentation, convolutional neural network models integrated with Transformer still face many challenges. First, the generalization ability of the model is limited. Most existing studies focus on the segmentation of single-modal datasets, and the effect is not ideal when segmenting multiple-modal datasets. Second, the lesion sizes in different-modal data vary significantly, and the frequency features are different. The model cannot comprehensively and effectively extract multi-scale feature information and make accurate segmentations of edge textures. Summary of the Invention

[0006] The purpose of the present invention is to overcome the deficiencies of the prior art and propose a medical image segmentation method based on spatial domain and frequency domain feature attention for accurately segmenting various different modality medical images.

[0007] The technical solution adopted by the present invention is as follows:

[0008] A medical image segmentation method based on spatial domain and frequency domain feature attention, and the specific implementation steps are as follows:

[0009] Step 1, select publicly available datasets of different modalities, divide the datasets into a training set, a test set, and a validation set, and perform preprocessing operations on the dataset images;

[0010] Step 2, create a medical image segmentation model based on an encoder, a decoder, a skip connection block, and a weighted loss function, where the decoder is composed of a multi-scale hybrid convolution module (MSHCB), a convolutional multi-head self-attention module (CMSA), a frequency enhancement channel attention module (FECA), and a dynamic grouped convolution spatial attention module (DGCSA);

[0011] Step 3, train the medical image segmentation model through the training set, verify the segmentation effect of the medical image segmentation model through the validation set, and test the verified medical image segmentation model through the test set;

[0012] Step 4, deploy the medical image segmentation model with good test results on a server, and use this medical image segmentation model to perform medical image segmentation tasks;

[0013] Further, the specific content of Step 1 is as follows:

[0014] The selected medical image segmentation datasets are respectively the cardiac dataset (ACDC), the multi-organ dataset (Synapse), the skin lesion datasets (ISIC2017 and ISIC2018), and the polyp segmentation datasets (CVC-ClinicDB, Kvasir, and ColonDB). During preprocessing, all dataset images are adjusted to images of size 256×256 and standardized operations such as image flipping and noise reduction are performed.

[0015] Further, in Step 2, the encoder is composed of a multi-axis vision Transformer (MaxViT) backbone network divided into four layers.

[0016] The encoder is used to perform downsampling operations on the dataset images layer by layer, gradually extract medical image features, and use block-based local attention and dilated global attention, enabling the model to perform global-local spatial interactions at any input resolution.

[0017] Furthermore, in step 2, a multi-scale hybrid convolutional module (MSHCB) is provided in the decoder, and the module processes the original features input from the skip connection block. This module combines depth convolution, dilated convolution, and deformable convolution, achieving flexible receptive field adjustment while significantly reducing the computational cost. The module extracts image features from multiple angles and levels, enhancing the model's adaptability to complex medical image scenarios;

[0018] Furthermore, in step 2, a convolutional multi-head self-attention module (CMSA) is provided in the decoder, and the module processes the original features input from the skip connection block. This module uses convolutional operations to replace the fully connected layer to generate queries (Q), keys (K), and values (V), and pre-extracts local pattern and structural information through the local receptive field characteristics of convolution, enabling the subsequent self-attention calculation to effectively fuse local and global information.

[0019] Furthermore, in step 2, a frequency-enhanced channel attention module (FECA) is provided in the decoder, and the module operates on the feature map that has been processed and fused by the multi-scale hybrid convolutional module (MSHCB) and the convolutional multi-head self-attention module (CMSA). This module processes the feature map at three different scales, comprehensively captures the frequency features using the discrete cosine transform (DCT) and discrete wavelet transform (DWT), calibrates the feature maps at different scales using channel attention, and finally aggregates the feature maps at the three scales in a cascaded addition manner. The FECA module comprehensively captures the key information of the image at low-frequency and high-frequency levels, improving the performance of the model.

[0020] Furthermore, in step 2, a dynamic group convolutional spatial attention module (DGCSA) is provided in the decoder, and the module operates on the feature map that has been processed by the frequency-enhanced channel attention module (FECA). This module divides the feature map into four parts in the H and W dimensions respectively for processing, and combines four dynamic convolutions using different-sized convolutional kernels to extract features from them respectively, capturing spatial information in different directions to enhance the understanding of the spatial structure and the model's adaptability to different types of features.

[0021] Furthermore, in step 2, the skip connection block first normalizes the features from the next decoder layer, then concatenates the normalized features with the skip connection from the encoder, and finally performs activation and normalization operations. The skip connection block inputs the processed features into the decoder block for decoding.

[0022] Furthermore, in step 2, the weighted loss function is obtained by adding the boundary difference union loss (BDoU), weighted binary cross-entropy loss (CE), and dice loss (Dice) multiplied by their respective weights.

[0023] Further, step 3 is specifically as follows:

[0024] The medical image segmentation model is trained using the training set. During the training process, the loss function, optimizer function, and learnable hyperparameters used by the medical image segmentation model are continuously optimized until the segmentation performance of the model reaches the best.

[0025] The trained medical image segmentation model is verified using the validation set. If the verification effect is not good, training continues. If the verification effect is good, the model is tested using the test set.

[0026] When using the test set to verify the true segmentation effect of the model, if the segmentation effect is good, training ends. If the segmentation effect is not good, training continues.

[0027] Further, step 4 is specifically as follows:

[0028] The medical image segmentation model that passes the test is deployed to the server, and the call interface and call permissions are set. The user inputs the preprocessed data to be segmented into the medical image segmentation model, and the medical image segmentation model outputs the segmentation result with detailed data annotations to complete the segmentation task.

[0029] The method proposed by the present invention mainly has the following advantages.

[0030] Select publicly available datasets of different modalities, divide the datasets into training sets, test sets, and validation sets, and perform preprocessing operations on the dataset images; then create a medical image segmentation model based on an encoder, a decoder, a skip connection block, and a weighted loss function. Train the medical image segmentation model using the training set, verify the segmentation effect of the medical image segmentation model using the validation set, and test the verified medical image segmentation model using the test set; deploy the medical image segmentation model that has passed the test and has good results on the server, and use this medical image segmentation model to perform medical image segmentation tasks; the multi-scale hybrid convolution module (MSHCB) and the convolutional multi-head self-attention module (CMSA) in the decoder capture the long-range dependencies and multi-scale local features of the feature map, strengthening the model's understanding of the multi-scale and diversity of the lesion area. The frequency enhancement channel attention module (FECA) in the decoder extracts the global and local frequency feature information of the feature map on this basis and dynamically weights each frequency component, enhancing the weight of important features and improving the model's noise suppression ability and robustness. The dynamic grouped convolution spatial attention module (DGCSA) in the decoder focuses on modeling the pixel-level spatial relationship in the feature map, enhancing the model's segmentation ability for boundary and detail regions. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is the flowchart of the medical image segmentation method based on spatial domain and frequency domain feature attention of the present invention.

[0032] Figure 2 This is the overall framework diagram of the medical image segmentation model based on spatial domain and frequency domain feature attention.

[0033] Figure 3 This is the structural diagram of the multi-scale hybrid convolution module (MSHCB) in the decoder.

[0034] Figure 4 This is the structural diagram of the frequency enhancement channel attention module (FECA) in the decoder.

[0035] Figure 5 This is the structural diagram of the internal module (FEM) of the frequency enhancement channel attention module (FECA) in the decoder.

[0036] Figure 6 This is the structural diagram of the dynamic grouped convolution spatial attention module (DGCSA) in the decoder. Detailed implementation manners

[0037] The following further describes the present invention with reference to the accompanying drawings, so that those skilled in the art can better understand the present invention. It should be noted that without departing from the core idea of the present invention, those skilled in the art can make some improvements to the present invention, and these all belong to the protection scope of the present invention.

[0038] Please refer to Figures 1 to 6 As shown, the detailed implementation manners of a medical image segmentation method based on spatial domain and frequency domain feature attention in the present invention include the following steps:

[0039] Step 1, select publicly available datasets of different modalities, divide the datasets into a training set, a test set, and a validation set, and perform preprocessing operations on the dataset images;

[0040] Step 2, create a medical image segmentation model based on an encoder, a decoder, a skip connection block, and a weighted loss function, where the decoder is composed of a multi-scale hybrid convolution module (MSHCB), a convolutional multi-head self-attention module (CMSA), a frequency enhancement channel attention module (FECA), and a dynamic grouped spatial attention module (DGCSA);

[0041] Step 3, train the medical image segmentation model through the training set, verify the segmentation effect of the medical image segmentation model through the validation set, and test the verified medical image segmentation model through the test set;

[0042] Step 4: Deploy the medically trained and effective medical image segmentation model on the server and use this medical image segmentation model to perform medical image segmentation tasks;

[0043] Specifically, step 1 is as follows:

[0044] The selected medical image segmentation datasets are the cardiac dataset (ACDC), the multi-organ dataset (Synapse), the skin lesion datasets (ISIC2017 and ISIC2018), and the polyp segmentation datasets (CVC-ClinicDB, Kvasir, and ColonDB). During preprocessing, all dataset images are resized to 256×256 and standardized operations such as image flipping and noise reduction are performed. Specifically, the dataset is divided into a training set, a validation set, and a test set at a ratio of 8:1:1.

[0045] In step 2, the encoder consists of a multi-axis vision Transformer (MaxViT) backbone network divided into four layers. The model framework is as Figure 2 shown.

[0046] The encoder is used to perform downsampling operations on the dataset images layer by layer, gradually extract medical image features, and use block-wise local attention and dilated global attention, enabling the model to perform global-local spatial interactions at any input resolution. When the encoder performs downsampling, the resolution of the feature map is reduced layer by layer to 1 / 2, 1 / 4, 1 / 8, and 1 / 16 of the initial input image, while the number of feature channels increases proportionally layer by layer.

[0047] In step 2, the decoder is equipped with a multi-scale hybrid convolution block (MSHCB), which processes the original features input from the skip connection block. The block structure is as Figure 3 shown. This block combines depth convolution, dilated convolution, and deformable convolution, achieving flexible receptive field adjustment while significantly reducing the computational cost. The block extracts image features from multiple angles and levels, enhancing the model's adaptability to complex medical image scenarios. The specific details are as follows:

[0048] In the initial stage, the block uses a standard convolution with a kernel size of 3×3 to preliminarily extract features, and then uses depthwise separable dilated convolutions with dilation rates of 1, 2, 3, and 5 and a kernel size of 3×3 to perform multi-scale extraction of features in parallel, obtaining the feature map X i , where i ∈ {1, 2, 3, 5} represents the dilation rate. Subsequently, the block performs batch normalization and SiLU activation operations on the features at each scale to make the feature data more stable. The formula is as follows:

[0049] X i= S(B(DWC i (Conv3(X)))),

[0050] where Conv3 represents a standard convolution with a convolutional kernel size of 3×3, DWC i represents depthwise separable dilated convolution, B represents batch normalization, and S represents the SiLU activation function.

[0051] After performing the addition and concatenation operations on the feature maps X2, X3, and X5, we obtain The module respectively uses deformable depthwise convolutions with convolutional kernel sizes of 3×3 and 5×5 to perform convolution operations on the features, obtaining more refined features, that is Finally, the module obtains the final feature Y through element-wise addition and residual connection, further enhancing the accuracy and robustness of the model. Its formula is expressed as follows:

[0052]

[0053] where Concat represents the feature concatenation operation. Conv1 represents a standard convolution with a convolutional kernel size of 1×1, and DDWC g represents deformable depthwise separable convolution, where g = {3, 5} represents the size of the convolutional kernel.

[0054] In step 2, a convolutional multi-head self-attention module (CMSA) is provided in the decoder, and the module processes the original features input from the skip connection block. This module uses convolution operations to replace the fully connected layer to generate queries (Q), keys (K), and values (V), and pre-extracts local pattern and structural information through the local receptive field characteristics of convolution, enabling the subsequent self-attention calculation to effectively fuse local and global information. The specific content is as follows:

[0055] The input feature map is where C is the number of channels, and H and W are the height and width of the feature map respectively. To generate the queries (Query), keys (Key), and values (Value) required by the self-attention mechanism, the module first converts the input feature map X into these three representations through convolution operations. The formula is expressed as follows:

[0056]

[0057] where N represents layer normalization, G represents the GELU activation function, and C x represents the convolutional layer for generating queries (Q), keys (K), and values (V), where x = {q, k, v}. The formula for the self-attention mechanism is expressed as follows:

[0058]

[0059] where Q is the Query matrix, K is the Key matrix, V is the Value matrix, and d k is the dimension of the key, which is usually equal to the number of channels C and is used for scaling. The final formula for multi-head self-attention is expressed as follows:

[0060] MHSA(Q, K, V) = Concat(head1, head2, …, head h )W O ,

[0061] where MHSA represents multi-head self-attention, h is the number of heads, and W O is the final linear transformation matrix. The number of heads in the bottleneck layer and the other four decoder blocks are 10, 8, 6, 4, and 2 respectively.

[0062] In step 2, a Frequency Enhancement Channel Attention module (FECA) is provided in the decoder. The module operates on the feature map that has been processed and fused by the Multi-Scale Hybrid Convolution Block (MSHCB) and the Convolutional Multi-Head Self-Attention module (CMSA). The module structure is as Figure 4 shown. This module processes the feature map at three different scales, uses the Discrete Cosine Transform (DCT) and the Discrete Wavelet Transform (DWT) to comprehensively capture the frequency features, uses channel attention to calibrate the feature maps at different scales, and finally aggregates the feature maps at the three scales by concatenation and addition. The FECA module comprehensively captures the key information of the image at low-frequency and high-frequency levels, improving the performance of the model. The specific content is as follows:

[0063] The FECA module divides the features into three different scales for processing, namely the original size, downsampled by a factor of 2, and downsampled by a factor of 4, named X1, X2, and X4. In each module branch, the module first performs preliminary sampling on the feature map using a depthwise separable convolution with a kernel size of 3×3, and then successively passes through batch normalization and ReLU function activation operations. Subsequently, the features enter the Frequency Enhancement Module (FEM) (as Figure 5 shown) for processing. The frequency features processed by DCT are denoted as DCT k . The DCT k formula is expressed as follows:

[0064]

[0065] where (m k , n k ) represents the index pair of the k-th frequency component, k ∈ [0, k - 1], X i represents the feature map processed by DCT, and i ∈ {1, 2, 4}. H and W respectively represent the height and width of the feature map Xi The height and width, represent the basis image of the DCT. The formula is expressed as follows:

[0066]

[0067] For a 2D image, wavelet transform decomposes it into low-frequency and high-frequency components. The low-frequency component represents the smooth part of the image, while the high-frequency component reflects the details and edges of the image. Specifically, the original image is decomposed into four sub-bands by wavelet transform: the low-frequency component (LL), the horizontal high-frequency component (LH), the vertical high-frequency component (HL), and the diagonal high-frequency component (HH). These sub-bands respectively preserve the low-frequency information and the high-frequency information in different directions in the original image. The low-frequency component (LL) is used as the low-frequency part of the image, and this part is named LF, while the high-frequency part is represented by the sum of the high-frequency components (LH, HL, HH) in the horizontal, vertical, and diagonal directions, and the high-frequency part is named HF. The module introduces two learnable weighting parameters α and β to perform weighted fusion on LF and HF, enabling the model to automatically adjust the contribution degrees of the low-frequency (LF) and high-frequency (HF) information to the final feature (WF). After the feature map undergoes DCT and DWT operations to generate the frequency feature DCT k and WF, after batch normalization is performed on them, the module respectively performs average pooling and max pooling on DCT k and WF and adds the pooled results to obtain P avg and P max . The module uses two fully connected layers with a reduction ratio of r and to process the frequency features, and finally uses sigmoid to obtain the channel attention map S. The formula is expressed as follows:

[0068] S = σ(M2(R(M1P avg )) + M2(R(M1P max ))),

[0069] where σ represents the sigmoid activation function, and R represents the Relu activation function. The attention maps generated by each scale branch calibrate the features to generate and The formula is expressed as follows:

[0070]

[0071] Among them, FEM represents the frequency enhancement module, R represents the ReLU activation function, B represents batch normalization, DWC represents depthwise separable convolution, and i ∈ {1, 2, 4}. The module gradually aggregates the features of three scales in a cascaded addition manner, avoiding the loss of details and information inconsistency that may be caused by drastic upsampling, thereby achieving smoother feature fusion. First, the module is upsampled by a factor of 2, and then added to to obtain Subsequently, the module upsamples by a factor of 2 and adds it to to obtain Finally, the module uses a residual connection so that is added to X to obtain the final feature map Y. The formula is expressed as follows:

[0072]

[0073] where Upsample represents upsampling by a factor of 2.

[0074] In step 2, the decoder is provided with a dynamic group convolution spatial attention module (DGCSA), and the module operates on the feature map processed by the frequency enhancement channel attention module (FECA). The module structure is as Figure 6 shown. This module divides the feature map into four parts in the H and W dimensions respectively for processing, and combines four dynamic convolutions using different kernel sizes to perform feature extraction on them respectively, capturing spatial information in different directions to enhance the understanding of the spatial structure and improve the adaptability of the model to different types of features. The specific content is as follows:

[0075] The module decomposes the feature map into two dimensions, H and W, for processing respectively, which not only reduces the computational complexity but also enables the module to capture spatial information in different directions to enhance the understanding of the spatial structure. Specifically, first, the module performs global average pooling operations on the input feature map in the H and W dimensions to obtain two 1D sequence features and Subsequently, the module divides the features into N parts along the channel dimension to obtain N independent sub-features and i ∈ [1, N], and the number of channels of each sub-feature is To balance the computational complexity and model performance, N is set to 4. Next, the module uses dynamic 1D convolutions with kernel sizes of 2, 3, 5, and 7 to perform feature extraction on the 4 sub-features respectively, obtaining and The formula is expressed as follows:

[0076]

[0077] where represents the i-th dynamic 1D convolution. In the dynamic convolution, K different convolutional kernels are defined, W = {W1, W2, …, W K}, and each convolutional kernel extracts different feature patterns. The module also designs a lightweight weight generation network weight_generator composed of global pooling, fully connected layer and Softmax, which dynamically calculates the weights of each convolutional kernel according to the input features , and the weights are defined as a = {a1, a2, …, a K}. Finally, the outputs of multiple convolutional kernels are weighted and summed to obtain the final feature representation The formula is expressed as follows:

[0078]

[0079] The module performs a concatenation operation on and to achieve the aggregation of sub-features, and uses group normalization (GroupNorm) with K groups for normalization. By independently normalizing each sub-feature, GroupNorm can effectively reduce the semantic interference between sub-features, thus avoiding the dilution effect of the attention mechanism. Finally, the module uses the Sigmoid activation function to generate the spatial attention maps S H and S W and recalibrates the original features. The formula is expressed as follows:

[0080]

[0081] Y = X * S H * S W .

[0082] where G K represents group normalization with K groups.

[0083] In step 2, the skip connection block, the module structure is shown in Figure 1 . The module first normalizes the features from the next layer, then concatenates the normalized features with the skip connection from the encoder, and finally performs activation and normalization operations. The skip connection block inputs the processed features into the decoder block for decoding.

[0084] In step 2, the weighted loss function is obtained by adding the boundary difference union loss (BDoU), weighted binary cross-entropy loss (CE) and dice loss (Dice) multiplied by their respective weights. The formula is expressed as follows:

[0085] Loss = 0.4 * loss CE + 0.6 * loss Dice+0.5*loss BDoU .

[0086] The specific steps of step 3 are as follows:

[0087] The medical image segmentation model is trained using the training set. During the training process, the loss function, optimizer function, and learnable hyperparameters used by the medical image segmentation model are continuously optimized until the segmentation performance of the model reaches the best;

[0088] The trained medical image segmentation model is verified using the validation set. If the verification effect is not good, continue training. If the verification effect is good, use the test set to test the model.

[0089] When using the test set to verify the true segmentation effect of the model, if the segmentation effect is good, end the training. If the segmentation effect is not good, continue training.

[0090] The specific steps of step 4 are as follows:

[0091] Deploy the medical image segmentation model that has passed the test to the server and set the call interface and call permissions. The user inputs the preprocessed data to be segmented into the medical image segmentation model, and the medical image segmentation model outputs the segmentation results with detailed data annotations to complete the segmentation task.

[0092] In summary, the advantages of the present invention are as follows:

[0093] Select public datasets of different modalities, divide the datasets into training sets, test sets, and validation sets, and perform preprocessing operations on the dataset images; then create a medical image segmentation model based on an encoder, a decoder, a skip connection block, and a weighted loss function, train the medical image segmentation model through the training set, verify the segmentation effect of the medical image segmentation model through the validation set, and test the verified medical image segmentation model through the test set; deploy the medical image segmentation model that has passed the test and has a good effect on the server, and use this medical image segmentation model to perform medical image segmentation tasks; the multi-scale hybrid convolution module (MSHCB) and the convolutional multi-head self-attention module (CMSA) in the decoder capture the long-range dependencies and multi-scale local features of the feature map, strengthening the model's understanding of the multi-scale and diversity of the lesion area. The frequency enhancement channel attention module (FECA) in the decoder extracts the global and local frequency feature information of the feature map on this basis and dynamically weights each frequency component, enhancing the weight of important features and improving the model's noise suppression ability and robustness. The dynamic grouped convolution spatial attention module (DGCSA) in the decoder focuses on modeling the pixel-level spatial relationship in the feature map, enhancing the model's segmentation ability for boundary and detail regions.

Claims

1. A medical image segmentation method based on spatial domain and frequency domain feature attention, and the specific implementation includes the following steps: Step 1: Select publicly available datasets of different modalities, divide the datasets into training sets, test sets, and validation sets, and perform preprocessing operations on the dataset images; Step 2: Create a medical image segmentation model based on an encoder, a decoder, a skip connection block, and a weighted loss function, where the decoder is composed of a multi-scale hybrid convolutional module (MSHCB), a convolutional multi-head self-attention module (CMSA), a frequency enhancement channel attention module (FECA), and a dynamic grouped convolutional spatial attention module (DGCSA). Step 3: Train the medical image segmentation model through the training set, verify the segmentation effect of the medical image segmentation model through the validation set, and test the verified medical image segmentation model through the test set; Step 4: Deploy the medical image segmentation model that has been tested and has good results on the server, and use this medical image segmentation model to perform medical image segmentation tasks.

2. The medical image segmentation method based on spatial domain and frequency domain feature attention according to claim 1, wherein: The medical image segmentation datasets in Step 1 are the cardiac dataset (ACDC), the multi-organ dataset (Synapse), the skin lesion datasets (ISIC2017 and ISIC2018), and the polyp segmentation datasets (CVC-ClinicDB, Kvasir, and ColonDB). During preprocessing, all dataset images are adjusted to images of size 256×256 and standardized operations such as image flipping and noise reduction are performed.

3. The medical image segmentation method based on spatial domain and frequency domain feature attention according to claim 1, characterized in that: In Step 2, the encoder is composed of a multi-axis vision Transformer (MaxViT) backbone network divided into four layers. The encoder is used to perform downsampling operations on the dataset images layer by layer, gradually extract medical image features, and use block-based local attention and dilated global attention, enabling the model to perform global-local spatial interactions at any input resolution.

4. The medical image segmentation method based on spatial domain and frequency domain feature attention according to claim 1, wherein: In Step 2, a multi-scale hybrid convolutional module (MSHCB) is provided in the decoder, and the module processes the original features input from the skip connection block. This module combines depth convolution, dilated convolution, and deformable convolution, achieving flexible receptive field adjustment while significantly reducing the computational amount. The module extracts image features from multiple angles and levels, enhancing the model's adaptability to complex medical image scenarios.

5. The medical image segmentation method based on spatial domain and frequency domain feature attention according to claim 1, characterized in that: In Step 2, a convolutional multi-head self-attention module (CMSA) is provided in the decoder, and the module processes the original features input from the skip connection block. This module uses convolutional operations to replace the fully connected layer to generate queries (Q), keys (K), and values (V), and pre-extracts local pattern and structural information through the local receptive field characteristics of convolution, enabling subsequent self-attention calculations to effectively fuse local and global information.

6. The medical image segmentation method based on spatial domain and frequency domain feature attention according to claim 1, characterized in that: In step 2, the decoder is equipped with a Frequency Enhancement Channel Attention Module (FECA). This module operates on the feature map that has been processed by the Multi-Scale Hybrid Convolution Block (MSHCB) and the Convolutional Multi-Head Self-Attention Block (CMSA) and then added together. This module processes the feature map at three different scales, comprehensively captures the frequency features using the Discrete Cosine Transform (DCT) and the Discrete Wavelet Transform (DWT), calibrates the feature maps at different scales using channel attention, and finally aggregates the feature maps at the three scales in a cascaded addition manner. The FECA module comprehensively captures the key information of the image at low-frequency and high-frequency levels, improving the performance of the model.

7. The medical image segmentation method based on spatial domain and frequency domain feature attention according to claim 1, characterized in that: In step 2, the decoder is equipped with a Dynamic Group Convolution Spatial Attention Module (DGCSA). This module operates on the feature map that has been processed by the Frequency Enhancement Channel Attention Module (FECA). This module divides the feature map into four parts in the H and W dimensions respectively for processing, combines four dynamic convolutions with different kernel sizes to extract features from them respectively, captures the spatial information in different directions to improve the understanding of the spatial structure, and enhances the adaptability of the model to different types of features.

8. The medical image segmentation method based on spatial domain and frequency domain feature attention according to claim 1, characterized in that: In step 2, the skip connection block first normalizes the features from the next layer, then concatenates the normalized features with the skip connection from the encoder, and finally performs activation and normalization operations. The skip connection block inputs the processed features into the decoder block for decoding.

9. The medical image segmentation method based on spatial domain and frequency domain feature attention according to claim 1, characterized in that: In step 2, the weighted loss function is obtained by adding the Boundary Difference Union Loss (BDoU), the Weighted Binary Cross-Entropy Loss (CE), and the Dice Loss (Dice) after multiplying them by their respective weights.

10. The medical image segmentation method based on spatial domain and frequency domain feature attention according to claim 1, characterized in that: The specific implementation of step 3 is as follows: The medical image segmentation model is trained using the training set. During the training process, the loss function, optimizer function, and learnable hyperparameters used by the medical image segmentation model are continuously optimized until the segmentation performance of the model reaches the best. The trained medical image segmentation model is verified using the validation set. If the verification effect is not good, training continues. If the verification effect is good, the model is tested using the test set. When verifying the true segmentation effect of the model using the test set, if the segmentation effect is good, the training ends. If the segmentation effect is not good, training continues.

11. The medical image segmentation method based on spatial domain and frequency domain feature attention according to claim 1, wherein: The specific implementation of step 4 is as follows: The medical image segmentation model that passes the test is deployed to the server, and the call interface and call permissions are set. The user inputs the preprocessed data to be segmented into the medical image segmentation model, and the medical image segmentation model outputs the segmentation result with detailed data annotations to complete the segmentation task.

Citation Information

Cited By

  • Anti-NMDAR encephalitis clinical prognosis evaluation method based on artificial intelligence

    CN120766939A

  • Artificial intelligence-based clinical prognosis evaluation method for anti-nmdar encephalitis

    CN120766939B

  • Medical image segmentation method based on multi-scale convolution bidirectional Mama

    CN120876871A