Colon polyp segmentation method based on multi-frequency guiding attention network
The multi-frequency guided attention network addresses the challenge of capturing both high-frequency edges and low-frequency structures in colorectal polyp segmentation, enhancing segmentation accuracy and stability through adaptive dilation spatial pooling and iterative attention mechanisms.
Patent Information
- Application Number
- CN202510443911.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-15
AI Technical Summary
Existing deep learning methods are difficult to capture high-frequency edge details and low-frequency global structures simultaneously in colon polyps segmentation, resulting in blurred boundaries and incomplete morphology of segmentation results, and insufficient segmentation accuracy and stability under large changes in polyps morphology and complex tissue structure.
A multi-frequency guided attention network is adopted, and multi-level features are extracted through the ResNet-50 backbone network, combined with the multi-frequency guided attention module and iterative content guided attention module, optimize high-frequency edge information and low-frequency global feature modeling, and improve segmentation accuracy and stability through multi-stage loss supervision training.
While capturing details, maintaining overall structural integrity is improved, and the accuracy and stability of colon polyps segmentation are especially in segmentation effects in blurred boundaries and complex morphological areas.
Smart Images

Figure CN120318252A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of deep learning and medical image technology, and particularly relates to a colon polyp segmentation method based on a multi-frequency guided attention network. Background Art
[0002] Colon polyp segmentation is a core task in computer vision and has important applications in computer-aided diagnosis (CAD) systems, which can help clinicians accurately identify and separate lesion areas in medical images. High-quality segmentation results can provide more accurate diagnostic basis for doctors, thereby improving the treatment effect and clinical prognosis of patients. Colorectal cancer (CRC) is the main early lesion of colorectal cancer, and its incidence is on the rise, becoming one of the major diseases endangering human health globally. Therefore, how to improve the early detection and precise treatment of colorectal cancer has become the core challenge in the cross-research of medicine and artificial intelligence. However, due to the limitations of imaging methods, colonoscopy images are often interfered by factors such as blurred boundaries, low contrast, uneven tissue structure, and large scale variations, making it extremely difficult to distinguish polyps from surrounding tissues.
[0003] In recent years, deep learning has made remarkable progress in the field of medical image segmentation. Among them, the encoder-decoder architecture, especially U-Net and its variants, has occupied a core position in medical image segmentation tasks due to its strong feature learning and reconstruction capabilities. To further improve the performance of U-Net in segmentation tasks, many researchers have dedicated to proposing various improvement schemes. For example, UNet++ adds densely connected skip connections and convolutional blocks to fuse features from different receptive fields. ResUNet++ combines residual blocks and SE attention mechanisms to enhance feature expression ability and improve segmentation accuracy. CPFNet further adds two pyramid feature extraction modules during the decoding process to fuse global and multi-scale context information, enhancing the model's ability to capture complex structures. Although these improvements have improved the segmentation accuracy to a certain extent, they mainly focus on local feature extraction and have insufficient ability to model long-range dependencies, making it difficult to effectively capture global semantic information. With the successful application of Transformer in the field of computer vision, some researchers have tried to introduce it into medical image segmentation tasks. For example, TransUNet, DS-TransUNet, Transfuse, BGDR-TransUNet, Trans-UNeter, and MedicalTransformer. Although the Transformer-based methods show good performance in modeling long-range dependencies, these methods often increase the computational complexity and rely strongly on large-scale datasets and high-computing-power devices, limiting their application in resource-constrained environments.
[0004] Boundary information plays a crucial role in medical image segmentation, especially when dealing with images with weak boundaries and noise interference. In recent years, researchers have proposed various segmentation methods that explicitly enhance boundary information. For example, CaraNet introduces a context axial reverse attention mechanism to strengthen the attention to polyp boundaries and optimizes for small object segmentation tasks, thereby improving the model's detection ability for fine-grained objects. DCRNet designs a memory mechanism that stores the feature embeddings of previous images through a queue, extracts co-occurring visual patterns during cross-image learning, and combines the internal and external context information within a single image to enhance the stability and consistency of segmentation. CCBANet adopts a cascaded context module and an attention balance module. The former combines the region information extracted from the current layer and the lower layer and passes it to the higher layer to fuse region and global information in a waterfall mode. The latter uses the predicted output of the adjacent lower layer as a guidance map to implement the attention mechanism for three regions: background, polyp, and boundary curve. M2UNet integrates MetaFormer and multi-scale information to enhance the utilization of context. MSAByNet proposes multi-scale subtraction attention modules between and within layers and sets different receptive fields for modules at different levels to avoid the loss of feature map resolution and edge detail features.
[0005] In summary, although deep learning has made remarkable progress in the field of medical image segmentation, existing methods still face two core challenges: (1) it is difficult to capture high-frequency edge details (edges, textures) and low-frequency global structures (semantic information) simultaneously, resulting in problems such as blurred boundaries and incomplete morphology in the segmentation results; (2) in the case of large morphological variations and complex tissue structures of polyps, it is difficult for the model to adaptively adjust the attention area, thereby affecting the segmentation accuracy and stability. These factors limit the performance of traditional image segmentation methods in polyp detection tasks and are difficult to meet the clinical application requirements. Therefore, constructing a robust and efficient automatic polyp segmentation method has important clinical value for the early prevention, diagnosis, and treatment of colorectal cancer. Summary of the Invention
[0006] To solve the above problems, through multi-frequency analysis, diverse features are extracted in different frequency ranges to enhance high-frequency edge information while optimizing low-frequency global feature modeling, ensuring that the model captures details while maintaining the integrity of the overall structure. The present invention adopts the following technical solutions:
[0007] A colon polyp segmentation method based on a multi-frequency guided attention network, comprising the following steps:
[0008] Step 1: Dataset construction, the dataset includes a number of polyp images;
[0009] Step 2: Use the ResNet-50 backbone network to extract features from the polyp images, obtaining five hierarchical features of different scales. Among them, the shallow feature with the largest scale is the low-level feature The remaining shallow features are high-level features For the high-level features Perform dimensionality reduction processing to obtain features
[0010] Step 3: According to the features And the corresponding hierarchical high-level prediction features Use the multi-frequency guided attention module to generate refined features Directly input the features Into the decoder to generate high-level features
[0011] Step 4: Through the iterative content-guided attention module, fuse the refined features With the high-level features Gradually generate high-level semantic features And the prediction feature map
[0012] Step 5: Perform multi-stage loss supervised training on the segmentation model to obtain an optimized polyp segmentation model.
[0013] Furthermore, the multi-level features of the images obtained in step 2 are specifically: Import the image I ∈ R of the dataset in step 1 3×H×w , Use the ResNet-50 backbone network to extract features from the polyp images to obtain features Then for Perform dimensionality reduction processing on the features to obtain Keep the original number of channels unchanged, where R represents the set of real numbers, N represents the image channels, and The N of i ∈ {64, 256, 512, 1024, 2048}, H represents the height of the image, W represents the width of the image, The N of i ∈ {64, 64, 128, 256, 512}.
[0014] Furthermore, in step 3, use the multi-frequency guided attention module to obtain refined features: The specific steps are as follows:
[0015] Step 3.1: Pass the dimensionality-reduced features Through the multi-scale frequency extraction in the multi-frequency guided attention module to obtain multi-frequency attention features In order to extract from the input dimensionality-reduced feature map To obtain features with different receptive fields, adaptive dilated spatial pyramid pooling with different dilation rates is adopted. The dilation rate r is as follows:
[0016]
[0017] The adaptive dilated spatial pyramid pooling consists of multiple parallel dilated convolutional layers. After each convolutional layer, batch normalization and the ReLU activation function are connected. Finally, the multi-scale features are integrated through 1×1 convolution, and the final output is passed to the subsequent attention module to further optimize the multi-frequency feature modeling.
[0018] The output of the adaptive dilated spatial pyramid pooling is passed through the multi-scale frequency extraction module in the multi-frequency guided attention module to obtain the multi-frequency attention features
[0019] Furthermore, the two-dimensional discrete cosine transform (2D DCT) is used to represent an image as a weighted sum of basis images generated by different frequency cosine functions:
[0020]
[0021] where (u k , v k ) represents the frequency index corresponding to . In addition, the two-dimensional discrete cosine transform (2D DCT) basis image is defined as follows:
[0022]
[0023] And the Top-K selection strategy is adopted to screen important frequency components. Subsequently, each obtains G avg , G max , G min through global average pooling, global max pooling, and global min pooling respectively. Then, through the fully connected layer W ∈ R C×C . As shown in the following formula:
[0024] M i = σ(∑ d∈{avg,max,min} W(δ(G d )))
[0025] where δ and σ represent the ReLU and Sigmoid activation functions respectively. Finally, the feature map is recalibrated using M i :
[0026] Furthermore, the multi-frequency attention features pass through the feature-enhanced spatial attention in the multi-frequency guided attention module to obtain the final output feature map
[0027] Calibrated feature map It is used to extract significant boundary information at different scales in the spatial domain. In this process, we introduce two learnable parameters (β i , α i ), which are used to control the background and foreground respectively. As shown in the following formula:
[0028]
[0029] Among them, the foreground attention map F i is calculated through 1×1 convolution and Sigmoid: The corresponding background attention map pair W i is obtained by element-wise subtraction, that is, (1 - W i ). Finally It is used as the multi-frequency feature input of MFGA.
[0030] Step 3.2: The high-level features obtained in Step 4 are convolved to obtain a higher-level prediction map Then, through a decomposition method, two attention maps are generated from the higher-level prediction map : the reverse attention map and the boundary attention map
[0031] Step 3.3: The multi-frequency feature, reverse attention map, boundary attention map and dimensionality-reduced feature are fused to generate the combined feature f i c ;
[0032]
[0033] Step 3.4: f i c is convolved and passed through the sigmoid function to obtain an attention mask, denoted as Ai. Multiply Ai with f i c and then add it to the dimensionality-reduced feature to obtain f i a ;
[0034] Among them, [·] represents the concatenation operation of features. Considering that the edge information may contain redundant details of noise that are not beneficial to segmentation, an attention mask, denoted as Ai, is introduced. This mask aims to guide the model to focus on important regions while suppressing background noise and redundant information. The attention feature map f i a of the i-th layer is defined as follows:
[0035]
[0036] Among them, the symbol σ represents the sigmoid function.
[0037] Step 3.5: Apply DCA (Dual Cross-Attention Module) to recalibrate the attention feature f i a to obtain a refined feature
[0038]
[0039] Furthermore, in the said Step 4, the iterative content-guided attention module gradually optimizes the attention weights for the refined feature obtained in Step 3 and the high-level feature of the previous layer to dynamically adjust the model's attention in different regions to obtain high-level semantic features and prediction feature maps. The specific steps are as follows:
[0040] Step 4.1: Add the refined feature and the high-level feature to generate a spatial importance map W through content-guided attention;
[0041] The operation of the output spatial importance map W of the content-guided attention module:
[0042] We calculate the corresponding W through the following formula C and W S .
[0043]
[0044]
[0045] Among them, max(0, x) represents the ReLU activation function, C k×k (·) represents a convolutional layer with a kernel size of k×k, and [·] represents a channel-level concatenation operation. and respectively represent the features processed by global average pooling across the spatial dimension, global average pooling across the channel dimension, and global max pooling across the channel dimension. To reduce the number of parameters and limit the model complexity, the first 1×1 convolution reduces the channel dimension from C to (r is the reduction ratio). Subsequently, the second 1×1 convolution expands it back to C. In our implementation, we choose to reduce the number of channels to a fixed value (i.e., 16) and set r to
[0046] We fuse W C and W S by element-wise addition to obtain a rough spatial importance map Wcoa ∈R C×H×W :
[0047] W coa =W C +W S
[0048] Since W C is channel-based, W coa and X are aligned in the channel dimension. To obtain the final refined spatial importance map W, we need to adjust it according to the content of the input features. Specifically, we adopt channel shuffle operation and combine it with group convolution:
[0049] W = σ(GC 7×7 (CS([X, W coa )))
[0050] where σ represents the sigmoid function, CS(·) represents the channel shuffle operation, and GC K×K (·) represents the group convolution layer with kernel size k×k. In our implementation, the number of groups is set to C.
[0051] Step 4.2: Multiply the feature by W and multiply (1 - W) by , and finally add them together:
[0052]
[0053] where the input of W is W is obtained by the first CGA calculation.
[0054] Step 4.3: Feed the result obtained in Step 4.2 into CGA again to get the spatial importance map W′:
[0055] W′ = CGA(F1),
[0056] Step 4.4: Multiply the feature by W′ and multiply (1 - W′) by , and finally add them together to get the high-level feature f i d :
[0057]
[0058] f i d is the output feature, and are the MFGA output feature and the high-level feature of the previous layer's Decoder respectively, and W′ is the weight map calculated by the second CGA.
[0059] Further, in step 5, a dedicated multi-stage loss and output aggregation strategy is adopted to supervise the model with binary cross-entropy loss (L BCE ) with weight decay and dice loss (L Dice ) to train the polyp image segmentation model, including the following steps:
[0060] Step 5.1: Input the predicted region and the ground truth region, calculate the binary cross-entropy loss to measure the classification error between the predicted region and the ground truth region at the pixel level, so as to ensure that the model can accurately distinguish the polyp region and the background region:
[0061]
[0062] Step 5.2: Calculate the dice loss to measure the overlap degree between the predicted region and the ground truth region, enhance the model's learning ability of the overall shape of the polyp region, and improve the segmentation performance of the model on polyps of different sizes:
[0063]
[0064] Step 5.3: During the training process, use the deep fusion features and boundary detail information as inputs to calculate the binary cross-entropy loss and dice loss respectively, enhance the model's perception ability of the polyp region boundary, and improve the segmentation accuracy;
[0065] Step 5.4: Accumulate the binary cross-entropy loss and dice loss corresponding to all deep fusion features and boundary detail information to obtain the total binary cross-entropy loss and total dice loss respectively;
[0066]
[0067] Perform global constraint and local constraint weighting on the calculated total binary cross-entropy loss and total dice loss to obtain the weighted binary cross-entropy loss and weighted dice loss, and further combine them to form the final loss to optimize the training process of the polyp image segmentation model.
[0068] Further, in step 1, the collected polyp image dataset is divided into a training set and a test set; in step 5, the polyp image segmentation model is trained through the training set. This method further includes:
[0069] Step 6: Use the best polyp segmentation model obtained after training to test the polyp images in the test set, obtain the polyp segmentation results after testing, and evaluate them;
[0070] Step 7: Input the polyp image to be segmented into the trained polyp image segmentation model, and output the polyp segmentation result.
[0071] The advantages and beneficial effects of the present invention are as follows:
[0072] A colon polyp segmentation method based on a multi-frequency guided attention network according to the present invention combines a multi-frequency guided attention module on the basis of a Res2Net network. The core component of this module is a multi-scale frequency extraction module, which combines an adaptive dilated spatial pyramid pooling mechanism to enhance the adaptability to scale changes in the target area by fusing dilated convolutions with different receptive fields. Subsequently, the multi-scale frequency extraction module extracts frequency statistical information through a multi-frequency channel attention module combined with 2D discrete cosine transform, thereby generating a channel attention map, and finally differentiates boundary features through feature-enhanced spatial attention. More specifically, through multi-frequency analysis, this method extracts diverse features in different frequency ranges to enhance high-frequency edge information while optimizing low-frequency global feature modeling, ensuring that the model captures details while maintaining the integrity of the overall structure. In addition, to further optimize the segmentation results, we also introduce an iterative content-guided attention module, which dynamically adjusts the model's attention in different regions by gradually optimizing the attention weights, especially for lesion regions with blurred boundaries and complex morphologies, and can effectively improve the ability to capture key anatomical structures, ensuring the accuracy and stability of polyp segmentation. Brief Description of the Drawings
[0073] Figure 1 Schematic diagram of the overall model structure for colon polyp segmentation according to the present invention
[0074] Figure 2 Schematic diagram of the frequency-guided attention module according to the present invention
[0075] Figure 3 Schematic diagram of the multi-scale frequency extraction attention module according to the present invention
[0076] Figure 4 Schematic diagram of the iterative content-guided attention module according to the present invention
[0077] Figure 5 Comparison chart of visualization effects of the present invention and other methods obtained through experiments Detailed Description of the Preferred Embodiments
[0078] The following detailed description of the specific embodiments of the present invention is provided in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not intended to limit the present invention. It should be noted that the diagrams provided in the following embodiments are only used to illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and the features in the embodiments can be combined with each other without conflict.
[0079] Among them, the accompanying drawings are only for illustrative purposes, showing only schematic diagrams rather than actual pictures, and should not be construed as a limitation on the present invention; for better illustration of the embodiments of the present invention, some components in the accompanying drawings will be omitted, enlarged or reduced, which do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.
[0080] Please refer to Figure 1 , one aspect of the present invention provides a colon polyp segmentation method based on a multi-frequency guided attention network, including
[0081] Step 1: Dataset construction, the dataset includes several polyp images;
[0082] Step 2: Use the ResNet-50 backbone network to extract features from the polyp images to obtain 5 hierarchical features of different scales, where the shallow feature with the largest scale is the low-level feature The remaining shallow features are high-level features For the high-level features Perform dimensionality reduction processing to obtain the feature
[0083] Step 3: According to the feature And the corresponding hierarchical high-level prediction features Use the multi-frequency guided attention module to generate refined features Input the feature Directly into the decoder to generate high-level features
[0084] Step 4: Fusion the refined features And the high-level features Through the iterative content-guided attention module to gradually generate high-level semantic features And the prediction feature map
[0085] Step 5: Perform multi-stage loss supervised training on the segmentation model to obtain an optimized polyp segmentation model.
[0086] Furthermore, the multi-level features of the images obtained in step 2 are specifically: import the image I∈R of the dataset in step 1 3×H×w , use the ResNet-50 backbone network to extract features from the polyp images to obtain the feature Then for The feature is processed by dimensionality reduction to obtain Keep the original number of channels unchanged, where R represents the set of real numbers, N represents the image channels, and The N of i∈ {64, 256, 512, 1024, 2048}, where H represents the height of the image and W represents the width of the image. of N i ∈ {64, 64, 128, 256, 512}.
[0087] Please refer to Figure 2 , Figure 3 , in step 3, a multi-frequency guided attention module is used to obtain refined features: The specific steps are as follows:
[0088] Step 3.1: The dimensionality-reduced feature obtains multi-frequency attention features through multi-scale frequency extraction in the multi-frequency guided attention module To obtain features with different receptive fields from the input dimensionality-reduced feature map , adaptive dilated spatial pyramid pooling with different dilation rates is adopted. The dilation rate r is:
[0089]
[0090] Adaptive dilated spatial pyramid pooling consists of multiple parallel dilated convolutional layers. After each convolutional layer, batch normalization and ReLU activation functions are connected. Finally, the multi-scale features are integrated through 1×1 convolution, and the final output is passed to the subsequent attention module to further optimize the multi-frequency feature modeling.
[0091] The output of the adaptive dilated spatial pyramid pooling is passed through the multi-scale frequency extraction module in the multi-frequency guided attention module to obtain multi-frequency attention features
[0092] Furthermore, the two-dimensional discrete cosine transform (2D DCT) is used to generate a weighted sum of basis images represented by different frequency cosine functions for the image:
[0093]
[0094] where (u k , v k ) represents the frequency index corresponding to . In addition, the two-dimensional discrete cosine transform (2D DCT) basis image is defined as follows:
[0095]
[0096] And the Top-K selection strategy is adopted to screen important frequency components. Subsequently, each obtains G avg , G max , G min . Then, through the fully connected layer W ∈ RC×C As shown in the following formula:
[0097] M i = σ(∑ d∈{avg,max,min} W(δ(G d )))
[0098] where δ and σ represent the ReLU and Sigmoid activation functions respectively. Finally, use M i to recalibrate the feature map:
[0099] Furthermore, the multi-frequency attention feature obtains the final output feature map by enhancing the spatial attention of the feature in the multi-frequency guided attention module
[0100] The calibrated feature map is used to extract significant boundary information at different scales in the spatial domain. In this process, we introduce two learnable parameters (β i , α i ), which are used to control the background and foreground respectively. As shown in the following formula:
[0101]
[0102] where the foreground attention map F i is calculated through 1×1 convolution and Sigmoid: The corresponding background attention map pair W i is obtained by element-wise subtraction, that is, (1 - W i ). Finally is used as the multi-frequency feature input of MFGA.
[0103] Step 3.2: Convolve the high-level features obtained in Step 4 to obtain a higher-level prediction map Then, use the decomposition method on the higher-level prediction map to generate two attention maps: the reverse attention map and the boundary attention map
[0104] Step 3.3: Fuse the multi-frequency feature, reverse attention map, boundary attention map with the dimensionality-reduced feature to generate the combined feature f i c ;
[0105]
[0106] Step 3.4: For f i cPerform convolution and sigmoid function to obtain an attention mask, denoted as Ai, and multiply Ai by f i c and then add the result to the dimensionality-reduced feature to obtain f i a ;
[0107] where [·] represents the feature concatenation operation. Considering that edge information may contain redundant details that are not beneficial to segmentation, an attention mask, denoted as A i . This mask aims to guide the model to focus on important regions while suppressing background noise and redundant information. The attention feature map f i a of the i-th layer is defined as follows:
[0108]
[0109] where the symbol σ represents the sigmoid function.
[0110] Step 3.5: Apply DCA (Dual Cross-Attention Module) to recalibrate the attention feature f i a to obtain a refined feature
[0111]
[0112] Please refer to Figure 4 , in step 4, the iterative content-guided attention module gradually optimizes the attention weights of the refined feature obtained in step 3 and the high-level feature of the previous layer to dynamically adjust the model's attention in different regions to obtain high-level semantic features and prediction feature maps. The specific steps are as follows:
[0113] Step 4.1: Add the refined feature and the high-level feature to generate a spatial importance map W through content-guided attention;
[0114] where the operation of the content-guided attention module to output the spatial importance map W:
[0115] We calculate the corresponding W C and W S .
[0116]
[0117]
[0118] Among them, max(0, x) represents the ReLU activation function, C k×k (·) represents a convolutional layer with a kernel size of k×k, and [·] represents a channel-level concatenation operation. and represent the features processed by global average pooling across the spatial dimension, global average pooling across the channel dimension, and global max pooling across the channel dimension, respectively. To reduce the number of parameters and limit the model complexity, the first 1×1 convolution reduces the channel dimension from C to (r is the reduction ratio). Subsequently, the second 1×1 convolution expands it back to C. In our implementation, we choose to reduce the number of channels to a fixed value (i.e., 16) and set r to
[0119] We fuse W C and W S by element-wise addition to obtain the rough spatial importance map W coa ∈R C×H×W :
[0120] W coa = W C + W S
[0121] Since W C is channel-based, W coa and X are aligned in the channel dimension. To obtain the final fine spatial importance map W, we need to adjust it according to the content of the input features. Specifically, we adopt the channel shuffle operation and combine it with group convolution:
[0122] W = σ(GC 7×7 (CS([X, W coa )))
[0123] where σ represents the sigmoid function, CS(·) represents the channel shuffle operation, and GC K×K (·) represents a group convolutional layer with a kernel size of k×k. In our implementation, the number of groups is set to C.
[0124] Step 4.2: Multiply the feature by W and multiply (1 - W) by and finally add them:
[0125]
[0126] where the input of W is W is obtained by the first CGA calculation.
[0127] Step 4.3: Feed the result obtained in Step 4.2 into the CGA again to obtain the spatial importance map W′:
[0128] W′ = CGA(F1),
[0129] Step 4.4: Multiply the feature by W′ and multiply (1 - W′) by and then sum them up to obtain the high-level feature f i d :
[0130]
[0131] f i d is the output feature, and are the MFGA output feature and the high-level feature of the previous layer of the Decoder respectively. W′ is the weight mapping calculated by the second CGA.
[0132] Furthermore, in the said Step 5, a dedicated multi-stage loss and output aggregation strategy is adopted to supervise the model with the binary cross-entropy loss with weight decay (L BCE ) and the dice loss (L Dice ) to train the polyp image segmentation model, including the following steps:
[0133] Step 5.1: Input the prediction region and the ground truth region, calculate the binary cross-entropy loss to measure the classification error between the prediction region and the ground truth region at the pixel level, so as to ensure that the model can accurately distinguish the polyp region and the background region:
[0134]
[0135] Step 5.2: Calculate the dice loss to measure the overlap degree between the prediction region and the ground truth region, enhance the model's learning ability of the overall shape of the polyp region, and improve the segmentation performance of the model on polyps of different sizes:
[0136]
[0137] Step 5.3: During the training process, use the depth fusion feature and the boundary detail information as inputs to calculate the binary cross-entropy loss and the dice loss respectively, enhance the model's perception ability of the polyp region boundary, and improve the segmentation accuracy;
[0138] Step 5.4: Accumulate all the binary cross-entropy losses and dice losses corresponding to the depth fusion features and the boundary detail information to obtain the total binary cross-entropy loss and the total dice loss respectively;
[0149]
[0140] Global and local constraint weighting is performed on the calculated total binary cross-entropy loss and total Dice loss to obtain a weighted binary cross-entropy loss and a weighted Dice loss, which are further combined to form a final loss to optimize the training process of the polyp image segmentation model.
[0141] Further, in step 1, the collected polyp image dataset is divided into a training set and a test set; in step 5, the polyp image segmentation model is trained using the training set. This method further includes:
[0142] Step 6: Using the best polyp segmentation model obtained after training, test the polyp images in the test set to obtain the polyp segmentation results after testing and evaluate them;
[0143] Step 7: Input the polyp image to be segmented into the trained polyp image segmentation model to output the polyp segmentation result.
[0144] The advantages and beneficial effects of the present invention are as follows:
[0145] A colon polyp segmentation method based on a multi-frequency guided attention network according to the present invention combines a multi-frequency guided attention module on the basis of a Res2Net network. The core component of this module is a multi-scale frequency extraction module, which combines an adaptive dilated spatial pyramid pooling mechanism to enhance the adaptability to scale changes in the target area by fusing dilated convolutions with different receptive fields. Subsequently, the multi-scale frequency extraction module extracts frequency statistical information through a multi-frequency channel attention module combined with a 2D discrete cosine transform to generate a channel attention map, and finally distinguishes boundary features through feature-enhanced spatial attention. More specifically, through multi-frequency analysis, this method extracts diverse features in different frequency ranges to enhance high-frequency edge information while optimizing low-frequency global feature modeling, ensuring that the model captures details while maintaining the integrity of the overall structure. In addition, to further optimize the segmentation results, we also introduce an iterative content-guided attention module, which dynamically adjusts the attention of the model in different regions by gradually optimizing the attention weights, especially for lesion regions with blurred boundaries and complex morphologies, and can effectively improve the ability to capture key anatomical structures, ensuring the accuracy and stability of polyp segmentation.
[0146] To illustrate the effectiveness of this method, some polyp images were selected from the classic polyp segmentation datasets CVC-300, Kvasir, ColonDB, ETIS, and ClinicDB to train the complete network. Then, the five datasets CVC-300, Kvasir, ColonDB, ETIS, and ClinicDB were used for testing, and the test results were compared with the mainstream polyp segmentation algorithms in the prior art. Pytorch and NVIDIA RTX 4060Ti were used. We trained with a batch size of 6, 200 batches, and used the Stochastic Gradient Descent (SGD) optimizer with a momentum of 0.9 and a weight decay of 1e-5. We adjusted the size of the input images to 352×352, and used this size in both the training and inference phases, and then adjusted back to the original size when calculating the evaluation metrics. To perform data augmentation, we used random horizontal and vertical flips, random rotations, and multi-scale training strategies.
[0147] The test results on the CVC-300, Kvasir, ColonDB, ETIS, and ClinicDB datasets are shown in Table 1:
[0148] Table 1 Test results of polyp segmentation of the present invention and 7 other methods
[0149]
[0150] The bold font is the optimal value for each row.
[0151] We adopted five common metrics: mean Dice (mDice), mean IoU (mIoU), weighted Dice metric structural metric (S a ), enhanced alignment metric mean squared error (MAE). These metrics have a dual purpose, that is, to evaluate the performance of our method compared with the ground truth labels (i.e., between the prediction and the ground truth f). The lower the MAE value, the better, and the higher the other values, the better.
[0152] As Figure 5 shown, from the results, it can be concluded that the proposed polyp segmentation method of the present invention can obtain more accurate boundaries and more complete semantic structures in the segmentation results.
[0153] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A colon polyp segmentation method based on a multi-frequency guided attention network, characterized in that, Including the following steps: Step 1: Dataset construction, where the dataset includes a number of polyp images; Step 2: Use the ResNet-50 backbone network to extract features from polyp images, obtaining five hierarchical features of different scales. Among them, the shallow feature with the largest scale is the low-level feature The remaining shallow features are high-level features For the high-level features perform dimensionality reduction processing to obtain features Step 3: According to the feature and the high-level prediction feature of the corresponding level Use the multi-frequency guided attention module to generate refined features Input the feature directly into the decoder to generate high-level features Step 4: Fuse the refined features through the iterative content-guided attention module with the high-level features to gradually generate high-level semantic features and the predicted feature map Step 5: Conduct multi-stage loss supervised training on the segmentation model to obtain an optimized polyp segmentation model.
2. The colon polyp segmentation method based on a multi-frequency guided attention network according to claim 1, wherein the feature lies in And: The multi-level features of the image obtained in step 2 are specifically as follows: Import the image I ∈ R of the dataset in step 1 3×H×W , and use the ResNet-50 backbone network to extract features from the polyp image to obtain features Then, for the features, perform dimensionality reduction processing to obtain Keep the original number of channels unchanged, where R represents the set of real numbers, N represents the image channels, and the N of i ∈ {64, 256, 512, 1024, 2048}, H represents the height of the image, W represents the width of the image, the N of i ∈ {64, 64, 128, 256, 512}.
3. The colon polyp segmentation method based on a multi-frequency guided attention network according to claim 1, wherein In step 3, use a multi-frequency guided attention module to obtain refined features: The specific steps are as follows: Step 3.1: The dimensionality-reduced features are used to obtain multi-frequency attention features through multi-scale frequency extraction in the multi-frequency guided attention module Step 3.2: The high-level features obtained in Step 4 are convolved to obtain a higher-level prediction map Then, through a decomposition method, the higher-level prediction map generates two attention maps: the reverse attention map and the boundary attention map Step 3.3: Fuse the multi-frequency features, reverse attention map, boundary attention map and dimensionality-reduced features to generate combined features Step 3.4: For perform convolution and apply the sigmoid function to obtain an attention mask, denoted as Ai. Multiply Ai with and then add the result to the dimensionality-reduced feature to get Step 3.5: For the attention feature perform DCA (Dual Cross-Attention Module) recalibration to obtain a refined feature 4. A colon polyp segmentation method based on a multi-frequency guided attention network according to claim 1, characterized in that: In the step 4, the iterative content-guided attention module refines the features obtained in step 3 and the high-level features of the previous layer to gradually optimize the attention weights, dynamically adjust the attention of the model in different regions, obtain high-level semantic features and prediction feature maps. The specific steps are as follows: Step 4.1: Add the refined features and the high-level features to generate a spatial importance map W by guiding attention through content; Step 4.2: Multiply feature by W and multiply (1 - W) by and then add them up; Step 4.3: Send the result obtained in step 4.2 into the CGA again to obtain the spatial importance map W'; Step 4.4: Multiply the feature by W′ and multiply (1 - W′) by and finally add them together to obtain the high-level feature 5. A colon polyp segmentation method based on a multi-frequency guided attention network according to claim 1, characterized in that: In step 5, a dedicated multi-stage loss and output aggregation strategy is adopted to supervise the model with binary cross-entropy loss (L BCE ) with weight decay and dice loss (L Dice ), and the polyp image segmentation model is trained, including the following steps: Step 5.1: Input the predicted region and the ground truth region, calculate the binary cross-entropy loss to measure the classification error between the predicted region and the ground truth region at the pixel level, so as to ensure that the model can accurately distinguish the polyp region and the background region; Step 5.2: Calculate the Dice loss to measure the overlap degree between the predicted region and the ground truth region, enhance the model's learning ability for the overall shape of the polyp region, and improve the segmentation performance of the model on polyps of different sizes; Step 5.3: During the training process, use the depth fusion features and boundary detail information as inputs, calculate the binary cross-entropy loss and the Dice loss respectively, enhance the model's perception ability of the polyp region boundary, and improve the segmentation accuracy; Step 5.4: Accumulate the binary cross-entropy losses and Dice losses corresponding to all the depth fusion features and boundary detail information to obtain the total binary cross-entropy loss and the total Dice loss respectively; Step 5.5: Perform global constraint and local constraint weighting on the calculated total binary cross-entropy loss and total Dice loss to obtain the weighted binary cross-entropy loss and the weighted Dice loss, and further combine them to form the final loss to optimize the training process of the polyp image segmentation model.
6. The colon polyp segmentation method based on a multi-frequency guided attention network according to claim 1, characterized in that: In step 1, divide the collected polyp image dataset into a training set and a test set; in step 5, train the polyp image segmentation model through the training set. This method further includes: Step 6: Use the best polyp segmentation model obtained after training to test the polyp images in the test set, obtain the polyp segmentation results after testing, and conduct an evaluation; Step 7: Input the polyp image to be segmented into the trained polyp image segmentation model to output the polyp segmentation result.
Citation Information
Cited By
Copper strip image defect segmentation method based on time-space domain
CN121392843A
A method for segmenting defects in copper strip images based on space-time domain
CN121392843B
Lightweight polyp segmentation method based on multi-scale differential features and spatial attention
CN121504959A
Lightweight polyp segmentation method based on multi-scale differential features and spatial attention
CN121504959B