A polyp image segmentation method based on boundary-guided multi-level attention network
By designing a multi-level attention network based on boundary guidance in polyp image segmentation, the segmentation problem caused by similar polyp shape and background in the prior art is solved, and higher segmentation accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202411414098.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-11
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-10-11
AI Technical Summary
The existing polyp image segmentation method is difficult to accurately segment when the shape, size and texture of polyp are different and highly similar to the background visually, and the multi-layer feature aggregation effect is poor, affecting the segmentation accuracy and robustness.
A multi-level attention network based on boundary guidance was designed, and multi-level features were acquired through the pyramid visual Transformer-PVT-v2 backbone and multi-scale feature extraction module. Combined with the boundary perception module and the multi-level attention module, the target polyp region is dynamically selected to enhance the model's refinement and segmentation effect on polyp boundaries.
Effectively fuse low-level features and high-level global features to generate fine polyp boundaries, improve segmentation accuracy and robustness, and can more accurately locate polyp areas and generate fine boundaries.
Smart Images

Figure CN118941587B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a polyp image segmentation method based on a boundary-guided multi-level attention network. Background Art
[0002] Colorectal cancer is a major health problem worldwide, ranking third in morbidity and second in mortality. Colonoscopy is the standard for diagnosing colorectal cancer, which can help doctors detect and remove potential cancerous polyps early. However, due to the complexity and subjectivity of the colonoscopy process, it may lead to missed diagnosis. Failure to detect polyps in time may lead to the diagnosis of advanced cancer, when the patient's survival rate is less than 10%. Therefore, developing a method that can automatically and accurately segment polyps from colonoscopic images will significantly improve diagnostic efficiency and reduce the burden of manual annotation for doctors. However, there are still several technical challenges in achieving automatic polyp segmentation, especially the large differences in shape, size and texture of polyps, the high visual similarity of polyps to surrounding tissues (such as mucosa), and the easy fusion with the background. These factors make it difficult to generalize the segmentation algorithm in different cases.
[0003] Early polyp segmentation methods relied on hand-crafted features such as color, shape, and texture, but due to the limited ability of manual feature extraction, these methods performed poorly when processing complex samples and had a high rate of missed detection. With the rapid development of deep learning technology, methods based on convolutional neural networks have been widely used in image segmentation tasks. Among them, U-Net, as a representative architecture, uses an encoder-decoder structure and jump connections to enable it to effectively integrate spatial information and achieve excellent performance in medical image segmentation. Inspired by U-Net, multiple variants have been applied to the field of polyp segmentation and have made significant progress. These deep learning models can better extract features, thereby more accurately segmenting polyps and surrounding tissues, significantly improving the accuracy of segmentation.
[0004] In recent years, attention mechanisms have been introduced into deep learning models to improve segmentation performance. For example, SANet reduces the influence of polyp color through color swapping operations and uses shallow attention mechanisms to filter background noise; CaraNet proposes an axial reverse attention module to enhance the model's sensitivity to salient information by analyzing position information and multi-scale features; LDNet further improves the segmentation effect by extracting global context features and introducing a self-attention mechanism to capture long-range contextual relationships. However, although these models have made some progress in segmentation accuracy, they often ignore the processing of boundary information, which is crucial for accurate segmentation of polyps. Ignoring boundary information may lead to missed detection or inaccurate segmentation boundaries, thus affecting the overall segmentation effect. Therefore, incorporating boundary-aware mechanisms into segmentation models may be the key to improving segmentation accuracy.
[0005] In order to improve segmentation performance, some models have begun to make greater use of boundary information. For example, CFANet proposed a two-stream structure in which low-level features are used to generate boundary-aware features, which are then gradually integrated into high-level features to form a cross-layer polyp feature fusion. MEGANet retains high-frequency boundary information by integrating classic edge detection techniques with attention mechanisms, avoiding the loss of boundary information during network deepening. However, these methods often perform poorly in dealing with the relationship between low-level details and high-level semantics, resulting in unsatisfactory multi-layer feature aggregation. For example, MEGANet retains high-frequency edge information through the Laplacian operator, thereby improving the model's ability to deal with weak boundaries. However, this high-frequency edge information includes not only polyp boundaries, but may also contain noise boundaries. Combining these noises with high-level features may limit the overall performance of the model. Similarly, CFANet generates boundary-aware features by merging low-level features and integrating them into high-level features, but due to the large amount of noise in low-level features, simply aggregating features may lead to blurred boundaries and affect segmentation performance.
[0006] In summary, although deep learning-based polyp segmentation methods have made significant progress, they still face multiple challenges. In particular, how to better utilize boundary information and the aggregation of multi-layer features has become the key to improving the segmentation accuracy and robustness of the model. Summary of the invention
[0007] In order to overcome the above-mentioned defects of the prior art, the implementation of the present invention provides a polyp image segmentation method based on a boundary-guided multi-level attention network to solve the problems raised in the above-mentioned background technology.
[0008] A polyp image segmentation method based on a boundary-guided multi-level attention network, comprising:
[0009] Step S1: constructing a data set, the data set including a number of polyp images;
[0010] Step S2: Construct a boundary-guided multi-level attention network. The boundary-guided multi-level attention network includes a pyramid vision Transformer-PVT-v2 backbone, a multi-scale feature extraction module, a parallel partial decoder, a boundary perception module, and a boundary-guided multi-level attention module. Import the image of the data set in step S1 into the pyramid vision Transformer-PVT-v2 backbone to obtain the multi-level features of the image, which are the first multi-level features , the second multi-level feature , the third multi-level feature and the fourth multi-level feature ;
[0011] Step S3: Import the multi-level features of the image in step S2 into the multi-scale feature extraction module to obtain the multi-scale features of the compression channel, which are the first multi-scale features , the second multi-scale feature , the third multi-scale feature and the fourth multi-scale feature ;
[0012] Step S4: Import the multi-scale features of step S3 into the parallel partial decoder to obtain the global feature map ;
[0013] Step S5: The global feature map of step S4 is successively processed by the boundary perception module And the multi-scale features of step S3 , and Use convolution and addition operations to obtain aggregate features, which are the second aggregate features , the third aggregate feature and the fourth aggregate feature ; The first multi-scale feature in step S3 and the second aggregate feature Fusion to obtain boundary perception features ;
[0014] Step S6: The multi-scale features of step S3 and the global feature map of step S4 are integrated through a multi-level attention module based on boundary guidance. , Boundary perception features of step S5 Perform multi-level attention feature enhancement to obtain segmentation feature maps that focus on different areas of the polyp image , segmentation feature map and segmentation feature map ;
[0015] Step S7: Construct a loss function and minimize the loss function to optimize the parameters of the boundary-guided multi-level attention network.
[0016] Furthermore, the multi-level features of the image are obtained in step S2, specifically: the image of the data set in step S1 is imported , obtain multi-level features through the pyramid vision Transformer-PVT-v2 backbone ,in, represents the set of real numbers, Indicates the height of the image. Indicates the width of the image. ,and is the number of channels in each layer.
[0017] Furthermore, in step S3, multi-scale features are obtained, specifically:
[0018] The receptive field block of the constructed multi-scale feature extraction module compresses the number of channels of the multi-level features generated in step S2 to 64 using a convolution operation with a convolution kernel size of 1;
[0019] Then the receptive field block uses three parallel convolution operations with a kernel size of 3 and dilation rates of 3, 5, and 7 for each multi-level feature to obtain features of different scales, and splices these multi-scale features in the channel dimension. The spliced features are further fused with multi-scale features through a convolution operation with a kernel size of 3.
[0020] For four multi-level features , , and Perform this operation separately to generate multi-scale features , , and .
[0021] Furthermore, step S5 is specifically as follows:
[0022] First, the global feature map With multi-scale features , and Gradually add to supplement the spatial features lost during the downsampling process, and finally generate the second aggregate feature ;
[0023] Then, in order to preserve the spatial information, the first multi-scale feature and the second aggregate feature The spliced features are spliced in the channel dimension, and the spliced features are subjected to a convolution operation with a convolution kernel size of 1 for feature interaction, which is expressed as:
[0024] ;
[0025] ;
[0026] in, represents concatenation in the channel dimension, represents a convolution operation with a convolution kernel size of 1. Represents the splicing feature, Represents convolutional splicing features;
[0027] After that, the attention map is calculated in the spatial dimension and channel dimension respectively;
[0028] In the spatial branch, three parallel convolutions with different convolution kernel sizes are used to enhance the multi-scale feature representation of the model and obtain enhanced convolution splicing features. Spatial attention is applied to further refine the boundary details provided by the enhanced convolution splicing features. In the channel branch, channel attention is used to suppress the noise introduced by the convolution splicing features.
[0029] The obtained spatial attention map and channel attention map are multiplied with the original features respectively to generate a spatial branch feature map and a channel branch feature map. The spatial branch feature map and the channel branch feature map are added and fused to generate a boundary perception feature. , expressed as:
[0030] ;
[0031] ;
[0032] ;
[0033] ;
[0034] in, represents a convolution operation with a convolution kernel size of 3. represents a convolution operation with a convolution kernel size of 5. represents the enhanced convolutional splicing feature, represents the spatial branch feature map, represents a convolution operation with a convolution kernel size of 7. represents the addition operation, represents spatial attention, represents channel attention, represents matrix multiplication, Represents the channel branch feature map.
[0035] Further, step S6 is specifically as follows: the multi-level attention module based on boundary guidance is Calculate the reverse feature :
[0036]
[0037] in, represents the sigmoid activation function, Represents the matrix The reverse operation of subtracting the input;
[0038] Then, the predicted features , reverse feature and boundary-aware features Upsampling to multi-scale features The same spatial resolution; then, the predicted features , reverse feature , boundary-aware features and multi-scale features Multiplying them together produces three attention maps focusing on different areas, expressed as:
[0039] ;
[0040] ;
[0041] ;
[0042] in, Represented as an upsampling operation, , ,and They are the attention maps of predicted features, reverse features, and boundary-aware features;
[0043] The three attention maps are concatenated along the channel dimension and subjected to a convolution operation with a convolution kernel size of 1 for feature interaction to generate a combined feature. , expressed as:
[0044] ;
[0045] Then, the binding characteristics are calculated The attention map is combined with the feature Multiply to further highlight the polyp features, and use a residual connection to fuse the features from the encoder to restore the lost spatial details and semantic information, expressed as:
[0046] ;
[0047] in, Represents the combined features Attention map of
[0048] Finally, a CBAM module and a convolution operation with a convolution kernel size of 1 are used to suppress noise in the features, reduce information redundancy, and enhance the model's attention to the polyp area, expressed as:
[0049] ;
[0050] ;
[0051] in, represents the predicted segmentation feature map, Represents the suppression operation of the CBAM module.
[0052] Furthermore, step S7 is specifically as follows: the total loss function is expressed as:
[0053] ;
[0054] in, Represents the true mask map of the polyp image, represents the ground truth mask of the boundary obtained by applying the Canny edge detection algorithm to each polyp image, is a combination of weighted intersection-over-union loss and weighted binary cross entropy loss for region segmentation of polyp images, is the weighted binary cross entropy loss for polyp boundary segmentation.
[0055] Compared with the prior art, the present invention has the following beneficial effects:
[0056] 1. The present invention designs a boundary perception module that effectively aggregates low-level features and high-level global features to generate fine polyp boundaries; the module first preliminarily fuses low-level features and high-level global features, and then dynamically selects target polyp regions through attention mechanisms in the spatial dimension and channel dimension, respectively. In the spatial dimension, in order to improve the model's perception of polyps of different sizes, multi-scale feature extraction is used; then, the spatial attention features and channel attention features are further fused to generate clear polyp boundaries. This method effectively fuses low-level features and high-level global features, realizes the complementarity of low-level and high-level features, and effectively solves the impact of fuzzy polyp boundaries on the model.
[0057] 2. The present invention provides a multi-level attention module based on boundary guidance, which performs multi-level attention feature enhancement and fusion on the multi-scale features from the encoder, the prediction results of the previous level and the fine boundary perception features, so as to improve the model's refinement of the polyp boundary and the segmentation effect of the polyp main area; the module uses reverse features to improve the model's attention to the polyp boundary, and introduces boundary perception features to refine the model's processing of the polyp boundary. By fusing multiple attention features, the relationship between the polyp main body and the polyp boundary can be further explored, and the segmentation accuracy of the model can be enhanced. Compared with the prior art, this method can more accurately locate the polyp area when processing polyp images, and generate polyp segmentation results with finer boundaries.
[0058] 3. Compared with the existing technology, this method has better learning ability and generalization performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 This is a flowchart of the steps of a polyp image segmentation method based on a boundary-guided multi-level attention network of the present invention. DETAILED DESCRIPTION
[0060] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0061] Reference Figure 1 , a polyp image segmentation method based on boundary-guided multi-level attention network, comprising:
[0062] Step S1: constructing a data set, the data set including a number of polyp images;
[0063] Step S2: Construct a boundary-guided multi-level attention network. The boundary-guided multi-level attention network includes a pyramid vision Transformer-PVT-v2 backbone, a multi-scale feature extraction module, a parallel partial decoder, a boundary perception module, and a boundary-guided multi-level attention module. Import the image of the data set in step S1 into the pyramid vision Transformer-PVT-v2 backbone to obtain the multi-level features of the image, which are the first multi-level features , the second multi-level feature , the third multi-level feature and the fourth multi-level feature ;
[0064] Step S3: Import the multi-level features of the image in step S2 into the multi-scale feature extraction module to obtain the multi-scale features of the compression channel, which are the first multi-scale features , the second multi-scale feature , the third multi-scale feature and the fourth multi-scale feature ;
[0065] Step S4: Import the multi-scale features of step S3 into the parallel partial decoder to obtain the global feature map ;
[0066] Step S5: The global feature map of step S4 is successively processed by the boundary perception module And the multi-scale features of step S3 , and Use convolution and addition operations to obtain aggregate features, which are the second aggregate features , the third aggregate feature and the fourth aggregate feature ; The first multi-scale feature in step S3 and the second aggregate feature Fusion to obtain boundary perception features ;
[0067] Step S6: The multi-scale features of step S3 and the global feature map of step S4 are integrated through a multi-level attention module based on boundary guidance. , Boundary perception features of step S5 Perform multi-level attention feature enhancement to obtain segmentation feature maps that focus on different areas of the polyp image , segmentation feature map and segmentation feature map ;
[0068] Step S7: Construct a loss function and minimize the loss function to optimize the parameters of the boundary-guided multi-level attention network.
[0069] Furthermore, the multi-level features of the image are obtained in step S2, specifically: the image of the data set in step S1 is imported , obtain multi-level features through the pyramid vision Transformer-PVT-v2 backbone ,in, represents the set of real numbers, Indicates the height of the image. Indicates the width of the image. ,and is the number of channels in each layer.
[0070] Furthermore, in step S3, multi-scale features are obtained, specifically:
[0071] The receptive field block of the constructed multi-scale feature extraction module compresses the number of channels of the multi-level features generated in step S2 to 64 using a convolution operation with a convolution kernel size of 1;
[0072] Then the receptive field block uses three parallel convolution operations with a kernel size of 3 and dilation rates of 3, 5, and 7 for each multi-level feature to obtain features of different scales, and splices these multi-scale features in the channel dimension. The spliced features are further fused with multi-scale features through a convolution operation with a kernel size of 3.
[0073] For four multi-level features , , and Perform this operation separately to generate multi-scale features , , and .
[0074] Furthermore, step S5 is specifically as follows:
[0075] First, the global feature map With multi-scale features , and Gradually add to supplement the spatial features lost during the downsampling process, and finally generate the second aggregate feature ;
[0076] Then, in order to preserve the spatial information, the first multi-scale feature and the second aggregate feature The spliced features are spliced in the channel dimension, and the spliced features are subjected to a convolution operation with a convolution kernel size of 1 for feature interaction, which is expressed as:
[0077] ;
[0078] ;
[0079] in, represents splicing in the channel dimension, represents a convolution operation with a convolution kernel size of 1. Represents the splicing feature, Represents convolutional splicing features;
[0080] After that, the attention map is calculated in the spatial dimension and channel dimension respectively;
[0081] In the spatial branch, three parallel convolutions with different convolution kernel sizes are used to enhance the multi-scale feature representation of the model and obtain enhanced convolution splicing features. Spatial attention is applied to further refine the boundary details provided by the enhanced convolution splicing features. In the channel branch, channel attention is used to suppress the noise introduced by the convolution splicing features.
[0082] The obtained spatial attention map and channel attention map are multiplied with the original features respectively to generate a spatial branch feature map and a channel branch feature map. The spatial branch feature map and the channel branch feature map are added and fused to generate a boundary perception feature. , expressed as:
[0083] ;
[0084] ;
[0085] ;
[0086] ;
[0087] in, represents a convolution operation with a convolution kernel size of 3. represents a convolution operation with a convolution kernel size of 5. represents the enhanced convolutional splicing feature, represents the spatial branch feature map, represents a convolution operation with a convolution kernel size of 7. represents the addition operation, represents spatial attention, represents channel attention, represents matrix multiplication, Represents the channel branch feature map.
[0088] Further, step S6 is specifically as follows: the multi-level attention module based on boundary guidance is Calculate the reverse feature :
[0089]
[0090] in, represents the sigmoid activation function, Represents the matrix The reverse operation of subtracting the input;
[0091] Then, the predicted features , reverse feature and boundary-aware features Upsampling to multi-scale features The same spatial resolution; then, the predicted features , reverse feature , boundary-aware features and multi-scale features Multiplying them together produces three attention maps focusing on different areas, expressed as:
[0092] ;
[0093] ;
[0094] ;
[0095] in, Represented as an upsampling operation, , ,and They are the attention maps of predicted features, reverse features, and boundary-aware features;
[0096] The three attention maps are concatenated along the channel dimension and subjected to a convolution operation with a convolution kernel size of 1 for feature interaction to generate a combined feature. , expressed as:
[0097] ;
[0098] Then, the binding characteristics are calculated The attention map is combined with the feature Multiply to further highlight the polyp features, and use a residual connection to fuse the features from the encoder to restore the lost spatial details and semantic information, expressed as:
[0099] ;
[0100] in, Represents the combined features Attention map of
[0101] Finally, a CBAM module and a convolution operation with a convolution kernel size of 1 are used to suppress noise in the features, reduce information redundancy, and enhance the model's attention to the polyp area, expressed as:
[0102] ;
[0103] ;
[0104] in, represents the predicted segmentation feature map, Represents the suppression operation of the CBAM module.
[0105] Furthermore, step S7 is specifically as follows: the total loss function is expressed as:
[0106] ;
[0107] in, Represents the true mask map of the polyp image, represents the ground truth mask of the boundary obtained by applying the Canny edge detection algorithm to each polyp image, is a combination of weighted intersection-over-union loss and weighted binary cross entropy loss for region segmentation of polyp images, is the weighted binary cross entropy loss for polyp boundary segmentation.
[0108] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A polyp image segmentation method based on a boundary-guided multi-level attention network, characterized in that: include: Step S1: constructing a data set, the data set including a number of polyp images; Step S2: Construct a boundary-guided multi-level attention network. The boundary-guided multi-level attention network includes a pyramid vision Transformer-PVT-v2 backbone, a multi-scale feature extraction module, a parallel partial decoder, a boundary perception module, and a boundary-guided multi-level attention module. Import the image of the data set in step S1 into the pyramid vision Transformer-PVT-v2 backbone to obtain the multi-level features of the image, which are the first multi-level features , the second multi-level feature , the third multi-level feature and the fourth multi-level feature ; Step S3: Import the multi-level features of the image in step S2 into the multi-scale feature extraction module to obtain the multi-scale features of the compression channel, which are the first multi-scale features , the second multi-scale feature , the third multi-scale feature and the fourth multi-scale feature ; Step S4: Import the multi-scale features of step S3 into the parallel partial decoder to obtain the global feature map ; Step S5: The global feature map of step S4 is successively processed by the boundary perception module And the multi-scale features of step S3 , and Use addition operation to obtain the second aggregate feature ; The first multi-scale feature in step S3 and the second aggregate feature Fusion to obtain boundary perception features ; Step S6: The boundary-aware features of step S5 are integrated through a boundary-guided multi-level attention module Perform multi-level attention feature enhancement to obtain segmentation feature maps that focus on different areas of the polyp image , segmentation feature map and segmentation feature map ; Step S7: construct a loss function, and minimize the loss function to optimize the parameters of the boundary-guided multi-level attention network; Step S5 is specifically as follows: First, the global feature map With multi-scale features , and Gradually add to supplement the spatial features lost during the downsampling process, and finally generate the second aggregate feature ; Then, in order to preserve the spatial information, the first multi-scale feature and the second aggregate feature The spliced features are spliced in the channel dimension, and the spliced features are subjected to a convolution operation with a convolution kernel size of 1 for feature interaction, which is expressed as: ; ; in, represents concatenation in the channel dimension, represents a convolution operation with a convolution kernel size of 1. Represents the splicing feature, Represents convolutional splicing features; After that, the attention map is calculated in the spatial dimension and channel dimension respectively; In the spatial branch, three parallel convolutions with different convolution kernel sizes are used to enhance the multi-scale feature representation of the model and obtain enhanced convolution splicing features. Spatial attention is applied to further refine the boundary details provided by the enhanced convolution splicing features. In the channel branch, channel attention is used to suppress the noise introduced by the convolution splicing features. The obtained spatial attention map and channel attention map are multiplied with the original features respectively to generate a spatial branch feature map and a channel branch feature map. The spatial branch feature map and the channel branch feature map are added and fused to generate a boundary perception feature. , expressed as: ; ; ; ; in, represents a convolution operation with a convolution kernel size of 3. represents a convolution operation with a convolution kernel size of 5. represents the enhanced convolutional splicing feature, represents the spatial branch feature map, represents a convolution operation with a convolution kernel size of 7. represents the addition operation, represents spatial attention, represents channel attention, represents matrix multiplication, Represents the channel branch feature map; Step S6 is as follows: first, the boundary-guided multi-level attention module predicts features Calculate the reverse feature : ; in, represents the sigmoid activation function, Represents the matrix The reverse operation of subtracting the input; Then, the predicted features , reverse feature and boundary-aware features Upsampling to multi-scale features The same spatial resolution; then, the predicted features , reverse feature , boundary-aware features and multi-scale features Multiplying them together produces three attention maps focusing on different areas, expressed as: ; ; ; in, Represented as an upsampling operation, , and They are the attention maps of predicted features, reverse features, and boundary-aware features; The three attention maps are concatenated along the channel dimension and subjected to a convolution operation with a convolution kernel size of 1 for feature interaction to generate a combined feature. , expressed as: ; Then, the binding characteristics are calculated The attention map is combined with the feature Multiply to further highlight the polyp features, and use a residual connection to fuse the features from the encoder to restore the lost spatial details and semantic information, expressed as: ; in, Represents the combined features Attention map of Finally, the multi-level attention module successively suppresses the noise in the features through the CBAM module and the convolution operation with a convolution kernel size of 1, reduces information redundancy, and enhances the model's attention to the polyp area, which is expressed as: ; ; in, represents the predicted segmentation feature map, Represents the suppression operation of the CBAM module.
2. The polyp image segmentation method based on boundary-guided multi-level attention network according to claim 1 is characterized in that: The multi-level features of the image in step S2 are as follows: import the image of the data set in step S1 , obtain multi-level features through the pyramid vision Transformer-PVT-v2 backbone ,in, represents the set of real numbers, Indicates the height of the image. Indicates the width of the image. ,and is the number of channels in each layer.
3. The polyp image segmentation method based on boundary-guided multi-level attention network according to claim 2 is characterized in that: In step S3, multi-scale features are obtained, specifically: The receptive field block of the constructed multi-scale feature extraction module compresses the number of channels of the multi-level features generated in step S2 to 64 using a convolution operation with a convolution kernel size of 1; Then the receptive field block uses three parallel convolution operations with a kernel size of 3 and dilation rates of 3, 5, and 7 for each multi-level feature to obtain features of different scales, and splices these multi-scale features in the channel dimension. The spliced features are further fused with multi-scale features through a convolution operation with a kernel size of 3. For four multi-level features , , and Perform this operation separately to generate multi-scale features , , and .
4. The polyp image segmentation method based on boundary-guided multi-level attention network according to claim 1, characterized in that: Step S7 is specifically: loss function It is expressed as: ; in, Represents the true mask map of the polyp image, represents the ground truth mask of the boundary obtained by applying the Canny edge detection algorithm to each polyp image, is a combination of weighted intersection-over-union loss and weighted binary cross entropy loss for region segmentation of polyp images, is the weighted binary cross entropy loss for polyp boundary segmentation.
Citation Information
Patent Citations
Medical image segmentation method and device, computer equipment and storage medium
CN114419020A
Medical image segmentation method based on boundary perception and attention mechanism
CN117078930A