A method, device and storage medium for colon polyp segmentation
Through the boundary guiding of the bidirectional attention residual network, the problem of camouflage and diversity in colon polyps segmentation is solved, high-precision colon polyps segmentation is achieved, and the automation and accuracy of colonoscopy detection is improved.
Patent Information
- Application Number
- CN202310481889.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-04-28
AI Technical Summary
The existing colon polyps segmentation method has high similarity with the surrounding mucosal tissue and blurred boundaries, which leads to strong camouflage, making it difficult to accurately identify and segment, and has large differences in size and shape, making it difficult to completely segment.
Using a method of guiding the bidirectional attention residual network based on the boundary, multi-level encoding features are extracted through feature encoder, and boundary attention maps are generated by combining the context enrichment module and the boundary extractor. The bidirectional attention residual module is used to enhance the foreground area and suppress background noise, and improve segmentation accuracy.
High-quality colon polyps segmentation has been achieved in a variety of challenging scenarios, improving the detection ability of polyps of different morphology and sizes, and reducing the misdiagnosis rate.
Smart Images

Figure CN116542921B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer-aided medicine, and particularly to a colon polyp segmentation method, device and storage medium based on a boundary-guided bidirectional attention residual network. Background Art
[0002] Currently, colorectal cancer has become the third most common cancer worldwide and the second most lethal cancer. More than 90% of colorectal cancer cases evolve from colon polyps, which form on the inner wall of the colon as small, non-cancerous benign cell masses that carry a risk of transforming into colorectal cancer over time. Therefore, the most effective way to prevent and treat colorectal cancer is to detect and remove colon polyps before they turn into colorectal cancer. Currently, colonoscopy is the most commonly used method for detecting colon polyps, but this process requires a large amount of expensive manual labor and has a high misdiagnosis rate. Therefore, an automatic and accurate polyp segmentation method has important practical value and broad application prospects.
[0003] Traditional polyp segmentation methods mainly rely on manually selected features such as color, texture, shape, etc. Although these methods have a certain effect on polyp segmentation, due to their lack of the ability to express high-level semantic information, they often lead to a high probability of missed and misjudged cases. In recent years, with the wide application of deep learning technology in the field of medical image analysis, many deep learning-based polyp segmentation methods have emerged. Early methods usually added modules based on FCN to improve the accuracy of colon polyp segmentation, but since FCN relied too much on high-level features and ignored the rich detailed features contained in low-level features, the segmentation results were relatively rough. With the great success of U-Net in the field of biomedical images, more methods have adopted the design idea of U-Net, that is, using a deep convolutional neural network as an encoder to extract multi-level features, and then designing a decoder according to the top-down principle to fuse multi-level features and predict the final segmentation result.
[0004] Although these methods have made great progress, the segmentation of colon polyps remains a challenging task, mainly for the following two reasons:
[0005] First, the color and texture of colon polyps are very similar to the surrounding mucosal tissue, and the boundary is very blurred, resulting in a high degree of camouflage of colon polyps and making it difficult to identify them; second, the size, shape, position, etc. of colon polyps vary greatly, making it difficult for the model to completely segment all polyps of different shapes. Due to the existence of the above two problems, the accuracy of colon polyp segmentation methods still needs to be improved. Therefore, the present invention aims to solve these two problems in the colon polyp segmentation task. Summary of the Invention
[0006] Aiming at the problems of camouflage and blurred boundaries caused by the high similarity between colon polyps and surrounding mucosal tissues, the present invention provides a colon polyp segmentation method, device and storage medium based on a boundary-guided bidirectional attention residual network. The bidirectional attention mechanism is used to synchronously enhance the foreground region and suppress the high-similarity noise in the background, thereby increasing the contrast between the foreground and the background. At the same time, a boundary guidance module is introduced to enhance the boundary region. In addition, the present invention also proposes a context-rich module, which adaptively extracts global information, local information, and multi-scale information to improve the expression ability of the initial encoded features, and further improve the detection ability of polyps of different shapes and sizes, and finally achieve high-quality colon polyp segmentation in different scenarios.
[0007] To this end, the present invention provides the following technical solutions:
[0008] The present invention provides a colon polyp segmentation method based on a boundary-guided bidirectional attention residual network, including the following steps:
[0009] A. Input the colonoscopy image into the feature encoder to obtain multi-level encoded features (F i , i = 2, 3, 4, 5). After the last layer of the feature encoder, an average pooling layer is concatenated to obtain the top-level feature F6;
[0010] B. Input the features obtained in step A into the context-rich module to obtain a richer feature representation (B i , i = 2, 3, 4, 5, 6), including:
[0011] B1. For each level of features (F i , i = 2, 3, 4, 5), a global-local feature extractor is used to synchronously extract global features and local features and fuse them to obtain (B i , i = 2, 3, 4, 5), thereby improving the polyp localization ability and detail description ability;
[0012] B2. For the top-level feature F6, a U-shaped dilated convolutional spatial pyramid pooling module is used to extract multi-scale features B6, thereby improving the detection ability of polyps of different sizes;
[0013] C. At the same time, the multi-level encoded features (F i , i = 2, 3, 4, 5) obtained in step A are input into a designed boundary extractor to generate a boundary attention map E, supervised by the boundary ground truth;
[0014] D. Through a series of boundary-guided bidirectional attention residual modules, the richer feature representation obtained in step B is gradually decoded according to the principle from high to low to obtain the fused feature (R i, where \(i = 2, 3, 4, 5\)), the fused feature and B6 are respectively passed through the prediction layer to obtain multi-stage prediction maps (S i , where \(i = 2, 3, 4, 5, 6\)) and are supervised with the ground truth, and S2 is selected as the final segmentation result. The boundary-guided bidirectional attention residual module includes:
[0015] D1. Adopt a bidirectional attention mechanism to complement and enhance the features of adjacent two layers, making the foreground regions of the high-level features more complete and suppressing the noise interference in the background regions of the low-level features as much as possible.
[0016] D2. First, the adjacent features processed by the bidirectional attention mechanism are processed by a multi-scale channel attention module for dynamic region weighting, and then are respectively input into the boundary enhancement module. The boundary clues collected in step C are used to enhance their boundary regions. Finally, the two-way features are fused to obtain the fused feature of the current stage.
[0017] D3. The fused feature is redistributed in channels and space through a shuffle attention module, and then the output of the current boundary-guided bidirectional attention residual module is obtained by adding the residual structure and the high-level feature.
[0018] Furthermore, step A includes:
[0019] The feature encoder is of Res2Net architecture, and the last two layers are discarded to retain the spatial structure; the feature encoder generates 4 feature maps (F i , where \(i = 2, 3, 4, 5\)) with different spatial resolutions and channel numbers for each image.
[0020] Furthermore, the Res2Net architecture is Res2Net-50 architecture. The encoder is initialized with the pre-trained weights of Res2Net-50 on ImageNet. While deleting the last fully connected layer, an average pooling layer with a pooling kernel of \(3\times3\) and a stride of 2 is added after the last convolutional layer, and F5 is downsampled to obtain the top-level feature F6.
[0021] Furthermore, step B1 includes:
[0022] The multi-level encoded features (F i , where \(i = 2, 3, 4, 5\)) are passed through a transformer branch to obtain global features. At the same time, the multi-level encoded features (F i , where \(i = 2, 3, 4, 5\)) are passed through a globally-guided CNN branch to obtain local features. Finally, the global features and the local features are fused through channel stacking and a \(3\times3\) convolutional layer to generate enhanced features (B i , where \(i = 2, 3, 4, 5\)).
[0023] Further, step B2 includes:
[0024] The U-shaped dilated convolutional spatial pyramid pooling introduces a complementary path from small receptive fields to large receptive fields and a guiding path from large receptive fields to small receptive fields. The former enables the top-level feature F6 to superimpose the output results of the dilated convolutional layer with a smaller dilation rate each time it passes through a dilated convolutional layer with a specific dilation rate; the latter enables the features obtained by processing through dilated convolutional layers with different dilation rates to be guided by the output results of the dilated convolutional layer with a larger dilation rate before fusion, and the guiding method is the combination of channel superposition and spatial attention mechanism.
[0025] Further, the dilation rates d of the dilated convolutional layers are taken as 1, 2, 3, and 4 respectively.
[0026] Further, step C includes:
[0027] For the encoded features (F i , i = 2, 3, 4, 5) of each level, they are transformed to the same spatial resolution by using a 1×1 convolutional layer and an upsampling operation, then superimposed in the channel direction, and then fused by using a 3×3 convolutional layer, and finally a boundary attention map is obtained through a 1×1 convolutional layer.
[0028] Further, step D1 includes:
[0029] First, the enhanced features of two adjacent layers are fused through a 3×3 convolutional layer, then a foreground attention map is generated through a 1×1 convolutional layer and a sigmoid activation function, and then the background attention map is obtained by subtracting the foreground attention map from 1 pixel by pixel. The foreground attention map and the background attention map are respectively used to enhance the foreground of the high-level features and filter out the background noise of the low-level features.
[0030] Further, the boundary guiding mechanism in step D2 is: using the boundary attention map to weight the feature to be enhanced, then adding it to the original unweighted feature pixel by pixel, and obtaining the final boundary-enhanced feature through a 3×3 convolutional layer.
[0031] The above technical solutions provided by the present invention have the following beneficial effects:
[0032] The present invention proposes a colon polyp segmentation method based on boundary-guided bidirectional attention residual network, which takes into account the camouflage and diversity of colon polyps in the colon polyp segmentation task. First, the colonoscopy image is input into the feature encoder to obtain multi-level encoding features; then, for the encoding features, a richer feature representation is obtained through the context enrichment module to adapt to the polyp segmentation scenarios of different sizes, shapes and positions, wherein for the top-level features, a U-shaped dilated convolution space pyramid pooling module is used to extract multi-scale features, thereby improving the detection ability of multi-size polyps, and for each level of features, a global-local feature extractor is used to synchronously extract global features and local features and fuse them, thereby improving the positioning ability and detail characterization ability of polyps; at the same time, the present invention designs a boundary extractor, which generates boundaries through multi-level encoding features for boundary enhancement in the subsequent boundary-guided bidirectional attention residual module; next, the enhanced features obtained by the context enrichment module are dynamically fused according to the criteria from high to low levels through a series of boundary-guided bidirectional attention residual modules, and the boundaries are enhanced at the same time, thereby enhancing the capture ability of camouflaged polyps and improving the segmentation accuracy of the boundary area. Experimental results show that the colon polyp segmentation method based on boundary-guided bidirectional attention residual network proposed in this paper can achieve accurate segmentation results in many challenging polyp segmentation scenarios.
[0033] Based on the above reasons, the present invention can be widely promoted in the field of computer-assisted medical treatment. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0035] Figure 1 is a schematic diagram of the polyp segmentation scenario;
[0036] Figure 2 It is a flow chart of a method for segmenting colon polyps based on a boundary-guided bidirectional attention residual network according to an embodiment of the present invention;
[0037] Figure 3 Schematic diagram of the structure of the global-local feature extractor in an embodiment of the present invention.
[0038] Figure 4 Schematic diagram of the structure of the U-shaped dilated convolution spatial pyramid pooling module in an embodiment of the present invention.
[0039] Figure 5It is a schematic structural diagram of the boundary-guided bidirectional attention residual module in the embodiments of the present invention. Detailed implementation manners
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] Refer to Figure 2 , which shows a flowchart of a colon polyp segmentation method based on a boundary-guided bidirectional attention residual network in the embodiments of the present invention. The method includes the following steps:
[0042] A. Input a colonoscopy image into a feature encoder to obtain multi-level encoded features (F i , i = 2, 3, 4, 5), and concatenate an average pooling layer after the last layer of the feature encoder to obtain a top-level feature F6;
[0043] The input colonoscopy image is as Figure 1 shown.
[0044] B. Input the features obtained in step A into a context enrichment module to obtain a richer feature representation (B i , i = 2, 3, 4, 5, 6).
[0045] In specific implementation, step B includes:
[0046] B1. For each level of features (F i , i = 2, 3, 4, 5), respectively use a global-local feature extractor to synchronously extract global features and local features and fuse them to obtain (B i , i = 2, 3, 4, 5), thereby improving the polyp localization ability and detail description ability. Specifically, as Figure 3 shown:
[0047] Pass the multi-level encoded features (F i , i = 2, 3, 4, 5) through a transformer branch to obtain global features, which are expressed as follows:
[0048]
[0049] where MHSA(·) represents the multi-head self-attention mechanism, LN(·) represents layer normalization, and MLP(·) represents the multi-layer perceptron. At the same time, pass the multi-level encoded features (F i, where \(i = 2, 3, 4, 5\), obtain local features through a globally-guided CNN branch, expressed as follows:
[0050]
[0051] where \(GSC(\cdot)\) represents group normalization, SiLU activation function, and a \(3\times3\) convolutional layer. is the global attention map, obtained by through a \(1\times1\) convolutional layer and a sigmoid activation function. Finally, the global features and local features are fused through channel stacking and a \(3\times3\) convolutional layer to generate enhanced features (\(B\) i , where \(i = 2, 3, 4, 5\)).
[0052] B2. For the top-level feature \(F6\), a U-shaped dilated convolutional spatial pyramid pooling module is used to extract multi-scale features \(B6\), thereby improving the detection ability for polyps of different sizes. Specifically, as Figure 4 shown:
[0053] The U-shaped dilated convolutional spatial pyramid pooling introduces two paths based on the classical dilated convolutional spatial pyramid pooling method: a complementary path from small receptive fields to large receptive fields and a guiding path from large receptive fields to small receptive fields. The former makes the top-level feature \(F6\) add the output result of the dilated convolutional layer with a dilation rate one level smaller when passing through the dilated convolutional layer with a certain dilation rate, which can be expressed as:
[0054]
[0055] where represents a \(3\times3\) dilated convolution with a dilation rate of \(d\), \(C1(\cdot)\) represents a \(1\times1\) convolution, and \(t\) represents the index of the branch. The latter makes the feature \(f'\) t guided by the output result of the dilated convolutional layer with a dilation rate one level larger before fusion, and the guiding method is the combination of channel stacking and spatial attention mechanism, which can be expressed as:
[0056]
[0057] where \(Cat(\cdot)\) represents concatenation in the channel dimension, \(SA(\cdot)\) represents the spatial attention mechanism, and \(f\) t represents the multi-branch feature generated through guidance. Finally, the multi-scale feature \(B\) b is generated through the following steps:
[0058] \(B6 = C1(Cat(f1, f2, f3, f4, f5)) + C1(F6)\);
[0059] C. The multi-level encoded features (\(F\) obtained in step A i, for \(i = 2, 3, 4, 5\), input a designed boundary extractor to generate a boundary attention map \(E\), supervised by the boundary ground truth.
[0060] In a specific implementation, step C includes:
[0061] For the encoded features of each level (\(F\) i , for \(i = 2, 3, 4, 5\)), use a \(1\times1\) convolutional layer and an upsampling operation to transform to the same spatial resolution, then stack them in the channel direction, and then use a \(3\times3\) convolutional layer for fusion, and finally obtain the boundary attention map through a \(1\times1\) convolutional layer, which can be expressed as:
[0062]
[0063] Among them, \(Up(\cdot)\) represents upsampling, and \(\sigma(\cdot)\) represents the Sigmoid activation function.
[0064] D. Through a series of boundary-guided bidirectional attention residual modules, gradually perform feature decoding on the richer feature representations obtained in step B according to the criterion from high to low to obtain fused features (\(R\) i , for \(i = 2, 3, 4, 5\)), pass the fused features and \(B6\) through the prediction layer to obtain multi-stage prediction maps (\(S\) i , for \(i = 2, 3, 4, 5, 6\)) and supervise them with the ground truth, and select \(S2\) as the final segmentation result.
[0065] Among them, the boundary-guided bidirectional attention residual module includes:
[0066] D1. Adopt a bidirectional attention mechanism to complement and enhance the features of adjacent two layers, making the foreground region of the high-level features more complete and suppressing the noise interference in the background region of the low-level features as much as possible, specifically as Figure 5 shown.
[0067] First, fuse the adjacent two-layer enhanced features (\(R\) i+1 and \(B\) i , where \(R\) i+1 is the output of the previous boundary-guided bidirectional attention residual module) through a \(3\times3\) convolutional layer to obtain a fused feature Then generate a foreground attention map through a \(1\times1\) convolutional layer and the Sigmoid activation function and subtract the foreground attention map from 1 pixel by pixel to obtain a background attention map The process can be expressed as:
[0068]
[0069]
[0070] Next, the foreground attention map and the background attention map are used respectively to enhance the high-level feature R i+1 and filter out the low-level feature B i of background noise. The process can be expressed as:
[0071]
[0072] D2. For the adjacent features processed by the bidirectional attention mechanism and First, dynamic region weighting is performed through the multi-scale channel attention module, and then they are respectively input into the boundary enhancement module, and the boundary clues E collected in step C are used to enhance their boundary regions.
[0073]
[0074]
[0075] where EG i (·) represents the boundary guidance mechanism, which is specifically as follows:
[0076]
[0077] D3. The features after boundary guidance and are fused and weighted adjustment in channels and space is performed through a shuffle attention module (SHA). Then, through the residual structure plus the high-level features, it is the output R of the current boundary-guided bidirectional attention residual module i , and the process can be expressed as:
[0078]
[0079] E. Training and optimization of the boundary-guided bidirectional attention residual network:
[0080] The overall method can be divided into two stages: training and inference. During training, the tensors of the training set are used as inputs to obtain the trained network parameters; during the inference stage, the parameters saved in the training stage are used for testing to obtain the final colon polyp segmentation results.
[0081] This embodiment of the present invention is implemented under the Pytorch framework. During the training stage, the Adam optimizer is used, the learning rate is 5e -5 , β1 = 0.9, β2 = 0.999, and the batch size is 8. During training, the spatial resolution of the images is 352×352, and the model must also be input at a resolution of 352×352 during testing.
[0082] The colon polyp segmentation method based on the boundary-guided bidirectional attention residual network proposed in the embodiments of the present invention uses a bidirectional attention mechanism to synchronously enhance the foreground region and suppress high-similarity noise in the background, thereby increasing the contrast between the foreground and the background. At the same time, a boundary-guided module is introduced to enhance the boundary region. In addition, the present invention also proposes a context enrichment module, which adaptively extracts global information, local information, and multi-scale information to improve the expression ability of the initial encoded features, and further improve the detection ability for polyps of different shapes and sizes. Experimental results show that the colon polyp segmentation method based on the boundary-guided bidirectional attention residual network proposed in the present invention can obtain accurate prediction results for many challenging colon polyp segmentation scenarios.
[0083] Corresponding to the colon polyp segmentation method based on the boundary-guided bidirectional attention residual network in the above embodiment, the embodiments of the present invention also provide a colon polyp segmentation device based on the boundary-guided bidirectional attention residual network, including:
[0084] An encoding unit, configured to input a colonoscopy image into a feature encoder to obtain multi-level encoded features, and concatenate an average pooling layer after the last layer of the feature encoder to obtain top-level features;
[0085] A feature enhancement unit, configured to perform context feature enhancement on the multi-level encoded features and top-level features obtained by the encoding unit to obtain rich feature representations and multi-scale features;
[0086] A boundary extraction unit, configured to input the multi-level encoded features obtained by the encoding unit into a boundary extractor to generate a boundary attention map supervised by the boundary ground truth;
[0087] A prediction unit, configured to gradually perform feature decoding on the rich feature representations obtained by the feature enhancement unit according to the principle from high to low through a series of boundary-guided bidirectional attention residual modules to obtain fused features, and respectively pass the fused features and multi-scale features through a prediction layer to obtain multi-stage prediction maps, and supervise them with the ground truth, and select the second-layer prediction map as the colon polyp segmentation result;
[0088] Among them, the boundary-guided bidirectional attention residual module includes: using a bidirectional attention mechanism to complementarily enhance the features of two adjacent layers; dynamically weighting the adjacent features processed by the bidirectional attention mechanism through a multi-scale channel attention mechanism, and then using the boundary attention map obtained by the boundary extraction unit to enhance the boundary region. Finally, the two-way features are fused to obtain the fused feature of the current stage; the fused feature is re-allocated weights in channels and spaces through a shuffling attention mechanism, and then added to the high-level features through a residual structure, which is the output of the current boundary-guided bidirectional attention residual module.
[0089] For the colon polyp segmentation device based on the boundary-guided bidirectional attention residual network in the embodiments of the present invention, since it corresponds to the colon polyp segmentation method based on the boundary-guided bidirectional attention residual network in the above embodiments, the description is relatively simple. For relevant similarities, please refer to the description of the colon polyp segmentation method based on the boundary-guided bidirectional attention residual network in the above embodiments, and details will not be elaborated here.
[0090] An embodiment of the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the colon polyp segmentation method based on the boundary-guided bidirectional attention residual network as described above. Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A colon polyp segmentation method based on a boundary-guided bidirectional attention residual network, characterized in that It includes the following steps: Input the colonoscopy image into the feature encoder to obtain multi-level encoded features, and concatenate an average pooling layer after the last layer of the feature encoder to obtain the top-level features; Input the multi-level encoded features and the top-level features into the context enrichment module to obtain rich feature representations and multi-scale features, including: for the multi-level encoded features, use a global-local feature extractor to synchronously extract global features and local features, and fuse the global features and the local features to obtain rich features corresponding to each level of the encoded features; for the top-level features, use a U-shaped dilated convolutional spatial pyramid pooling module to extract multi-scale features; in the U-shaped dilated convolutional spatial pyramid pooling module, a supplementary path from a small receptive field to a large receptive field and a guiding path from a large receptive field to a small receptive field are introduced, including: the former enables different dilated convolutional layers with different dilation rates to receive the supplement of the output results of the dilated convolutional layer with a lower level during feature extraction, and the supplement method is pixel-wise addition; the latter enables the features processed by different dilated convolutional layers with different dilation rates to be guided by the output results of the dilated convolutional layer with a higher dilation rate before fusion, and the guiding method is the combination of channel stacking and spatial attention mechanism; Input the multi-level encoded features into the boundary extractor to generate a boundary attention map supervised by the boundary ground truth; Gradually perform feature decoding on the rich feature representation according to the high-to-low criterion through a series of boundary-guided bidirectional attention residual modules to obtain the fused features, pass the fused features and the multi-scale features through the prediction layer respectively to obtain multi-stage prediction maps, and supervise them with the ground truth, and select the second-layer prediction map as the colon polyp segmentation result; Among them, the boundary-guided bidirectional attention residual module includes: Use a bidirectional attention mechanism to complement and enhance the features of adjacent layers; Dynamically weight the adjacent features processed by the bidirectional attention mechanism through a multi-scale channel attention mechanism, then use the boundary attention map to enhance the boundary region, and finally fuse the two-way features to obtain the fused features of the current stage; Re-distribute the weights of the fused features in channels and space through the shuffle attention mechanism, and then add the high-level features through the residual structure, which is the output of the current boundary-guided bidirectional attention residual module.
2. The colon polyp segmentation method based on the boundary-guided bidirectional attention residual network according to claim 1, wherein The feature encoder is of the Res2Net architecture, and the last two layers are discarded to retain the spatial structure; the feature encoder generates 4 feature maps with different spatial resolutions and channel numbers for each image.
3. The colon polyp segmentation method based on a boundary-guided bidirectional attention residual network according to claim 1, characterized in that The global-local feature extractor includes: the multi-level encoded features pass through a transformer branch to obtain global features; The multi-level encoded features pass through a globally guided CNN branch to obtain local features; The global features and the local features are fused through channel stacking and a 3×3 convolutional layer to generate enhanced features.
4. The colon polyp segmentation method based on a boundary-guided bidirectional attention residual network according to claim 1, wherein The boundary extractor generates a boundary attention map supervised by the boundary ground truth, including: The encoded features at each level are transformed to the same spatial resolution using a 1×1 convolutional layer and an upsampling operation, then stacked in the channel direction, followed by fusion using a 3×3 convolutional layer, and finally a boundary attention map is obtained through a 1×1 convolutional layer.
5. The colon polyp segmentation method based on the boundary-guided bidirectional attention residual network according to claim 1, wherein A bidirectional attention mechanism is used to complementarily enhance the features of adjacent layers, including: First, fuse adjacent two layers of enhanced features through a 3×3 convolutional layer, and then generate a foreground attention map through a convolutional layer and a sigmoid activation function, and then subtract the foreground attention map from 1 pixel by pixel to obtain a background attention map. Use the foreground attention map and the background attention map respectively to enhance the foreground of the high-level features and filter out the background noise of the low-level features.
6. The colon polyp segmentation method based on a boundary-guided bidirectional attention residual network according to claim 1, wherein Enhancing the boundary region using the boundary attention map, including: weighting the features to be enhanced using the boundary attention map, then performing pixel-wise addition with the original features, and finally obtaining the final boundary-enhanced features through a 3×3 convolutional layer.
7. A colon polyp segmentation device based on a boundary-guided bidirectional attention residual network, characterized in that, Including: An encoding unit for inputting a colonoscopy image into a feature encoder to obtain multi-level encoded features, and concatenating an average pooling layer after the last layer of the feature encoder to obtain top-level features; A feature enhancement unit for contextually enhancing the multi-level encoded features and the top-level features obtained by the encoding unit to obtain rich feature representations and multi-scale features, including: for the multi-level encoded features, globally-local feature extractors are respectively used to synchronously extract global features and local features, and the global features and the local features are fused to obtain rich features corresponding to the encoded features at each level; for the top-level features, a U-shaped dilated convolutional spatial pyramid pooling module is used to extract multi-scale features; in the U-shaped dilated convolutional spatial pyramid pooling module, a complementary path from a small receptive field to a large receptive field and a guiding path from a large receptive field to a small receptive field are introduced, including: the former enables different dilated convolutional layers with different dilation rates to receive the output results of the dilated convolutional layer with a lower level during feature extraction, and the complementary method is pixel-wise addition; the latter enables the features processed by different dilated convolutional layers with different dilation rates to be guided by the output results of the dilated convolutional layer with a higher dilation rate before fusion, and the guiding method is the combination of channel stacking and spatial attention mechanism; A boundary extraction unit for inputting the multi-level encoded features obtained by the encoding unit into a boundary extractor to generate a boundary attention map supervised by the boundary ground truth; A prediction unit for gradually decoding the rich feature representations obtained by the feature enhancement unit according to the principle from high to low through a series of boundary-guided bidirectional attention residual modules to obtain fused features, passing the fused features and the multi-scale features through a prediction layer respectively to obtain multi-stage prediction maps, and supervising with the ground truth, and selecting the second prediction map as the colon polyp segmentation result; Among them, the boundary-guided bidirectional attention residual module includes: using a bidirectional attention mechanism to complementarily enhance the features of adjacent layers; dynamically weighting the adjacent features after being processed by the bidirectional attention mechanism through a multi-scale channel attention mechanism, then enhancing the boundary region using the boundary attention map obtained by the boundary extraction unit, and finally fusing the two-way features to obtain the fused features at the current stage; passing the fused features through a shuffling attention mechanism to re-distribute the weights in the channel and spatial dimensions, and then adding the high-level features through a residual structure, which is the output of the current boundary-guided bidirectional attention residual module.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to execute the colon polyp segmentation method based on a boundary-guided bidirectional attention residual network according to any one of claims 1-6.
Citation Information
Patent Citations
Polyp segmentation method combining attention U-shaped network and multi-scale feature fusion
CN114820635A
Convolutional neural network polyp segmentation method fusing channel and space attention
CN114842029A