A colorectal polyp segmentation method integrating multiple attention mechanism neural networks
Through a neural network that integrates multiple attention mechanisms, the problem of low accuracy of colorectal polyps segmentation is solved, and more accurate polyps recognition and segmentation is achieved, supporting the accuracy of clinical diagnosis.
Patent Information
- Application Number
- CN202111276806.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-10-29
AI Technical Summary
Due to different size, morphology, color and texture and similar to the surrounding tissue mucosa, color and texture, color and are similar to the segmentation accuracy, which affects the diagnostic accuracy.
A neural network that integrates multiple attention mechanisms is adopted, including channel packet space enhancement module, axial self-attention combined with receptive field enhancement module and reverse attention boundary enhancement module. Through these modules, the feature importance of characteristics are adjusted, the context-dependent model is established, and the boundary information is mined, and the feature representation ability of polyp segmentation is improved.
The segmentation accuracy of colorectal polyps is significantly improved, especially when processing polyps of different sizes and morphology and blurred boundaries, it enhances the recognition ability of polyps and supports more accurate clinical diagnosis.
Smart Images

Figure CN113989301B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of digital images, image segmentation, computer vision and deep learning, and in particular to a colorectal polyp segmentation method based on a neural network integrating multiple attention mechanisms. Background Art
[0002] Colorectal cancer (CRC) is a common malignant tumor of the digestive tract and the third most common cancer in the world. Most colorectal cancers evolve from adenomatous polyps, so early diagnosis of colorectal cancer is crucial to improving the survival rate of colorectal cancer patients. In fact, the survival rate of colorectal cancer in the first stage exceeds 95%, while in the fourth and fifth stages, it drops to less than 35%. Currently, colonoscopy has been widely used in clinical practice and has become a standard method for screening colorectal cancer. In clinical practice, colonoscopy relies heavily on the doctor's experience, and the misdiagnosis rate is very high due to the different sizes, shapes, colors, textures of polyps, and their high similarity to the surrounding tissue mucosa, unclear boundaries of abnormal tissues, and low contrast at the boundaries of the environment.
[0003] In early studies, learning-based methods mainly rely on manually extracted features such as color, shape, texture, appearance or a combination thereof. Such methods usually train a classifier to separate polyps from colonoscopy images. However, due to the limited representation ability of manually extracted features in describing heterogeneous polyps and the similarities between polyps and difficult samples, there is usually a problem of low detection accuracy, which is not conducive to clinical diagnosis. Therefore, it is of great significance to segment colorectal polyps with high precision. Summary of the invention
[0004] Based on the above, the purpose of the present invention is to provide a colorectal polyp segmentation method based on a neural network that integrates multiple attention mechanisms, so as to solve the problem of low segmentation accuracy caused by polyps having different sizes, shapes, colors, textures, and being highly similar to surrounding tissue mucosa and unclear boundaries of abnormal tissue.
[0005] In order to achieve the above-mentioned invention object, the present invention provides the following technical solutions:
[0006] A colorectal polyp segmentation method integrating multiple attention mechanism neural networks, characterized in that it comprises the following steps:
[0007] 1) Collect various types of endoscopic colorectal polyp images and perform image enhancement to form training samples;
[0008] 2) Construct a neural network with multiple attention mechanisms, including a feature extraction module, a channel grouping spatial enhancement module, an axial self-attention combined with a receptive field enhancement module, and a reverse attention boundary enhancement module;
[0009] 3) using the training samples to train the neural network, optimizing the network parameters, and after the network parameters are determined, forming a model together with the neural network;
[0010] 4) When applied, endoscopic colorectal images are collected and input into the model, and the colorectal polyp segmentation image is output after calculation.
[0011] The colorectal polyp segmentation method integrating multiple attention mechanism neural networks is characterized in that the specific process of step 1) is as follows:
[0012] Step 1.1) Collecting an endoscopic colorectal polyp image dataset;
[0013] Step 1.2) The image resolution is adjusted to 352×352 by linear interpolation method, and the dataset is divided into training data and test data;
[0014] Step 1.3) The images in the training data are randomly horizontally flipped, randomly contrast enhanced, randomly scaled up and down by 0.75 to 1.25 times, and randomly rotated by 0 to 360 degrees.
[0015] The colorectal polyp segmentation method integrating multiple attention mechanism neural networks is characterized in that the specific process of constructing the channel grouping space enhancement module, the axial self-attention combined with the receptive field enhancement module and the reverse attention boundary enhancement module in step 2) is as follows:
[0016] Step 2.1) Construct channel grouping spatial enhancement module: The feature vector of each independent grouping position in space is represented by x i express, Where C represents the number of feature map channels output by the backbone network, H×W represents the size of the feature map, and G represents the number of groups. Indicates the spatial domain where the pixel is located; according to the batch size set in the experiment, combined with the number of feature map channels obtained by each layer of the backbone network, while preventing the loss of feature details caused by too small grouping, when the number of feature map channels obtained is 512, 1024, and 2048 respectively, the number of groups G is set to 32, 64, and 128 respectively; firstly, a global maximum average pooling is performed on each group to approximate the learned feature semantic vector of each group and suppress possible noise, that is:
[0017]
[0018] Then use this global feature GF and local feature xi Perform dot product operation to generate the attention coefficient c corresponding to each sub-feature i , c i =GF·x i , that is, to compare the similarity between each sub-feature and the global feature; at the same time, in order to prevent the coefficients between different samples from deviating too much, the coefficient c i Normalized to get the attention feature map a i ,Right now:
[0019]
[0020] in is a constant 1e -5 ,c j represents the attention coefficient of another sub-feature in the space, j represents another sub-feature in the space, and finally a i The final enhanced feature grouping y is obtained through the sigmoid activation function i ;
[0021] Step 2.2) Construct an axial self-attention receptive field enhancement module: First, the input feature map passes through the receptive field paths of the 1×3, 1×5, 1×7, and 1×1 convolution layers respectively, and then enters the axial self-attention module after expanding the receptive field through the dilated convolution layers with dilation rates of 3, 5, 7, and 1, and finally outputs after channel splicing with the feature map of the 1×1 receptive field path;
[0022] Step 2.3) Construct the reverse attention boundary enhancement module: regard the feature map obtained from the decoder as m d , and the corresponding foreground feature map m is obtained f and background feature map m b ,in
[0023] m f =max(m d -0.5,0),m b =max(0.5-m d ,0)
[0024] Then calculate the feature vector of each feature map:
[0025]
[0026] v f , v b , v d Represent the feature vectors corresponding to the feature maps from the foreground, background and decoder respectively, i represents each pixel in the spatial dimension, e represents a natural constant, and then calculate the similarity between each of the above feature vectors and the feature map from the encoder;
[0027]
[0028]
[0029] Then the similarity score k of each feature vector is obtained by weighted summation f ,k b ,k d ; Then, each pixel p in the semantic feature map is obtained by weighted average i ;
[0030] p i =δ(k fi ρ(v f )+k bi ρ(v b )+k di ρ(v d ))
[0031] ψ(·), φ(·), ρ(·), and δ(·) all represent convolution operations; finally, the semantic feature map p and the input features of the encoder are channel-wise concatenated to obtain the final semantic feature map.
[0032] The colorectal polyp segmentation method integrating multiple attention mechanism neural networks is characterized in that the specific process of step 3) is as follows:
[0033] Step 3.1) Construct a colorectal polyp segmentation model that integrates multiple attention mechanism neural networks, and use the Adam optimizer combined with the cosine annealing strategy for optimization. The initial learning rate is set to 0.0001, and the network loss function is Set to:
[0034]
[0035] Among them, i represents each pixel in the image, y represents the polyp label image, Represents the network prediction output image;
[0036] Step 3.2) Fine-tune the model by loading the model with the highest accuracy in the test set to obtain the polyp segmentation image result;
[0037] Step 3.3) Test the segmentation performance of the trained model and use the test set to verify the segmentation effect of the trained model; by comparing it with the original label image, evaluate the segmentation effect of the model from the perspective of subjective visual effects and objective evaluation indicators.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] The colorectal polyp segmentation method based on a neural network integrating multiple attention mechanisms provided by the present invention, the channel grouping space enhancement module, by generating an attention factor for each spatial position in each semantic feature group to adjust the importance of each sub-feature, so that each group can autonomously adjust the correlation of the features learned by each group, thereby enhancing the features of the backbone network. The axial self-attention module, combined with the receptive field strategy, can establish a rich context dependency model for local features, model remote dependencies, and obtain low-level features with edge details and global semantic high-level features to improve the feature representation of polyp segmentation. The reverse attention boundary enhancement module can generate a reverse saliency feature map, and calculate the feature vectors of the foreground feature map and the background feature map by aggregating pixel features, so as to mine boundary information and enhance the network's segmentation performance for smaller polyps and polyps with blurred tissue boundaries. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 This is a schematic diagram of a neural network structure integrating multiple attention mechanisms provided by an embodiment of the present invention;
[0041] Figure 2 is a schematic diagram of the structure of a channel grouping space enhancement module provided by an embodiment of the present invention;
[0042] Figure 3 It is a schematic diagram of the structure of the axis self-attention receptive field enhancement module provided by an embodiment of the present invention;
[0043] Figure 4 is a structural schematic diagram of a reverse attention boundary enhancement module provided by an embodiment of the present invention;
[0044] Figure 5 This is a graph showing the result of polyp segmentation by the polyp segmentation model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0045] To make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific implementation methods described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.
[0046] The embodiment of the present invention provides a schematic diagram of colorectal polyp segmentation integrating multiple attention mechanism neural networks, and the specific steps are as follows:
[0047] Step 1: Various types of endoscopic colorectal polyp images were collected and enhanced by random horizontal flipping, random contrast enhancement, 0.75-1.25 times random multi-scale zooming, and 0-360 degree random rotation to form training samples.
[0048] The specific process is:
[0049] Step 1.1: Collect endoscopic colorectal polyp image dataset;
[0050] Step 1.2: Adjust the image resolution to 352×352 by linear interpolation method, and divide the dataset into training data and test data;
[0051] Step 1.3: Perform random horizontal flipping, random contrast enhancement, 0.75-1.25 times random multi-scale upscaling and downscaling, and 0-360 degree random rotation on the images in the training data.
[0052] Step 2: Construct a neural network with multiple attention mechanisms, including a feature extraction module, a channel grouping spatial enhancement module, an axial self-attention combined with a receptive field enhancement module, and a reverse attention boundary enhancement module.
[0053] like Figure 1 As shown in the figure, Res2Net is used as the feature extraction backbone network to initially extract polyp image features, and then the channel grouping spatial enhancement module is used to adjust the feature correlation of each group and enhance the features extracted by the backbone network. After the channel grouping spatial enhancement module, an encoding path based on the axial self-attention receptive field module is added, which can not only establish a rich context dependency model for local features and model remote dependencies, but also obtain low-level features with edge details and global semantic high-level features. After that, the feature map will be sent to the partial decoder to obtain the global feature map. After the global feature map is channel-joined with the feature map output by the axial self-attention, it is sent to the reverse attention boundary enhancement module. By aggregating the pixel features of each feature map from the encoder, the foreground feature map, background feature map and feature vector of the decoder are calculated respectively, so as to mine the boundary information and enhance the network's segmentation ability for smaller polyps and those with blurred tissue boundaries. Finally, the feature map output from the reverse attention boundary enhancement module will be added to the global feature map output by the decoder, and finally the final prediction map will be obtained through the sigmoid function.
[0054] The specific process of building the network is as follows:
[0055] Step 2.1: Construct channel grouping spatial enhancement module. Figure 2 As shown, the feature vector of each independent group in space is represented by x i express, Where C represents the number of feature map channels output by the backbone network, H×W represents the size of the feature map, and G represents the number of groups. According to the batch size set in the experiment, combined with the number of feature map channels obtained by each layer of the backbone network, and to prevent the loss of feature details caused by too small groups, when the number of feature map channels obtained is 512, 1024, and 2048, we set the number of groups G to 32, 64, and 128 respectively. We first perform global maximum average pooling on each group to approximate the learned feature semantic vector of each group and suppress possible noise, that is:
[0056]
[0057] Then we use this global feature GF and local feature x i Perform dot product operation to generate the attention coefficient c corresponding to each sub-feature i , that is, compare the similarity between each sub-feature and the global feature. i =GF·x i At the same time, in order to prevent the coefficients of different samples from deviating too much, the coefficient c i Normalized to get the attention feature map a i ,Right now:
[0058] in is a constant 1e -5 , finally a i The final enhanced feature grouping y is obtained through the sigmoid activation function i .
[0059] Step 2.2: Construct the axial self-attention receptive field enhancement module. Figure 3 As shown in the figure, the input feature map first passes through the receptive field paths of the 1×3, 1×5, 1×7, and 1×1 convolutional layers respectively, and then enters the axial self-attention module after expanding the receptive field through the dilated convolutional layers with dilation rates of 3, 5, 7, and 1, and finally is output after channel concatenation with the feature map of the 1×1 receptive field path.
[0060] Step 2.3: Construct the reverse attention boundary enhancement module. Figure 4 As shown, the feature map obtained from the decoder is regarded as m d , and the corresponding foreground feature map m is obtained f and background feature map m b ,in
[0061] m f =max(m d -0.5,0),m b =max(0.5-m d ,0)
[0062] Then calculate the feature vector of each feature map:
[0063]
[0064] v f ,v b ,v d Represent the feature vectors corresponding to the feature maps from the foreground, background and decoder respectively, i represents each pixel in the spatial dimension, and then we calculate the similarity between each of the above feature vectors and the feature map from the encoder.
[0065]
[0066]
[0067] Then the similarity score k of each feature vector is obtained by weighted summation f ,k b ,k d Then, each pixel p in the semantic feature map is obtained by weighted average. i .
[0068] p i =δ(k fi ρ(v f )+k bi ρ(v b )+k di ρ(v d ))
[0069] ψ(·), φ(·), ρ(·), and δ(·) all represent convolution operations. Finally, the semantic feature map p and the input features of the encoder are channel-wise concatenated to obtain the final semantic feature map.
[0070] Step 3: Use training samples to train the neural network, optimize network parameters, and determine the network parameters to form a model together with the neural network.
[0071] The specific process is:
[0072] Step 3.1: Construct a colorectal polyp segmentation model that integrates multiple attention mechanism neural networks. The Adam optimizer is used in combination with the cosine annealing strategy for optimization. The initial learning rate is set to 0.0001, and the network loss function is Set to:
[0073]
[0074] Among them, i represents each pixel in the image, y represents the polyp label image, Represents the network predicted output image.
[0075] Step 3.2: Fine-tune the model by loading the model with the highest accuracy in the test set to obtain the polyp segmentation image results.
[0076] Step 3.3: Test the segmentation performance of the trained model. Use the test set (not in the training set) to verify the segmentation effect of the trained model. By comparing it with the original label map, evaluate the segmentation effect of the model from the subjective visual effect and objective evaluation indicators.
[0077] Step 4: When applied, endoscopic colorectal images are collected and input into the model, and the colorectal polyp segmentation image is output after calculation.
[0078] Specific experimental examples
[0079] 1) Select experimental data
[0080] The datasets selected in the experiment are from CVC-300, CVC-ClinicDB, CVC-ColonDB, Kvasir and ETIS-LaribPolypDB, all of which are open source datasets. 900 and 550 images were extracted from Kvasir and CVC-ClinicDB respectively, a total of 1450 images as training sets, and the rest of the data was used as test sets.
[0081] 2) Experimental results
[0082] The network is trained according to the steps in the neural network colorectal polyp segmentation method integrating multiple attention mechanisms. After the model is constructed, the model with the highest accuracy among the training models is loaded, and then the model performance is verified with the images in the test set to obtain the polyp segmentation result. Figure 5 shown.
[0083] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A colorectal polyp segmentation method integrating multiple attention mechanism neural networks, It is characterized in that The following steps are involved: 1) Collect various types of endoscopic colorectal polyp images and perform image enhancement to form training samples; 2) Construct a neural network with multiple attention mechanisms, including a feature extraction module, a channel grouping spatial enhancement module, an axial self-attention combined with a receptive field enhancement module, and a reverse attention boundary enhancement module; The specific process of constructing the channel grouping spatial enhancement module, the axial self-attention combined receptive field enhancement module and the reverse attention boundary enhancement module in step 2) is as follows: Step 2.1) Construct channel grouping spatial enhancement module: The feature vector of each independent grouping position in space is represented by x i It means that X={x 1...m }, m = H × W, where C represents the number of feature map channels output by the backbone network, H × W represents the size of the feature map, and G represents the number of groups. Indicates the spatial domain where the pixel is located; according to the batch size set in the experiment, combined with the number of feature map channels obtained by each layer of the backbone network, while preventing the loss of feature details caused by too small grouping, when the number of feature map channels obtained is 512, 1024, and 2048 respectively, the number of groups G is set to 32, 64, and 128 respectively; firstly, a global maximum average pooling is performed on each group to approximate the learned feature semantic vector of each group and suppress possible noise, that is: Then use this global feature GF and local feature x i Perform dot product operation to generate the attention coefficient c corresponding to each sub-feature i , c i =GF·x i , that is, to compare the similarity between each sub-feature and the global feature; at the same time, in order to prevent the coefficients between different samples from deviating too much, the coefficient c i Normalized to get the attention feature map a i ,Right now: in is a constant 1e -5 ,c j represents the attention coefficient of another sub-feature in the space, j represents another sub-feature in the space, and finally a i The final enhanced feature grouping y is obtained through the sigmoid activation function i ; Step 2.2) Construct an axial self-attention receptive field enhancement module: First, the input feature map passes through the receptive field paths of the 1×3, 1×5, 1×7, and 1×1 convolution layers respectively, and then enters the axial self-attention module after expanding the receptive field through the dilated convolution layers with dilation rates of 3, 5, 7, and 1, and finally outputs after channel splicing with the feature map of the 1×1 receptive field path; Step 2.3) Construct the reverse attention boundary enhancement module: regard the feature map obtained from the decoder as m d , and the corresponding foreground feature map m is obtained f and background feature map m b ,in m f =max(m d -0.5,0),m b =max(0.5-m d ,0) Then calculate the feature vector of each feature map: v f , v b , v d Represent the feature vectors corresponding to the feature maps from the foreground, background and decoder respectively, i represents each pixel in the spatial dimension, e represents a natural constant, and then calculate the similarity between each of the above feature vectors and the feature map from the encoder; Then the similarity score k of each feature vector is obtained by weighted summation f ,k b ,k d ; Then, each pixel p in the semantic feature map is obtained by weighted average i ; p i =δ(k fi p(v f )+k bi p(v b )+k di p(v d )) ψ(·), φ(·), ρ(·), δ(·) all represent convolution operations; finally, the semantic feature map p and the input features of the encoder are concatenated to obtain the final semantic feature map; 3) using the training samples to train the neural network, optimizing the network parameters, and after the network parameters are determined, forming a model together with the neural network; 4) When applied, endoscopic colorectal images are collected and input into the model, and the colorectal polyp segmentation image is output after calculation.
2. The colorectal polyp segmentation method integrating multiple attention mechanism neural networks as claimed in claim 1, It is characterized in that Step 1) The specific process is as follows: Step 1.1) Collecting an endoscopic colorectal polyp image dataset; Step 1.2) The image resolution is adjusted to 352×352 by linear interpolation method, and the dataset is divided into training data and test data; Step 1.3) The images in the training data are randomly horizontally flipped, randomly contrast enhanced, randomly scaled up and down by 0.75 to 1.25 times, and randomly rotated by 0 to 360 degrees.
3. The colorectal polyp segmentation method integrating multiple attention mechanism neural networks as claimed in claim 1, It is characterized in that Step 3) The specific process is as follows: Step 3.1) Construct a colorectal polyp segmentation model that integrates multiple attention mechanism neural networks, and use the Adam optimizer combined with the cosine annealing strategy for optimization. The initial learning rate is set to 0.0001, and the network loss function is Set to: Among them, i represents each pixel in the image, y represents the polyp label image, Represents the network prediction output image; Step 3.2) Fine-tune the model by loading the model with the highest accuracy in the test set to obtain the polyp segmentation image result; Step 3.3) Test the segmentation performance of the trained model and use the test set to verify the segmentation effect of the trained model; by comparing it with the original label image, evaluate the segmentation effect of the model from the perspective of subjective visual effects and objective evaluation indicators.
Citation Information
Patent Citations
Liver segmentation method based on residual-attention deep neural network
CN110889852A
Liver pathological image segmentation model establishment and segmentation method based on attention mechanism
CN112017191A