Colorectal Polyp Segmentation Method Combining Convolution and Multilayer Perceptron Neural Network

By fusion convolution and multi-layer perceptron neural network, a neural network including multi-layer perceptron encoder, parallel self-attention module, cascaded semantic feature aggregation module and channel-guided grouping reverse attention module was built, which solved the problem of low colorectal polyps segmentation accuracy and achieved higher polyps feature extraction and segmentation accuracy.

CN114511508BActive Publication Date: 2025-06-10ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210028910.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-11
Publication Date
2025-06-10
Estimated Expiration
2042-01-11

AI Technical Summary

Technical Problem

Due to different size, morphology, color and texture and similar to the surrounding tissue mucosa, color and texture, color and are similar to the segmentation accuracy, which affects the diagnostic accuracy.

Method used

Using the method of fusion convolution and multi-layer perceptron neural network, a neural network including multi-layer perceptron encoder, parallel self-attention module, cascaded semantic feature aggregation module and channel-guided packet reverse attention module is constructed. The network parameters are optimized by training samples to improve the accuracy of polyp feature extraction and segmentation.

Benefits of technology

This method can extract polyp features more comprehensively and robustly, reduce irrelevant feature interference, effectively aggregate high-level semantic features and low-level edge information, and improve the segmentation accuracy of small polyps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114511508B_ABST
    Figure CN114511508B_ABST
Patent Text Reader

Abstract

The present invention discloses a colorectal polyp segmentation method integrating convolutional and multi-layer perceptron neural networks, comprising: 1) collecting various types of endoscopic colorectal polyp images, and performing image enhancement by methods such as randomly horizontally flipping the images, randomly enhancing the contrast, randomly magnifying and shrinking by 0.75 to 1.25 times in multiple scales, and randomly rotating by 0 to 360 degrees to form training samples; 2) constructing a neural network integrating convolution and multi-layer perceptron, which includes a convolutional and multi-layer perceptron encoder, a parallel self-attention module, a cascaded semantic feature aggregation module, and a channel-guided grouped inverse attention module; 3) training the neural network with the training samples, optimizing the network parameters, and after determining the network parameters, jointly forming a model with the neural network; 4) during application, collecting endoscopic colorectal images and inputting them into the model, and calculating to output colorectal polyp segmentation images. This method improves the segmentation effect of colorectal polyps.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of digital images, image segmentation, computer vision, and deep learning, and particularly relates to a colorectal polyp segmentation method based on a neural network that fuses convolution and multi-layer perceptrons. Background Art

[0002] Colorectal cancer (CRC) is a common malignant tumor of the digestive tract and is the third most common cancer globally. Most colorectal cancers evolve from adenomatous polyps. Therefore, early diagnosis of colorectal cancer is crucial for improving the survival rate of patients with colorectal cancer. In fact, the survival rate of colorectal cancer in the first stage exceeds 95%, while in the fourth and fifth stages, it drops to less than 35%. Currently, colonoscopy has been widely used in clinical practice and has become the standard method for screening colorectal cancer. In clinical practice, colonoscopy largely depends on the experience of doctors. And due to the characteristics of polyps such as different sizes, shapes, colors, textures, high similarity to the surrounding tissue mucosa, unclear boundaries of abnormal tissues, and low contrast at the boundary of the surrounding environment, the missed diagnosis rate is very high.

[0003] In early studies, learning-based methods mainly relied on manually extracted features such as color, shape, texture, appearance, or combinations thereof. Such methods usually trained a classifier to separate polyps from colonoscopy images. However, due to the limited representational ability of manually extracted features in describing heterogeneous polyps and the similarity between polyps and difficult samples, there is usually a problem of low detection accuracy, which is not conducive to clinical diagnosis. Therefore, it is of great significance to accurately segment colorectal polyps. Summary of the Invention

[0004] Based on the above, the purpose of the present invention is to provide a colorectal polyp segmentation method based on a neural network that fuses convolution and multi-layer perceptrons, and solve the problem of low segmentation accuracy caused by polyps having different sizes, shapes, colors, textures, high similarity to the surrounding tissue mucosa, and unclear boundaries of abnormal tissues.

[0005] To achieve the above object of the invention, the present invention provides the following technical solutions:

[0006] A colorectal polyp segmentation method based on a neural network that fuses convolution and multi-layer perceptrons, characterized by comprising the following steps:

[0007] 1) Collect various types of endoscopic colorectal polyp images, and perform image enhancement to form training samples;

[0008] 2) Construct a neural network that combines convolution and multi-layer perceptron, which includes a multi-layer perceptron encoder, a parallel self-attention module, a cascaded semantic feature aggregation module, and a channel-guided grouped reverse attention module;

[0009] 3) Use the training samples to train the neural network, optimize the network parameters, and after determining the network parameters, jointly form a model with the neural network;

[0010] 4) During application, collect endoscopic colorectal images and input them into the model, and after calculation, output colorectal polyp segmentation images.

[0011] The colorectal polyp segmentation method using a neural network that combines convolution and multi-layer perceptron is characterized in that the specific process of step 1) is as follows:

[0012] Step 1.1) Collect an endoscopic colorectal polyp image dataset;

[0013] Step 1.2) Adjust the image resolution through linear interpolation method, and divide the dataset into two parts: training data and test data;

[0014] Step 1.3) Randomly horizontally flip, randomly enhance the contrast, randomly magnify and reduce by 0.75 - 1.25 times, and randomly rotate by 0 - 360 degrees for the images in the training data.

[0015] Furthermore, the colorectal polyp segmentation method using a neural network that combines convolution and multi-layer perceptron is characterized in that the specific construction processes of the multi-layer perceptron encoder, the parallel self-attention module, the cascaded semantic feature aggregation module, and the channel-guided grouped reverse attention module in step 2) are as follows:

[0016] Step 2.1) Construct a multi-layer perceptron encoder: First, construct a pure convolutional layer, which includes a 1×1 convolution, a 3×3 convolution, and a 1×1 convolution, and repeat this sequence; then construct a convolution and multi-layer perceptron hybrid layer, which sequentially includes a channel perceptron, a 3×3 depthwise separable convolution, and a channel perceptron, and repeat this sequence respectively, so as to form 3 convolution and perceptron hybrid layers;

[0017] Step 2.2) Construct a parallel attention module: The parallel self-attention module includes a channel attention branch and a spatial attention branch, where the channel attention branch can be described as

[0018] A tt ch (X) = M z (σ 1 (M v (X)) × δ(σ 2 (M q (X))))

[0019] is the real number field, M v , M q and M z is a 1×1 convolution, σ 1 , σ 2 is an integer operation, δ is a softmax operation; the output of the last channel attention branch is ⊙ is the Hadamard product; then construct the spatial attention branch:

[0020]

[0021] C represents the number of convolution channels, H represents the height of the feature map, and W represents the width of the feature map; σ 1 , σ 2 and σ 3 are integer operations, M v , M q is a 1×1 convolution, ρ is a global average pooling operation, and δ is a softmax operation. The final output of the spatial attention branch is Finally, the output of the entire module is

[0022] Step 2.3) Construct the cascaded semantic feature aggregation module: First, multiply the high-level feature H 3 and H 2 and then add the result to the low-level feature H 1 in channels. At the same time, multiply the high-level feature H 1 and the low-level feature L 1 and then add the result to the high-level feature H 3 and H 2 in channels after fusion; finally, reduce the channel dimension through 3×3 and 1×1 convolution operations; the output of the final module is G:

[0023] G = M 1 M 3 (Concat(M 3 M 3 (Concat(M 3 (H 3 )) × M 3 (H 2 ))M 3 (H 1 ))(M 3 (L 1 )) × M 3 (H 1 ))))

[0024] is a 1×1 convolution, M 3It is a 3×3 convolution, and Concat is an operation of adding channels;

[0025] Step 2.4) Construct a channel-guided grouped inverse attention module: First, the global feature map from the cascaded semantic feature aggregation module forms a saliency feature map through the sigmoid operation; then, the inverse attention-guided map R is obtained through the inverse attention operation:

[0026] R = φ[σ(μ(S)), E]

[0027] μ is a bilinear interpolation operation, σ(x) = 1 / (1 + e -x ) is the sigmoid function, φ is the inverse attention operation, and E is a matrix of all 1s.

[0028] Then the high-level feature H 1 , H 2 , H 3 will be divided into multiple groups in the channel dimension, and the number of groups is 1, 8, and 32 respectively; this process can be described as:

[0029] F s (H i ) = {H i,1 ,..., H i,m}

[0030] Then the inverse attention-guided map R is cyclically inserted into each group. This process can be described as:

[0031] Y i = F c ({H i,1 , R}..., {H i,m , R})

[0032] where i represents the serial number of the high-level feature, which corresponds one-to-one with the number of grouped channels m, F s represents the channel separation operation, F c represents the channel addition, and finally, the output map of the last module is obtained through H i and Y i .

[0033] Furthermore, the colorectal polyp segmentation method using the fusion convolution and multi-layer perceptron neural network is characterized in that the specific process of step 3) is as follows:

[0034] Step 3.1) Construct a colorectal polyp segmentation model using the fusion convolution and multi-layer perceptron neural network, and optimize it using the Adam weight decay optimizer combined with the polynomial learning rate decay strategy. The initial learning rate is set to 0.0002, and the loss function L of the network is set to:

[0035] L = L BCE + LIoU

[0036]

[0037]

[0038] Among them, i represents each pixel in the image, and y represents the polyp label image, represents the network prediction output image;

[0039] Step 3.2) Fine-tune the model, and obtain the polyp segmentation image result by loading the model with the highest accuracy in the test set;

[0040] Step 3.3) Test the segmentation performance of the trained model, and use the test set to verify the segmentation effect of the trained model; by comparing with the original label map, evaluate the segmentation effect of the model from subjective visual effects and objective evaluation indicators.

[0041] Compared with the prior art, the beneficial effects of the present invention at least include:

[0042] The colorectal polyp segmentation method based on the fusion of convolutional and multi-layer perceptron neural networks provided by the present invention, the convolutional and multi-layer perceptron encoder can extract more comprehensive and robust polyp features. The parallel attention module can focus more attention on the polyp area in the low-level features and reduce the interference of irrelevant features. The cascaded semantic feature aggregation module can effectively aggregate high-level semantic features and low-level edge information features. The channel-guided grouped reverse attention module can establish the connection between the polyp area and the boundary area in the image, so as to better segment small polyps. Brief Description of the Drawings

[0043] Figure 1 is a schematic diagram of the neural network integrating convolution and multi-layer perceptron provided by an embodiment of the present invention;

[0044] Figure 2 is a schematic diagram of the structure of the convolutional and multi-layer perceptron encoder provided by an embodiment of the present invention;

[0045] Figure 3 is a schematic diagram of the structure of the parallel attention module provided by an embodiment of the present invention;

[0046] Figure 4 is a schematic diagram of the structure of the cascaded semantic feature aggregation module provided by an embodiment of the present invention;

[0047] Figure 5 is a schematic diagram of the structure of the channel-guided grouped reverse attention module provided by an embodiment of the present invention;

[0048] Figure 6It is the result diagram of polyp segmentation by the polyp segmentation model provided by the embodiment of the present invention. Detailed implementation manners

[0049] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific implementation manners described herein are only used to explain the present invention and do not limit the protection scope of the present invention.

[0050] The embodiment of the present invention provides a schematic diagram of colorectal polyp segmentation that integrates a convolutional neural network and a multi-layer perceptron neural network. The specific steps are as follows:

[0051] Step 1: Collect various types of endoscopic colorectal polyp images, and perform image enhancement through methods such as random horizontal flipping, random contrast enhancement, random multi-scale zooming in and out by 0.75 to 1.25 times, and random rotation by 0 to 360 degrees to form training samples.

[0052] The specific process is as follows:

[0053] Step 1.1: Collect an endoscopic colorectal polyp image dataset;

[0054] Step 1.2: Adjust the image resolution to 352×352 through linear interpolation, and divide the dataset into two parts: training data and test data;

[0055] Step 1.3: Perform random horizontal flipping, random contrast enhancement, random multi-scale zooming in and out by 0.75 to 1.25 times, and random rotation by 0 to 360 degrees on the images in the training data.

[0056] Step 2: Build a neural network that integrates a convolutional neural network and a multi-layer perceptron, which includes a convolutional and multi-layer perceptron encoder, a parallel self-attention module, a cascaded semantic feature aggregation module, and a channel-guided grouped reverse attention module.

[0057] As Figure 1 shown, use the convolutional and multi-layer perceptron encoder as the feature extraction backbone network to initially extract the features of the polyp image. The obtained low-level features obtain more edge detail information through the parallel attention module. After the number of channels of the high-level features is readjusted through the convolutional layer, they are sent to the cascaded semantic aggregation module together with the low-level features to obtain a global feature map. Then, each high-level feature and the global feature map are sent to the channel-guided grouped reverse attention module to obtain 3 prediction maps. After the 3 prediction maps are optimized in pairs, the final prediction map is finally obtained through the sigmoid function.

[0058] The specific construction process of the network is as follows:

[0059] Step 2.1: Build a convolutional and multi-layer perceptron encoder: AsFigure 2 As shown in the figure, first construct a pure convolutional layer, which includes a 1×1 convolution, a 3×3 convolution, and a 1×1 convolution, and repeat this sequence 3 times. Then construct a hybrid layer of convolution and multi-layer perceptron, which successively includes a channel perceptron, a 3×3 depthwise separable convolution, and a channel perceptron, and repeat this sequence 4 times, 8 times, and 3 times respectively, so as to form 3 hybrid layers of convolution and perceptron. Among them, the scaling ratio of the hidden layer in the perceptron in each layer is 3;

[0060] Step 2.2: Construct a parallel self-attention module: As Figure 3 shown, the parallel self-attention module includes a channel attention branch and a spatial attention branch, where the channel attention branch can be described as

[0061] A tt ch (X) = M z (σ 1 (M v (X)) × δ(σ 2 (M q (X))))

[0062] is the real number field, M v , M q and M z are 1×1 convolutions, σ 1 , σ 2 is an integer operation, and δ is a sofmax operation. Finally, the output of the channel attention branch is ⊙ is the Hadamard product. Then construct the spatial attention branch:

[0063]

[0064] σ 1 , σ 2 and σ 3 are integer operations, M v , M q are 1×1 convolutions, ρ is the global average pooling operation, and δ is the softmax operation. The final output of the spatial attention branch is Finally, the output of the entire module is

[0065] Step 2.3: Construct a cascaded semantic feature aggregation module: As Figure 4 shown, first multiply the high-level features H 3 and H 2 and then add the channels to the low-level feature H 1 , and at the same time, the high-level feature H 1 and the low-level feature L1 After performing matrix multiplication, it will be added to the high-level feature H 3 and H 2 The results after fusion are added in channels. Finally, through 3×3 and 1×1 convolution operations, the channel dimension is reduced to 64. The output of the last module is G:

[0066] G = M 1 M 3 (Concat(M 3 M 3 (Concat(M 3 (H 3 ) × M 3 (H 2 ))M 3 (H 1 ))(M 3 (L 1 ) × M 3 (H 1 ))))

[0067] is a 1×1 convolution, M 3 is a 3×3 convolution, and Concat is an operation of adding in channels;

[0068] Step 2.4: Construct a channel-guided grouped reverse attention module: As Figure 5 shown, first, the global feature map from the cascaded semantic feature aggregation module forms a saliency feature map through the sigmoid operation. Then, a reverse attention guidance map R is obtained through the reverse attention operation:

[0069] R = φ[σ(μ(S)), E]

[0070] μ is a bilinear interpolation operation, σ(x) = 1 / (1 + e -x ) is the sigmoid function, φ is the reverse attention operation, and E is a matrix of all 1s.

[0071] Then the high-level feature H 1 , H 2 , H 3 will be divided into multiple groups in the channel dimension, and the number of groups is 1, 8, and 32 respectively. This process can be described as:

[0072] F s (H i ) = {H i,1 ,..., H i,m}

[0073] Then R is cyclically inserted into each group. This process can be described as:

[0074] Y i = F c ({Hi,1 , R}..., {H i,m , R})

[0075] where \(i\in\{1, 2, 3\}\), \(m\in\{1, 8, 32\}\), \(F s represents the channel separation operation, \(F c represents the channel addition. Finally, through \(H i and \(Y i the output image of the final module is obtained.

[0076] Step 3: Use the training samples to train the neural network, optimize the network parameters, and after determining the network parameters, jointly form a model with the neural network.

[0077] The specific process is as follows:

[0078] Step 3.1: Construct a colorectal polyp segmentation model that combines convolution and a multi-layer perceptron neural network, and optimize it using the Adam weight decay optimizer combined with the polynomial learning rate decay strategy. The initial learning rate is set to 0.0002, and the loss function \(L\) of the network is set as:

[0079] \(L = L BCE + L IoU

[0080]

[0081] where \(i\) represents each pixel in the image, \(y\) represents the polyp label image, represents the network predicted output image.

[0082] Step 3.2: Fine-tune the model, and obtain the polyp segmentation image result by loading the model with the highest accuracy in the test set.

[0083] Step 3.3: Test the segmentation performance of the trained model, and use the test set (not appearing in the training set) to verify the segmentation effect of the trained model; by comparing with the original label map, evaluate the segmentation effect of the model from the subjective visual effect and objective evaluation indicators.

[0084] Step 4: During application, collect the endoscopic colorectal images and input them into the model, and calculate the output colorectal polyp segmentation image. Specific embodiments:

[0086] 1) Select experimental data

[0087] The datasets selected in the experiment come from CVC-300, CVC-ClinicDB, CVC-ColonDB, Kvasir, and ETIS, all of which are open-source datasets. 900 images and 550 images are extracted from Kvasir and CVC-ClinicDB respectively, for a total of 1450 images as the training set, and the remaining data as the test set.

[0088] 2) Experimental results

[0089] Train the network according to the steps in the colorectal polyp segmentation method that combines convolution and multi-layer perceptron neural network. After constructing the model, load the model with the highest accuracy in the training model, and then use the images in the test set to verify the model performance to obtain the polyp segmentation result. As Figure 6 shown.

[0090] The specific embodiments described above have elaborated on the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A colorectal polyp segmentation method integrating convolution and multi-layer perceptron neural network, characterized in that, it includes the following steps: 1) Collect various types of endoscopic colorectal polyp images and perform image enhancement to form training samples; 2) Construct a neural network containing convolution and multi-layer perceptron, which includes a convolution and multi-layer perceptron encoder, a parallel self-attention module, a cascaded semantic feature aggregation module, and a channel-guided grouped reverse attention module; The specific construction processes of the convolution and multi-layer perceptron encoder, the parallel self-attention module, the cascaded semantic feature aggregation module, and the channel-guided grouped reverse attention module in step 2) are as follows: Step 2.1) Construct a convolution and multi-layer perceptron encoder: First, construct a pure convolution layer, including a 1×1 convolution, a 3×3 convolution, a 1×1 convolution, and a repeating sequence; then construct a convolution and multi-layer perceptron hybrid layer, which sequentially includes a channel perceptron, a 3×3 depthwise separable convolution, and a channel perceptron, and repeat the sequence respectively, so as to form 3 convolution and perceptron hybrid layers; Step 2.2) Construct a parallel self-attention module: The parallel self-attention module includes a channel attention branch and a spatial attention branch, where the channel attention branch is described as A tt ch (X) = M z (σ 1 (M v (X)) × δ(σ 2 (M q (X)))) is the real number field, M v , M q and M z is a 1×1 convolution, σ 1 , σ 2 is an integer operation, δ is a softmax operation; the output of the last channel attention branch is ⊙ is the Hadamard product; then construct the spatial attention branch: C represents the number of convolutional channels, H represents the height of the feature map, and W represents the width of the feature map; σ 1 , σ 2 and σ 3 are integer operations, M v , M q is a 1×1 convolution, ρ is a global average pooling operation, and δ is a softmax operation; the final output of the spatial attention branch is The final output of the entire module is Step 2.3) Construct a cascaded semantic feature aggregation module: First, multiply the high-level features H 3 and H 2 through matrix multiplication and then add the result to the low-level feature L 1 along the channels. At the same time, multiply the high-level feature H 1 and the low-level feature L 1 through matrix multiplication and then add the result to the high-level feature H 3 and H 2 along the channels after fusion. Finally, reduce the channel dimension through 3×3 and 1×1 convolution operations. The output of the final module is G: G = M 1 M 3 (Concat(M 3 M 3 (Concat(M 3 (H 3 ) × M 3 (H 2 ))M 3 (H 1 ))(M 3 (L 1 ) × M 3 (H 1 )))) M 1 is a 1×1 convolution, M 3 is a 3×3 convolution, and Concat is an operation of channel addition; Step 2.4) Construct a channel-guided grouped reverse attention module: First, the global feature map from the cascaded semantic feature aggregation module forms a saliency feature map through a sigmoid operation; then a reverse attention guidance map R is obtained through a reverse attention operation: R = φ[σ(μ(S)), E] μ is a bilinear interpolation operation, σ(x) = 1 / (1 + e -x ) is the sigmoid function, φ is the reverse attention operation, and E is a matrix of all 1s; Then the high-level feature H 1 , H 2 , H 3 will be divided into multiple groups in the channel dimension; this process can be described as: F s (H i ) = {H i,1 ,..., H i,m} Then the reverse attention guidance map R is cyclically inserted into each group, and this process can be described as: Y i = F c ({H i,1 , R}..., {H i,m , R}) where i represents the serial number of the high-level feature, corresponding one-to-one with the number of grouped channels m; F s represents the channel separation operation, F c represents the channel addition; finally, through H i and Y i the output graph of the last module is obtained; 3) Use the training samples to train the neural network, optimize the network parameters, and after determining the network parameters, form a model together with the neural network; 4) During application, collect endoscopic colorectal images and input them into the model, and calculate and output colorectal polyp segmentation images.

2. A colorectal polyp segmentation method integrating convolution and multi-layer perceptron neural network according to claim 1, characterized in that, the specific process of step 1) is as follows: Step 1.1) Collect an endoscopic colorectal polyp image dataset; Step 1.2) Adjust the image resolution by linear interpolation method and divide the dataset into two parts: training data and test data; Step 1.3) Randomly horizontally flip, randomly enhance the contrast, randomly magnify and reduce by 0.75 - 1.25 times, and randomly rotate by 0 - 360 degrees the images in the training data.

3. A colorectal polyp segmentation method integrating convolution and multi-layer perceptron neural network according to claim 1, characterized in that, the specific process of step 3) is as follows: Step 3.1) Construct a colorectal polyp segmentation model integrating multiple attention mechanism neural networks, and optimize it using the Adam weight decay optimizer combined with the polynomial learning rate decay strategy. Set the initial learning rate, and the loss function of the network is set as: Among them, i represents each pixel in the image, and y represents the polyp label image, indicating the network prediction output image; Step 3.2) Fine-tune the model, by loading the model with the highest accuracy in the training model, and then use the images in the test set to verify the model performance to obtain the polyp segmentation image result; Step 3.3) Test the segmentation performance of the trained model, use the test set to verify the segmentation effect of the trained model; by comparing with the original label map, evaluate the segmentation effect of the model from subjective visual effects and objective evaluation indicators.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method based on attention multi-scale feature fusion

    CN111127493A

  • Semantic segmentation method based on feature pyramid attention and mixed attention cascading

    CN112651973A