An Intestinal Polyp Segmentation Method and Device Based on Deep Fusion of Endoscopic Images
By combining the fusion strategy of gating axial attention mechanism and sliding window attention mechanism, the problem of intestinal polyp image segmentation method dependence on label data is solved, segmentation accuracy and robustness are improved, and training costs are reduced.
Patent Information
- Application Number
- CN202211464842.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-22
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-11-22
AI Technical Summary
The prior art intestinal polyp image segmentation method relies on a large amount of labeled data, and the training cost and difficulty are high, making it difficult to obtain an accurate segmentation model.
A fusion mechanism combining the gated axial attention mechanism module and the sliding window attention mechanism module is adopted to build a local-global learning strategy, feature learning is performed through shallow global branches and deep local branches, and data augmentation and adaptive threshold extraction are performed, and the model is optimized using the joint loss function.
It reduces the training cost and difficulty, obtains richer feature information, improves the robustness and accuracy of the segmentation network, reduces spatial information loss, and realizes efficient intestinal polyp image segmentation.
Smart Images

Figure CN115965630B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep learning, computer vision, and medical image processing, and more particularly to an intestinal polyp segmentation method and apparatus based on deep fusion of endoscopic images. Background Art
[0002] Currently, endoscopic examination has been widely used in clinical practice and has become one of the important standard methods for colorectal cancer diagnosis. There are many clinical outcome programs for colorectal cancer diagnosis. Among them, endoscopic images are one of the main ways for intestinal polyp detection. However, there are significant differences in the size, color, etc. of intestinal polyps among individuals, and many intestinal polyps do not protrude from the surrounding mucosa. Therefore, accurate segmentation of intestinal polyps remains a difficult challenge.
[0003] In early studies, the segmentation of intestinal polyps mainly relied on manual segmentation by professional doctors with rich clinical experience, which was easily interfered by subjective factors, high similarity between heterogeneous intestinal polyps and samples, etc., resulting in low segmentation accuracy. Currently, deep learning-based segmentation methods have improved the accuracy of intestinal polyp segmentation to a certain extent and saved a large amount of manpower and material resources. In particular, methods based on Transformer have been widely applied to computer vision tasks and achieved satisfactory performance. However, this highly depends on a large amount of labeled data, which is difficult to meet in the field of medical image processing. In addition, due to the high complexity of the intestinal polyp structure, there is an instance imbalance in endoscopic images, making it difficult to distinguish between diseased and non-diseased bodies, and even affecting treatment decisions.
[0004] Chinese Patent Publication No. CN107146229A discloses a method for segmenting colon polyp images based on a cellular automaton model, which mainly solves the problems of low segmentation efficiency, poor repeatability, and low segmentation accuracy in existing colon polyp image segmentation technologies. The technical solution is as follows: (1) Read a color colonoscopy image containing polyps; (2) Repair the highlight area in the image; (3) Initially detect the polyp area in the image; (4) Mark the seed pixels; (5) Construct a cellular automaton model; (6) Initialize the cellular automaton model; (7) Perform image segmentation; (8) Output the segmented image. By using the prior knowledge that the shape of polyps is approximately elliptical, the seed pixels are automatically marked, a cellular automaton model is constructed, and image segmentation is performed through the formulated local transformation rules, making full use of the local information of the image, and having the advantages of high segmentation efficiency and accuracy, and can be used for automatic segmentation of colon polyp images. However, in order to improve the segmentation accuracy, this patent application also needs to extract a large amount of labeled data, and the difficulty is relatively high. Summary of the Invention
[0005] The technical problem to be solved by the present invention is that the training of the existing intestinal polyp image segmentation method relies on a large amount of labeled data, resulting in high training costs and difficulties, making it difficult to obtain a relatively accurate segmentation model.
[0006] The present invention solves the above technical problems through the following technical means: An intestinal polyp segmentation method based on deep fusion of endoscopic images, the method comprising:
[0007] Step 1: Collect intestinal polyp images under the endoscope and perform preprocessing to obtain a training set and a test set;
[0008] Step 2: Use a gated axial attention mechanism module and a sliding window attention mechanism module to construct a deep feature fusion module;
[0009] Step 3: Use a gated axial attention mechanism module, a deep feature fusion module, and an attention gate module to construct an intestinal polyp image segmentation model;
[0010] Step 4: Use the training set to train the intestinal polyp image segmentation model to obtain an optimal intestinal polyp image segmentation model;
[0011] Step 5: Input the real-time collected intestinal polyp image under the endoscope into the optimal intestinal polyp image segmentation model to obtain a predicted segmentation image.
[0012] The present invention designs a fusion mechanism that combines a gated axial attention mechanism module and a sliding window-based attention mechanism module to form a local-global learning strategy. It uses a shallow global branch and a deep local branch to perform feature learning on intestinal polyp image patches, solves the problem of lack of a large amount of labeled data, has low training costs and difficulties. At the same time, it obtains richer feature information, reduces the loss of spatial information, improves the robustness of the segmentation network, and enhances the model accuracy.
[0013] Further, the step 1 includes:
[0014] S11. Collect intestinal polyp images under the endoscope;
[0015] S12. Use the linear interpolation method to adjust the resolution of all obtained images to 224×224;
[0016] S13. Perform random horizontal rotation, random vertical rotation, random deformation, random contrast, and random brightness changes on the data images, as well as 0.70 - 1.20 times random multi-scale scaling. Use the high-hat transformation in morphology to enhance the contrast between the intestinal polyp and the background. The formula for the high-hat transformation is:
[0017]
[0018] f b_hat= f - (f · b) = f - f d ,
[0019] f enchance = f t_hat - f b_hat ,
[0020] In the formula, f represents the intestinal polyp image, and f op represents the opening operation, and f d represents the closing operation, b represents the structural element in morphology, and f enchance represents the enhanced endoscopic intestinal polyp image;
[0021] S14. Use the method of adaptive threshold to extract the area where the intestinal polyp is located in the above processed image, and divide the adjusted image into two parts: a training set and a test set according to 5:1.
[0022] Furthermore, the second step includes:
[0023] S21. Construct a gated axial attention mechanism module: Split the original self-attention mechanism module into two modules. One module performs self-attention calculation on the height axis, and the other module performs calculation on the width axis. The calculation formula is as follows
[0024]
[0025]
[0026] Among them, W represents the width, H represents the height, q and k respectively represent the query vector and the key vector. q i,j represents the query vector at any position in i ∈ {1,...H}, j ∈ {1,...W}, and k i,W , u i,w represents the key vector and value vector at any position in i ∈ {1,...H} on a certain width axis, and represent the position biases corresponding to the query vector, key vector, and value vector, Use the gated mechanism to control the weight of the position information and update the self-attention calculation formula on the height axis, that is:
[0027]
[0028] Among them, G q , G k , are all control parameters;
[0029] S22. Construct a window-based attention mechanism module. The sliding window attention mechanism module is expressed as:
[0030] Z′i = W_MSA(Norm(Z i-1 )) + Z i-1
[0031] Z i = FFN(Norm(Z`i)) + Z′ i
[0032] Z′ i+1 = SW_MSA(Norm(Z i )) + Z i
[0033] Z i+1 = FFN(Norm(Z`i)) + Z′ i+1
[0034] Wherein, W_MSA represents a window-based attention module through which the input features pass, SW_MSA represents a sliding window attention mechanism module, Z i-1 represents the input features of W_MSA in the i-th layer, and Z′ i represents the output features of W_MSA in the i-th layer, and also the input features of SW_MSA in the i-th layer; Norm represents normalization, i is an identifier of a certain intermediate block, and FFN represents a fully connected network of a hidden layer;
[0035] S23. Construct a deep feature fusion module: The outputs of the gated axial attention mechanism module and the sliding window attention mechanism module are reshaped to obtain feature maps, the sizes of the feature maps are adjusted using convolution operations, the feature maps of each path are fused through feature concatenation, and then the fused features are obtained through convolution operations
[0036] Furthermore, the step three includes:
[0037] Input the intestinal polyp features of the endoscopic image into the first convolutional layer and multiple gated axial attention mechanism modules to obtain different depth features. The outputs of the gated axial attention mechanism modules are respectively connected to a second convolutional layer, the first convolutional layer is cascaded in sequence with all the second convolutional layers, the output ends of the first convolutional layer and the first n - 1 second convolutional layers are respectively connected to a deep feature fusion module, n ≥ 2, the output end of the n-th second convolutional layer is connected to a third convolutional layer, the output ends of each deep feature fusion module are respectively connected to an attention gating module, the output ends of each attention gating module are respectively connected to a fourth convolutional layer, the third convolutional layer and all the fourth convolutional layers are cascaded in sequence, and the output of the last fourth convolutional layer outputs the image segmentation result, completing the construction of the deep fusion neural network, and using this neural network as the intestinal polyp image segmentation model.
[0038] Furthermore, the attention gating module includes:
[0039] Perform a linear transformation on the input tensor using a 1×1 convolution kernel for convolution operation, then sequentially use the PReLU non-linear activation function and the Sigmoid activation function as two sets of attention coefficients, and finally adjust the feature dimension to jointly fuse the input features to obtain the attention features.
[0040] Furthermore, the fourth step includes:
[0041] S41. Input the training set into the intestinal polyp image segmentation model, and use the Adam optimizer to optimize the intestinal polyp image segmentation model. The number of training epochs is defaulted to 200, and the initial learning rate is set to 0.001;
[0042] S42. Set the loss function as: L total = αL BCE + βL blob ,
[0043] L BCE (p i , g i ) = L = {l1,…, l N} T ,
[0044] l N = -w n [g n logp n + (1 - g n )log(1 - p n )]
[0045]
[0046] where α and β both represent constraint weights, N represents the number of instances, p n , g n respectively represent the predicted output value for intestinal polyps and the true result of the intestinal polyp image under the nth instance, L BCE (p i , g i ) represents the binary cross-entropy loss, p i , g i respectively represent the prediction and label in the output result, l N represents the loss corresponding to the nth sample, w n represents the set hyperparameter, Ω represents the image domain, and Ω n represents the image domain under the nth instance;
[0047] S43. Continuously update the training parameters. When the value of the loss function is the smallest, stop training to obtain the optimal intestinal polyp image segmentation model.
[0048] Further, after the step four, the following steps are further included:
[0049] Input the test set into the optimal intestinal polyp image segmentation model to obtain the intestinal polyp segmentation result map;
[0050] Compare the intestinal polyp segmentation result map with the corresponding label to evaluate the segmentation performance of the intestinal polyp image segmentation model.
[0051] The present invention also provides an intestinal polyp segmentation system based on deep fusion of endoscopic images, and the system includes:
[0052] An image preprocessing device, configured to collect intestinal polyp images under an endoscope and perform preprocessing to obtain a training set and a test set;
[0053] A feature fusion device, configured to construct a deep feature fusion module by using a gated axial attention mechanism module and a sliding window attention mechanism module;
[0054] A model construction device, configured to construct an intestinal polyp image segmentation model by using a gated axial attention mechanism module, a deep feature fusion module, and an attention gating module;
[0055] A model training device, configured to train the intestinal polyp image segmentation model by using the training set to obtain the optimal intestinal polyp image segmentation model;
[0056] A result output device, configured to input the real-time collected intestinal polyp image under the endoscope into the optimal intestinal polyp image segmentation model to obtain a predicted segmentation image.
[0057] Further, the image preprocessing device is further configured to:
[0058] S11. Collect intestinal polyp images under an endoscope;
[0059] S12. Use the linear interpolation method to adjust the resolution of all acquired images to 224×224;
[0060] S13. For the data images, through random horizontal rotation, random vertical rotation, random deformation, random contrast and random brightness change, and random multi-scale scaling by 0.70 to 1.20 times, use the high-hat transformation in morphology to enhance the contrast between the intestinal polyp and the background. The formula for the high-hat transformation is:
[0061]
[0062] f b_hat = f - (f·b) = f - f d ,
[0063] f enchance = f t_hat - f b_hat ,
[0064] In the formula, f represents the intestinal polyp image, and f op represents the opening operation, and f d represents the closing operation. b represents the structural element in morphology, and f enchance represents the enhanced endoscopic intestinal polyp image;
[0065] S14. Use the method of adaptive threshold to extract the area where the intestinal polyp is located in the processed image above, and divide the adjusted image into two parts: a training set and a test set according to 5:1.
[0066] Furthermore, the feature fusion device is also used for:
[0067] S21. Construct a gated axial attention mechanism module: Split the original self-attention mechanism module into two modules. One module performs self-attention calculation on the height axis, and the other module performs calculation on the width axis. The calculation formula is as follows
[0068]
[0069]
[0070] Among them, W represents the width, H represents the height, q and k respectively represent the query vector and the key vector. q i,j represents the query vector at any position in i ∈ {1,...H}, j ∈ {1,...W}, and k i,W , u i,w represents the key vector and value vector at any position in i ∈ {1,...H} on a certain width axis, and represent the position biases corresponding to the query vector, key vector, and value vector, Use the gated mechanism to control the weight of the position information and update the self-attention calculation formula on the height axis, that is:
[0071]
[0072] Among them, G q , G k , are all control parameters;
[0073] S22. Construct a window-based attention mechanism module. The sliding window attention mechanism module is expressed as:
[0074] Z′ i = W_MSA(Norm(Z i-1 )) + Z i-1
[0075] Z i= FFN(Norm(Z`i)) + Z′ i
[0076] Z′ i+1 = SW_MSA(Norm(Z i )) + Z i
[0077] Z i+1 = FFN(Norm(Z`i)) + Z′ i+1
[0078] Wherein, W_MSA represents a window-based attention module through which the input features pass, SW_MSA represents a sliding window attention mechanism module, Z i-1 represents the input features of W_MSA in the i-th layer, and Z′ i represents the output features of W_MSA in the i-th layer, and also the input features of SW_MSA in the i-th layer; Norm represents normalization, i is an identifier of a certain intermediate block, and FFN represents a fully connected network of a hidden layer;
[0079] S23. Construct a deep feature fusion module: The outputs of the gated axial attention mechanism module and the sliding window attention mechanism module are reshaped to obtain feature maps, the sizes of the feature maps are adjusted by convolution operations, the feature maps of each path are fused by feature splicing, and then the fused features are obtained through convolution operations
[0080] Furthermore, the model construction device is further configured to:
[0081] Input the intestinal polyp features of the endoscopic image into the first convolutional layer and multiple gated axial attention mechanism modules to obtain different depth features. The outputs of the gated axial attention mechanism modules are respectively connected to a second convolutional layer. The first convolutional layer and all the second convolutional layers are cascaded in sequence. The output ends of the first convolutional layer and the first n - 1 second convolutional layers are respectively connected to a deep feature fusion module, where n ≥ 2. The output end of the n-th second convolutional layer is connected to a third convolutional layer. The output ends of each deep feature fusion module are respectively connected to an attention gating module. The output ends of each attention gating module are respectively connected to a fourth convolutional layer. The third convolutional layer and all the fourth convolutional layers are cascaded in sequence. The output of the last fourth convolutional layer outputs the image segmentation result, completing the construction of the deep fusion neural network, and using this neural network as the intestinal polyp image segmentation model.
[0082] Furthermore, the attention gating module includes:
[0083] Perform a linear transformation on the input tensor using a 1×1 convolutional kernel, then sequentially use the PReLU non-linear activation function and the Sigmoid activation function as two sets of attention coefficients, and finally adjust the feature dimensions to jointly fuse the input features to obtain the attention features.
[0084] Furthermore, the model training device is also used for:
[0085] S41. Input the training set into the intestinal polyp image segmentation model, and use the Adam optimizer to optimize the intestinal polyp image segmentation model. The number of training epochs is defaulted to 200, and the initial learning rate is set to 0.001;
[0086] S42. Set the loss function as: L total =αL BCE +βL blob ,
[0087] L BCE (p i ,g i )=L={l1,…,l N} T ,
[0088] l N =-w n [g n logp n +(1 - g n )log(1 - p n )]
[0089]
[0090] where α and β both represent constraint weights, N represents the number of instances, p n ,g n respectively represent the predicted output value for intestinal polyps and the true result of the intestinal polyp image under the nth instance, L BCE (p i ,g i ) represents the binary cross-entropy loss, p i ,g i respectively represent the prediction and label in the output result, l N represents the loss corresponding to the nth sample, w n represents the set hyperparameter, Ω represents the image domain, and Ω n represents the image domain under the nth instance;
[0091] S43. Continuously update the training parameters. When the value of the loss function is the smallest, stop training to obtain the optimal intestinal polyp image segmentation model.
[0092] Furthermore, after the model training device, it also includes:
[0093] Input the test set into the optimal intestinal polyp image segmentation model to obtain the intestinal polyp segmentation result map;
[0094] The intestinal polyp segmentation result image is compared with the corresponding label to evaluate the segmentation performance of the intestinal polyp image segmentation model.
[0095] The advantages of the present invention are:
[0096] (1) The present invention designs a fusion mechanism that combines a gated axial attention mechanism module with a sliding window-based attention mechanism module to form a local-global learning strategy. It uses shallow global branches and deep local branches to perform feature learning on intestinal polyp image blocks, which solves the problem of lack of a large amount of labeled data. The training cost and difficulty are low. At the same time, it obtains richer feature information, reduces the loss of spatial information, improves the robustness of the segmentation network, and improves the model accuracy.
[0097] (2) The present invention performs certain preprocessing on the endoscopic intestinal polyp image, including data enhancement. By adding changes in the intestinal polyp image, the robustness of the segmentation model is improved and overfitting is reduced. The high and low hat transformation in morphology is used to enhance the contrast between the intestinal polyp target and the background. The adaptive threshold method is used to extract the area where the intestinal polyp is located in the endoscopic image to mine the intestinal polyp boundary information.
[0098] (3) The present invention uses a joint loss function to guide the segmentation model to enhance the learning of intestinal polyp morphology and texture features, and uses training samples to train the neural network to obtain the optimal segmentation model to segment intestinal polyp images. The test set is used to optimize the multi-model prediction results to obtain the final intestinal polyp segmentation results, thereby achieving good segmentation performance.
[0099] (4) The present invention constructs an attention gating module to dynamically and implicitly generate target regions, highlighting features that are useful for characterizing intestinal polyps, thereby suppressing irrelevant feature responses and guiding the model to pay more attention to the extraction of target features. BRIEF DESCRIPTION OF THE DRAWINGS
[0100] Figure 1 This is a flow chart of a method for intestinal polyp segmentation based on deep fusion of endoscopic images disclosed in Example 1 of the present invention;
[0101] Figure 2 A schematic diagram of an intestinal polyp segmentation model in an intestinal polyp segmentation method based on deep fusion of endoscopic images disclosed in Example 1 of the present invention;
[0102] Figure 3 A schematic diagram of a gate-space axial attention mechanism model in a method for intestinal polyp segmentation based on deep fusion of endoscopic images disclosed in Example 1 of the present invention;
[0103] Figure 4 Schematic diagram of the attention mechanism module based on a sliding window in a method for segmenting intestinal polyps by depth fusion of endoscopic images disclosed in Embodiment 1 of the present invention;
[0104] Figure 5 Schematic diagram of the attention gating module in a method for segmenting intestinal polyps by depth fusion of endoscopic images disclosed in Embodiment 1 of the present invention;
[0105] Figure 6 Schematic diagram of the depth feature fusion module in a method for segmenting intestinal polyps by depth fusion of endoscopic images disclosed in Embodiment 1 of the present invention. Detailed implementation manners
[0106] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0107] Embodiment 1
[0108] As Figure 1 shown, a method for segmenting intestinal polyps by depth fusion of endoscopic images, the method includes:
[0109] S1: Collect endoscopic images of intestinal polyps and perform preprocessing to obtain a training set and a test set; the specific process is as follows:
[0110] S11. Collect abdominal organ images of multiple modalities and collect endoscopic images of intestinal polyps;
[0111] S12. Use the linear interpolation method to adjust the resolution of all acquired images to 224×224;
[0112] S13. Perform random horizontal rotation, random vertical rotation, random deformation, random contrast, and random brightness change on the data images, as well as random multi-scale scaling by 0.70 to 1.20 times. Use the high-hat transformation in morphology to enhance the contrast between the intestinal polyps and the background. The formula for the high-hat transformation is:
[0113]
[0114] f b_hat = f - (f·b) = f - f d ,
[0115] f enchance= f t_hat - f b_hat ,
[0116] where f represents the intestinal polyp image, f op represents the opening operation, f d represents the closing operation, b represents the structure element in morphology, f enchance represents the enhanced endoscopic intestinal polyp image;
[0117] S14. Use the method of adaptive threshold to extract the area where the intestinal polyp is located in the above processed image, and divide the adjusted image into two parts: a training set and a test set according to 5:1.
[0118] S2: Use the gated axial attention mechanism module and the sliding window attention mechanism module to construct a deep feature fusion module; the specific process is as follows:
[0119] S21. Construct the gated axial attention mechanism module: As Figure 3 shown, split the original self-attention mechanism module into two modules. One module performs self-attention calculation on the height axis, and the other module performs calculation on the width axis, effectively simulating the working mechanism of the original self-attention mechanism module. At the same time, in order to make the module sensitive to position information, a relative position encoding is added, and the calculation formula is as follows
[0120]
[0121]
[0122] where W represents the width, H represents the height, q and k represent the query vector and the key vector respectively, q i,j represents the query vector at any position in i ∈ {1,...H}, j ∈ {1,...W}, k i,W , u i,w represents the key vector and the value vector at any position in i ∈ {1,...H} on a certain width axis, and represent the position biases corresponding to the query vector, the key vector and the value vector, In addition, in order to be able to effectively learn the standard position information in the low-scale feature map, a gated mechanism is used to control the weight of the position information and control the influence of the position bias in the non-local context encoding. Therefore, a gated mechanism is used to control the weight of the position information and update the self-attention calculation formula on the height axis, that is:
[0123]
[0124] where G q , G k , They are all control parameters;
[0125] S22. Construct a window attention mechanism module: As Figure 4 shown, the sliding window attention mechanism module is mainly composed of two consecutive attention mechanism modules and a feed-forward network module. Among them, the two consecutive attention mechanism modules are both window-based multi-head attention mechanism modules and window-based multi-head attention mechanism modules. The sliding window attention mechanism module is expressed as:
[0126] Z′ i = W_MSA(Norm(Z i-1 )) + Z i-1
[0127] Z i = FFN(Norm(Z`i)) + Z′ i
[0128] Z′ i+1 = SW_MSA(Norm(Z i )) + Z i
[0129] Z i+1 = FFN(Norm(Z`i)) + Z′ i+1
[0130] Among them, W_MSA represents the window-based attention module through which the input features pass, SW_MSA represents the sliding window attention mechanism module, Z i-1 represents the input features of W_MSA in the i-th layer, Z′ i represents the output features of W_MSA in the i-th layer, and also the input features of SW_MSA in the i-th layer. Norm represents normalization, i is an identifier for a certain intermediate block, and FFN represents a fully connected network with a hidden layer;
[0131] S23. Construct a deep feature fusion module: As Figure 6 shown, this module includes the above-mentioned gated axial attention mechanism module, sliding window attention mechanism module, segmentation backbone convolution module, and post-processing module. In the deep feature fusion module, the outputs of the gated axial attention module and the Swin Transformer module (sliding window attention mechanism module) are reshaped to obtain feature maps, and convolution operations are used to adjust the sizes of the feature maps to match the feature maps input to the deep fusion module. The three-way features output by the gated axial attention mechanism module, sliding window attention mechanism module, and segmentation backbone convolution module are fused through feature concatenation, and the required fusion features are obtained through convolution operations
[0132] S3: Construct an intestinal polyp image segmentation model using a gated axial attention mechanism module, a deep feature fusion module, and an attention gating module; as shown in Figure 2 the figure, input the intestinal polyp features of the endoscopic image into the first convolutional layer and multiple gated axial attention mechanism modules to obtain different depth features. The outputs of the gated axial attention mechanism modules are respectively connected to a second convolutional layer. The first convolutional layer is cascaded in sequence with all the second convolutional layers. The output ends of the first convolutional layer and the first n - 1 second convolutional layers are respectively connected to a deep feature fusion module, where n ≥ 2. The output end of the nth second convolutional layer is connected to a third convolutional layer. The output end of each deep feature fusion module is respectively connected to an attention gating module. The output end of each attention gating module is respectively connected to a fourth convolutional layer. The third convolutional layer and all the fourth convolutional layers are cascaded in sequence. The output of the last fourth convolutional layer is the image segmentation result, completing the construction of the neural network with deep fusion, and using this neural network as the intestinal polyp image segmentation model. The following details the construction process of the intestinal polyp image segmentation model:
[0133] Step S31: As shown in Figure 2 the figure, use the Unet segmentation framework as the backbone network for intestinal polyp feature extraction. Extract preliminary features through encoder downsampling, and use the decoder to fuse the learned features and restore the original size. The above-mentioned first convolutional layer and second convolutional layer are the convolutional layers in the encoder, while the third convolutional layer and the fourth convolutional layer are the convolutional layers in the decoder. The encoder is mainly composed of multiple convolutional layers and pooling layers. Each convolutional layer in the encoder includes 2 convolutional blocks with a convolutional kernel size of 2×2. The formula for the convolutional block is as follows
[0134]
[0135] where N represents the number of feature maps in the lth layer, represents the weight matrix mapping from the nth feature map in the lth layer to the mth feature map in the (l + 1)th layer, * represents a 2D convolutional operation, represents the nth feature map in the lth layer, represents the corresponding bias term, and correspondingly represents the mth feature map in the (l + 1)th layer. A batch normalization function pooling block (BatchNormalization) and a non-linear activation function ReLU are added after the convolutional block. The pooling block contains a max pooling layer with a pooling window size of 2×2;
[0136]
[0137] where n and m represent the area covered by the pooling window.
[0138] The decoder is mainly composed of upsampling and skip connections, that is, the connection method shown in the third convolutional layer and the fourth convolutional layer. By using the convolutional layer of upsampling and skip connections, the initial pixels are gradually restored. Figure 2 Using the convolutional layer of upsampling and skip connections, the initial pixels are gradually restored.
[0139] Step S32: Construct a multi-scale input module, and send the input data to multiple gated axial attention mechanism modules to obtain features of different depths, improving the integrity and richness of the extraction of global and local semantic information of the image. At the same time, making the segmentation network more robust and generalizable, and also alleviating the problem of lack of a large amount of data to a certain extent. As shown in Figure 2 Send the input picture to gated axial attention mechanism modules at different levels. The outputs of the ×3, ×6, ×9, and ×12 modules are respectively used as the input of a branch of the first convolutional layer and the second convolutional layer in the encoder.
[0140] Step S33: The feature output extracted by the encoder is divided into four branches in the output direction. Three branches are input to the depth feature fusion module. The first branch is used as the input of the gated axial attention mechanism module, the second branch is used as the input of the sliding window attention mechanism module, the features of the third branch will be fused with the outputs of the first and second branch modules to obtain a new feature representation and sent to the decoder to complete the decoding process, and the fourth branch will continue to perform encoding to learn the target features.
[0141] Step S34: Dynamically and implicitly generate the target area by using the attention gating module to highlight the features useful for intestinal polyp features, thereby suppressing the response of irrelevant features. As shown in Figure 5 The attention method is to use the method of vector concatenation. A 1×1 convolutional operation is used to perform a linear transformation on the input tensor. Two sets of concatenated features are linearly mapped into a high dimension. The PReLU non-linear activation function and the Sigmoid activation function are used as the two sets of attention coefficients. Finally, the feature dimension is adjusted and combined with the input features to fuse and obtain the attention features. It can be expressed as:
[0142]
[0143] Among them, represents the attention coefficient, represents each pixel vector.
[0144] Through four skip connections, it is fused with the features obtained by the decoder upsampling, and the feature map gradually restores the initial pixel size.
[0145]
[0146] Among them, represents the feature map obtained after passing through the i th layer of attention gating module, Denotes the feature map obtained after upsampling the (i - 1)th layer, Denotes the feature map obtained after the ith skip connection.
[0147] Use the Sigmoid activation function to obtain the final output
[0148] S(x) = 1 / (1 + e -x )
[0149] x is the feature map input into the Sigmoid activation function.
[0150] S4: Use the training set to train the intestinal polyp image segmentation model to obtain the optimal intestinal polyp image segmentation model; the specific process is as follows:
[0151] S41. Input the training set into the intestinal polyp image segmentation model, and use the Adam optimizer to optimize the intestinal polyp image segmentation model. The number of training epochs is defaulted to 200, and the initial learning rate is set to 0.001;
[0152] S42. Set the loss function to improve the situation of instance imbalance in the endoscopic image. The loss function is: L total = αL BCE + βL blob ,
[0153] L BCE (p i ,g i ) = L = {l1,…,l N} T ,
[0154] l N = -w n [g n logp n +(1 - g n )log(1 - p n )]
[0155]
[0156] where α and β both represent constraint weights, N represents the number of instances, p n ,g n respectively represent the predicted output value of intestinal polyps and the true result of the intestinal polyp image under the nth instance, L BCE (p i ,g i ) represents the binary cross - entropy loss, p i ,g i respectively represent the prediction and label in the output result, l N represents the loss corresponding to the nth sample, w nLet \(\theta\) denote the set of set hyperparameters, and \(\Omega\) denote the image domain. \(\Omega_{n}\) n denotes the image domain under the \(n\)th instance;
[0157] S43. Continuously update the training parameters. When the value of the loss function is minimized, stop training to obtain the optimal intestinal polyp image segmentation model.
[0158] After that, use the test set to test the trained intestinal polyp image segmentation model. The specific process is as follows:
[0159] Input the test set into the optimal intestinal polyp image segmentation model to obtain the intestinal polyp segmentation result map;
[0160] Use the level set method to optimize the fusion model for image segmentation;
[0161] Compare the intestinal polyp segmentation result map with the corresponding label to evaluate the segmentation performance of the intestinal polyp image segmentation model.
[0162] S5: Input the endoscope intestinal polyp image collected in real time into the optimal intestinal polyp image segmentation model to obtain the predicted segmentation image.
[0163] Through the above technical solutions, the traditional Transformer structure lacks some innate inductive biases of CNN (prior experience brought by the convolutional structure), such as translational invariance and inclusion of local relationships. Therefore, it does not perform as well on datasets with insufficient scale. Thus, the present invention adds convolutional blocks, designs a fusion mechanism combining a gated axial attention mechanism module and a sliding window attention mechanism module, constructs a local-global learning strategy, uses a shallow global branch and a deep local branch to perform feature learning on intestinal polyp image patches, solves the problem of lack of a large amount of labeled data, has lower training cost and training difficulty. At the same time, it obtains richer feature information, reduces the loss of spatial information, improves the robustness of the segmentation network, and enhances the model accuracy.
[0164] Embodiment 2
[0165] Based on Embodiment 1, Embodiment 2 of the present invention further provides an intestinal polyp segmentation system based on deep fusion of endoscopic images. The system includes:
[0166] An image preprocessing device for collecting endoscope intestinal polyp images and performing preprocessing to obtain a training set and a test set;
[0167] A feature fusion device for constructing a deep feature fusion module using a gated axial attention mechanism module and a sliding window attention mechanism module;
[0168] A model construction device for constructing an intestinal polyp image segmentation model using a gated axial attention mechanism module, a deep feature fusion module, and an attention gating module;
[0169] A model training device for training an intestinal polyp image segmentation model using a training set to obtain an optimal intestinal polyp image segmentation model;
[0170] A result output device for inputting a real-time acquired endoscopic intestinal polyp image into the optimal intestinal polyp image segmentation model to obtain a predicted segmentation image.
[0171] Specifically, the image preprocessing device is further configured to:
[0172] S11. Acquire endoscopic intestinal polyp images;
[0173] S12. Use the linear interpolation method to adjust the resolution of all acquired images to 224×224;
[0174] S13. For the data images, through random horizontal rotation, random vertical rotation, random deformation, random contrast and random brightness changes, and random multi-scale scaling by 0.70 to 1.20 times, use the high-hat transformation in morphology to enhance the contrast between the intestinal polyp and the background. The formula for the high-hat transformation is:
[0175]
[0176] f b_hat = f - (f·b) = f - f d ,
[0177] f rnchance = f t_hat - f b_hat ,
[0178] In the formula, f represents the intestinal polyp image, f op represents the opening operation, f d represents the closing operation, b represents the structure element in morphology, and f enchance represents the enhanced endoscopic intestinal polyp image;
[0179] S14. Use the method of adaptive threshold to extract the region where the intestinal polyp is located in the processed image, and divide the adjusted image into two parts: a training set and a test set according to 5:1.
[0180] More specifically, the feature fusion device is further configured to:
[0181] S21. Construct a gated axial attention mechanism module: Split the original self-attention mechanism module into two modules. One module performs self-attention calculation on the height axis, and the other module performs calculation on the width axis. The calculation formula is as follows
[0182]
[0183]
[0184] Among them, W represents the width, H represents the height, q and k represent the query vector and the key vector respectively, and q i,j represents the query vector at any position where i ∈ {1,...H}, j ∈ {1,...W}, and k i,W , u i,w represent the key vector and the value vector at any position where i ∈ {1,...H} on a certain width axis, and represent the position biases corresponding to the query vector, the key vector, and the value vector, The gating mechanism is used to control the weight of the position information, and the self-attention calculation formula on the height axis is updated, that is:
[0185]
[0186] Among them, G q , G k , are all control parameters;
[0187] S22. Construct a module based on the window attention mechanism, and the module based on the sliding window attention mechanism is expressed as:
[0188] Z′ i = W_MSA(Norm(Z i-1 )) + Z i-1
[0189] Z i = FFN(Norm(Z`i)) + Z′ i
[0190] Z′ i+1 = SW_MSA(Norm(Z i )) + Z i
[0191] Z i+1 = FFN(Norm(Z`i)) + Z′ i+1
[0192] Among them, W_MSA represents the window-based attention module through which the input features pass, SW_MSA represents the module based on the sliding window attention mechanism, Z i-1 represents the input features of W_MSA in the i-th layer, and Z′ i represents the output features of W_MSA in the i-th layer, and is also the input features of SW_MSA in the i-th layer; Norm represents normalization, i is an identifier of a certain intermediate block, and FFN represents a fully connected network of a hidden layer;
[0193] S23. Construct a deep feature fusion module: The outputs of the gated axial attention mechanism module and the sliding window attention mechanism module are reshaped to obtain feature maps, the sizes of the feature maps are adjusted using convolutional operations, the feature maps of each path are fused through feature concatenation, and then fused features are obtained through convolutional operations.
[0194] More specifically, the model construction device is further configured to:
[0195] Input the intestinal polyp features of the endoscopic image into the first convolutional layer and multiple gated axial attention mechanism modules to obtain different-depth features. The outputs of the gated axial attention mechanism modules are respectively connected to a second convolutional layer, the first convolutional layer is cascaded in sequence with all the second convolutional layers, the output ends of the first convolutional layer and the first n - 1 second convolutional layers are respectively connected to a deep feature fusion module, n≥2, the output end of the nth second convolutional layer is connected to a third convolutional layer, the output ends of each deep feature fusion module are respectively connected to an attention gating module, the output ends of each attention gating module are respectively connected to a fourth convolutional layer, the third convolutional layer and all the fourth convolutional layers are cascaded in sequence, and the image segmentation result is output by the last fourth convolutional layer, completing the construction of the deep fusion neural network, and using this neural network as the intestinal polyp image segmentation model.
[0196] More specifically, the attention gating module includes:
[0197] Perform a linear transformation on the input tensor using a 1×1 convolutional operation with a convolutional kernel, then use the PReLU non-linear activation function and the Sigmoid activation function in sequence as two sets of attention coefficients, and finally adjust the feature dimensions and jointly fuse the input features to obtain the attention features.
[0198] More specifically, the model training device is further configured to:
[0199] S41. Input the training set into the intestinal polyp image segmentation model, optimize the intestinal polyp image segmentation model using the Adam optimizer, the default number of training epochs is 200, and the initial learning rate is set to 0.001;
[0200] S42. Set the loss function as: L total =αL BCE +βL blob ,
[0201] L BCE (p i ,g i ) = L = {l1,…,l N} T ,
[0202] l N =-w n[g n logp n +(1 - g n )log(1 - p n )]
[0203]
[0204] where α and β both represent constraint weights, N represents the number of instances, p n , g n respectively represent the predicted output value for intestinal polyp prediction and the true result of the intestinal polyp image under the nth instance, L BCE (p i , g i ) represents the binary cross - entropy loss, p i , g i respectively represent the prediction and label in the output result, l N represents the loss corresponding to the nth sample, w n represents the set hyperparameter, Ω represents the image domain, Ω n represents the image domain under the nth instance;
[0205] S43. Continuously update the training parameters. When the value of the loss function is the smallest, stop the training to obtain the optimal intestinal polyp image segmentation model.
[0206] More specifically, after the model training device, there is further included:
[0207] Input the test set into the optimal intestinal polyp image segmentation model to obtain the intestinal polyp segmentation result map;
[0208] Compare the intestinal polyp segmentation result map with the corresponding label to evaluate the segmentation performance of the intestinal polyp image segmentation model.
[0209] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An intestinal polyp segmentation method based on deep fusion of endoscopic images, characterized in that, The method includes: Step 1: Collect endoscopic intestinal polyp images and preprocess them to obtain a training set and a test set; Step 2: Construct a deep feature fusion module using a gated axial attention mechanism module and a sliding window attention mechanism module; The Step 2 includes: S21. Construct a gated axial attention mechanism module: Split the original self-attention mechanism module into two modules. One module performs self-attention calculation on the height axis, and the other module performs calculation on the width axis. The calculation formula is as follows Among them, represents the width, represents the height, and represent the query vector and the key vector respectively, represents the query vector at any position in i {1,...H}, j {1,...W}, and represent the key vector and the value vector at any position in i {1,...H} on a certain width axis, and represent the position biases corresponding to the query vector, the key vector, and the value vector; a gating mechanism is used to control the weight of the position information, and the self-attention calculation formula on the height axis is updated, that is: Among them, , , , are all control parameters; S22. Construct a window-based attention mechanism module. The sliding window attention mechanism module is expressed as: = W_MSA(Norm( )) + = FFN(Norm(Z`i)) + = SW_MSA(Norm( )) + = FFN(Norm(Z`i)) + Among them, W_MSA represents the window-based attention module through which the input features pass, and SW_MSA represents the sliding window attention mechanism module. represents the input features of W_MSA in the i-th layer. represents the output features of W_MSA in the i-th layer, and also the input features of SW_MSA in the i-th layer; Norm represents normalization, i is an identifier for a certain intermediate block, and FFN represents a fully connected network with one hidden layer. S23. Construct a deep feature fusion module: The outputs of the gated axial attention mechanism module and the sliding window attention mechanism module are reshaped to obtain feature maps. The sizes of the feature maps are adjusted using convolutional operations. The feature maps of each path are fused through feature concatenation, and then the fused features are obtained through convolutional operations. ; Step 3: Construct an intestinal polyp image segmentation model using the gated axial attention mechanism module, the deep feature fusion module, and the attention gating module; Step 4: Train the intestinal polyp image segmentation model using the training set to obtain the optimal intestinal polyp image segmentation model; Step 5: Input the real-time collected endoscopic intestinal polyp image into the optimal intestinal polyp image segmentation model to obtain a predicted segmentation image.
2. The intestinal polyp segmentation method based on depth fusion of endoscopic images according to claim 1, wherein The Step 1 includes: S11. Collect endoscopic intestinal polyp images; S12. Use the linear interpolation method to adjust the resolution of all acquired images to 224×224; S13. For the data images, through random horizontal rotation, random vertical rotation, random deformation, random contrast and random brightness change, and random multi-scale scaling by 0.70 to 1.20 times, use the high-hat transformation in morphology to enhance the contrast between the intestinal polyp and the background. The formula for the high-hat transformation is: = -( ) = - , = -( ) = - , = - , In the formula, represents the intestinal polyp image, represents the opening operation, represents the closing operation, represents the structural element in morphology, represents the enhanced endoscopic intestinal polyp image; S14. Use the method of adaptive threshold to extract the area where the intestinal polyp is located in the processed images above, and divide the adjusted images into a training set and a test set in a ratio of 5:
1.
3. The method for segmenting intestinal polyps based on depth fusion of endoscopic images according to claim 1, wherein, The Step 3 includes: Input the intestinal polyp features of the endoscopic image into the first convolutional layer and multiple gated axial attention mechanism modules to obtain different depth features. The outputs of the gated axial attention mechanism modules are respectively connected to a second convolutional layer. The first convolutional layer and all the second convolutional layers are cascaded in sequence. The output ends of the first convolutional layer and the first n - 1 second convolutional layers are respectively connected to a deep feature fusion module, n≥2. The output end of the nth second convolutional layer is connected to a third convolutional layer. The output ends of each deep feature fusion module are respectively connected to an attention gating module. The output ends of each attention gating module are respectively connected to a fourth convolutional layer. The third convolutional layer and all the fourth convolutional layers are cascaded in sequence. The output of the last fourth convolutional layer outputs the image segmentation result, completing the construction of the deep fusion neural network. This neural network is used as the intestinal polyp image segmentation model.
4. A method for segmenting intestinal polyps based on depth fusion of endoscopic images according to claim 3, characterized in that, The attention gating module includes: Perform a linear transformation on the input tensor using a 1×1 convolutional operation with a convolutional kernel, then use the PReLU non-linear activation function and the Sigmoid activation function in sequence as two sets of attention coefficients, and finally adjust the feature dimension to jointly fuse the input features to obtain the attention features.
5. A method for segmenting intestinal polyps based on depth fusion of endoscopic images according to claim 3, characterized in that, The Step 4 includes: S41. Input the training set into the intestinal polyp image segmentation model, and use the Adam optimizer to optimize the intestinal polyp image segmentation model. The default number of training epochs is 200, and the initial learning rate is set to 0.001; S42. Set the loss function as: , ( )=L={ ,…, , where α and β both represent constraint weights, N represents the number of instances, , respectively represent the predicted output value for intestinal polyp prediction and the true result of the intestinal polyp image under the th instance, ( )represents the binary cross-entropy loss, respectively represent the prediction and label in the output result, represents the loss corresponding to the nth sample, represents the set hyperparameter, Ω represents the image domain, represents the th image domain under the instance; S43. Continuously update the training parameters. When the value of the loss function is the smallest, stop the training to obtain the optimal intestinal polyp image segmentation model.
6. The method for segmenting intestinal polyps based on depth fusion of endoscopic images according to claim 5, wherein After the above step four, it further includes: Input the test set into the optimal intestinal polyp image segmentation model to obtain the intestinal polyp segmentation result map; Compare the intestinal polyp segmentation result map with the corresponding label to evaluate the segmentation performance of the intestinal polyp image segmentation model.
7. An intestinal polyp segmentation device based on deep fusion of endoscopic images, characterized in that, The device includes: An image preprocessing device, configured to collect intestinal polyp images under an endoscope and perform preprocessing to obtain a training set and a test set; A feature fusion device, configured to construct a deep feature fusion module by using a gated axial attention mechanism module and a sliding window attention mechanism module; the feature fusion device is further configured to: S21. Construct a gated axial attention mechanism module: Split the original self-attention mechanism module into two modules. One module performs self-attention calculation on the height axis, and the other module performs calculation on the width axis. The calculation formula is as follows Among them, represents the width, represents the height, and represent the query vector and the key vector respectively, represents the query vector at any position in i {1,...H}, j {1,...W} of the query vector at any position, and represent the key vector and the value vector at any position in i {1,...H} on a certain width axis, and represent the position biases corresponding to the query vector, the key vector, and the value vector; a gating mechanism is used to control the weight of the position information, and the self-attention calculation formula on the height axis is updated, that is: Among them, , , , are all control parameters; S22. Construct a window-based attention mechanism module. The sliding window attention mechanism module is expressed as: = W_MSA(Norm( )) + = FFN(Norm(Z`i)) + = SW_MSA(Norm( )) + = FFN(Norm(Z`i)) + Among them, W_MSA represents the window-based attention module through which the input features pass, and SW_MSA represents the sliding window attention mechanism module. represents the input features of W_MSA in the i-th layer. represents the output features of W_MSA in the i-th layer, which are also the input features of SW_MSA in the i-th layer; Norm represents normalization, i is an identifier for a certain intermediate block, and FFN represents a fully connected network with one hidden layer. S23. Construct a deep feature fusion module: The outputs of the gated axial attention mechanism module and the sliding window attention mechanism module are reshaped to obtain feature maps. The sizes of the feature maps are adjusted using convolutional operations. The feature maps of each path are fused through feature concatenation, and then the fused features are obtained through convolutional operations. ; A model construction device, configured to construct an intestinal polyp image segmentation model by using a gated axial attention mechanism module, a deep feature fusion module, and an attention gating module; A model training device, configured to train the intestinal polyp image segmentation model by using the training set to obtain the optimal intestinal polyp image segmentation model; A result output device, configured to input the real-time collected intestinal polyp image under an endoscope into the optimal intestinal polyp image segmentation model to obtain a predicted segmentation image.
8. An intestinal polyp segmentation device based on depth fusion of endoscopic images according to claim 7, characterized in that, The image preprocessing device is further configured to: S11. Collect intestinal polyp images under an endoscope; S12. Use the linear interpolation method to adjust the resolution of all acquired images to 224×224; S13. Perform random horizontal rotation, random vertical rotation, random deformation, random contrast, and random brightness change on the data images, as well as random multi-scale scaling by a factor of 0.70 to 1.
20. Use the high-hat and low-hat transformation in morphology to enhance the contrast between the intestinal polyp and the background. The formula for the high-hat and low-hat transformation is: = -( ) = - , = -( ) = - , = - , In the formula, represents the intestinal polyp image, represents the opening operation, represents the closing operation, represents the structural element in morphology, represents the enhanced endoscopic intestinal polyp image; S14. Use the method of adaptive threshold to extract the area where the intestinal polyp is located in the above processed image, and divide the adjusted image into two parts, namely a training set and a test set, according to a ratio of 5:
1.
9. The intestinal polyp segmentation device based on depth fusion of endoscopic images according to claim 7, characterized in that, The model construction device is further configured to: Input the intestinal polyp features of the endoscopic image into the first convolutional layer and multiple gated axial attention mechanism modules to obtain features of different depths. The outputs of the gated axial attention mechanism modules are respectively connected to a second convolutional layer. The first convolutional layer is cascaded sequentially with all the second convolutional layers. The output ends of the first convolutional layer and the first n-1 second convolutional layers are respectively connected to a depth feature fusion module, where n≥2. The output end of the nth second convolutional layer is connected to a third convolutional layer. The output end of each depth feature fusion module is respectively connected to an attention gating module. The output end of each attention gating module is respectively connected to a fourth convolutional layer. The third convolutional layer and all the fourth convolutional layers are cascaded sequentially. The output of the last fourth convolutional layer is the image segmentation result, completing the construction of the neural network for depth fusion. Use this neural network as the intestinal polyp image segmentation model.
10. The intestinal polyp segmentation device based on depth fusion of endoscopic images according to claim 9, characterized in that, The attention gating module includes: Perform a linear transformation on the input tensor using a 1×1 convolutional operation with a convolutional kernel, then sequentially use the PReLU non-linear activation function and the Sigmoid activation function as two sets of attention coefficients, and finally adjust the feature dimensions to jointly fuse the input features to obtain the attention features.
Citation Information
Patent Citations
Colonic polyp image segmentation method based on cellular automaton model
CN107146229A
Retinal blood vessel image segmentation method based on double-channel U-shaped improved Transform network
CN114820632A
Colon polyp segmentation method based on cascade structure attention mechanism network
CN114897870A