Polyp image segmentation method

Through the improved polyp image segmentation model, combined with CSEBlock, GCT module, PA module and Self-Attention mechanism, the problem of low polyp segmentation accuracy in small and medium-sized goals and complex backgrounds of the existing technology is solved, and higher segmentation accuracy and robustness are achieved.

CN120374978APending Publication Date: 2025-07-25SHENYANG UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510460077.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing polyp image segmentation method has low accuracy when processing small target data sets and complex backgrounds, making it difficult to effectively segment polyp regions in complex backgrounds.

Method used

Using an improved polyp image segmentation model, including a PVT module, an improved cascade fusion module (CFM*), an improved camouflage identification module (CIM*) and an improved similar aggregation module (SAM*), an improved global and local feature extraction capabilities, reduce background interference, and improve segmentation accuracy by introducing CSEBlock, GCT module, PA module and Self-Attention mechanism.

Benefits of technology

It improves the accuracy of polyp image segmentation, enhances the model's ability to recognize complex backgrounds and small targets, and improves segmentation accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374978A_ABST
    Figure CN120374978A_ABST
Patent Text Reader

Abstract

The invention provides a polyp image segmentation method, and relates to the technical field of medical image processing. The polyp image segmentation method comprises the following steps: establishing a polyp image segmentation model, wherein the polyp image segmentation model comprises a PVT module, an improved cascade fusion module, an improved camouflage identification module and an improved similar aggregation module; the method comprises the following steps: replacing Basic Conv2d in charge of extracting high-level features in an original cascade fusion module with CSEBlock to obtain an improved cascade fusion module; the sigmoid (out) of the original CA module of the original camouflage identification module is replaced by sigmoid (out) * x to obtain an improved CA module, and the sigmoid (out) of the original SA module of the original camouflage identification module is replaced by sigmoid (out) * x to obtain an improved SA module; a GCT module is introduced in front of the improved CA module, a PA module is introduced between the improved CA module and the improved SA module, and an improved camouflage identification module is obtained; a Self-Attention module is added between Wz operation and addition operation of an original similar aggregation module, an improved similar aggregation module is obtained, and the recognition accuracy is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and particularly to a polyp image segmentation method. Background Art

[0002] Polyp segmentation is a key task in medical image analysis, aiming to automatically detect and accurately segment polyp regions from colonoscopy images to assist doctors in early diagnosis and treatment. Its main characteristics are that the morphology, size, and color of polyps vary greatly, the boundaries may be blurred and easily confused with surrounding tissues, making the segmentation task highly challenging. Specific difficulties include that the color of polyps is similar to that of normal tissues, resulting in unclear boundaries, the effects of illumination changes, motion blur, and noise on the segmentation effect, and small polyps are easily overlooked or misdetected. Currently, the research on polyp segmentation tasks at home and abroad is mainly based on traditional polyp segmentation methods and deep learning-based segmentation methods.

[0003] Traditional polyp segmentation methods rely on handcrafted features and classical image processing techniques, such as threshold segmentation, region growing, and active contour models (ACM). Usually, image preprocessing is first performed to extract color, texture, and edge features, and the Otsu threshold, K-means clustering, or level set method is used to segment the polyp region, and finally the result is optimized through morphological operations. However, these methods have poor adaptability, are easily interfered by illumination and noise, are difficult to handle polyp morphological changes, and have limited detection ability for small polyps.

[0004] Polyp segmentation has made significant progress driven by deep learning. Early methods based on traditional image processing have limited effects in complex backgrounds and low contrast situations, while CNNs, especially U-Net and its variants (such as ResUNet, AttentionU-Net), perform excellently in automatic feature extraction. However, CNNs mainly rely on local convolutions and are prone to ignoring global information.

[0005] In recent years, the introduction of Transformer architectures (such as TransUNet, Swin-UNet) has enhanced the global feature extraction ability using self-attention mechanisms and improved the recognition effect for polyps of different sizes. In addition, the application of multi-modal learning and transfer learning has enhanced the generalization ability of the model on small-sample datasets.

[0006] Vision Transformer (ViT) has a stronger global feature capture ability compared to CNNs. The self-attention mechanism can model long-distance pixel relationships. In addition, ViT has stronger generalization ability after large-scale pre-training, is suitable for small-sample learning, and the attention mechanism improves the model interpretability, which is particularly important in fields such as medical image analysis.

[0007] However, although the ViT model shows advantages in global feature extraction, it has some limitations. For example, ViT divides the input image into image patches of a fixed size for processing. This fixed-granularity feature extraction method may ignore the detailed features of the polyp region, especially the key information capture of small target polyps is insufficient. In addition, the self-attention mechanism of ViT may cause the loss of low-level visual features in deep networks, further weakening the model's segmentation ability for complex boundaries and detailed information. These problems significantly affect the model's segmentation performance in complex backgrounds. The problems existing in the above ViT model affect the final polyp segmentation accuracy.

[0008] In the existing polyp image segmentation methods, the segmentation models used have low accuracy when dealing with small target datasets and complex backgrounds. Therefore, it is necessary to further optimize the segmentation model to provide a new polyp image segmentation method to improve the segmentation accuracy. Summary of the Invention

[0009] The present invention proposes a polyp image segmentation method, aiming to solve the problem of low accuracy of existing segmentation models when dealing with small target datasets and complex backgrounds.

[0010] The present invention provides a polyp image segmentation method, including the following steps: establishing a polyp image segmentation model, the polyp image segmentation model includes a PVT module, an improved cascaded fusion module (CFM*), an improved camouflage identification module (CIM*), and an improved similarity aggregation module (SAM*); wherein, replacing the BasicConv2d responsible for extracting high-level features in the original cascaded fusion module (CFM) with a CSEBlock to obtain the improved cascaded fusion module (CFM*); replacing sigmoid(out) of the original CA module in the original camouflage identification module (CIM) with sigmoid(out)×x to obtain an improved CA module, and replacing sigmoid(out) of the original SA module in the original camouflage identification module (CIM) with sigmoid(out)×x to obtain an improved SA module, where out represents the vector output through a convolutional or fully connected layer, x represents the input feature map, and sigmoid(out) is used as the weighting coefficient for each channel; introducing a GCT module before the improved CA module, and introducing a PA module between the improved CA module and the improved SA module to obtain the improved camouflage identification module (CIM*); adding a Self-Attention module between the W z operation and the addition operation in the original similarity aggregation module (SAM) to obtain the improved similarity aggregation module (SAM*).

[0011] The improved cascaded fusion module (CFM*) replaces part of the BasicConv2d with the CSEBlock, enhances the global information through global pooling and channel attention, and enhances the spatial features through 1×1 convolution, improving the polyp segmentation accuracy and reducing background interference. The improved camouflage identification module (CIM*) introduces the channel-level gated context transformation mechanism (GCT module) and pixel attention (PA module). The GCT module enhances the global perception ability, and the PA module enhances the local information, enabling the model to focus more precisely on the polyp area, reducing background interference, and thus improving the segmentation accuracy. In addition, the channel attention (CA module) and spatial attention (SA module) mechanisms are improved, strengthening the attention to key regions. The improved CA module dynamically adjusts the weights of each channel by multiplying sigmoid(out) with the input feature map. The improved SA module also introduces this idea of dynamic weighting, calculates the weighted coefficients of spatial positions, and combines them with the spatial dimension of the input feature map. The improved similarity aggregation module (SAM*) introduces the Self-Attention mechanism, effectively capturing the long-range dependencies of the polyp area and further improving the segmentation accuracy.

[0012] Furthermore, the polyp image segmentation method further includes the following steps: Input the image to be segmented into the polyp image segmentation model. The PVT module is used to extract the multi-scale long-range dependency features in the image to be segmented, and the multi-scale long-range dependency features include high-level features and low-level features. The improved cascaded fusion module (CFM*) is used to fuse the high-level features to obtain a high-level feature map. The improved camouflage identification module (CIM*) is used to process the low-level features, remove noise, and enhance the low-level representation information of the polyp to obtain a low-level feature map. The improved similarity aggregation module (SAM*) aligns and fuses the high-level feature map and the low-level feature map to obtain the segmentation prediction feature map as the output.

[0013] Furthermore, the PVT module includes an MLP, Window Attention, a Block, Overlap Patch Embed, Pyramid Vision Transformer Impr, and Depthwise Separable Convolution (DWConv); in the PVT module, the MLP is used for feature transformation, Window Attention is used to provide an efficient local attention mechanism, the Block is used to gradually extract features through multi-head attention and MLP layers, Overlap Patch Embed is used for patch embedding, Pyramid Vision Transformer Impr is used for multi-stage feature extraction, and Depthwise Separable Convolution (DWConv) is used to improve efficiency through depthwise separable convolution; the image to be segmented is passed into the PVT model of the PVT module to extract low-level feature x1; then the output low-level feature x1 is passed into the PVT model to extract high-level feature x2; then the output high-level feature x2 is passed into the PVT model to extract high-level feature x3; finally, the output high-level feature x3 is passed into the PVT model to extract high-level feature x4.

[0014] Furthermore, the improved Cascaded Fusion Module (CFM*) includes a CSEBlock and BasicConv2d; the improved Cascaded Fusion Module (CFM*) consists of two cascaded parts; in the first part, the high-level feature x4 is first upsampled to the same size as the high-level feature x3, and then processed by two convolutional units CSEBlock to obtain a smoothed feature map and perform a Hadamard product on the high-level feature x3, and then concatenate the result with along the channel dimension, and then process it through the convolutional unit BasicConv2d to obtain a fused feature map

[0015] In the second part, the high-level feature x4, the high-level feature x3, and are first upsampled to the same size as the high-level feature x2, and then the convolutional unit BasicConv2d is used to smooth these upsampled feature maps; then perform a Hadamard product operation on the smoothed x4 and x3 with x2, and concatenate the resulting map with the upsampled and smoothed ; input the concatenated feature map into two convolutional units BasicConv2d for dimensionality reduction processing to obtain the final output feature map T1.

[0016] ​Furthermore, the CSEBlock includes a 1x1 convolutional layer, batch normalization, ReLU activation function, GlobalPooling, fully connected layers, sigmoid activation function, and context convolutional layer; the input features are passed through a 1x1 convolutional layer to convert the number of input channels to the number of output channels, and batch normalization and ReLU activation function are applied to extract and enhance features; then Global Pooling is used to compress the feature map to obtain a global description of each channel; then, the features pass through two fully connected layers, where the first fully connected layer reduces the number of channels and then restores to the original number of channels through the second fully connected layer to learn the weights of the channels; the weight coefficients are obtained through the sigmoid activation function, and the weight coefficients represent the importance of each channel and are used to weight the input features; the context convolutional layer uses 1x1 depth convolution to perform fine-grained context modeling on the input features; the weighted features are combined with the original features through the Hadamard product and then fused with the features obtained by the context convolution to obtain the final fused feature map.

[0017] Furthermore, the improved camouflage identification module (CIM*) includes a GCT module, an improved CA module, a PA module, and an improved SA module; the GCT module includes learnable parameters for controlling the scaling and offset of feature maps; according to the L1 mode or L2 mode, the module calculates the embedded features of the input features and performs corresponding processing; a gating value is generated through the tanh activation function to adjust the weights of each channel of the input features; the improved CA module includes the following structure: two adaptive pooling layers, namely average pooling and max pooling, two convolutional layers, ReLU activation function, and sigmoid activation function; the PA module includes a convolutional layer, a batch normalization layer, and a sigmoid activation function; the convolutional layer is used for feature learning within the channel, the batch normalization layer helps to stabilize the training process, and the sigmoid activation function is used to generate pixel-level attention weights; the improved SA module includes a convolutional layer with two input channels and a sigmoid activation function, the convolutional layer adopts an adjustable convolutional kernel size, and corresponding padding is set according to the convolutional kernel size; by concatenating the results of average pooling and max pooling and inputting them into the convolutional layer, spatial attention weights are finally generated through the sigmoid activation function; the low-level features x1 extracted from the PVT module enter the embedding calculation layer of the GCT module, and through feature normalization and weighting operations based on the L1 mode or L2 mode norm, a weighted feature map is obtained, and the output weighted feature map is fed into the convolutional layer of the improved CA module for channel attention weighting processing to obtain a dynamically weighted feature map; the dynamically weighted feature map processed by the improved CA module enters the convolutional layer of the PA module for pixel-level modulation processing to generate a pixel-level modulated modulation feature map, and the modulation feature map processed by the PA module enters the convolutional layer of the improved SA module for spatial attention weighting processing to adjust the weights in the spatial dimension, and a low-level feature map after spatial weighting adjustment is output.

[0018] Furthermore, the improved similarity aggregation module (SAM*) includes an adaptive average pooling layer, three convolutional layers, GCN, and Self-Attention; the specific operation method of the improved similarity aggregation module (SAM*) is as follows: the high-level feature map output by CFM* is input into the convolutional operation of SAM* for linear transformation processing to obtain feature maps Q and K; the low-level feature map output by CIM* is input into the spatial attention module of SAM*, and through convolutional operation, interpolation, and Softmax, an attention map T2' is generated, and then it is multiplied by K through Hadamard product and subjected to adaptive pooling and cropping, and finally a feature map V is obtained; the SAM* module performs weighting, GCN, Self-Attention, and inner product calculations on the feature maps K and V, and finally obtains a segmentation prediction feature map Z.

[0019] Further, the polyp image segmentation method further includes the following steps: obtaining a polyp image dataset, selecting a part of the images in the polyp image dataset as the training set, and another part of the images as the test set; training the polyp image segmentation model using the training set, continuously monitoring the recognition accuracy during training, and saving the weight parameters with the highest recognition accuracy during the training process as the optimal weights; loading the optimal weights into the polyp image segmentation model, testing the polyp image segmentation model using the test set, and calculating the accuracy of the test set to verify the recognition effect of the polyp image segmentation model.

[0020] Further, preprocess five public datasets, namely Kvasir-SEG, ClinicDB, ColonDB, Endoscene, and ETIS, to obtain a polyp image dataset.

[0021] Further, uniformly adjust the sizes of the images in the five public datasets of Kvasir-SEG, ClinicDB, ColonDB, Endoscene, and ETIS to 352*352, and the size of the sample quantity used during training is 16.

[0022] The polyp image segmentation method provided by the present invention optimizes the polyp image segmentation model. Through experimental verification, the optimized polyp image segmentation model of the present invention has a higher recognition accuracy compared with the existing models.

[0023] Based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become easily understandable. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:

[0025] Figure 1 is a flowchart of the polyp image segmentation method according to an optional embodiment of the present invention;

[0026] Figure 2 is an overall structure diagram of the polyp image segmentation model according to an optional embodiment of the present invention;

[0027] Figure 3 is a structure diagram of the CFM module before improvement;

[0028] Figure 4 is a structure diagram of the improved CFM* module of the present invention;

[0029] Figure 5 is Figure 4The structural diagram of the CSEBlock in;

[0030] Figure 6 The structural diagram of the CIM module before improvement;

[0031] Figure 7 The structural diagram of the improved CIM* module of the present invention;

[0032] Figure 8 For Figure 7 The structural diagram of the GCT module in;

[0033] Figure 9 The structural diagram of the CA module;

[0034] Figure 10 For Figure 7 The structural diagram of the PA module in;

[0035] Figure 11 The structural diagram of the SA module;

[0036] Figure 12 The structural diagram of the SAM module before improvement;

[0037] Figure 13 The structural diagram of the improved SAM* module of the present invention;

[0038] Figure 14 The comparison diagram before and after polyp image segmentation processing. Detailed implementation manners

[0039] Hereinafter, the exemplary embodiments disclosed in the present application will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be completely conveyed to those skilled in the art. Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art.

[0040] In combination with Figures 1 to 14 As shown, an alternative embodiment of the present invention provides a polyp image segmentation method, which is applicable to segmenting colon polyp images based on complex backgrounds and small targets. In this embodiment, the specific steps of the method include:

[0041] Step 1: Data preprocessing: Five challenging public datasets, namely Kvasir-SEG, ClinicDB, ColonDB, Endoscene, and ETIS, are used to preprocess the images in the dataset, and the preprocessed images are used as training and test data;

[0042] Specifically, Kvasir-SEG was collected from the polyp category in the Kvasir dataset, including a total of 1000 polyp images; ClinicDB contains 612 images, all of which were extracted from 31 colonoscopy videos. In addition, ETIS contains 196 images, the ColonDB dataset contains 380 images, and EndoScene contains 60 images. In the experiment, 900 and 548 images from the Kvasir-SEG and ClinicDB datasets were used as the training set, and 100 images and 64 images were selected from Kvasir-SEG and ClinicDB respectively as the test set. In addition, to verify the generalization ability of the model, the ETIS, ColonDB, and EndoScene datasets were additionally used as additional test sets to evaluate the model performance, that is, the model was tested on these three datasets of ETIS, ColonDB, and EndoScene.

[0043] In the data preprocessing stage, in order to improve the training efficiency and generalization ability of the model, the input data was processed in multiple aspects. First, considering the differences in the sizes of polyp images, images of different polyp sizes were used for training, which can enhance the robustness of the model to polyps of different sizes. Second, the sizes of all input images were uniformly adjusted to 352×352 pixels, thereby standardizing the data input and reducing the computational complexity and training instability caused by size differences. In addition, the invention sets the amount of data processed each time to 16 images to optimize the training efficiency and ensure that the model can learn efficiently within the limited video memory space.

[0044] Step 2: Training: Design an improved polyp image segmentation model, and use the training set to train the improved polyp image segmentation model, and save the weight parameters when the recognition accuracy is the highest during the training process.

[0045] Specifically, the present invention optimizes and improves based on the Polyp-PVT model. Compared with the original Polyp-PVT model, the present invention designs an improved model to solve the problem of poor segmentation effect of the Polyp-PVT model when facing complex backgrounds and small targets. As Figure 2As shown in the figure, the improved model consists of four main parts: the PVT module (a), CFM* (b), CIM* (c), and SAM* (d). The pre-processed dataset images are first input into the PVT module for feature extraction to extract multi-scale long-range dependence features. The convolutional units in the PVT module are used to adjust the number of channels of the high-level features. The extracted high-level features are sent to CFM* for feature fusion to obtain feature maps. At the same time, the extracted low-level features are sent to CIM* to remove noise and enhance the low-level feature information of polyps. Finally, the feature maps generated by CIM* and CFM* are sent to SAM* for alignment and fusion to obtain the final feature maps, thereby obtaining the segmentation prediction (this process is one round of model training). By using the preset weights and hyperparameters, combined with the accuracy evaluation of multiple rounds of training and the validation set, the optimal model weights are dynamically saved to ensure that the model can efficiently and accurately capture the key features of the polyp area and perform precise segmentation.

[0046] In this application, as Figure 2 shown, the PVT module adopts an existing structure.

[0047] As Figure 3 and Figure 4 shown, in this application, an improved CFM* module is designed to improve the original cascaded fusion module CFM module. It is proposed to replace the BasicConv2d responsible for extracting high-level features with CSEBlock to enhance the global information modeling ability and feature expression ability, thus forming CFM*. Replacing the BasicConv2d convolution in the high-level feature part of CFM with CSEBlock provides stronger global information modeling and feature enhancement capabilities. Ordinary 3×3 convolutions mainly focus on local feature extraction, while CSEBlock draws on the idea of the channel attention mechanism. It extracts channel-level global information, that is, the overall information of each feature map, through global average pooling, and calculates the attention weights using fully connected layers (FC1 and FC2), and then normalizes through sigmoid, enabling the model to enhance the feature expression of key channels while suppressing irrelevant features.

[0048] Although both the CSEBlock and the SE (Squeeze-and-Excitation) mechanism adopt the idea of channel attention mechanism, different from the SE mechanism which only focuses on channel weighting, the CSEBlock further introduces a context enhancement module to effectively model spatial features through 1×1 convolution. That is, through the context enhancement module, 1×1 convolution is applied to the feature map, enabling the spatial features to be effectively modeled, making up for the deficiency that the SE mechanism only focuses on channel information and the limitation that the SE mechanism only focuses on channel information while ignoring the spatial structure. This makes the CFM* module not only able to process channel information but also enhance the expressive ability of spatial features. Finally, through the fusion of channel weighting and context enhancement, the CSEBlock makes the model more accurate when focusing on the polyp region, reduces background interference, and improves the segmentation accuracy.

[0049] Design an improved CFM* module, and propose to replace the BasicConv2d responsible for extracting high-level features with CSEBlock, combined with Figure 2 part b in Figure 4 and Figure 5 the CSEBlock structure, the specific operation method of the CFM* module is as follows:

[0050] That is, the three high-level features x2, x3, and x4 extracted by PVT are input into the CFM*. The highest-level feature map is upsampled to the same size as the lower-level feature map and processed through two convolutional units to obtain two sets of feature maps. One set of feature maps is multiplied by the lower-level feature map through the Hadamard product and concatenated with the other set of feature maps, and then smoothed through a convolutional unit to obtain a fused feature map. Then multiple feature maps are upsampled and processed by convolution. After the Hadamard product and smoothing processing, the feature maps are concatenated, and finally, the final fused feature map is obtained by using convolution for dimensionality reduction.

[0051] Among them, x4 extracts features through the CSEBlock, and the specific operation method of the CSEBlock structure is as follows: The input feature x4 has a shape of [16, 32, 8, 8], and the number of channels of the feature map is adjusted through a 1×1 convolution operation. The output shape of the convolution operation is [16, 32, 8, 8]. After this layer, the feature map is normalized through batch normalization, and then non-linearly transformed through an activation function (ReLU), and the output shape remains [16, 32, 8, 8].

[0052] Next, the feature x4 after ReLU activation is passed into the Global Pooling layer. This pooling operation averages the spatial dimensions (height and width) of each channel, and finally compresses each channel into a single value. The output shape after pooling is [16, 32, 1, 1], and at this time, each channel only retains its global average value.

[0053] Then, the pooling result is flattened into a vector [16, 32] and passed through two fully connected layers. In the first fully connected layer, the feature dimension is compressed, then a non-linear transformation is performed through the ReLU activation function, and then the dimension is restored to the original feature dimension through the second fully connected layer to learn the importance of channels.

[0054] Next, through the sigmoid activation function, a weight map ranging from 0 to 1 is generated, with a shape of [16, 32, 1, 1]. The weight map is then expanded to [16, 32, 8, 8], and then multiplied element-wise with the input feature x to complete channel-level weighting. At the same time, the input feature x is fed into a 1×1 convolutional layer, and this convolutional operation processes each channel independently. The output shape after convolution is [16, 32, 8, 8], which represents the extracted context information. Finally, the features after context extraction and the weighted features are added element-wise to obtain the final output feature map with a shape of [16, 32, 8, 8].

[0055] Compared with the original CIM module, in order to avoid the influence of the complex background of polyp images and improve the segmentation accuracy, the present invention proposes an improved designed CIM* module, as Figure 6 and Figure 7 shown, this application designs an improved CIM* module, improves the channel attention CA (Channel Attention) module and the spatial attention SA (Spatial Attention) module, and newly adds a pixel attention PA module (PixelAttention). At the same time, the gated context transformation mechanism GCT (GatedContext Transformation) is introduced into the camouflage recognition module CIM (Camouflage IdentificationModule) to form CIM*, so as to further enhance its performance.

[0056] As Figure 8As shown, the GCT module adopts an existing structure. GCT enhances the self - adaptability of feature expression through global channel feature modeling and gating regulation, significantly improving the robustness and accuracy of the model. First, GCT calculates the global features of each channel through L1 or L2 normalization, which helps capture the global dependencies between channels. Through the global channel features, GCT provides the model with stronger global context information, helping the model focus on key features and improving the segmentation accuracy. On this basis, the gating mechanism introduced by GCT further enhances the self - adaptability of the model. Through the tanh activation function, GCT generates gating weights that can dynamically adjust the intensity of channel features, determining which channels should be enhanced and which should be suppressed. For polyp segmentation, the gating mechanism can automatically enhance the features of key regions according to the content of the input image, while suppressing the interference of background noise and redundant information. In this way, the model can focus more on the polyp region, avoid the influence of the background, and improve the segmentation accuracy.

[0057] Compared with the original CA and SA, the present invention improves the application method of the attention mechanism. By multiplying the attention weights with the input feature map element - by - element, it emphasizes the model's attention to key regions and suppresses unimportant regions. In the improved CA and SA modules, the core improvement is the introduction of a dynamic weighting mechanism, which enables the weights of each channel and spatial position to be adaptively adjusted according to task requirements. The present invention uses the sigmoid function to generate the weighting coefficients and combines them with the input feature map, thereby dynamically adjusting their contributions according to the importance of each channel or spatial position. As Figure 9 and Figure 11 shown, the CA module adopts an existing structure and only optimizes the calculation formula.

[0058] In the CA module, the way of channel reduction is optimized. CA provides higher flexibility by introducing a reduction parameter, which can adjust to reduce redundant information according to the number of input channels, thus more effectively reducing the computational amount and optimizing feature extraction. This mechanism helps the model automatically enhance the key channel features related to polyps, especially the channels related to the edges and textures of polyps.

[0059] In the CA mechanism, the original method only uses sigmoid(out) as the static channel weighting coefficient, while the improved CA dynamically adjusts the weight of each channel by multiplying sigmoid(out) with the input feature map. The formula is:

[0060] y = sigmoid(out)×x

[0061] Among them, x represents the input feature map, out represents the vector output through a convolutional or fully connected layer, and sigmoid(out) is the weighting coefficient for each channel, which can adaptively adjust each channel according to the task requirements, enhance the contribution of important channels, and suppress the influence of irrelevant channels. Similarly, in the SA mechanism, the idea of dynamic weighting is also introduced after improvement to calculate the weighting coefficient of spatial positions and combine it with the spatial dimension of the input feature map.

[0062] This improvement enables the model to adaptively adjust the weighting coefficients of each channel and spatial position according to the input data and task objectives, so as to better focus on task-related features and reduce the interference of noise and unnecessary features. In this way, the model not only improves flexibility but also enhances its adaptability to changes, ultimately improving accuracy and generalization ability.

[0063] Such as Figure 2 、 Figure 6 、 Figure 7 and Figure 10 As shown, the PA module is added to the CIM module. As Figure 10 shown, the PA module adopts an existing structure. The PA module generates the weighting value for each pixel position through 1×1 convolution and batch normalization, and then restricts these weights between 0 and 1 through the sigmoid activation function. Subsequently, these weighting values are multiplied by the input feature map to perform pixel-by-pixel weighting operations, thereby enhancing the model's attention to key pixels. In the polyp segmentation task, the polyp area usually has a low contrast with the background and blurred boundaries, and traditional convolution operations may be difficult to capture these subtle features. PA can make the model more focused on the detailed features of the polyp area through refined pixel-level weighting, especially enhancing the edges and local textures, thus effectively improving the segmentation accuracy. In this way, PA helps the model provide higher resolution and more accurate segmentation results for small and difficult-to-segment targets such as polyps.

[0064] Such as Figure 7 shown, the improved CIM* module consists of four main parts: namely, GCT, the improved CA module, PA, and the improved SA module. Four pyramid features are extracted through the PVT module. First, the low-level feature x1 in the feature enters the GCT module, as Figure 8As shown, in the GCT module, the input is processed according to the selected mode L1 or L2, and weighted adjustment is performed in the channel dimension. If the L2 mode is selected, the input tensor x1 will first be squared, and the sum of each channel will be calculated in the spatial dimension. Then, the embedding will be calculated and normalized by the learned parameters. In the L1 mode, the absolute value of x1 will be taken to generate the corresponding embedding, and the normalization factor will be calculated. After that, the tanh activation function is used to generate a gate for weighted adjustment of the channel dimension of x1. Next, the processed x1 enters the channel attention module. In this module, first, average pooling and max pooling operations are performed on x1 to extract the average and maximum features of each channel respectively. Then, the fully connected layer is used to further process these pooling results, and a channel attention weight map is generated through sigmoid activation. Finally, it is multiplied element-wise with x1 to adjust the weights of different channels. Then, the output weighted feature map is fed into the improved CA module, and by dynamically adjusting the weighted coefficients of each channel, the expression effect of task-related features is enhanced, thus obtaining the dynamically weighted feature map. Subsequently, the feature map processed by the improved CA module enters the PA module, where the feature map is weighted and adjusted at the pixel level to generate the pixel-level modulated feature map. Finally, the feature map processed by the PA module enters the improved SA module, and by adjusting the weights in the spatial dimension, the spatially weighted and adjusted feature map is output.

[0065] In the improved CA, sigmoid(out) is not only used as the weighted coefficient of the channel but also used to weight the input feature map x, so that the contribution of each channel can be dynamically adjusted according to its importance in the current task. The core idea of this improvement is to weight each channel of the feature map according to the relative importance of each channel in the task, so as to enhance the expression effect of task-related features. Specifically, the improved weighting mechanism dynamically adjusts the weight of each channel by using sigmoid(out), enabling the model to adaptively allocate the importance of each channel according to the input data and task objectives, which is different from the original formula. The original weighting mechanism uses a fixed channel weighting coefficient, usually calculated by a vector output by a convolutional or fully connected layer. The formula is as follows:

[0066] y = sigmoid(out)

[0067] Among them, out represents the vector output through a convolutional or fully connected layer, and y represents the returned output. In the original formula, the weighting coefficients of the channels are static and fixed, and are only determined according to the calculation result of out. This process does not consider the influence of the input feature map x, nor does it dynamically adjust each channel according to the task requirements. The improved formula forms a dynamically weighted output by multiplying sigmoid(out) with the input feature map x, and the formula is as follows:

[0068] y = sigmoid(out) × x

[0069] Among them, x represents the input feature map, out represents the vector output through a convolutional or fully connected layer, and sigmoid(out) is the weighting coefficient of each channel, which assigns weights according to the importance of each channel in the current task. When the weight sigmoid(out) of a certain channel is greater than 1, the contribution of this channel will be amplified, enhancing its influence on the task. When the weight is close to 0, it means that the contribution of this channel to the task is small, so its influence is suppressed, thereby reducing the interference of this channel on the model output. In this way, the model can automatically select the features closely related to the task and strengthen them, while suppressing the irrelevant or redundant features.

[0070] Compared with the static weighting coefficients in the original code, the improved weighting mechanism significantly improves the flexibility of the model. The weighting coefficients of the original method are fixed and not adaptively adjusted according to the data and task requirements, while the improved method enhances the adaptive ability of the model through dynamic weighting. For different input data and tasks, the model can dynamically adjust the weighting coefficients of each channel, thereby optimizing the learning process and improving the generalization ability. This mechanism not only effectively avoids the overfitting problem, but also improves the robustness of the model to different tasks and datasets. During the training process, the model can focus on the features that contribute more to the task, while suppressing the noise and redundant features, thereby improving the overall performance.

[0071] In this way, the network can better focus on the features related to the task, reduce the dependence on irrelevant features, thereby improving the accuracy and robustness of the network, and finally showing stronger adaptability in different task and data scenarios.

[0072] Next, the feature map processed by the CA module enters the PA module, and the structure of the PA module is as Figure 2 and 10As shown. The feature map after being processed by the CA module is first processed through a convolutional layer to extract features. Next, through the effects of batch normalization and activation functions, a pixel-level attention map is generated, which reflects the relative importance of each pixel as key information in the image. Then, this generated pixel-level attention map is weighted by element-wise multiplication with the input feature map x1, thereby modulating the feature map at the pixel level. In this way, the network can pay more attention to key pixel positions, enhance the perception ability of important regions, and thus improve the performance of the model in the task.

[0073] Finally, the feature map after being processed by PA enters the SA module. This module calculates average pooling and max pooling along the channel dimension to obtain two single-channel feature maps, and after splicing these two maps, they are processed through a convolutional layer. Through sigmoid activation, a spatial attention map is generated and multiplied with x1, thereby performing weighted adjustment on the input in the spatial dimension.

[0074] This series of operations enables the CIM* module to perform multi-level weighted adjustments on the input tensor x1, integrating the attention mechanisms in the channel, pixel, and spatial dimensions, allowing the network to focus on the most critical features, and thus improving the performance and robustness of the model.

[0075] To address the problem that the original SAM model lacks global perception ability and it is difficult to distinguish polyp regions with blurred boundaries, resulting in low polyp segmentation accuracy, the present invention proposes combining the self-attention mechanism with the original SAM model to obtain an improved module SAM*.

[0076] As Figure 12 and Figure 13 shown, an improved SAM* module is designed, improving the original Similarity Aggregation Module (SAM) module by introducing the self-attention mechanism into SAM to form SAM*.

[0077] In the improved SAM* module, a channel-based self-attention mechanism is introduced to enhance the interaction between local and global information. Specifically, SAM* adds a new Self-Attention class that maps the input features to query (Q), key (K), and value (V) matrices through QKV linear transformation, calculates the attention scores, and finally updates the feature representation through weighted summation. This method effectively captures long-range dependencies, making the feature representation more accurate and enabling better understanding of the global information in the image. By combining GCN (Graph Convolutional Network) and the self-attention mechanism, the feature interaction method is optimized. GCN is still responsible for capturing local connection relationships, while Self-Attention further enhances the learning ability of global dependencies, enabling the model to have stronger feature representation ability when dealing with complex structures. This combination method avoids the local information limitations that may be caused by relying solely on GCN, thus improving the usability of the overall features.

[0078] As Figure 13 shown, the part marked by the dashed box is the key to the improvement. The specific operation method of the SAM* module is as follows:

[0079] First, input T1 output by CFM* and T2 output by CIM* into SAM*. Apply two linear mapping functions to T1 to reduce the channel dimension and extract key feature information. These two mapping functions respectively generate query feature (Q) and key feature (K), and their dimensions are both In the implementation process, use 1×1 convolution operation as the linear mapping operation to project T1 into a lower-dimensional feature space to reduce the computational complexity while maintaining its global semantic information. This process can be represented by the following formula:

[0080] Q = W θ (T1), K = W φ (T1)

[0081] where W θ (·) and W φ (·) are linear mapping functions, and T1 represents high-level features.

[0082] Next, process T2 to extract low-level visual features and provide guiding information for attention calculation. First, use the convolutional unit W g (·) to compress the channels of T2 and reduce the number of channels to 32. Subsequently, T2 undergoes bilinear interpolation to adjust to the same spatial size as T1 to ensure alignment during calculation. To further enhance the important features of T2, SAM* applies Softmax normalization in the channel dimension and selects the second channel as the attention map T2', and its dimension is This attention map T2' reflects the importance of each pixel in T2, enabling subsequent calculations to focus more on significant regions. This series of operations can be represented by F(·). After obtaining T2', we calculate the Hadamard product between K and T2' to assign different attention weights to different pixels. The purpose of this operation is to enhance the weights of edge pixels so that the detailed information of T2 can be better injected into T1. To prevent feature drift, we use adaptive pooling operation to globally downsample the features after calculation and perform central cropping, finally obtaining the feature map V with a dimension of 4×4×16. This process can be expressed as follows:

[0083] V = AP(K⊙F(W g (T2)))

[0084] where AP(·) represents adaptive pooling and central cropping, ⊙ represents the Hadamard product (element-wise multiplication) for feature weighting of different pixels, K represents the key feature, W g (·) represents the convolutional unit, and T2 represents the high-level feature.

[0085] After obtaining V, we calculate the similarity between V and K to establish global dependencies between different positions. Specifically, we use the matrix inner product to calculate the correlation between V T and K:

[0086]

[0087] where, represents the matrix inner product, δ(·) represents Softmax normalization, V T represents the transpose of V, and f represents the correlation attention map.

[0088] After obtaining the correlation attention map f, we apply it to Q and input the resulting features into GCN for global feature modeling to obtain the enhanced feature g with a dimension of 4×4×16:

[0089]

[0090] where f T represents the transpose of f, Q represents the feature map obtained above, and GCN(·) represents the graph convolutional layer, whose role is to model global relationships through the graph convolutional network, enabling the attention relationships calculated by f to be not only based on the pixel level but also have higher-level structural information.

[0091] After obtaining the global feature g processed by GCN, we further calculate the inner product between f and g to remap the graph domain features back to the original structural feature space:

[0092]

[0093] Among them, Y' represents the original structural feature.

[0094] Relying solely on GCN for global feature propagation is restricted by local adjacency relationships, resulting in insufficient information interaction between remote pixels. Therefore, the invention further introduces a Self-Attention mechanism to make up for this deficiency, enabling the model to perform global dependency modeling across the entire image range. To adapt to Self-Attention calculations, the features processed by GCN are transformed and flattened into a sequence format, such that the feature vectors of each pixel point are regarded as independent tokens. In this way, Self-Attention can calculate the correlations between different pixel points, rather than relying solely on local neighborhood information. During the Self-Attention calculation process, the model first projects the improved features into the Q, K, and V spaces, representing the query, index, and feature information between different pixel points respectively. Subsequently, the dot product of Q and K is calculated to obtain the attention distribution between pixel points, and normalized by Softmax to ensure a smoother attention distribution between different pixel points. After that, the value features are weighted and summed using the attention weights, such that the final feature of each pixel point depends not only on its own information but also incorporates key information across the entire image. This process ensures the effective propagation of global information, enabling the model to focus on the interactions between distant pixels, especially showing significant improvements in object segmentation and boundary region processing.

[0095] After the Self-Attention calculation is completed, the model needs to restore the enhanced representation that incorporates global features to a CNN-compatible format. First, the channel order is adjusted to rearrange the number of channels back to the original format. Subsequently, the features enhanced by the attention mechanism are reshaped into a four-dimensional tensor to ensure that they can continue to be used for subsequent CNN processing. Finally, the model further adjusts the feature scale through channel expansion and fuses the enhanced features with the original input through residual connections.

[0096] After the model is established, continue to complete the training of the improved model, and continuously monitor the recognition accuracy during training, saving the model weights with the highest recognition accuracy during the training process.

[0097] The hardware platform used in this experiment is an NVIDIA RTX 4090 server. The network model adopts the PyTorch framework and uses Python 3.8. The weights of the pre-trained model are the model weights with the highest accuracy saved during the training process. The images are sliced with a patch size of 16, the batch size is set to 16, the number of training epochs is set to 100, and the optimizer uses AdamW. If SGD is used, the momentum is 0.9. The initial learning rate (lr) is set to 0.0001, and the cosine annealing strategy is used for learning rate decay, with a decay of 0.1 every 50 epochs. The gradient clipping clip is set to 0.5 to prevent gradient explosion. The model is cyclically trained on the pre-processed dataset images according to the set pre-trained weight file and hyperparameters, and the model weights with the highest recognition accuracy during the training process are saved.

[0098] Step 3: Testing: Load the model with the trained weight file obtained in Step 2, that is, the best weight obtained from training, and perform a test evaluation on the test set. Calculate the accuracy of the test set to evaluate and verify the recognition performance and recognition effect of the improved model.

[0099] The beneficial effects of the present invention:

[0100] (1) Looking at the CFM module, ordinary 3×3 convolutions mainly focus on local features and are difficult to capture global information, resulting in insufficient attention of the model to key regions in the polyp segmentation task. At the same time, relying only on channel-level features may ignore spatial information, making it impossible to effectively extract edge details and affecting the segmentation accuracy. The proposed CFM* replaces the BasicConv2d convolution in the high-level feature part of CFM with the CSEBlock module. CSEBlock enhances the model's attention to key features through channel-level global information modeling. First, it uses global average pooling to extract channel-level features and calculates global context information, enabling the model to understand feature relationships in a larger range. Then, through fully connected layers FC1, FC2, and sigmoid, channel attention weights are generated to adaptively adjust the weights of each channel, enhancing key channels and weakening the influence of irrelevant channels, thereby improving the segmentation effect.

[0101] In addition to optimizing channel information, CSEBlock also introduces a spatial context enhancement (ContextualExcitation) mechanism to make up for the defect that the SE mechanism only focuses on the channel dimension and ignores spatial information. Applying a 1×1 convolution on the feature map enables the model to further enhance the expression of local information and improve the ability to capture details. Especially in the polyp segmentation task, it can more clearly identify boundary and texture information.

[0102] (2)Regarding the CIM module, CIM lacks the ability to model global channel dependencies and adaptively adjust features, making it vulnerable to background noise and redundant information interference. As a result, it cannot accurately highlight key region features, affecting the accuracy and robustness of the model. A CIM* is proposed to capture the global dependencies between channels and enhance the model's global context awareness.

[0103] (3)Regarding the CIM module, when the original CA and SA process small target tasks, their feature extraction ability is limited, and it is difficult to adaptively adjust the dimensionality reduction ratio, resulting in a large amount of computation. At the same time, the enhancement or suppression of key regions may be too extreme, affecting the segmentation accuracy. The CIM* module is proposed to improve the application method of the attention mechanism. By calculating the attention weights through element-wise multiplication, the model can focus more on key regions and suppress unimportant background regions.

[0104] (4)Regarding the CIM module, in the polyp segmentation task, the contrast between the polyp region and the background in the original CIM module is low, and the boundary is blurred, making it difficult for traditional convolution operations to effectively capture subtle features, thus affecting the segmentation accuracy. PA is added to the CIM* module. There are two main improvements in PA: one is to calculate the weighted value at each pixel position using 1×1 convolution and batch normalization, and limit the weight to the range of [0,1] through sigmoid activation to make the weight adjustment smooth. The other is to multiply the calculated weighted value with the input feature map pixel by pixel to enhance the model's attention to key pixels.

[0105] (5)Regarding the SAM module, in the traditional feature interaction process, only local information is relied on, making it difficult to effectively capture global dependencies, resulting in limited feature expression and affecting the understanding of complex structures. The proposed SAM* introduces the self-attention mechanism, combines GCN and the self-attention mechanism, further models global dependencies, and improves the overall feature expression ability.

[0106] To verify the effectiveness and rationality of the algorithm proposed in the present invention, ablation experiments are carried out under the same experimental environment and conditions. The results of the ablation experiments are shown in Table 1.

[0107] Table 1 mDice results of ablation experiments

[0108]

[0109] The model method of the present invention is based on the recognition research of polyps with complex backgrounds and small targets. Therefore, to verify the effectiveness of the overall improvement, the network of the present invention is compared with the original Polyp-PVT model and other mainstream polyp segmentation models in recent years. The comparison results are shown in Tables 2, 3, 4, 5, and 6.

[0110] According to the experimental results in Table 2, the recognition accuracy of the method proposed in the present invention on the CVC-300 dataset exceeds that of other methods. Compared with SANet, the present invention is 1.4% higher; compared with MSNet, it is 3.3% higher; compared with PPNet, it is 0.3% higher; compared with Polyp-PVT, it is 0.2% higher. Generally speaking, the present invention performs best among all the listed models and has the highest mDice coefficient.

[0111] Table 2 Comparison of mDice between the improved model and other methods on the CVC-300 dataset

[0112]

[0113]

[0114] According to the experimental results in Table 3, the recognition accuracy of the method proposed in the present invention on the ClinicDB dataset is better than that of other methods. Compared with SANet, the present invention is 3.1% higher; compared with MSNet, it is 2.6% higher; compared with PPNet, it is 2.6% higher; compared with Polyp-PVT, it is 1.2% higher. Generally speaking, the present invention performs best among all the listed models and has the highest mDice coefficient.

[0115] Table 3 Comparison of mDice between the improved model and other methods on the ClinicDB dataset

[0116]

[0117] According to the experimental results in Table 4, the recognition accuracy of the method proposed in the present invention on the Kvasir dataset is better than that of other methods. Compared with SANet, the present invention is 2.7% higher; compared with MSNet, it is 2.4% higher; compared with PPNet, it is 1.1% higher; compared with Polyp-PVT, it is 0.5% higher. Generally speaking, the present invention performs best among all the listed models and has the highest mDice coefficient.

[0118] Table 4 Comparison of mDice between the improved model and other methods on the Kvasir dataset

[0119]

[0120] According to the experimental results in Table 5, the method proposed by the present invention has a higher recognition accuracy on the ColonDB dataset than other methods. Compared with SANet, the present invention is 6.7% higher; compared with MSNet, it is 6.5% higher; compared with PPNet, it is 2.9% higher; compared with Polyp-PVT, it is 1.6% higher. Generally speaking, the present invention performs best among all the listed models and has the highest mDice coefficient.

[0121] Table 5 Comparison of mDice between the improved model and other methods on the ColonDB dataset

[0122]

[0123] According to the experimental results in Table 6, the method proposed by the present invention has a higher recognition accuracy on the ETIS dataset than other methods. Compared with SANet, the present invention is 5.7% higher; compared with MSNet, it is 8.8% higher; compared with PPNet, it is 2.3% higher; compared with Polyp-PVT, it is 2.2% higher. Generally speaking, the present invention performs best among all the listed models and has the highest mDice coefficient.

[0124] Table 6 Comparison of mDice between the improved model and other methods on the ETIS dataset

[0125]

[0126] The polyp image segmentation method provided by the present invention optimizes the polyp image segmentation model. According to the above experimental comparison results, the optimized polyp image segmentation model of the present invention has a higher recognition accuracy compared with the existing models.

[0127] As Figure 14 shown, the polyp image segmentation method provided by the present application can achieve polyp image segmentation processing and has a good segmentation effect.

[0128] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, and all should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claimed rights.

Claims

1. A polyp image segmentation method, characterized in that, It includes the following steps: Build a polyp image segmentation model, which includes a PVT module, an improved cascaded fusion module (CFM*), an improved camouflage identification module (CIM*), and an improved similarity aggregation module (SAM*); Among them, the BasicConv2d responsible for extracting high-level features in the original cascaded fusion module (CFM) is replaced with a CSEBlock to obtain the improved cascaded fusion module (CFM*); Replace sigmoid(out) in the original CA module of the original camouflage identification module (CIM) with sigmoid(out)×x to obtain an improved CA module, and replace sigmoid(out) in the original SA module of the original camouflage identification module (CIM) with sigmoid(out)×x to obtain an improved SA module. out represents the vector output through a convolutional or fully connected layer, x represents the input feature map, and sigmoid(out) is used as the weighting coefficient for each channel; introduce a GCT module before the improved CA module, and introduce a PA module between the improved CA module and the improved SA module to obtain the improved camouflage identification module (CIM*); Add a Self-Attention module between the W operation and the addition operation of the original Similarity Aggregation Module (SAM) to obtain an improved Similarity Aggregation Module (SAM*). z ​ 2. The polyp image segmentation method according to claim 1, wherein The polyp image segmentation method further includes the following steps: Input the image to be segmented into the polyp image segmentation model. The PVT module is used to extract multi-scale long-range dependence relationship features in the image to be segmented, and the multi-scale long-range dependence relationship features include high-level features and low-level features; The improved cascaded fusion module (CFM*) is used to fuse the high-level features to obtain a high-level feature map; the improved camouflage identification module (CIM*) is used to process the low-level features, remove noise and enhance the low-level representation information of the polyp to obtain a low-level feature map; the improved similarity aggregation module (SAM*) aligns and fuses the high-level feature map and the low-level feature map to obtain a segmentation prediction feature map as the output.

3. The polyp image segmentation method according to claim 2, wherein The PVT module includes an MLP, WindowAttention, a Block, an OverlapPatchEmbed, an improved PyramidVisionTransformerImpr, and a Depthwise Separable Convolution (DWConv); In the PVT module, the MLP is used for feature transformation, WindowAttention is used to provide an efficient local attention mechanism, the Block is used to gradually extract features through multi-head attention and MLP layers, the OverlapPatchEmbed is used for image patch embedding, the improved PyramidVisionTransformerImpr is used for multi-stage feature extraction, and the Depthwise Separable Convolution (DWConv) is used to improve efficiency through depthwise separable convolution; The image to be segmented is passed into the PVT model of the PVT module to extract the low-level feature x1; then the output low-level feature x1 is passed into the PVT model to extract the high-level feature x2; then the output high-level feature x2 is passed into the PVT model to extract the high-level feature x3; finally, the output high-level feature x3 is passed into the PVT model to extract the high-level feature x4.

4. The polyp image segmentation method according to claim 3, wherein the improved cascaded fusion module (CFM*) includes a CSEBlock and a BasicConv2d; the improved cascaded fusion module (CFM*) consists of two cascaded parts; In the first part, the high-level feature x4 is first upsampled to the same size as the high-level feature x3, and then processed by two convolutional units CSEBlock to obtain a smoothed feature map and Perform the Hadamard product with the high-level feature x3, and then concatenate the result with along the channel dimension, and then process it through the convolutional unit BasicConv2d to obtain a fused feature map In the second part, first upsample the high-level features x4, x3, and to the same size as the high-level feature x2 respectively, and then use the convolutional unit BasicConv2d to smooth these upsampled feature maps; then perform the Hadamard product operation on the smoothed x4 and x3 with x2, and combine the resulting map with the upsampled and smoothed for concatenation; input the concatenated feature map into two convolutional units BasicConv2d for dimensionality reduction to obtain the final output feature map T1.

5. The polyp image segmentation method according to claim 4, wherein the CSEBlock includes a 1x1 convolutional layer, batch normalization, a ReLU activation function, global pooling, a fully connected layer, a sigmoid activation function, and a context convolutional layer; The input features pass through a 1x1 convolutional layer to convert the number of input channels into the number of output channels, and batch normalization and the ReLU activation function are applied to extract and enhance the features; Then global pooling is used to compress the feature map to obtain the global description of each channel; Then, the features pass through two fully connected layers, where the first fully connected layer reduces the number of channels and then restores to the original number of channels through the second fully connected layer to learn the weights of the channels; The weight coefficients are obtained through the sigmoid activation function. The weight coefficients represent the importance of each channel and are used to weight the input features; The context convolutional layer uses 1x1 depth convolution to perform fine-grained context modeling on the input features; The weighted features are combined with the original features through the Hadamard product, and then fused with the features obtained by the context convolution to obtain the final fused feature map.

6. The polyp image segmentation method according to claim 3, wherein the improved camouflage identification module (CIM*) includes a GCT module, an improved CA module, a PA module, and an improved SA module; The GCT module includes learnable parameters for controlling the scaling and offset of the feature map; according to the L1 mode or the L2 mode, the module calculates the embedded features of the input features and performs corresponding processing; The gating value is generated through the tanh activation function and is used to adjust the weights of each channel of the input features; The improved CA module includes the following structure: two adaptive pooling layers, namely average pooling and max pooling, two convolutional layers, a ReLU activation function, and a sigmoid activation function; The PA module includes a convolutional layer, a batch normalization layer, and a sigmoid activation function; the convolutional layer is used to perform feature learning within the channel, the batch normalization layer helps to stabilize the training process, and the sigmoid activation function is used to generate pixel-level attention weights; The improved SA module includes a convolutional layer with two input channels and a sigmoid activation function. The convolutional layer uses an adjustable convolutional kernel size and sets corresponding padding according to the convolutional kernel size. By concatenating the results of average pooling and max pooling and inputting them into the convolutional layer, spatial attention weights are finally generated through the sigmoid activation function. The low-level features x1 extracted from the PVT module enter the embedding calculation layer of the GCT module. Through feature normalization and weighting operations based on the L1 mode or L2 mode norm, a weighted feature map is obtained. The output weighted feature map is sent to the convolutional layer of the improved CA module for channel attention weighting processing to obtain a dynamically weighted feature map. The dynamically weighted feature map processed by the improved CA module enters the convolutional layer of the PA module for pixel-level modulation processing to generate a pixel-level modulated feature map. The modulated feature map processed by the PA module enters the convolutional layer of the improved SA module for spatial attention weighting processing to adjust the weights in the spatial dimension and output the low-level feature map after spatial weighting adjustment.

7. The polyp image segmentation method according to any one of claims 2 to 6, characterized in that The improved similarity aggregation module (SAM*) includes an adaptive average pooling layer, three convolutional layers, GCN, and Self-Attention. The specific operation mode of the improved similarity aggregation module (SAM*) is as follows: The high-level feature map output by CFM* is input into the convolutional operation of SAM* for linear transformation processing to obtain feature maps Q and K. The low-level feature map output by CIM* is input into the spatial attention module of SAM*. Through convolutional operation, interpolation, and Softmax, an attention map T2' is generated, which is then subjected to Hadamard product with K, followed by adaptive pooling and cropping to finally obtain feature map V. The SAM* module performs weighting, GCN, Self-Attention, and inner product calculations on feature map K and feature map V to finally obtain the segmentation prediction feature map Z.

8. The polyp image segmentation method according to claim 1, characterized in that, The polyp image segmentation method further includes the following steps: Obtain a polyp image dataset, select a part of the images in the polyp image dataset as the training set, and another part of the images as the test set. Use the training set to train the polyp image segmentation model. Continuously monitor the recognition accuracy during training and save the weight parameters with the highest recognition accuracy during the training process as the optimal weights. Load the optimal weights into the polyp image segmentation model and use the test set to test the polyp image segmentation model, and calculate the accuracy of the test set to verify the recognition effect of the polyp image segmentation model.

9. The polyp image segmentation method according to claim 8, characterized in that Preprocess five public datasets, namely Kvasir-SEG, ClinicDB, ColonDB, Endoscene, and ETIS, to obtain the polyp image dataset.

10. The polyp image segmentation method according to claim 8 or 9, characterized in that The sizes of the images in the five public datasets of Kvasir-SEG, ClinicDB, ColonDB, Endoscene, and ETIS were uniformly adjusted to 352 * 352, and the size of the number of samples used during training was 16.