Complex polyp image segmentation method based on multi-scale feature fusion

By designing a cross-scale fusion module, combining an encoder and a decoder, the multi-scale feature fusion of polyp images is achieved, which solves the problem of poor segmentation effect of polyp images in the prior art, and improves segmentation accuracy and model adaptability.

CN120298420APending Publication Date: 2025-07-11ANHUI UNIVERSITY OF TECHNOLOGY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510358514.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing deep learning methods are difficult to fuse features of different scales in polyp image segmentation, resulting in poor segmentation of small polyps or polyp detail areas, and traditional methods are inefficient in complex polyp image processing.

Method used

A cross-scale fusion module is designed, and multi-scale feature fusion is carried out through the combination of encoder and decoder, and efficient feature integration of polyp images is achieved using global average pooling, full-connection layer mapping, multi-head attention layer fusion and attention weight mechanisms.

Benefits of technology

It significantly improves the accuracy and robustness of polyp image segmentation, can capture multi-scale features, adapt to polyp images of different types and complexities, and improves the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298420A_ABST
    Figure CN120298420A_ABST
Patent Text Reader

Abstract

The invention discloses a complex polyp image segmentation method based on multi-scale feature fusion, and belongs to the technical field of medical image processing. The method comprises the following steps: inputting a preprocessed polyp image into a lower sampling layer of an encoder, sequentially extracting feature maps of different scales, retaining, carrying out global average pooling, full connection layer mapping, splicing, multi-head attention layer fusion, splitting and dimension reduction operations, and outputting the feature maps fused with different scales; and inputting the feature map output by the last down-sampling layer into a decoder for up-sampling, splicing, fusing and storing the up-sampled feature map and a feature map which is output by an encoder and has the same scale as the feature map, unifying the size, splicing in the channel dimension, and finally applying the attention weight to the feature map to obtain the attention of the user. According to the method, the detection capability of complex pathological changes and tiny targets of the polyp is remarkably improved, and the segmentation precision of the polyp image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image processing, and more specifically, relates to a complex polyp image segmentation method based on multi-scale feature fusion. Background Art

[0002] Polyps are one of the common lesions in the digestive tract. Early detection and accurate segmentation of the polyp region are crucial for the diagnosis and treatment of diseases. Precise image segmentation can clearly separate the lesion area in the image, providing intuitive and accurate pathological information for doctors. This helps doctors more accurately judge the severity of the condition, develop personalized treatment plans, and effectively monitor and evaluate the changes in the condition during the treatment process, thereby improving the treatment effect.

[0003] Currently, polyp segmentation mainly relies on endoscopic image analysis. However, the segmentation of polyp images faces the following challenges:

[0004] (1) The shapes, sizes, and positions of polyps vary greatly, and it is difficult for traditional methods to adapt to this diversity;

[0005] (2) The contrast between polyps and surrounding tissues is low, and the boundaries are blurred, resulting in low segmentation accuracy;

[0006] (3) Large and small polyps may coexist in polyp images, and it is difficult for traditional methods to capture multi-scale features simultaneously.

[0007] In the field of medical image processing, existing deep learning methods (such as U-Net) are widely used in medical image segmentation. However, when applied to polyp image processing, due to their difficulty in fusing features of different scales, the segmentation effect on small polyps or detailed regions of polyps is relatively poor.

[0008] After retrieval, the Chinese patent application number is 202210289442.7, the application date is March 23, 2022, and the invention creation name is: A polyp segmentation method and system for colon polyp images based on multi-scale feature fusion enhancement. The polyp processing method disclosed in this application case includes: inputting the image to be segmented into an encoder network for continuous convolutional pooling to generate feature maps of multiple different scales; inputting the feature maps of different scales into a multi-scale feature fusion enhancement module to obtain feature maps of multiple different scales rich in multi-scale information, and inputting them into a decoder; in the decoder network, the feature maps generated after multiple iterations of the feature maps of different scales rich in multi-scale information are used to complete the polyp segmentation of the colon polyp image, and the segmentation result map is output. In this application case, a multi-scale feature fusion enhancement technology is proposed to improve the efficiency and accuracy of polyp segmentation. However, it uses feature maps of all scales in the decoder part, increasing the computational complexity and resulting in relatively low efficiency of polyp segmentation, especially when dealing with complex polyp images, the efficiency is even lower.

[0009] Therefore, there is an urgent need for a segmentation method for complex polyp images to solve the above problems existing in the segmentation of complex polyp images in the prior art. Summary of the Invention

[0010] 1. Problems to be Solved

[0011] The purpose of the present invention is to provide a complex polyp image segmentation method based on multi-scale feature fusion, which realizes the efficient fusion of multi-scale features of polyp images by designing a cross-scale fusion module, comprehensively captures the feature information of polyps at different scales, improves the accuracy of polyp segmentation, and provides more accurate and reliable information for medical diagnosis and treatment.

[0012] 2. Technical Solutions

[0013] In order to solve the above problems, the technical solutions adopted by the present invention are as follows:

[0014] The present invention provides a complex polyp image segmentation method based on multi-scale feature fusion, including the following steps:

[0015] S1. Preprocessing of polyp images;

[0016] S2. Input the preprocessed polyp image into an encoder, and sequentially extract features of different scales through a downsampling layer, and retain the feature maps obtained by each downsampling;

[0017] S3. Perform global average pooling, fully connected layer mapping, splicing, multi-head attention layer fusion, splitting and dimension reduction operations on the feature maps obtained in step S2, and output multiple feature maps that fuse different scales;

[0018] S4. Input the feature map output from the last downsampling layer of the encoder in step S2 into the decoder for upsampling, and splice and fuse the upsampled feature map with the feature map of the same scale output in step S3.

[0019] S5. Sequentially save the feature maps output after each upsampling, and unify the scale of the obtained feature maps to the scale of the feature map output by the last upsampling. Then, splice all the feature maps in the channel dimension.

[0020] S6. Process the feature map obtained in step S5, apply the attention weight to this feature map to obtain the final feature map, and finally process the final feature map through convolution and activation functions to obtain the segmentation result.

[0021] Furthermore, in step S1, the preprocessing of the polyp image includes cropping and scaling the polyp image.

[0022] Furthermore, step S2 specifically includes the following steps:

[0023] S2.1. Input the polyp image X with size C in ×H×W into the convolution-batch normalization-activation module, where C in is the number of input channels, H and W are the height and width of the image respectively, and the activation function uses the ReLu function.

[0024] S2.2. Use the downsampling module to downsample the image processed in step S2.1, output the feature map F1, then input the obtained feature map F1 into the convolution-batch normalization-activation module for processing, and then perform downsampling again to obtain the feature map F2, and so on. The feature map output after each downsampling is saved, and finally a set of feature maps of different scales {F1, F2,.., F j} is obtained.

[0025] Furthermore, step S3 specifically includes the following steps:

[0026] S3.1. Perform global average pooling on each feature map in the set of feature maps {F1, F2,.., F j} to obtain a one-dimensional vector V j , that is, V j =GPA(F j ).

[0027] S3.2. Use the set of fully connected layers {FC 1,j} to map the one-dimensional vector V j to the same dimension D, and obtain Concatenate the mapped vectors in order to form a combined vector

[0028] S3.3. Use the multi-head attention layer to perform feature fusion on the combined vector V C The calculation formula is (V A , A) = MHA(V C , V C , V C ), where V A is the attention output vector and A is the attention weight matrix;

[0029] S3.4. Split the attention output vector V A into vectors V Sj corresponding to different scales, and restore it to the original dimension through the fully connected layer set {FC 2,j}, obtaining V bj = FC 2,j (V Sj );

[0030] S3.5. Restore the vector V bj mapped back to the original dimension to the feature map shape F fj , that is where is a matrix of all 1s with the same spatial size as the feature map F j , and finally obtain the set of fused feature maps {F f1 , F f2 ,..., F fj}.

[0031] Furthermore, step S4 specifically includes the following steps:

[0032] S4.1. Use the upsampling module to upsample the output feature map F j of the last downsampling layer. The l-th UB module first upsamples the feature map F l through the bilinear interpolation upsampling layer Upsample j to obtain F uj ;

[0033] S4.2. Concatenate F uj with the feature map F fj that fuses features of different scales at its corresponding scale in the channel dimension to obtain F C ;

[0034] S4.3. Finally, after being processed by m l CBN modules, output the feature map Then upsample again, and so on, retaining the output after each upsampling to form the output set {F d1 , F d2 ,... F dj}.

[0035] Further, in step S5, the spatial dimensions H×W of the feature map F output in the last upsampling are used as the target dimensions, and bilinear interpolation operations are performed on the feature maps output by other upsamplings to unify their dimensions to the target dimensions, and then all the feature maps with unified dimensions are concatenated in the channel dimension. dj Further, step S6 includes flattening and averaging the concatenated feature maps in the spatial dimension, and then calculating the attention weights through the attention module AM.

[0036] Further, the attention module AM includes a fully connected layer FC

[0037] and FC a1 and FC a2 as well as the activation functions ReLu and Sigmoid. The output dimension of FC a1 is half of the input dimension, and the output dimension of FC a2 is 1.

[0038] Further, the attention weights are extended to the same dimension as the concatenated feature maps, and then applied to the feature maps to obtain the weighted feature maps.

[0039] Further, the weighted feature maps are processed by a 1×1 convolutional layer, and the number of channels of the feature maps is adjusted to the number of target categories, and then processed using an activation function to obtain the final segmentation result.

[0040] 3. Beneficial Effects

[0041] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0042] (1) In the image segmentation method of the present invention, by using an encoder combined with a decoder to perform cross-scale fusion processing on polyp images, it can effectively integrate feature information of different scales, fully exploit the complementarity between multi-scale features, which enables the model to not only capture the detailed features of polyp micro-lesions but also take into account the overall information of large-scale organ structures, thereby significantly improving the detection ability for complex polyp lesions and small targets and enhancing the segmentation accuracy of polyp images;

[0043] (2) In the image segmentation method of the present invention, by introducing an attention enhancement mechanism, it can automatically learn the importance of features and perform weighted processing on key regions. This method enhances the focusing ability of the model on the target region, reduces the interference of irrelevant information, and enables the segmentation result to more accurately reflect the actual situation of the target tissue and the lesion region, further improving the accuracy and robustness of the segmentation;

[0044] (3) The image segmentation method of the present invention. The feature selection and fusion module enables the features obtained from each upsampling to be retained and concatenated in the channel dimension. By using the attention mechanism, the decoding process can be automatically adjusted according to the characteristics of the input image, making full use of the feature information at different levels, enabling the model to adapt to polyp images of different types and complexities, improving the generalization ability of the model, and achieving good segmentation results on images with different datasets and large individual differences. Description of the Drawings

[0045] Figure 1 It is a model framework diagram of the image segmentation method of the present invention; Detailed Embodiment

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0047] To achieve accurate cutting of complex polyp images, in combination with Figure 1 , this embodiment proposes a complex polyp image segmentation method based on multi-scale feature fusion, and the specific steps are as follows:

[0048] S1. Preprocessing of polyp images;

[0049] Perform cropping and scaling processing on the input polyp images to unify the size of the polyp images into RGB images of 224x224.

[0050] S2. Input the preprocessed polyp images into the encoder, and sequentially extract features of different scales through the downsampling layer, and retain the feature maps obtained from each downsampling;

[0051] Specifically, input the polyp image X with a size of C in ×H×W into the convolution-batch normalization-activation module (CBN). Among them, C in is the number of input channels, and H and W are the height and width of the image respectively. Assume that the convolution layer of the i-th CBN module is C i , the batch normalization layer is BN i , and the activation function is σ i , then the output feature map F i = σ i (BN i (C i (X))). In this embodiment, the convolution layer C iThe convolution kernel size is 3, the padding is 1, and the ReLu function is used by default as the activation function.

[0052] Then, the downsampling module (DB) is used to downsample F1, and the obtained feature map is input into the CBN. Then, downsampling is performed again to obtain the second feature map F2, and so on. That is: the j-th downsampling module first downsamples the input feature map through the max pooling layer MaxPool j (the pooling kernel size is 2), and then processes it through n j CBN modules. Let the input feature map be F in , then the output feature map After multiple downsamplings are performed in sequence, a set of feature maps {F1, F2,.., F j} with different scales is obtained, and the corresponding number of channels are {C1, C2,..., C j}.

[0053] S3. Perform global average pooling, fully connected layer mapping, concatenation, multi-head attention layer fusion, splitting, and dimensionality reduction operations on the feature maps obtained in step S2, and output multiple feature maps that fuse different scales;

[0054] First, perform global average pooling on each feature map in the set of feature maps {F1, F2,.., F j} to obtain a one-dimensional vector V j , that is, V j = GPA(F j );

[0055] Then, use the set of fully connected layers {FC 1,j} to map the one-dimensional vector V j to the same dimension D, and obtain Concatenate the mapped vectors in order to form a combined vector

[0056] Then, use the multi-head attention layer to perform feature fusion on the combined vector V C , and the calculation formula is (V A , A) = MHA(V C , V C , V C ), where V A is the attention output vector, and A is the attention weight matrix;

[0057] Finally, split the attention output vector V A into vectors V Sj corresponding to different scales, and restore it to the original dimension through the set of fully connected layers {FC 2,j}, and obtain V bj = FC 2,j (VSj ), the vector V mapped back to the original dimension bj is restored to the feature map shape F fj , that is where is a matrix of all 1s with the same spatial dimensions as the feature map F j , and finally the fused set of feature maps {F f1 , F f2 ,..., F fj} is obtained.

[0058] S4. Input the feature map output by the last downsampling layer of the encoder in step S2 into the decoder for upsampling, and splice and fuse the upsampled feature map with the feature map of the same scale output in step S3;

[0059] S4.1. Use the upsampling module to upsample the output feature map F j of the last downsampling layer. The l-th UB module first upsamples the feature map F l through the bilinear interpolation upsampling layer Upsample j to obtain F uj ;

[0060] S4.2. Splice F uj with the feature map F fj that has fused features of different scales at its corresponding scale in the channel dimension to obtain F C ;

[0061] S4.3. Finally, after being processed by m l CBN modules, the output feature map is then upsampled again, and so on. The output after each upsampling is retained to form the output set {F d1 , F d2 ,... F dj}.

[0062] S5. Save the feature maps output after each upsampling in turn. The spatial dimensions H×W of the feature map F dj output by the last upsampling are used as the target dimensions, and the bilinear interpolation operation (Resample) is used for the other upsampled output feature maps to make their dimensions unified to the target dimensions, that is The outputs with unified dimensions are spliced in the channel dimension to obtain the spliced feature map

[0063] S6. Process the spliced feature map F stack in step S5. Assume its shape is first flattened in the spatial dimension and the mean value is calculated to obtain V flat = mean(Fstack , with dim=(2, 3)), and then, calculate the attention weights through the attention module AM, which includes a fully connected layer FC a1 and FC a2 as well as activation functions ReLu and Sigmoid. The calculation process is W = Sigmoid(FC a2 (ReLu(FC a1 (V flat )))). Here, the output dimension of FC a1 is half of the input dimension, and the output dimension of FC a2 is 1.

[0064] Use the broadcast mechanism to expand the attention weight W to the same dimension as F stack and apply it to the feature map to obtain the weighted feature map F W = W · F stack . Finally, process F W through a 1×1 convolutional layer to adjust the number of channels of the feature map to the number of target classes, and then use the activation function for processing to obtain the final segmentation result.

[0065] In addition, train and test the model of the present invention, collect polyp image datasets, including public datasets such as CVC-300 and ColonDB, and divide them into a training set and a validation test set. Use the method of the present invention to construct a polyp image segmentation model, train the model using the training set, and use the validation test set for testing after training. The test results show that the average Dice on five datasets, namely ETIS, ColonDB, ClinicDB, CVC-300, and Kvasir, are 58.02%, 72.98%, 92.03%, 79.11%, and 88.52% respectively.

[0066] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A complex polyp image segmentation method based on multi-scale feature fusion, characterized in that: It includes the following steps: S1. Preprocessing of polyp images; S2. Input the preprocessed polyp images into the encoder, and sequentially extract features of different scales through the downsampling layer, and retain the feature maps obtained by each downsampling; S3. Perform global average pooling, fully connected layer mapping, splicing, multi-head attention layer fusion, splitting and dimensional reduction operations on the feature maps obtained in step S2, and output multiple feature maps that fuse features of different scales; S4. Input the feature map output by the last downsampling layer of the encoder in step S2 into the decoder for upsampling, and splice and fuse the upsampled feature map with the feature map of the same scale output in step S3; S5. Sequentially save the feature maps output after each upsampling, and unify the scale of the obtained feature maps to the scale of the feature map output by the last upsampling, and then splice all the feature maps in the channel dimension; S6. Process the feature map obtained in step S5, apply the attention weight to the feature map to obtain the final feature map, and finally process the final feature map through convolution and activation functions to obtain the segmentation result.

2. The image segmentation method according to claim 1, wherein In step S1, the preprocessing of the polyp images includes cropping and scaling the polyp images.

3. The image segmentation method according to claim 1, characterized in that, Step S2 specifically includes the following steps: S2.

1. Input the polyp image X with dimensions C in ×H×W into the convolution-batch normalization-activation module, where C in is the number of input channels, H and W are the height and width of the image respectively, and the ReLu function is used as the activation function; S2.

2. Downsample the image processed in step S2.1 using the downsampling module to output the feature map F1, then input the obtained feature map F1 into the convolution-batch normalization-activation module for processing, and then perform downsampling again to obtain the feature map F2, and so on. The feature map output after each downsampling is saved, and finally a set of feature maps with different scales {F1, F2,.., F j} is obtained.

4. The image segmentation method according to claim 3, wherein Step S3 specifically includes the following steps: S3.

1. Perform global average pooling on each feature map in the set of feature maps {F1, F2,.., F j} to obtain a one-dimensional vector V j , that is, V j = GPA(F j ); S3.

2. Use the fully connected layer set {FC 1,j} to map the one-dimensional vector V j to the same dimension D, obtaining V mapj = FC 1,j (V j ). Concatenate the mapped vectors in order to form a combined vector V C = {V map1 , V map2 ,..., V mapj}; S3.

3. Use the multi-head attention layer to perform feature fusion on the combined vector V C The calculation formula is (V A , A) = MHA(V C , V C , V C ), where V A is the attention output vector and A is the attention weight matrix; S3.

4. Split the attention output vector V A into vectors V corresponding to different scales Sj , and restore it to the original dimension through the fully connected layer set {FC 2,j}, obtaining V bj = FC 2,j (V Sj ); S3.

5. Restore the vector V mapped back to the original dimension to the feature map shape F bj , that is fj , where is a matrix of all 1s with the same spatial dimensions as the feature map F . Finally, the set of fused feature maps {F j , F f1 ,..., F f2} is obtained fj .

5. The image segmentation method according to claim 4, wherein Step S4 specifically includes the following steps: S4.

1. Use the upsampling module to upsample the output feature map F of the last downsampling layer j For upsampling, the l-th UB module first uses the bilinear interpolation upsampling layer Upsample l to upsample the feature map F j to obtain F uj ; S4.

2. Merge F uj with the feature map F that has fused different-scale features at its corresponding scale fj by concatenating them in the channel dimension to obtain F C ; S4.

3. Finally, after passing through m l CBN modules to process and output the feature map Then, perform upsampling again, and so on, retaining the output after each upsampling to form an output set {F d1 , F d2 ,... F dj}.

6. The image segmentation method according to claim 5, characterized in that In step S5, the spatial size H×W of the feature map F output by the last upsampling is used as the target size, and bilinear interpolation is performed on the feature maps output by other upsamplings to make their sizes unified to the target size, and then all the feature maps with unified sizes are concatenated in the channel dimension. dj ​ 7. The image segmentation method according to claim 6, characterized in that, Step S6 includes flattening and averaging the spliced feature maps in the spatial dimension, and then calculating the attention weight through the attention module AM.

8. The image segmentation method according to claim 7, wherein The described Attention Module AM includes a fully connected layer FC a1 and FC a2 as well as activation functions ReLu and Sigmoid, FC a1 outputs a dimension that is half of the input dimension, FC a2 outputs a dimension of 1.

9. The image segmentation method according to claim 7, wherein Expand the attention weight to the same dimension as the spliced feature map, and then apply it to the feature map to obtain the weighted feature map.

10. The image segmentation method according to claim 9, wherein, Process the weighted feature map with a 1×1 convolutional layer, adjust the number of channels of the feature map to the number of target categories, and then use the activation function for processing to obtain the final segmentation result.

Citation Information

Patent Citations

  • Colon polyp image polyp segmentation method and system based on multi-scale feature fusion enhancement

    CN116863132A