A polyp image segmentation method based on a dual-branch feature progressive fusion network

Through the two-branch feature progressive fusion network, the problem that polyp image segmentation method in the prior art is difficult to extract boundary information, and a more accurate and complete polyp segmentation effect is achieved.

CN119360031BActive Publication Date: 2025-07-08CHINA JILIANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411919763.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-07-08
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

The existing deep learning-based polyp image segmentation method is difficult to effectively extract and utilize the boundary information of polyp regions, resulting in insufficient performance when dealing with polyp segmentation tasks of different sizes and shapes.

Method used

Using a method based on a two-branch feature progressive fusion network, multi-level features are extracted through the PVTv2 backbone network, combining a single-line gating mechanism, boundary information perception branch, misalignment fusion module, misalignment single-layer enhancement module and multi-level residual decoding module to achieve global fusion and enhancement of polyp boundary features.

Benefits of technology

It significantly improves the accuracy and reliability of polyp segmentation, enhances the extraction and utilization of polyp region boundary information, and improves segmentation performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360031B_ABST
    Figure CN119360031B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of image processing, and discloses a polyp image segmentation method based on a dual-branch feature progressive fusion network. This method constructs a dual-branch network for progressive fusion training based on the extraction of feature and boundary information, mainly including a dual-branch backbone network based on point convolution visual transformation and reverse perception information layer, a dislocation fusion module, a dislocation single-layer fusion module, a perception information fusion module, and a multi-level residual decoding module. The trained optimal model is used to segment polyp images and evaluate the results. The method of the present invention effectively extracts boundary information and hierarchical features through a dual-branch structure, adopts a progressive fusion method to retain and refine global semantic information, overcomes the limitations of traditional algorithms in the utilization of polyp boundary features, and uses a boundary difference joint loss to optimize the model training process during training, realizing high-precision recognition and segmentation of polyp regions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a polyp image segmentation method based on a dual-branch feature progressive fusion network. Background Art

[0002] With the rapid development of deep learning methods, the efficiency of feature extraction and utilization in polyp image segmentation has been significantly improved. Architectures based on convolutional neural networks (CNNs) have been widely used in the field of image segmentation. Among them, typical network architectures represented by U-Net have greatly promoted the application of CNNs in medical image segmentation. Many studies have improved on the basis of U-Net to overcome its limitations in practical applications, resulting in significant improvements in the segmentation performance of various segmentation models based on U-Net. In addition, methods based on Transformer achieve dynamic weight allocation and global receptive field feature extraction capabilities through multi-head self-attention layers. Vision Transformer (ViT) proposes a strategy of dividing an image into multiple regions for encoding and classification respectively, improving the training efficiency through parallel computing and effectively enhancing the ability to extract global information.

[0003] However, current deep learning-based segmentation methods mainly focus on the extraction of texture information features in the polyp region and are difficult to effectively extract and utilize the boundary information of the polyp region. Due to the lack of deep fusion of global features and boundary information, these methods often show deficiencies when dealing with polyp segmentation tasks of different sizes and shapes. Therefore, it is necessary to explore new methods to improve the accuracy and reliability of polyp segmentation, so as to achieve a more complete semantic structure expression. Summary of the Invention

[0004] The purpose of the present invention is to provide a polyp image segmentation method based on a dual-branch feature progressive fusion network to solve the above technical problems.

[0005] To solve the above technical problems, the specific technical solution of a polyp image segmentation method based on a dual-branch feature progressive fusion network of the present invention is as follows:

[0006] A polyp image segmentation method based on a dual-branch feature progressive fusion network includes the following steps:

[0007] Step 1: Collect a colonoscopy polyp image and polyp mask label dataset, and divide the dataset into a training set and a test set;

[0008] Step 2: Construct a polyp image segmentation model based on a dual-branch feature progressive fusion network;

[0009] Step 3: Perform data augmentation on the training data obtained in Step 1 and input it into the dual-branch feature progressive fusion network training model;

[0010] Step 4: Use the best model obtained after training to test the colonoscopy images in the test set, obtain the polyp segmentation results after testing, and evaluate the segmentation results of the model.

[0011] Furthermore, the dataset in Step 1 contains five sub-datasets. Among them, 1450 pictures selected from two sub-datasets are used as the training set, 798 pictures selected from five datasets are used as the test set, and finally, all five datasets are used to evaluate the segmentation performance of the final model.

[0012] Furthermore, the dual-branch feature progressive fusion network proposed in Step 2 is used to perform the following steps:

[0013] Step 2.1: Use the PVTv2 backbone network as the multi-level feature extraction branch to extract four-stage multi-level features from the original input image T ={ T 1 , T 2 , T 3 , T 4};

[0014] Step 2.2: Use a single-line gating mechanism to perform feature weighted extraction on the original input image for preliminary feature filtering;

[0015] Step 2.3: Use the boundary information perception branch to extract high-fidelity boundary information from the features obtained in Step 2.2 through hierarchical feature enhancement and cascaded fusion strategies, construct a polyp boundary segmentation mask, and extract boundary feature information;

[0016] Step 2.4: Use the misalignment fusion module to perform pairwise fusion on adjacent hierarchical features in Step 2.1 to generate a new comprehensive feature representation. At the same time, use the misalignment single-layer fusion module to extract the features of the bottom layer T 4 and refine and enhance the high-level semantic feature information with the help of the channel and spatial attention fusion mechanism to obtain a new four-stage multi-level feature F ={ F 1 , F 2 , F 3 , F 4};

[0017] Step 2.5: Further mine effective semantic information from the boundary feature information extracted in Step 2.3. Through a multi-level cascaded perception information fusion module, fuse the mined effective semantic information with multi-level features F for progressive global fusion to obtain boundary aggregation perception features;

[0018] Step 2.6: Use a multi-level residual decoding module to cross-level fuse multi-level features F and boundary feature information. The features obtained from each level of fusion are dimensionally reduced by 1×1 convolution and participate in deep supervision during the training process.

[0019] Furthermore, the single-line gating mechanism proposed in Step 2.2 performs preliminary feature extraction and filtering on the original image input to the network X First, generate a weight tensor through convolution and normalization operations W as , and then use the weight tensor after channel division , ( i = 1, 2) to perform weighted summation on the input features to obtain the fused feature G f . The specific expression method is as follows:

[0020] ,

[0021] where Softmax represents the normalization exponential function, div c represents the operation of dividing the tensor by channels, represents element-wise multiplication, represents element-wise addition.

[0022] Furthermore, the boundary information perception branch proposed in Step 2.3 includes the following steps:

[0023] Step 2.3.1: The boundary information perception branch contains three groups of cascaded reverse perception information layers. As a parallel branch to the multi-level feature extraction branch, the boundary information perception branch uses the original image input to the network as the perception target for boundary feature information extraction. Each group of reverse perception information layers first performs 3×3 convolution on the feature E j ( j = 1, 2, 3) to obtain, and then performs activation and inversion operations to obtain the inverted feature E rj , and multiplies E rj and element-wise to obtain the original boundary feature E cj . The specific expression method is as follows:

[0024] ,

[0025] In the formula represents Sigmoid function, E is an all - one matrix;

[0026] Step 2.3.2: Input together with E rj and E cj into the residual inverse enhancement module, and fuse features by combining residual connection and convolution operations E cj , and E rj , and finally obtain enhanced boundary - aware features with strong semantic information E xj , and the specific expression method is as follows:

[0027] ,

[0028] In the formula, || represents the concatenation operation along the channel dimension, Conv 1 represents the 1×1 convolution operation, Conv 3 represents the 3×3 convolution operation.

[0029] Furthermore, the misalignment fusion module proposed in step 2.4 takes adjacent - level features T i and T i+1 ( i = 1, 2, 3, 4) as inputs. After convolutions of two scales, the concatenation operation is used to enrich the feature information, and the channel attention operation and the spatial attention operation are used to obtain features with good information expression ability M 2 . Concatenate M 2 with itself in the channel dimension to construct a new feature representation, doubling the number of channels of the output tensor, expanding the representation ability of the feature space while maintaining the integrity of the original feature information, and finally performing local feature extraction through 3×3 convolution to obtain features with strong spatial context expression ability F i , and the specific expression method is as follows:

[0030] ,

[0031] In the formula U represents the upsampling operation, Conv 5Indicates a 5×5 convolution, sa represents spatial attention, ca represents channel attention.

[0032] Furthermore, the misaligned single-layer enhancement module proposed in step 2.4 consists of channel attention and spatial attention. When i = 4, the misaligned single-layer enhancement module uses a dual-branch splicing method of parallel attention operation and cascaded attention operation to perform semantic feature enhancement on the hierarchical features T 4 and finally further extracts the enhanced feature information after splicing through a 3×3 convolution to obtain the feature F 4 , and the specific expression method is as follows:

[0033]

[0034] Furthermore, the perceptual information fusion module proposed in step 2.5 uses the hierarchical feature F i and the boundary perception feature E xj as inputs, and globally fuses the hierarchical feature and the boundary perception feature through a multi-level cascading method. The perceptual information fusion module first uses channel attention operation to reduce redundancy and noise in the boundary perception feature, adopts a dual-branch parallel structure, splices the two features in different orders to obtain the features and , thus forming two feature combination branches. In the two branches, two different scale convolution operations of 3×3 and 5×5 are respectively used, and a global average pooling layer and a residual connection are used to further enhance the fused feature to obtain and , and finally element-wise addition is performed to obtain the perceptual feature The specific expression method is as follows:

[0035] ,

[0036] wherein GAP represents the global average pooling layer.

[0037] Furthermore, the multi-level residual decoding module proposed in step 2.6 includes the following steps:

[0038] Step 2.6.1: The multi-level residual decoding module uses a cascading method to take the output of the hierarchical perceptual information fusion module, the boundary feature E x3 and the output feature of the previous-level multi-level residual decoding module as inputs, globally fuses the features between different levels and the boundary information feature, and finally obtains the fused feature Dk ( k = 1, 2, 3, 4);

[0039] Step 2.6.2: In the multi-level residual decoding module, first, for and D k+1 , ( k = 2, 3), perform an element-wise addition operation to obtain a new fused feature. Subsequently, concatenate this fused feature with the two original features to form a joint feature representation containing three features , thereby enhancing the expressive ability of the features. The specific expression method is as follows:

[0040] ,

[0041] Step 2.6.3: Input the enhanced feature into two cascaded residual convolution modules, and use a two-level residual connection method to improve the semantic expression ability of the features. The specific expression method is as follows:

[0042] ,

[0043] When k = 1, the top-level perception information fusion module uses the output feature D 2 of the previous-level multi-level residual decoding module and the output feature E x3 of the boundary information perception branch as inputs; when k = 4, the bottom-level perception information fusion module uses the output feature F P32 and F P42 as inputs.

[0044] Furthermore, in step 3, the boundary difference joint loss is used to optimize the model training process, which specifically includes the following steps:

[0045] Step 3.1: Take the real mask G and the predicted mask P as inputs, and extract the neighborhood features of the real mask by customizing the edge convolution kernel to generate a tensor T E containing boundary information. The specific expression method is as follows:

[0046] ,

[0047] where Pad 01 represents padding one layer of edges with 0,Conv e represents a custom edge convolution kernel, and there is:

[0048] ,

[0049] Step 3.2: Calculate the true mask G and the boundary information tensor T E to determine the parameter ɑ for adjusting the loss calculation, and the specific expression method is as follows:

[0050] ,

[0051] where N Tnz and N Gnz are the numbers of non-zero elements in the boundary information tensor T E and the true mask G respectively, and a very small constant is used to prevent errors caused by the denominator being zero;

[0052] Step 3.3: Calculate the intersection and sum of squares of the true mask and the predicted mask, and thus obtain the difference measure between the predicted value and the true value as the numerator. Use ɑ as a variable parameter to weight the true mask G and the boundary information tensor T E as the denominator, and calculate the overall loss. The specific expression method is as follows:

[0053] ,

[0054] where sum ele represents the sum of the internal elements of the tensor.

[0055] A polyp image segmentation method based on a dual-branch feature progressive fusion network of the present invention has the following advantages:

[0056] (1) The present invention adopts a dual-branch structure, which can capture multi-level boundary information features while obtaining hierarchical information. The features captured by the dual-branch structure achieve deep fusion of global semantic information, strengthen the relevance of context information, and thus effectively improve the performance of polyp segmentation.

[0057] (2) The present invention proposes an inverse perception information layer as the core of the boundary information perception branch, which combines two parts: feature inversion and residual enhancement. The boundary features are extracted by cascading three groups of inverse perception information layers, and each group of boundary features is combined with the hierarchical features, significantly improving the semantic expression ability of the boundary features.

[0058] (3) The present invention proposes a misalignment fusion module and a misalignment single-layer enhancement module, which use a multi-branch attention mechanism to perform pairwise fusion and filtering on adjacent hierarchical features, and perform separate enhancement on the highest hierarchical features, thereby effectively retaining key information while optimizing the segmentation process.

[0059] (4) The present invention proposes a perceptual information fusion module, which realizes the progressive fusion of features by further filtering multi-level boundary information and adopting a multi-level cascade method, and generates multi-level perceptual features that fully fuse boundary information.

[0060] (5) The present invention proposes a multi-level residual decoding module, which uses a cascaded residual structure to fully perform global fusion on multi-level perceptual features to obtain an accurate and complete segmentation result.

[0061] (6) The present invention uses a joint boundary difference loss to supervise the training process, effectively improving the training effect of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 It is a schematic flow chart of a polyp image segmentation method based on a dual-branch feature progressive fusion network of the present invention;

[0063] Figure 2 It is the overall network structure diagram of the present invention;

[0064] Figure 3 It is the structure diagram of the single-line gating mechanism;

[0065] Figure 4 It is the structure diagram of the reverse perception information layer;

[0066] Figure 5 It is the structure diagram of the misalignment fusion module;

[0067] Figure 6 It is the structure of the misalignment single-layer fusion module;

[0068] Figure 7 It is the structure diagram of the perceptual information fusion module;

[0069] Figure 8 It is the structure diagram of the multi-level residual decoding module;

[0070] Figure 9 It is the experimental visualization comparison diagram. DETAILED DESCRIPTION OF THE INVENTION

[0071] In order to better understand the purpose, structure and function of the present invention, the following further describes in detail a polyp image segmentation method based on a dual-branch feature progressive fusion network of the present invention with reference to the accompanying drawings.

[0072] AsFigure 1 As shown in the figure, a polyp image segmentation method based on a dual-branch feature progressive fusion network of the present invention includes the following steps:

[0073] Step 1: Collect a colonoscopy polyp image and polyp mask label data set, and divide the data set into a training set and a test set;

[0074] The data set contains five sub-data sets, where 1450 pictures selected from two sub-data sets are used as the training set, 798 pictures selected from five data sets are used as the test set, and finally all five data sets are used to evaluate the segmentation performance of the final model.

[0075] Step 2: Construct a polyp image segmentation model based on a dual-branch feature progressive fusion network;

[0076] As Figure 2 shown, in Step 2, the proposed dual-branch feature progressive fusion network includes the following steps:

[0077] Step 2.1: Use the PVTv2 backbone network as a multi-stage feature extraction branch to extract four-stage multi-level features from the original input image T ={ T 1 , T 2 , T 3 , T 4}.

[0078] Step 2.2: Use a single-line gating mechanism to perform feature weighted extraction on the original input image for preliminary feature filtering.

[0079] As Figure 3 shown, the single-line gating mechanism proposed in Step 2.2 performs preliminary feature extraction and filtering on the original image input to the network X First, generate a weight tensor W as through convolution and normalization operations, and then use the weight tensor after channel division ( i = 1, 2) to perform weighted summation on the input features to obtain the fused feature G f , and the specific expression method is as follows:

[0080]

[0081] where Softmax represents the normalization exponential function, div c represents the operation of dividing the tensor by channels, Denotes element-wise multiplication, Denotes element-wise addition.

[0082] Step 2.3: Using the boundary information perception branch, extract high-fidelity boundary information from the features obtained in Step 2.2 through a hierarchical feature enhancement and cascaded fusion strategy, construct a polyp boundary segmentation mask, and extract boundary feature information.

[0083] As Figure 4 shown, the boundary information perception branch proposed in Step 2.3 includes the following steps:

[0084] Step 2.3.1: The boundary information perception branch contains three cascaded reverse perception information layers. As a parallel branch to the multi-level feature extraction branch, the boundary information perception branch uses the original image input to the network as the perception target to extract boundary feature information. Each group of reverse perception information layers first performs a 3×3 convolution on the feature E j ( j = 1, 2, 3) to obtain , then performs an activation and inversion operation to obtain the inverted feature E rj , and multiplies E rj and element-wise to obtain the original boundary feature E cj , and the specific expression method is as follows:

[0085]

[0086] In the formula, denotes Sigmoid function, E is a matrix of all ones.

[0087] Step 2.3.2: Input together with E rj and E cj into the residual inverse enhancement module, and fuse the features E cj , and E rj through residual connection and convolution operations, and finally obtain the enhanced boundary perception feature E xj with stronger semantic information, and the specific expression method is as follows:

[0088]

[0089] In the formula, || denotes the concatenation operation along the channel dimension, Conv1 Represents a 1×1 convolution operation, Conv 3 Represents a 3×3 convolution operation.

[0090] Step 2.4: Use the dislocation fusion module to fuse adjacent hierarchical features in Step 2.1 pairwise to generate a new comprehensive feature representation. At the same time, use the dislocation single-layer fusion module to extract the bottom-layer features T 4 and, with the help of the channel and spatial attention fusion mechanism, refine and enhance the high-level semantic feature information, thereby obtaining a new four-stage multi-level feature F ={ F 1 , F 2 , F 3 , F 4}.

[0091] As Figure 5 shown, the dislocation fusion module proposed in Step 2.4 takes adjacent hierarchical features T i and T i+1 ( i = 1, 2, 3, 4) as inputs. After convolutions of two scales, the splicing operation is used to enrich the feature information, and the channel attention operation and the spatial attention operation are used to obtain features with good information expression ability M 2 . Then, M 2 is self-spliced in the channel dimension to construct a new feature representation, doubling the number of channels of the output tensor, expanding the representation ability of the feature space while maintaining the integrity of the original feature information, and finally, local feature extraction is performed through 3×3 convolution to obtain features with strong spatial context expression ability F i . The specific expression method is as follows:

[0092]

[0093] In the formula U represents the upsampling operation, Conv 5 represents a 5×5 convolution, sa represents the spatial attention, ca represents the channel attention.

[0094] As Figure 6 shown, the dislocation single-layer enhancement module proposed in Step 2.4 is mainly composed of channel attention and spatial attention. When i= 4, the staggered single-layer enhancement module uses a dual-branch splicing method of parallel attention operation and cascade attention operation to perform hierarchical features. T 4 The semantic features are enhanced, and finally the enhanced feature information after splicing is further extracted through 3×3 convolution to obtain the feature F 4 , the specific expression method is as follows:

[0095]

[0096] Step 2.5: Further mine effective semantic information from the boundary feature information extracted in step 2.3, and combine the mined effective semantic information with the multi-level feature information through a multi-level cascade perception information fusion module. F Perform progressive global fusion to obtain boundary aggregation-aware features.

[0097] like Figure 7 As shown, in step 2.5, the proposed perceptual information fusion module uses hierarchical features F i and boundary-aware features E xj As input, the hierarchical features are globally fused with the boundary perception features through a multi-level cascade. The perception information fusion module first uses the channel attention operation to reduce the redundancy and noise in the boundary perception features, and adopts a dual-branch parallel structure to splice the two features in different orders to obtain the feature and , thus forming two feature combination branches. In the two branches, two convolution operations of different scales, 3×3 and 5×5, are used respectively, and the global average pooling layer and residual connection are used to further enhance the fusion features to obtain and , and finally the elements are added to obtain the perceptual features The specific expression method is as follows:

[0098]

[0099] In the formula GAP Represents a global average pooling layer.

[0100] Step 2.6: Cross-level fusion of multi-level features using a multi-level residual decoding module F The features obtained at each level are fused and reduced in dimension by 1×1 convolution to participate in deep supervision during the training process.

[0101] like Figure 8 As shown, in step 2.6, the proposed multi-level residual decoding module includes the following steps:

[0102] Step 2.6.1: The multi-level residual decoding module combines the output of the hierarchical perception information fusion module in a cascade manner. , boundary features E x3 and the output features of the upper-level multi-level residual decoding module as inputs, globally fuse the features between different levels and the boundary information features, and finally obtain the fused features D k ( k = 1, 2, 3, 4).

[0103] Step 2.6.2: In the multi-level residual decoding module, first perform element-wise addition on and D k+1 , ( k = 2, 3) to obtain a new fused feature, and then concatenate the fused feature with the two original features to form a joint feature representation containing three features , thereby enhancing the expression ability of the features. The specific expression method is as follows:

[0104]

[0105] Step 2.6.3: Input the enhanced features into two cascaded residual convolution modules, and use the two-level residual connection method to improve the semantic expression ability of the features. The specific expression method is as follows:

[0106]

[0107] When k = 1, the top-level perceptual information fusion module uses the output features D 2 of the upper-level multi-level residual decoding module E x3 and the output features of the boundary information perception branch k = 4, the bottom-level perceptual information fusion module uses the output features F P32 and F P42 as inputs.

[0108] Step 3: Perform data augmentation on the training data obtained in Step 1, and input it into the dual-branch feature progressive fusion network training model;

[0109] In Step 3, the boundary difference joint loss is used to optimize the model training process, which specifically includes the following steps:

[0110] Step 3.1: The real mask G and the predicted mask PAs the input, the neighborhood features of the real mask are extracted by customizing the edge convolution kernel to generate a tensor containing boundary information. T E , and the specific expression method is as follows:

[0111]

[0112] Among them Pad 01 represents padding the edge with 0 for one layer, Conv e represents the custom edge convolution kernel, and there is:

[0113]

[0114] Step 3.2: Calculate the relationship between the real mask G and the boundary information tensor T E to determine the parameter ɑ for adjusting the loss calculation. The specific expression method is as follows:

[0115]

[0116] Among them N Tnz and N Gnz are the numbers of non-zero elements in the boundary information tensor T E and the real mask G respectively. A very small constant is used to prevent division by zero from causing errors.

[0117] Step 3.3: Calculate the intersection and sum of squares of the real mask and the predicted mask, and use the difference measure between the predicted value and the real value obtained thereby as the numerator. Use ɑ as a variable parameter to weight the real mask G and the boundary information tensor T E as the denominator to calculate the overall loss. The specific expression method is as follows:

[0118]

[0119] Among them sum ele represents summing the elements inside the tensor.

[0120] Step 4: Use the best model obtained after training to test the colonoscopy images in the test set, obtain the polyp segmentation results after testing, and evaluate the segmentation results of the model.

[0121] In step 4, the polyp images in the test set are input into the trained polyp image segmentation model, and the output segmentation mask is evaluated.

[0122] Five classic polyp segmentation datasets, namely CVC-300, Kvasir, ColonDB, ETIS, and ClinicDB, are used to train and test the network proposed in this method. The test results are compared with the mainstream polyp segmentation algorithms in the prior art to illustrate the effectiveness of this method. The proposed network model is implemented using PyTorch and trained on an NVIDIA GeForce RTX 2080 SUPER GPU with 8GB of memory. The input of all images is uniformly set to 352×352, and the Adam optimization algorithm is used to optimize the overall parameters, with the learning rate set to 0.25e-4. The test results on the CVC-ClinicDB and Kvasir datasets are shown in Table 1:

[0123] Table 1 Test results of polyp segmentation on the CVC-ClinicDB and Kvasir datasets for the present invention and 12 other methods

[0124]

[0125] The experimental results on the CVC-ClinicDB and CVC-ColonDB datasets are shown in Table 2:

[0126] Table 2 Test results of polyp segmentation on the CVC-300 and CVC-ColonDB datasets for the present invention and 12 other methods

[0127]

[0128] The experimental results on the ETIS dataset are shown in Table 3:

[0129] Table 3 Test results of polyp segmentation on the ETIS dataset for the present invention and 12 other methods

[0130]

[0131] The bold font represents the optimal value for each row, and the underline indicates the second-best result.

[0132] Among the 5 evaluation indicators, the first 2 are commonly used evaluation indicators in semantic segmentation tasks, and the closer the value of the indicator is to 1, the better the segmentation result.

[0133] Figure 9 This is the visualization comparison chart of the experimental results of the method in the present invention and other methods. From the results, it can be seen that the polyp segmentation method proposed in the present invention can obtain a segmentation result with more accurate boundaries and more complete semantic structures.

[0134] It will be understood that the present invention is described by way of some embodiments, and those skilled in the art will be aware that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the present invention. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.

Claims

1. A polyp image segmentation method based on a dual-branch feature progressive fusion network, characterized in that It includes the following steps: Step 1: Collect a dataset of colonoscopy polyp images and polyp mask labels, and divide the dataset into a training set and a test set; Step 2: Construct a polyp image segmentation model based on a dual-branch feature progressive fusion network; The dual-branch feature progressive fusion network proposed in Step 2 is used to perform the following steps: Step 2.1: Use the PVTv2 backbone network as a multi-level feature extraction branch to extract four-stage multi-level features T = {T1, T2, T3, T4} from the original input image; Step 2.2: Use a single-line gating mechanism to perform feature weighted extraction on the original input image for preliminary feature filtering; Step 2.3: Utilize the boundary information perception branch to extract high-fidelity boundary information from the features obtained in Step 2.2 through hierarchical feature enhancement and cascaded fusion strategies, construct a polyp boundary segmentation mask, and extract boundary feature information; Step 2.4: Use a misaligned fusion module to perform pairwise fusion on adjacent hierarchical features in Step 2.1 to generate a new comprehensive feature representation. At the same time, use a misaligned single-layer fusion module to extract the bottom-layer feature T4, and refine and enhance the high-level semantic feature information with the channel and spatial attention fusion mechanism, thereby obtaining a new four-stage multi-level feature F = {F1, F2, F3, F4}; Step 2.5: Further mine effective semantic information from the boundary feature information extracted in Step 2.3, and perform progressive global fusion of the mined effective semantic information with the multi-level feature F through a multi-level cascaded perception information fusion module to obtain boundary aggregation perception features; The proposed perceptual information fusion module in step 2.5 uses the hierarchical feature F i and the boundary-aware feature E xj as inputs, and globally fuses the hierarchical feature and the boundary-aware feature through a multi-level cascading method. The perceptual information fusion module first uses channel attention operations to reduce redundancy and noise in the boundary-aware feature, adopts a double-branch parallel structure, and concatenates the two features in different orders to obtain features and thus forming two feature combination branches. In the two branches, convolution operations with two different scales of 3×3 and 5×5 are used respectively, and global average pooling layers and residual connections are adopted to further enhance the fused features to obtain and Finally, element-wise addition is performed to obtain the perceptual feature F Pij The specific expression method is as follows: Where GAP represents the global average pooling layer, ca represents channel attention, Conv1 represents a 1×1 convolution operation, Conv3 represents a 3×4 convolution operation, and Conv5 represents a 5×5 convolution operation; Step 2.6: Use a multi-level residual decoding module to cross-level fuse the multi-level feature F and boundary feature information. The features obtained from each level of fusion are reduced in dimension by a 1×1 convolution and participate in deep supervision during training; The multi-level residual decoding module proposed in Step 2.6 includes the following steps: Step 2.6.1: The multi-level residual decoding module, through a cascading method, fuses the output F of the hierarchical perception information fusion module Pi2 , the boundary feature E x3 and the output feature of the previous-level multi-level residual decoding module as inputs, globally fuses the features between different levels and the boundary information feature, and finally obtains the fused feature D k , k = 1, 2, 3, 4; Step 2.6.2: In the multi-level residual decoding module, first perform an element-wise addition operation on F P(k-1)2 and D k+1 , where k = 2, 3, to obtain a new fused feature. Subsequently, concatenate this fused feature with the two original features to form a joint feature representation containing three features to enhance the expressive power of the features. The specific expression method is as follows: Step 2.6.3: Input the enhanced features into two cascaded residual convolution modules, and use a two-level residual connection method to improve the semantic expression ability of the features. The specific expression method is as follows: When k = 1, the topmost perception information fusion module takes the output feature D2 of the upper-level multi-level residual decoding module and the output feature E of the boundary information perception branch as inputs; when k = 4, the bottommost perception information fusion module takes the output feature F of the bottommost perception information fusion module and F as inputs; x3 When k = 4, the bottommost perception information fusion module takes the output feature F of the bottommost perception information fusion module and F P32 and F P42 as inputs; Step 3: Perform data augmentation on the training data obtained in Step 1 and input it into the dual-branch feature progressive fusion network training model; Step 4: Use the best model obtained after training to test the colonoscopy images in the test set, obtain the polyp segmentation results after testing, and evaluate the segmentation results of the model.

2. The polyp image segmentation method based on the dual-branch feature progressive fusion network according to claim 1, wherein The dataset in Step 1 contains five sub-datasets. Among them, 1450 pictures selected from two sub-datasets are used as the training set, 798 pictures selected from five datasets are used as the test set, and finally all five datasets are used to evaluate the segmentation performance of the final model.

3. The polyp image segmentation method based on the dual-branch feature progressive fusion network according to claim 1, wherein The single-line gating mechanism proposed in step 2.2 performs preliminary feature extraction and filtering on the original image of the input network. First, a weight tensor W is generated through convolution and normalization operations. as , and then the weight tensor after channel division is used to perform weighted summation on the input features to obtain the fused feature G. f , and the specific expression method is as follows: W as = Softmax(Conv1(X)) Among them, Solftmax represents the softmax function, and div c represents the operation of dividing the tensor by channels, represents element-wise multiplication, represents element-wise addition.

4. The polyp image segmentation method based on the dual-branch feature progressive fusion network according to claim 1, wherein The boundary information perception branch proposed in Step 2.3 includes the following steps: Step 2.3.1: The boundary information perception branch contains three cascaded reverse perception information layers. As a parallel branch to the multi-level feature extraction branch, the boundary information perception branch takes the original image input into the network as the perception target for boundary feature information extraction. Each group of reverse perception information layers first performs 3×3 convolution on the feature E j , where j = 1, 2, 3, then performs activation and inversion operations to obtain the inverted feature E rj , and multiplies the elements of E rj and E j to obtain the original boundary feature E cj . The specific expression method is as follows: E rj = E - σ(Conv3(E j )) Where σ represents the Sigmoid function and E is a matrix of all ones; Step 2.3.2: Input E j together with E rj and E cj into the residual inverse enhancement module, and fuse the features E cj , E j and E rj by combining residual connection and convolution operations, and finally obtain the enhanced boundary-aware feature E xj with strong semantic information. The specific expression method is as follows: Where || represents the concatenation operation along the channel dimension, Conv1 represents a 1×1 convolution operation, and Conv3 represents a 3×3 convolution operation.

5. The polyp image segmentation method based on the dual-branch feature progressive fusion network according to claim 1, characterized in that The dislocation fusion module proposed in the step 2.4 takes the adjacent hierarchical features T i and T i+1 , where i = 1, 2, 3, 4 as the input. After convolutions of two scales, the splicing operation is used to enrich the feature information, and the channel attention operation and the spatial attention operation are used to obtain the feature M2 with good information expression ability. M2 is self-spliced in the channel dimension to construct a new feature representation, doubling the number of channels of the output tensor, expanding the representation ability of the feature space while maintaining the integrity of the original feature information. Finally, local feature extraction is performed through 3×3 convolution to obtain the feature F i with strong spatial context expression ability. The specific expression method is as follows: M1 = U(T i+1 ) || T i M2 = sa(ca(Conv3(M1)||Conv5(M1))) F i = Conv3(M2 || M2) Where U represents the upsampling operation, Conv5 represents the 5×5 convolution, sa represents the spatial attention, and ca represents the channel attention.

6. The polyp image segmentation method based on the dual-branch feature progressive fusion network according to claim 1, wherein The misaligned single-layer enhancement module proposed in step 2.4 is composed of channel attention and spatial attention. When i = 4, the misaligned single-layer enhancement module uses the dual-branch splicing method of parallel attention operation and cascaded attention operation to enhance the semantic features of the hierarchical feature T4, and finally further extracts the enhanced feature information after splicing through a 3×3 convolution to obtain the feature F4. The specific expression method is as follows: F' t1 = ca(T4) + sa(T4), F' t2 = sa(ca(T4)), F4 = Conv3(F' t1 || F' t2 )。 7. The polyp image segmentation method based on the dual-branch feature progressive fusion network according to claim 1, wherein In step 3, the boundary difference joint loss is used to optimize the model training process, which specifically includes the following steps: Step 3.1: Take the real mask G and the predicted mask P as inputs, and extract the neighborhood features of the real mask by means of a custom edge convolution kernel to generate a tensor T containing boundary information E , and the specific expression method is as follows: T E = Conv e (Pad 01 (G)) where Pad 01 denotes one - layer edge padding with 0, and Conv e denotes a custom edge convolution kernel, and there is: Step 3.2: Calculate the connection between the true mask G and the boundary information tensor T, so as to determine the parameter ɑ for adjusting loss calculation. The specific expression method is as follows: E to determine the parameter ɑ for adjusting loss calculation. The specific expression method is as follows: where N Tnz and N Gnz are the numbers of non-zero elements in the boundary information tensor T E and the true mask G respectively, and a very small constant is used to prevent errors caused by a zero denominator; Step 3.3: Calculate the intersection and sum of squares of the true mask and the predicted mask, and thus obtain the difference measure between the predicted value and the true value as the numerator. Using ɑ as a variable parameter, the true mask G and the predicted mask P are weighted as the denominator to calculate the overall loss. The specific expression method is as follows: I INT = sum ele (G × P) G MS = sum ele (G × G) P MS = sum ele (P × P) Among which sum ele represents the summation of the internal elements of the tensor.

Citation Information

Patent Citations

  • Polyp segmentation method and system based on double-branch feature fusion network

    CN117237641A