Image Segmentation Method Based on Multi-Scale Contourlet Parallel Inverse Attention Network
By paying attention to the network in parallel and reverse direction, combining the airspace and frequency domain characteristics, the problems of easy missing objects in the existing technology and under-digging of contour edge details are solved, achieving more accurate image segmentation and improving network approximation capabilities.
Patent Information
- Application Number
- CN202310678492.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-08
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-06-08
AI Technical Summary
The existing image segmentation method based on convolutional neural networks is prone to miss small-sized objects when dealing with large differences in shape, size, color and texture, and does not fully mine the outline edges and detailed information of the image, lacks a unified segmentation model, and relies heavily on large labeled data sets.
A multi-scale contour wave parallel reverse attention network is adopted. The feature aggregation module combines the airspace and frequency domain features, and a parallel partial decoder and reverse attention module are used to perform feature aggregation and attention calculation, thereby enhancing the ability to capture edge details.
It improves the accuracy of image segmentation and the approximation ability of the network, and can more effectively capture multi-scale and multi-directional contour information, reducing dependence on large labeled data sets.
Smart Images

Figure CN116703934B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to an image segmentation method based on a multi-scale contourlet parallel reverse attention network. Background Art
[0002] Segmentation is a pixel-level classification process that plays an important role in computer vision for various applications such as remote sensing images, medical images, and autonomous driving. Among them, semantic segmentation has limitations due to the variability of appearance, the ambiguity of object boundaries, and the lack of a unified segmentation model, especially in segmenting polyp and building images; polyp images have various shapes, sizes, colors, and textures, and in a clinical environment, small-sized flat polyps are easily missed; therefore, finding discriminative and effective feature representations is crucial.
[0003] Traditional feature extraction methods mainly focus on manually extracted feature descriptors such as color, shape, scale, and texture. Color is the most intuitive feature descriptor, and commonly used color spaces include CIE Lab, HSV, etc.; shape features are usually composed of geometric elements including points, lines, planes, etc., and due to the diversity of object shapes, existing methods consider prominent features such as angles, edges, and curvatures as features suitable for segmentation; multi-scale feature extraction usually uses traditional multi-resolution analysis methods (MRA), which describe the energy distribution in the frequency domain and mainly include wavelets, Shearlets, Ridgelets, curvelets, and contourlets; texture features describe the smoothness of the target surface through structural arrangements and relationships with neighboring pixels, and gray-level co-occurrence matrices (GLCM) and local binary patterns (LBP) are often used as texture features.
[0004] Since manually crafted feature descriptors rely to a large extent on certain types of fixed patterns, methods based on neural networks (CNNs) have received increasing attention. CNN-based image segmentation methods can be generally classified into two categories: patch-based image segmentation methods and full-image-based image segmentation methods. Among them, patch-based image segmentation methods regard segmentation as a per-pixel classification process, that is, the pixel point to be classified is used as the center of the patch. The features of the patch are extracted as the features of the target pixel, and then pattern classification is performed. Full-image-based image segmentation methods achieve end-to-end image segmentation, which is represented by FCN and SegNet. Later, Ronneberger et al. proposed the U-Net network and many variants have emerged, such as U-Net++, UNet 3+, ResUNet, and DenseUNet, etc.
[0005] However, the image segmentation method based on convolutional neural network still has the following limitations: 1. Due to the differences in the shape, size, color, and texture of the targets, the intra-class variation is very large. Therefore, small-sized objects are easily missed, reducing the segmentation accuracy. 2. The CNN-based method does not fully exploit the contour edges and detailed information of the image, ignoring the multi-scale, multi-directional, and multi-resolution representation of features in the frequency domain. 3. There is a lack of a unified segmentation model; most methods rely heavily on large labeled datasets and only achieve good performance on specific datasets.
[0006] Therefore, aiming at the limitations of the above methods, there is an urgent need to provide an efficient and accurate image segmentation method. Summary of the Invention
[0007] To solve the above problems existing in the prior art, the present invention provides an image segmentation method based on a multi-scale contourlet parallel reverse attention network. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0008] In a first aspect, the present invention provides an image segmentation method based on a multi-scale contourlet parallel reverse attention network, including:
[0009] Obtain the image to be segmented;
[0010] Use the feature aggregation module in the multi-scale contourlet parallel reverse attention network to extract the spatial domain features, frequency domain features, and features after frequency domain enhancement of the image to be segmented. Aggregate the features after frequency domain enhancement through the parallel part decoder in the multi-scale contourlet parallel reverse attention network to obtain a global mapping graph. Among them, the feature aggregation module includes a neural network module, a multi-scale contourlet sparse representation module, and a contourlet knowledge-guided learning module, and performs feature-guided learning on the spatial domain features obtained through the neural network module and the frequency domain features obtained through the multi-scale contourlet sparse representation module to obtain the features after frequency domain enhancement.
[0011] According to the reverse attention module in the multi-scale contourlet parallel reverse attention network, layer by layer obtain the reverse attention features by combining the global mapping graph with the spatial domain features, frequency domain features, and features after frequency domain enhancement, and process the reverse attention features through an activation function to obtain a prediction graph.
[0012] Advantages of the present invention:
[0013] An image segmentation method based on a multi-scale contourlet parallel reverse attention network provided by the present invention. First, a neural network module and a multi-scale contourlet sparse representation module are set up, which can combine spatial domain features and frequency domain features to achieve multi-scale and multi-directional contourlet sparse representation. Secondly, a contourlet knowledge-guided learning module is also set up, which can enhance the attention of the multi-scale contourlet parallel reverse attention network to edge learning, capture more edge details, so as to achieve more accurate segmentation and improve the approximation ability of the network. In addition, the present invention integrates frequency domain features and spatial domain features in the same framework, can utilize the advantages of both, avoid prematurely falling into local optimum in network learning, and at the same time the contourlet transform kernel improves the interpretability of the network model.
[0014] The following will further elaborate on the present invention in conjunction with the accompanying drawings and embodiments. Brief Description of the Drawings
[0015] Figure 1 is a flowchart of an image segmentation method based on a multi-scale contourlet parallel reverse attention network provided by an embodiment of the present invention;
[0016] Figure 2 is another flowchart of an image segmentation method based on a multi-scale contourlet parallel reverse attention network provided by an embodiment of the present invention;
[0017] Figure 3 is a schematic diagram of the segmentation results of some images in the CVC-ColonDB polyp dataset by the method provided by an embodiment of the present invention and the existing method;
[0018] Figure 4 is a schematic diagram of the segmentation results of some images in the CVC-ClinicDB polyp dataset by the method provided by an embodiment of the present invention and the existing method;
[0019] Figure 5 is a schematic diagram of the segmentation results of some images in the Kvasir-SEG polyp dataset by the method provided by an embodiment of the present invention and the existing method;
[0020] Figure 6 is a schematic diagram of the segmentation results of some images in the ETIS-Larib polyp dataset by the method provided by an embodiment of the present invention and the existing method;
[0021] Figure 7 is a schematic diagram of the segmentation result map of some images in the Massachusetts building dataset by the method provided by an embodiment of the present invention and the existing method;
[0022] Figure 8 is a schematic diagram of the segmentation result map of some images in the WHU building dataset by the method provided by an embodiment of the present invention and the existing method. Specific Embodiments
[0023] The following further describes the present invention in detail with reference to specific embodiments, but the embodiments of the present invention are not limited thereto. Please refer to Figures 1 - 2 as shown Figure 1 is a flowchart of an image segmentation method based on a multi-scale contourlet parallel inverse attention network provided by an embodiment of the present invention, Figure 2 is another flowchart of an image segmentation method based on a multi-scale contourlet parallel inverse attention network provided by an embodiment of the present invention. An image segmentation method based on a multi-scale contourlet parallel inverse attention network provided by the present invention includes:
[0024] S101. Obtain the image to be segmented.
[0025] S102. Use the feature aggregation module in the multi-scale contourlet parallel inverse attention network to extract the spatial domain features, frequency domain features, and features after frequency domain enhancement of the image to be segmented, and aggregate the features after frequency domain enhancement through the parallel part decoder in the multi-scale contourlet parallel inverse attention network to obtain a global mapping graph; wherein, the feature aggregation module includes a neural network module, a multi-scale contourlet sparse representation module, and a contourlet knowledge-guided learning module, and performs feature-guided learning on the spatial domain features obtained through the neural network module and the frequency domain features obtained through the multi-scale contourlet sparse representation module to obtain the features after frequency domain enhancement.
[0026] S103. According to the inverse attention module in the multi-scale contourlet parallel inverse attention network, combine the global mapping graph with the spatial domain features, frequency domain features, and features after frequency domain enhancement, and obtain the inverse attention features layer by layer. Process the inverse attention features through an activation function to obtain a prediction graph.
[0027] Specifically, please continue to refer to Figure 2As shown in the figure, in this embodiment, the multi-scale contour wave parallel inverse attention network includes a feature aggregation module, a parallel partial decoder, and an inverse attention module. Among them, the feature aggregation module includes a neural network module, a multi-scale contour wave sparse representation module, and a contour wave knowledge-guided learning module. In this embodiment, the neural network module is Res2Net. Through the processing of the neural network module, the spatial domain feature f is obtained. Optionally, the spatial domain features obtained in this embodiment are f1, f2, f3, and f4. In this embodiment, the multi-scale contour wave sparse representation module included in the feature aggregation module includes a scale filter and a direction filter. The essence of the contour wave transform is the process of iteratively convolving the signal with the Laplacian pyramid filter (LP) and the directional filter banks (DFB) to achieve independent decomposition in multiple scales and multiple directions. The frequency domain feature C obtained through the multi-scale contour wave sparse representation module l has the following expression:
[0028]
[0029] where F LP is the Laplacian pyramid filter, F DFB is the directional filter banks, I is the input image, is the four-level decomposition of the contour wave, l is the low-pass component, h is the high-pass component, h_bds is the band-pass direction sub-band in the frequency domain feature, and 4 is the four-level decomposition.
[0030] In this embodiment, L = 4, indicating 4 decomposition scales, and the direction filters from coarse to fine are nlevels = [0, 3, 3, 3], where 0 represents two-dimensional wavelet decomposition. The 4 obtained frequency domain features are cascaded with the 4 spatial domain features, that is, and
[0031] In this embodiment, the size of the feature coefficient matrix after decomposition by the multi-scale contour wave sparse representation module is adjusted to be equal to M 2 , and the expression is:
[0032]
[0033] where M is the adjusted length or width, s is the scale index of the contour wave transform, k is the direction sub-band index of the contour wave transform, and x and y are the coordinates of the feature map.
[0034] It should be noted that the coefficient matrix is the characteristic coefficient matrix after contourlet decomposition. The purpose of adjustment is that the characteristic map after contourlet decomposition is not necessarily a square matrix, so it cannot be cascaded with the characteristic map corresponding to the neural network. s is the scale index of the contourlet transform, k is the directional subband index of the contourlet transform, and x and y are the coordinates of the characteristic map.
[0035] In this embodiment, the feature aggregation module includes a contourlet knowledge-guided learning module for guiding the learning of spatial domain features based on frequency domain features. The specific process is as follows:
[0036] Perform morphological dilation operation on the frequency domain features to obtain the enhanced information of the frequency domain features; where is the real number domain, and W, H, and C1 are the width, height, and number of channels of the frequency domain feature map, respectively;
[0037] Calculate the binary mask according to the enhanced information to obtain the high-frequency information M i,j , and its expression is:
[0038]
[0039] where E i,j is the enhanced feature map corresponding to the pixel positions i and j in the feature map, T mid is the threshold, and i and j are the pixel positions in the feature map;
[0040] Divide the spatial domain features into s groups, where C2 is the number of channels of the spatial domain features, is rounding up; each group of spatial domain features has the following expression:
[0041]
[0042] Multiply the obtained frequency domain features F spe with high-frequency information by each of the s groups of spatial domain features one by one to obtain s groups of enhanced features, and cascade the s groups of enhanced features to obtain frequency domain enhanced features of the same size as the spatial domain features. Its expression is:
[0043]
[0044] where is the i-th group of spatial domain features, is the multiplication operator.
[0045] Cascade the obtained enhanced features with the corresponding spatial domain features and frequency domain features, that is, and
[0046] In this embodiment, the obtained enhanced features, cascaded frequency-domain features, spatial-domain features, and global mapping map are used as the inputs of the reverse attention module (RA), and the reverse attention features are obtained layer by layer. The entire multi-scale contourlet parallel reverse attention network model performs parameter learning through the error backpropagation algorithm to obtain the reverse attention features.
[0047] Finally, the reverse attention features are processed through an activation function, and the processing process is as follows:
[0048]
[0049] Among them, Pr(·) is the prediction operation, and p ij is the value of the prediction map at positions i and j, and exp(·) i,j is the exponential function with the natural constant e as the base. is the upsampling operation, and S3 is the image input to the activation function.
[0050] In summary, an image segmentation method based on a multi-scale contourlet parallel reverse attention network provided by the present invention first sets a neural network module and a multi-scale contourlet sparse representation module, which can combine spatial-domain features and frequency-domain features to achieve multi-scale and multi-directional contourlet sparse representation; secondly, a contourlet knowledge-guided learning module is also set, which can enhance the attention of the multi-scale contourlet parallel reverse attention network to edge learning, capture more edge details to achieve more accurate segmentation, and improve the approximation ability of the network; in addition, the present invention integrates frequency-domain features and spatial-domain features in the same framework, can utilize the advantages of both, avoid premature convergence to local optima in network learning, and at the same time the contourlet transform kernel improves the interpretability of the network model.
[0051] In an optional embodiment of the present invention, the multi-scale contourlet parallel reverse attention network is trained through the following process, specifically:
[0052] Obtain the dataset of images to be segmented and its corresponding dataset of label images, and perform normalization processing on the images to be segmented according to channels; shuffle the dataset of images to be segmented, use 80% of the images to be segmented as the training set, 10% of the images to be segmented as the validation set, and 10% of the images to be segmented as the test set.
[0053] Among them, normalizing the image to be segmented by channels includes normalizing the image to be segmented; optionally, the methods of normalization include min-max normalization, zero-mean normalization, and non-linear normalization, etc.; in this embodiment, but not limited to, the transforms.Normalize(mean, std) function in PyTorch is used to normalize the image to be segmented by channels, so as to accelerate the convergence speed of the model; among them, the parameters mean and std respectively represent the mean and variance sequences of each channel of the image; in this embodiment, the training dataset is X = {x n |n = 1, 2, …, N}, x n is the image to be segmented, N is the number of samples, and the corresponding label image dataset is is the label image, i = 1, …, K.
[0054] Input the training set into the preset multi-scale contourlet parallel reverse attention network for training; among them, obtain the spatial domain feature f according to the neural network module in the feature aggregation module; obtain the frequency domain feature C according to the multi-scale contourlet sparse representation module in the feature aggregation module; the contourlet knowledge-guided learning module in the feature aggregation module guides the spatial domain feature to learn according to the frequency domain feature, and obtains the feature f g after frequency domain enhancement; the parallel partial decoders aggregate the features after frequency domain enhancement to obtain the global mapping map; the reverse attention module obtains the reverse attention features layer by layer by combining the global mapping map with the features after frequency domain enhancement, the frequency domain features and the spatial domain features, and processes the reverse attention features through the activation function to obtain the predicted map set Train the multi-scale contourlet parallel reverse attention network according to the backpropagation algorithm, update the parameters, and obtain the trained multi-scale contourlet parallel reverse attention network. Use the trained multi-scale contourlet parallel reverse attention network to segment the images to be segmented in the test set to verify the performance of the trained multi-scale contourlet parallel reverse attention network.
[0055] It should be noted that in this embodiment, the neural network module in the feature aggregation module is based on Res2Net, combined with the multi-scale contourlet sparse representation module and the contourlet knowledge-guided learning module, including 1 parallel partial decoder (PPD) and 3 reverse attention modules (RA), and the main process is expressed as [f1’, f2’, f3’, f4’, f5’, PD, RA4, RA3, RA2, F sig, where f1’ is the low-level feature obtained after being processed by the neural network module, f2’ is the low-level feature obtained after being processed by the neural network module, the multi-scale contourlet sparse representation module, and the contourlet knowledge-guided learning module, and f3’, f4’, f5’ are the enhanced high-level features obtained after being processed by the neural network module, the multi-scale contourlet sparse representation module, and the contourlet knowledge-guided learning module.
[0056] It should be noted that when constructing the multi-scale contourlet sparse representation module, the parameters of the multi-scale contourlet sparse representation module need to be set, that is, the types of the size filter and the direction filter, as well as the decomposition scale and the number of directions; it should be noted that the size of the frequency domain feature after decomposition of the multi-scale contourlet sparse representation module is adjusted so as to be cascaded with the spatial feature output by the neural network module.
[0057] It should be noted that in the multi-scale Contourlet sparse representation module, the most discriminative features are extracted in a multi-scale and multi-directional manner. In the Contourlet knowledge-guided learning, the contourlet features are used to guide the learning, and the frequency domain features are used as the guidance map to strengthen the spatial features so as to fully capture the frequency domain information.
[0058] In an optional embodiment of the present invention, the effectiveness of the method provided by the present invention is verified through a simulation experiment.
[0059] I. Simulation Conditions
[0060] The simulation test platform of this embodiment is carried out on an HP-Z840 high-performance graphics workstation with the operating system Ubuntu16.04 LTS, having 2 NVIDIA GeForce GTX 1080 graphics cards and an Intel Xeon E5 processor, with a video memory of 64G, and the computer software configuration is PyTorch.
[0061] The data used in the simulation of this embodiment are five polyp data sets with different boundary features and two remote sensing building data sets.
[0062] II. Simulation Content
[0063] Simulation 1: Please refer to Figure 3 as shown Figure 3 is a schematic diagram of the segmentation results of some images in the CVC-ColonDB polyp data set by the method provided by the embodiment of the present invention and the existing method. For the simulation experiment of segmenting the polyp segmentation data set CVC-ColonDB using the method provided by the present invention and the existing method, the results can be seen in Figure 3 .
[0064] From Figure 3It can be intuitively seen from the segmentation results that for the sample images with low contrast between polyps and surrounding mucosa, this embodiment can accurately segment the polyp boundaries, as shown in Figure 3 the first and fourth lines of
[0065] Simulation 2: Please refer to Figure 4 as shown in Figure 4 which is a schematic diagram of the segmentation results of some images in the CVC-ClinicDB polyp dataset by the method provided in the embodiment of the present invention and the existing method. For the simulation experiment of segmenting the polyp segmentation dataset CVC-ClinicDB using the method provided by the present invention and the existing method, the results can be seen in Figure 4 .
[0066] Figure 4 shows the accurate segmentation results of this embodiment, which are consistent with the standard segmentation results at the boundary contours. The model proposed by the present invention can fully aggregate multi-scale contour information in different domains, and adaptively learn through contour waves to enhance important information and suppress irrelevant information, making the prediction output closer to the standard segmentation results.
[0067] Simulation 3: Please refer to Figure 5 as shown in Figure 5 which is a schematic diagram of the segmentation results of some images in the Kvasir-SEG polyp dataset by the method provided in the embodiment of the present invention and the existing method. For the simulation experiment of segmenting the polyp segmentation dataset Kvasir-SEG using the method of the present invention and the existing method, the results can be seen in Figure 5 .
[0068] As can be seen from Figure 5 due to the overexposure of colonoscopy images, the white areas will seriously affect the segmentation accuracy. This embodiment can effectively integrate multi-scale features and perform segmentation effectively without generating artifacts in this area. As shown in the 2-4 lines of Figure 5 .
[0069] Simulation 4: Please refer to Figure 6 as shown in Figure 6 which is a schematic diagram of the segmentation results of some images in the ETIS-Larib polyp dataset by the method provided in the embodiment of the present invention and the existing method. For the simulation experiment of segmenting the polyp segmentation dataset ETIS-Larib using the method of the present invention and the existing method, the results can be seen in Figure 6 .
[0070] Figure 6 illustrates that the present invention has good performance in the segmentation of small targets and "flat" polyps. Since the present invention can fully aggregate multi-resolution features, it can accurately segment small target areas, as shown in Figure 6As shown in lines 1 - 2, by mutually enhancing the features between different domains, rich high - frequency detailed features are obtained, thus effectively segmenting the boundaries of "flat" regions, such as Figure 6 shown in line 4.
[0071] Simulation Five: Please refer to Figure 7 as shown, Figure 7 is a schematic diagram of the segmentation result graphs of some images of the Massachusetts building dataset by the method provided in the embodiment of the present invention and the existing method. For the simulation experiment of segmenting the Massachusetts building dataset to be segmented with the present invention and the existing method, the results can be seen in Figure 7 .
[0072] As can be seen from Figure 7 , the present invention obtains more accurate boundaries and shapes in an environment with complex backgrounds and diverse building layouts. Due to the complementarity between the spatial domain and the frequency domain, and the frequency - domain features can further enhance the spatial - domain features as a guide, the perception ability and segmentation accuracy of the network are improved.
[0073] Simulation Six: Please refer to Figure 8 as shown, Figure 8 is a schematic diagram of the segmentation result graphs of some images of the WHU building dataset by the method provided in the embodiment of the present invention and the existing method. For the simulation experiment of segmenting the WHU building dataset to be segmented with the present invention and the existing method, the results can be seen in Figure 8 .
[0074] In the WHU remote - sensing building dataset, due to the diversity of the shapes, sizes, and spatial distribution patterns of buildings in its test set, the buildings can be divided into four groups according to the pixel ratio of the buildings in the whole image. As Figure 8 shown, the ratios from top to bottom are 0 - 25%, 25 - 50%, 50 - 75%, and 75 - 100% respectively. For the buildings with the smallest ratio in the first row, the present invention can clearly identify the buildings in the gaps. For the buildings with relatively larger ratios in the second to fourth rows, more accurate edges and shapes can be extracted. Generally speaking, the model proposed by the present invention has achieved good segmentation results on WHU building images with different ratios.
[0075] III. Evaluation of Segmentation Results
[0076] In the evaluation of the segmentation simulation experiment of polyp images, the Dice coefficient and mIOU are commonly used for evaluation. The calculation formula of the Dice coefficient is as follows:
[0077]
[0078]
[0079] Among them, \(G\) and \(P\) are the prediction map and the segmentation label map respectively, the symbols \(\cap\) and \(\cup\) represent the intersection and union operations respectively, and \(K\) is the total number of categories.
[0080] In the segmentation simulation experiment evaluation of remote sensing buildings, the overall accuracy (OA), precision, recall, and F1-score are used for evaluation, and their calculation formulas are as follows:
[0081]
[0082]
[0083]
[0084]
[0085] Among them, TP, TN, FN, and FP represent the number of true positive, true negative, false negative, and false positive samples respectively.
[0086] Calculate the Dice coefficient and mIOU of the above simulation on the CVC-ColonDB dataset, and the results are shown in Table 1.
[0087] Table 1 Results of the present invention and other comparison methods on the CVC-ColonDB dataset
[0088] Method PraNet Dice 0.709 mIoU 0.640 U - Net 0.759 - ResUNet+++TTA 0.847 0.847 DeepLabv3+ 0.862 - CDED - Net 0.896 - MED - Net 0.908 - The present invention 0914 0859
[0089] As can be seen from Table 1, the present invention is superior to other methods in both indicators; among them, the Dice coefficient (0.914) and mIoU (0.859) are respectively improved by 20.5% and 21.9% compared with PraNet; it shows that the contour wave has great potential in describing polyp images with multiple scales and multiple directions.
[0090] As can be seen from Table 2, the Dice coefficient (0.926) of the present invention on the CVC-ClinicDB dataset is superior to other comparison algorithms, which shows that by integrating frequency domain features and guidance information, aggregating multi-scale and multi-directional information, an accurate segmentation result can be effectively obtained on polyp images with a relatively sparse representation.
[0091] Table 2 Results of the present invention and other comparison methods on the CVC-ClinicDB dataset
[0092]
[0093]
[0094] As can be seen from Table 3, the results of the present invention in terms of the two evaluation metrics of Dice coefficient (0.865) and mIoU (0.798) are better than those of other comparison methods; due to the small number of samples and diverse polyp types in the ETIS-Larib dataset, it shows that the present invention has certain potential on small sample datasets.
[0095] Table 3 Results of the present invention and other comparison methods on the ETIS-Larib dataset
[0096] Method U - Net Dice 0.398 mIoU 0.335 U - Net++ 0.401 0.344 ResUNet+++TTA + CRF 0.602 0.743 ResUNet+++TTA 0.614 0.746 ResUNet+++CRF 0.622 0.752 PraNet 0.628 0.567 ResUNet++ 0.636 0.753 The present invention 0.865 0.798
[0097] Table 4 shows the segmentation results when the mixed dataset composed of Kvasir-SEG and CVC-ClinicDB is used as the training set and CVC-ColonDB is used as the test dataset; compared with other algorithms, the present invention obtains the highest Dice coefficient (0.734) on the CVC-ColonDB dataset, with a more accurate segmentation result.
[0098] Table 4 Results of the present invention and other comparison methods on the mixed dataset (Kvasir-SEG + CVC-ClinicDB)
[0099]
[0100]
[0101] Table 5 presents the cross-dataset performance of the present invention on Kvasir-SEG. Among them, Kvasir-SEG is used as the training set, and CVC-ClinicDB, ETIS-Larib, and CVC-ColonDB are used as the test sets. It can be seen that the present invention obtains the highest Dice coefficient on all three test sets. Thus, to a certain extent, it demonstrates the generalization performance of the present invention.
[0102] Table 5 Cross-dataset segmentation results of the present invention and other comparison methods with Kvasir-SEG as the training set
[0103]
[0104] As can be seen from Table 6, the present invention achieves the best performance in all four evaluation metrics on the Massachusetts dataset, which indicates that the multi-scale features in the frequency domain contribute to the improvement of segmentation performance, and the contourlet knowledge-guided learning module is effective for building segmentation in remote sensing images.
[0105] Table 6 Results of the present invention and other comparison methods on the Massachusetts dataset
[0106] Method FCN_8 Overall accuracy 0.878 Precision 0.798 Recall 0.787 F1 - score 0.793 U - Net 0.886 0.821 0.771 0.795 SegNet 0.889 0.883 0.732 0.801 PSPNet 0.844 0.741 0.643 0.688 PraNet 0.879 0.828 0.789 0.812 The present invention 0.901 0.889 0.791 0.838
[0107] As can be seen from Table 7, the present invention exceeds PraNet by approximately 1.8%, 1.5%, and 1.2% in terms of accuracy, F1 score, and recall rate, respectively. This verifies the effectiveness of the present invention in the remote sensing building extraction task.
[0108] Table 7 Results of the present invention and other comparison methods on the WHU dataset
[0109] Method FCN_8 Overall accuracy 0.961 Precision 0.892 Recall 0.887 F1 - score 0.890 U - Net 0.978 0.912 0.915 0.913 SegNet 0.972 0.914 0.900 0.907 PSPNet 0.947 0.871 0.861 0.866 PraNet 0.980 0.920 0.924 0.922 The present invention 0.981 0.938 0.936 0.937
[0110] In summary, the image segmentation method based on the multi-scale contourlet parallel reverse attention network provided by the present invention first sets up a neural network module and a multi-scale contourlet sparse representation module, which can combine spatial domain features and frequency domain features to achieve multi-scale and multi-directional contourlet sparse representation. Secondly, a contourlet knowledge-guided learning module is also set up, which can enhance the attention of the multi-scale contourlet parallel reverse attention network to edge learning and capture more edge details; it can enhance spatial features to achieve more accurate segmentation and improve the approximation ability of the network. In addition, the present invention integrates frequency domain features and spatial domain features in the same framework, can utilize the advantages of both, avoid premature convergence to local optima in network learning, and at the same time the contourlet transform kernel improves the interpretability of the network model.
[0111] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant is intended to cover non-exclusive inclusion, so that an article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of another identical element in the article or device including the said element. "Connection" or "connected" and other similar words are not limited to physical or mechanical connection, but may include electrical connection, whether direct or indirect. The orientation or positional relationship indicated by "up", "down", "left", "right", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation to the present invention.
[0112] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0113] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. An image segmentation method based on a multi-scale contourlet parallel reverse attention network, characterized in that, Including: Obtain the image to be segmented; Use the feature aggregation module in the multi-scale contourlet parallel reverse attention network to extract the spatial domain features, frequency domain features, and features after frequency domain enhancement of the image to be segmented, and aggregate the features after frequency domain enhancement through the parallel part decoder in the multi-scale contourlet parallel reverse attention network to obtain a global mapping graph; wherein, the feature aggregation module includes a neural network module, a multi-scale contourlet sparse representation module, and a contourlet knowledge-guided learning module, and performs feature-guided learning on the spatial domain features obtained through the neural network module and the frequency domain features obtained through the multi-scale contourlet sparse representation module to obtain the features after frequency domain enhancement; According to the reverse attention module in the multi-scale contourlet parallel reverse attention network, layer by layer obtain reverse attention features by combining the global mapping graph with the spatial domain features, the frequency domain features, and the features after frequency domain enhancement, and process the reverse attention features through an activation function to obtain a prediction graph.
2. The image segmentation method based on a multi-scale contourlet parallel reverse attention network according to claim 1, characterized in that, Also including: Perform normalization processing on the spatial domain features and the frequency domain features according to channels.
3. The image segmentation method based on a multi-scale contourlet parallel reverse attention network according to claim 1, characterized in that, The multi-scale contourlet sparse representation module includes a scale filter and a direction filter.
4. The image segmentation method based on a multi-scale contourlet parallel reverse attention network according to claim 3, characterized in that, The four-level decomposition expression of the multi-scale contourlet sparse representation module is: Among them, F LP is the Laplacian pyramid filter, F DFB is the directional filter bank, I is the input image, is the four-level decomposition of the contourlet, l is the low-pass component, h is the high-pass component, h_bds is the band-pass direction sub-band in the frequency domain feature, and 4 is the four-level decomposition.
5. The image segmentation method based on a multi-scale contourlet parallel reverse attention network according to claim 4, characterized in that, Adjust the size of the feature coefficient matrix after decomposition by the multi-scale contourlet sparse representation module to make it equal to M 2 , and the expression is as follows: where M is the adjusted length or width, s is the scale index of the contourlet transform, k is the direction sub-band index of the contourlet transform, and x and y are the coordinates of the feature map.
6. The image segmentation method based on a multi-scale contourlet parallel reverse attention network according to claim 1, characterized in that, The performing feature-guided learning on the spatial domain features obtained through the neural network module and the frequency domain features obtained through the multi-scale contourlet sparse representation module to obtain the features after frequency domain enhancement includes: Perform a morphological dilation operation on the frequency-domain feature to obtain enhanced information of the frequency-domain feature; where is the real number domain, and W, H, and C1 are the width, height, and number of channels of the frequency-domain feature map, respectively; Calculate a binary mask according to the enhanced information to obtain high-frequency information M i,j , and its expression is: Among them, E i,j is the feature enhancement map corresponding to the pixel positions i and j in the feature map, and T mid is the threshold, and i and j are the pixel positions in the feature map respectively; Divide the airspace feature into s groups, where C2 is the number of channels of the airspace feature, and is rounded up; the expression of each group of airspace features is: The frequency-domain feature F with high-frequency information will be obtained. spe Multiply it with each of the s groups of spatial-domain features one by one to obtain s groups of enhanced features, and cascade the s groups of enhanced features to obtain the frequency-domain enhanced features with the same size as the spatial-domain features. The expression is as follows: Among them, is the i-th group of airspace features, is the multiplication operator.
7. The image segmentation method based on a multi-scale contourlet parallel reverse attention network according to claim 1, characterized in that, The expression for processing the reverse attention features through an activation function is: where Pr(·) is the prediction operation, and p ij is the value of the prediction map at positions i and j, and exp(·) i,j is the exponential function with the natural constant e as the base, is the upsampling operation, and S3 is the image of the input activation function.
Citation Information
Patent Citations
Polyp segmentation method combining attention U-shaped network and multi-scale feature fusion
CN114820635A
AUPN727295A0