An Edge-Enhanced Salient Object Detection Network and Algorithm
By designing an edge-enhanced significance target detection network, combining feature fusion and edge detection of multiple modules, the problem of insufficient refinement of traditional algorithms when processing image edge details is solved, and high-precision detection and edge segmentation of significance targets are achieved.
Patent Information
- Application Number
- CN202210785259.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-05
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-07-05
AI Technical Summary
Traditional significance object detection algorithms have insufficient refinement when processing image edge details, especially when processing complex edges such as hair and animal hair.
An edge-enhanced significance target detection network is designed, including a Backbone network module, an incremental feature fusion module, a core change edge detection module and a pixel-by-pixel addition output operation module. Through the combination of these modules, structured detailed features of the image edge are captured and enhanced.
It has achieved the supplementary information of structured details such as the edge of the significant graph, improved the precise capture ability of the significance target, and can accurately segment the edge of the subject in complex environments, which has strong robustness.
Smart Images

Figure CN115205643B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of salient object detection, and in particular, to an edge-enhanced salient object detection network and algorithm. Background Art
[0002] In computer vision, a saliency map is an image that shows the unique quality of each pixel. The goal of a saliency map is to simplify or change the image representation into a more meaningful and easier-to-analyze image. For example, if a pixel has a high gray level or other unique color quality in a color image, the quality of that pixel will be shown in a more prominent way in the saliency map. Saliency detection can be regarded as an instance of image segmentation.
[0003] Salient object detection is mainly applied to image foreground segmentation. It can quickly design creative pictures, or replace the background for pictures or video frames, integrating foreground figures into different scenes to produce creative applications. However, the traditional manual processing method has certain requirements for the professional skills of personnel, and has problems such as huge workload, slow speed, and poor effect. In recent years, with the development of deep learning algorithms, image semantic segmentation algorithms have gradually matured, and segmentation algorithms based on salient objects have been widely used. However, there are also problems in the algorithm itself, such as the failure to refine the edges of the main body, such as the failure to refine the edges of hair, animal hair structures, etc.
[0004] Therefore, providing a new technical solution to improve the above problems is an urgent problem for those skilled in the art. Summary of the Invention
[0005] In view of this, the present invention provides an edge-enhanced salient object detection algorithm to solve the above technical problems.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] An edge-enhanced salient object detection network includes: a Backbone network module, a progressive feature fusion module, a kernel change edge detection module, and a pixel-by-pixel addition output operation module.
[0008] In the above solution, the Backbone network module is used to extract features from the input three-channel image to be detected and obtain coarse-grained salient features of multiple sizes.
[0009] In the above solution, the progressive feature fusion module is used to fuse the coarse-grained salient features obtained from the Backbone network module to obtain fused features.
[0010] In the above solution, the nuclear change edge detection module is used to capture the structured detail features lost by the salient object during the upsampling process from the coarse-grained saliency features obtained from the Backbone network module.
[0011] In the above solution, the pixel-by-pixel addition output operation module is used to perform pixel-by-pixel addition fusion on the fusion features obtained from the progressive feature fusion module and the structured features obtained from the nuclear change edge detection module to obtain a saliency map.
[0012] In the above solution, the Backbone network module includes a ResNet50 sub-network unit and a VGG16 sub-network unit. The ResNet50 sub-network unit includes a first sampling module, a second sampling module, a third sampling module, and a fourth sampling module. The first sampling module is used to perform 2-fold downsampling on the three-channel image to be processed to obtain a low-resolution feature map with a size of 1 / 2 of the three-channel image to be processed. The second sampling module is used to perform 4-fold downsampling on the three-channel image to be processed to obtain a low-resolution feature map with a size of 1 / 4 of the three-channel image to be processed. The third sampling module is used to perform 8-fold downsampling on the three-channel image to be processed to obtain a low-resolution feature map with a size of 1 / 8 of the three-channel image to be processed. The fourth sampling module is used to perform 8-fold downsampling on the three-channel image to be processed to obtain a low-resolution feature map with a size of 1 / 16 of the three-channel image to be processed. The VGG16 sub-network unit includes a first feature extraction module, a second feature extraction module, a third feature extraction module, and a fourth feature extraction module. The first feature extraction module is used to extract features from the low-resolution feature map output by the first sampling module to obtain a coarse-grained saliency feature map B_2 with a size of 1 / 2 of the three-channel image to be processed. The second feature extraction module is used to extract features from the low-resolution feature map output by the second sampling module to obtain a coarse-grained saliency feature map B_4 with a size of 1 / 4 of the three-channel image to be processed. The third feature extraction module is used to extract features from the low-resolution feature map output by the third sampling module to obtain a coarse-grained saliency feature map B_8 with a size of 1 / 8 of the three-channel image to be processed. The fourth feature extraction module is used to extract features from the low-resolution feature map output by the fourth sampling module to obtain a coarse-grained saliency feature map B_16 with a size of 1 / 16 of the three-channel image to be processed.
[0013] In the above solution, the progressive feature fusion module includes a feature channel reduction operation unit, a deconvolution upsampling operation unit, a simple pixel addition operation unit, a feature fusion operation unit, and a first backpropagation optimization unit. The feature channel reduction operation unit is used to perform a reduction operation on the channels of the saliency feature map B_4 output by the Backbone network module, the channels of the saliency feature map B_8, and the channels processed by the saliency feature map B_16 through a 1*1 convolutional kernel, so that the number of channels of the saliency feature map B_4, the number of channels of the saliency feature map B_8, and the number of channels of the saliency feature map B_16 are all the same as the number of channels of the saliency feature map B_2 output by the Backbone network module; the deconvolution upsampling operation unit is used to perform deconvolution upsampling operations on the saliency feature map B_4, the saliency feature map B_8, and the saliency feature map B_16 processed by the channel reduction operation, so that the sizes of the saliency feature map B_4, the saliency feature map B_8, and the saliency feature map B_16 processed by the channel reduction operation are all the same as the size of the saliency feature map B_2 output by the Backbone network module, obtaining the saliency feature map P_1_4, the saliency feature map P_1_8, and the saliency feature map P_1_16; the simple pixel addition operation unit is used to perform pixel-by-pixel addition operations on the saliency feature map B_2 output by the Backbone network module and the saliency feature map P_1_4, the saliency feature map P_1_8, and the saliency feature map P_1_16 output by the deconvolution upsampling operation unit, obtaining three feature maps with the same number of channels and a resolution of 1 / 2 the size of the three-channel image to be processed; the feature fusion operation unit is used to process the three feature maps output by the simple pixel addition operation unit using a convolution group containing different-sized convolutional kernels and fuse the three feature maps into one feature map F P ; the first backpropagation optimization unit is used to pass through the feature map F P The corresponding mask label is used to perform loss calculation and backpropagation optimization processing on the feature map F P output by the feature fusion operation unit.
[0014] In the above solution, the nuclear change edge detection module includes a structured loss label acquisition unit, an edge detection unit, and a second backpropagation optimization unit. The structured loss label acquisition unit is used to perform downsampling processing on the real label image corresponding to the three-channel image to be detected, and perform upsampling processing on the label image obtained by the downsampling processing in multiple ways. Subtracting the label image obtained by the upsampling processing from the real label image corresponding to the three-channel image to be detected pixel by pixel to obtain a structured loss label; the edge detection unit is used to process the different-sized coarse-grained saliency feature maps output by the Backbone network module using a convolutional group including convolutional kernels of different sizes, so that the processed feature maps are respectively the same size as the respective feature maps output by the Backbone network module, and adding adjacent processed feature maps pixel by pixel to obtain two intermediate edge feature maps, and performing a channel reduction operation on the two obtained intermediate edge feature maps so that the number of channel outputs of the two intermediate edge feature maps becomes 1, and adding the two intermediate edge maps obtained by the channel reduction operation pixel by pixel to obtain a structured feature map F E ; the second backpropagation optimization unit is used to perform loss calculation and backpropagation optimization processing on the structured feature map F E through the structured loss label obtained by the structured loss label acquisition unit.
[0015] The present invention also provides an edge-enhanced saliency target detection algorithm, which uses the edge-enhanced saliency target detection network described above for saliency target detection, including:
[0016] Performing feature extraction on the input three-channel image to be detected through the Backbone network module to obtain coarse-grained saliency features of multiple sizes;
[0017] Fusing the coarse-grained saliency features obtained from the Backbone network module through a progressive feature fusion module to obtain a fused feature;
[0018] Capturing the structured detail features lost by the saliency target during the upsampling process from the coarse-grained saliency features obtained from the Backbone network module through the nuclear change edge detection module;
[0019] Performing pixel-by-pixel addition fusion on the fused feature obtained by the progressive feature fusion module and the structured feature obtained by the nuclear change edge detection module through a pixel-by-pixel addition output operation module to obtain a saliency map.
[0020] In the above solution, performing feature extraction on the input three-channel image to be detected through the Backbone network module to obtain coarse-grained saliency features of multiple sizes:
[0021] The three-channel image to be processed is respectively downsampled by 2 times, 4 times, 8 times and 16 times through the ResNet50 sub-network unit to obtain low-resolution feature maps of 1 / 2 size, 1 / 4 size, 1 / 8 size and 1 / 16 size of the three-channel image to be processed;
[0022] The VGG16 sub-network unit is used to respectively extract features from the low-resolution feature maps of 1 / 2 size, 1 / 4 size, 1 / 8 size and 1 / 16 size output by the ResNet50 sub-network unit, and respectively obtain the coarse-grained saliency feature map B_2 of 1 / 2 size of the three-channel image to be processed, the coarse-grained saliency feature map B_4 of 1 / 4 size of the three-channel image to be processed, the coarse-grained saliency feature map B_8 of 1 / 8 size of the three-channel image to be processed, and the coarse-grained saliency feature map B_16 of 1 / 16 size of the three-channel image to be processed.
[0023] In the above solution, the fusing of the coarse-grained saliency features obtained from the Backbone network module through the progressive feature fusion module to obtain the fused features includes:
[0024] The feature channel reduction operation unit uses a 1*1 convolutional kernel to perform a reduction operation on the channels of the saliency feature map B_4 output by the Backbone network module, the channels of the saliency feature map B_8, and the channels processed by the saliency feature map B_16;
[0025] The deconvolution upsampling operation unit performs deconvolution upsampling operations on the saliency feature map B_4, the saliency feature map B_8, and the saliency feature map B_16 processed by the channel reduction operation to obtain the saliency feature map P_1_4, the saliency feature map P_1_8, and the saliency feature map P_1_16;
[0026] The simple pixel addition operation unit performs a per-pixel addition operation on the saliency feature map B_2 output by the Backbone network module and the saliency feature map P_1_4, the saliency feature map P_1_8, and the saliency feature map P_1_16 output by the deconvolution upsampling operation unit to obtain three feature maps with the same number of channels and a resolution of 1 / 2 size of the three-channel image to be processed;
[0027] The feature fusion operation unit uses a convolution group containing different-sized convolutional kernels to process the three feature maps output by the simple pixel addition operation unit and fuses the three feature maps into one feature map F P ;
[0028] The first backpropagation optimization unit uses the feature map F P The corresponding mask label for the feature map F output by the feature fusion operation unit PPerform loss calculation and backpropagation optimization processing.
[0029] In the above solution, capturing the structured detail features lost by the saliency target during the upsampling process from the coarse-grained saliency features obtained by the kernel change edge detection module from the Backbone network module includes:
[0030] The structured loss label acquisition unit performs downsampling processing on the true label image corresponding to the three-channel image to be detected, and performs various upsampling processes on the label image obtained by taking the difference of the downsampling processing, and subtracts the label image obtained by the upsampling processing from the true label image corresponding to the three-channel image to be detected pixel by pixel to obtain a structured loss label;
[0031] The edge detection unit processes the coarse-grained saliency feature maps of different sizes output by the Backbone network module using a convolution group including convolution kernels of different sizes, and adds adjacent processed feature maps pixel by pixel to obtain two intermediate edge feature maps, and performs a channel reduction operation on the two obtained intermediate edge feature maps, and adds the two intermediate edge maps obtained by the channel reduction operation pixel by pixel to obtain a structured feature map F E ;
[0032] The second backpropagation optimization unit uses the structured loss label obtained by the structured loss label acquisition unit to perform loss calculation and backpropagation optimization processing on the structured feature map F E Perform loss calculation and backpropagation optimization processing.
[0033] In summary, the beneficial effects of the present invention are: it can supplement structured detail information such as the edges of the saliency map, achieve precise capture of the saliency target, and then segment the saliency subject from the background. In various complex environments such as low foreground and background contrast, complex background, and complex subject shape, accurate segmentation of the subject edge can be obtained, and it has strong robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The drawings constituting a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0035] Figure 1 It is a schematic diagram of the network structure of the edge-enhanced saliency target detection network in the present invention.
[0036] Figure 2 It is a schematic diagram of the network structure of the progressive feature fusion module in the present invention.
[0037] Figure 3 It is a schematic diagram of the network structure of the structured loss label acquisition unit in the present invention.
[0038] Figure 4 This is a schematic diagram of the network structure of the edge detection unit in the present invention.
[0039] Figure 5 This is a step diagram of the edge-enhanced saliency object detection algorithm in the present invention.
[0040] Figure 6 This is a step diagram for feature extraction of the input three-channel image to be detected in the present invention.
[0041] Figure 7 This is a step diagram for obtaining the fused feature in the present invention.
[0042] Figure 8 This is a step diagram for obtaining the structured detail feature in the present invention.
[0043] Figure 9 This is a schematic diagram of the results of saliency object detection in different scenarios in the present invention. Detailed implementation manners
[0044] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the implementation manners and the accompanying drawings. Here, the illustrative implementation manners of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.
[0045] As Figure 1 shown, an edge-enhanced saliency object detection network of the present invention includes: a Backbone network module, a progressive feature fusion module, a kernel change edge detection module, and a pixel-by-pixel addition output operation module.
[0046] The connection relationships between the above-mentioned modules of the present invention will be further described in detail below in conjunction with the accompanying drawings.
[0047] The Backbone network module is used for feature extraction of the input three-channel image to be detected to obtain coarse-grained saliency features of multiple sizes; the progressive feature fusion module is used for fusing the coarse-grained saliency features obtained from the Backbone network module to obtain a fused feature; the kernel change edge detection module is used for capturing the structured detail features lost by the saliency object during the upsampling process from the coarse-grained saliency features obtained from the Backbone network module; the pixel-by-pixel addition output operation module is used for performing pixel-by-pixel addition fusion on the fused feature obtained from the progressive feature fusion module and the structured feature obtained by the kernel change edge detection module to obtain a saliency map.
[0048] Further, the Backbone network module includes a ResNet50 sub-network unit and a VGG16 sub-network unit. The ResNet50 sub-network unit includes a first sampling module, a second sampling module, a third sampling module, and a fourth sampling module. The first sampling module is used to perform 2-fold downsampling on the three-channel image to be processed, and obtain a low-resolution feature map with a size of 1 / 2 of the three-channel image to be processed. The second sampling module is used to perform 4-fold downsampling on the three-channel image to be processed, and obtain a low-resolution feature map with a size of 1 / 4 of the three-channel image to be processed. The third sampling module is used to perform 8-fold downsampling on the three-channel image to be processed, and obtain a low-resolution feature map with a size of 1 / 8 of the three-channel image to be processed. The fourth sampling module is used to perform 8-fold downsampling on the three-channel image to be processed, and obtain a low-resolution feature map with a size of 1 / 16 of the three-channel image to be processed. The VGG16 sub-network unit includes a first feature extraction module, a second feature extraction module, a third feature extraction module, and a fourth feature extraction module. The first feature extraction module is used to extract features from the low-resolution feature map output by the first sampling module, and obtain a coarse-grained saliency feature map B_2 with a size of 1 / 2 of the three-channel image to be processed. The second feature extraction module is used to extract features from the low-resolution feature map output by the second sampling module, and obtain a coarse-grained saliency feature map B_4 with a size of 1 / 4 of the three-channel image to be processed. The third feature extraction module is used to extract features from the low-resolution feature map output by the third sampling module, and obtain a coarse-grained saliency feature map B_8 with a size of 1 / 8 of the three-channel image to be processed. The fourth feature extraction module is used to extract features from the low-resolution feature map output by the fourth sampling module, and obtain a coarse-grained saliency feature map B_16 with a size of 1 / 16 of the three-channel image to be processed.
[0049] Such as Figure 2As shown, the progressive feature fusion module includes a feature channel reduction operation unit, a transposed convolution upsampling operation unit, a simple pixel addition operation unit, a feature fusion operation unit, and a first backpropagation optimization unit. The feature channel reduction operation unit is used to perform a reduction operation on the channels of the saliency feature map B_4 output by the Backbone network module, the channels of the saliency feature map B_8, and the channels processed by the saliency feature map B_16 through a 1*1 convolutional kernel, so that the number of channels of the saliency feature map B_4, the number of channels of the saliency feature map B_8, and the number of channels of the saliency feature map B_16 are all the same as the number of channels of the saliency feature map B_2 output by the Backbone network module. The transposed convolution upsampling operation unit is used to perform a transposed convolution upsampling operation on the saliency feature map B_4, the saliency feature map B_8, and the saliency feature map B_16 processed by the channel reduction operation, so that the sizes of the saliency feature map B_4, the saliency feature map B_8, and the saliency feature map B_16 processed by the channel reduction operation are all the same as the size of the saliency feature map B_2 output by the Backbone network module, obtaining the saliency feature map P_1_4, the saliency feature map P_1_8, and the saliency feature map P_1_16. The simple pixel addition operation unit is used to perform a pixel-by-pixel addition operation on the saliency feature map B_2 output by the Backbone network module and the saliency feature map P_1_4, the saliency feature map P_1_8, and the saliency feature map P_1_16 output by the transposed convolution upsampling operation unit, obtaining three feature maps with the same number of channels and a resolution of 1 / 2 the size of the three-channel image to be processed. The feature fusion operation unit is used to process the three feature maps output by the simple pixel addition operation unit using a convolution group containing different-sized convolutional kernels and fuse the three feature maps into one feature map F P ; The first backpropagation optimization unit is used to perform loss calculation and backpropagation optimization processing on the feature map F output by the feature fusion operation unit through the mask label corresponding to the feature map F P P
[0050] In this embodiment, the progressive feature fusion module is used to fuse four-resolution saliency features obtained from the Backbone network module. The features with a larger resolution contain more structured detail information, and the features with a smaller resolution contain more overall semantic information. Therefore, the cascading strategy adopted by the proposed module is from the features with a smaller resolution to the features with a larger resolution, so that the cascaded feature map contains more detail information.
[0051] In this embodiment, the simple pixel addition operation unit uses the add function in the PyTorch framework to perform a pixel-by-pixel addition operation on the saliency feature map B_2 with the saliency feature maps P_1_4, P_1_8, and P_1_16 output by the transposed convolution upsampling operation unit, respectively, to obtain three feature maps with 128 channels and a resolution of 1 / 2 the size of the three-channel image to be processed; the feature fusion operation unit processes the three feature maps obtained by the simple pixel addition operation unit using convolutional kernels with k = 3, k = 5, and k = 7, respectively, and uses the add function in the PyTorch framework to fuse the above three feature maps into one feature map.
[0052] As Figure 3 and Figure 4 shown, the kernel change edge detection module includes a structured loss label acquisition unit, an edge detection unit, and a second backpropagation optimization unit. The structured loss label acquisition unit is used to perform downsampling processing on the true label image corresponding to the three-channel image to be detected, and perform upsampling processing on the label image obtained by the downsampling processing in various ways, and subtract the label image obtained by the upsampling processing from the true label image corresponding to the three-channel image to be detected pixel by pixel to obtain a structured loss label; the edge detection unit is used to process the coarse-grained saliency feature maps of different sizes output by the Backbone network module using a convolutional group including convolutional kernels of different sizes, so that the sizes of the processed feature maps are respectively the same as the sizes of the respective feature maps output by the corresponding Backbone network module, and add adjacent processed feature maps pixel by pixel to obtain two intermediate edge feature maps, and perform a channel reduction operation on the two obtained intermediate edge feature maps so that the number of channel outputs of the two intermediate edge feature maps becomes 1, and perform a pixel-by-pixel addition operation on the two intermediate edge maps obtained by the channel reduction operation to obtain a structured feature map F E ; the second backpropagation optimization unit is used to perform loss calculation and backpropagation optimization processing on the structured feature map F E using the structured loss label obtained by the structured loss label acquisition unit.
[0053] In this embodiment, the ways of performing upsampling processing on the label image obtained by the downsampling processing include: nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation.
[0054] In this embodiment, the edge detection unit uses convolution kernels with k = 3, k = 5, and k = 7 to process the four feature maps output by the Backbone network module. Specifically, a convolution kernel with k = 3 is used to process the saliency feature map B_2, two convolution kernels with k = 5 are used to process the saliency feature maps B_4 and B_8 respectively, and a convolution kernel with k = 7 is used to process the saliency feature map B_16. At the same time, a convolution kernel with k = 7 can also be used to process the saliency feature map B_2, two convolution kernels with k = 5 are used to process the saliency feature maps B_4 and B_8 respectively, and a convolution kernel with k = 3 is used to process the saliency feature map B_16. The above processing process constitutes a two-way feature calculation process, and the add function in PyTorch is used to add the adjacent feature maps output by the edge detection unit pixel by pixel to obtain two intermediate edge features.
[0055] In this embodiment, the pixel-by-pixel addition output operation module is used to fuse the features obtained by the progressive feature fusion module and the features obtained by the kernel change edge detection module. The progressive feature fusion module obtains coarse-grained saliency target features, which contain more semantic information but lack some structural information. The kernel change edge detection module obtains detailed information such as edges lost due to non-learnable upsampling operations, but contains very little semantic information about saliency targets. Therefore, finally, the features obtained by the progressive feature fusion module and the kernel change edge detection module are added together to obtain a full-resolution saliency map that has both saliency target semantic information and more structural detail information.
[0056] The pixel-by-pixel addition output operation module uses the add function in PyTorch to add the fused feature map F that contains more semantic information P and the structured feature map F that contains more structural detail information E pixel by pixel for fusion. The saliency map finally obtained by this operation takes into account both the semantic information and the detail information of the saliency target.
[0057] As Figure 5 shown, the present invention also provides an edge-enhanced saliency target detection algorithm. Using the edge-enhanced saliency target detection network described above for saliency target detection, it includes:
[0058] Step S1: Extract features from the input three-channel image to be detected through the Backbone network module to obtain coarse-grained saliency features of various sizes;
[0059] Step S2: Fuse the coarse-grained saliency features obtained from the Backbone network module through the progressive feature fusion module to obtain fused features;
[0060] Step S3: Capture the structural detail features lost by the significant target during the upsampling process from the coarse-grained significant features obtained by the nuclear change edge detection module from the Backbone network module;
[0061] Step S4: Through the pixel-by-pixel addition output operation module, perform pixel-by-pixel addition fusion on the fusion features obtained by the progressive feature fusion module and the structural features obtained by the nuclear change edge detection module to obtain a saliency map.
[0062] As Figure 6 shown, the Backbone network module performs feature extraction on the input three-channel image to be detected to obtain coarse-grained significant features of multiple sizes:
[0063] Step S11: Respectively perform 2-fold, 4-fold, 8-fold, and 16-fold downsampling on the three-channel image to be processed through the ResNet50 sub-network unit to respectively obtain low-resolution feature maps of 1 / 2 size, 1 / 4 size, 1 / 8 size, and 1 / 16 size of the three-channel image to be processed;
[0064] Step S12: Respectively perform feature extraction on the low-resolution feature maps of 1 / 2 size, 1 / 4 size, 1 / 8 size, and 1 / 16 size output by the ResNet50 sub-network unit through the VGG16 sub-network unit to respectively obtain the coarse-grained significant feature map B_2 of 1 / 2 size of the three-channel image to be processed, the coarse-grained significant feature map B_4 of 1 / 4 size of the three-channel image to be processed, the coarse-grained significant feature map B_8 of 1 / 8 size of the three-channel image to be processed, and the coarse-grained significant feature map B_16 of 1 / 16 size of the three-channel image to be processed.
[0065] As Figure 7 shown, the progressive feature fusion module fuses the coarse-grained significant features obtained from the Backbone network module, and the obtained fusion features include:
[0066] Step S21: Through the feature channel reduction operation unit, perform a reduction operation on the channels of the significant feature map B_4 output by the Backbone network module, the channels of the significant feature map B_8, and the channels processed by the significant feature map B_16 using a 1*1 convolutional kernel;
[0067] Step S22: Through the deconvolution upsampling operation unit, perform deconvolution upsampling operations on the significant feature map B_4, the significant feature map B_8, and the significant feature map B_16 processed by the channel reduction operation to obtain the significant feature map P_1_4, the significant feature map P_1_8, and the significant feature map P_1_16;
[0068] Step S23: Perform pixel - by - pixel addition operations on the saliency feature map B_2 output by the Backbone network module with the saliency feature maps P_1_4, P_1_8, and P_1_16 output by the transposed convolution upsampling operation unit respectively, to obtain three feature maps with the same number of channels and a resolution of 1 / 2 the size of the three - channel image to be processed;
[0069] Step S24: Use a convolution group containing convolution kernels of different sizes in the feature fusion operation unit to process the three feature maps output by the simple pixel addition operation unit, and fuse the three feature maps into one feature map F P ;
[0070] Step S25: Use the mask label corresponding to the feature map F in the first backpropagation optimization unit P to perform loss calculation and backpropagation optimization processing on the feature map F output by the feature fusion operation unit P ;
[0071] As Figure 8 shown, the structured detail features lost by the saliency target during the upsampling process captured from the coarse - grained saliency features obtained by the kernel - variation edge detection module from the Backbone network module include:
[0072] Step S31: The structured loss label acquisition unit downsamples the true label image corresponding to the three - channel image to be detected, and performs upsampling processing on the label image obtained by taking the difference of the downsampling processing in multiple ways, and subtracts the label image obtained by the upsampling processing from the true label image corresponding to the three - channel image to be detected pixel by pixel to obtain the structured loss label;
[0073] Step S32: The edge detection unit uses a convolution group containing convolution kernels of different sizes to process the coarse - grained saliency feature maps of different sizes output by the Backbone network module, and adds adjacent processed feature maps pixel by pixel to obtain two intermediate edge feature maps, and performs channel reduction operation processing on the two obtained intermediate edge feature maps, and adds the two intermediate edge maps obtained by the channel reduction operation processing pixel by pixel to obtain the structured feature map F E ;
[0074] Step S33: The second backpropagation optimization unit uses the structured loss label obtained by the structured loss label acquisition unit to perform loss calculation and backpropagation optimization processing on the structured feature map F E ;
[0075] As Figure 9As shown, the present invention performs high-precision saliency target detection on targets in complex scenarios including multi-targets, target occlusion, small targets, etc., and has good detection effects.
[0076] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An edge-enhanced saliency object detection network, characterized in that, Including: Backbone network module, progressive feature fusion module, kernel change edge detection module, and pixel-by-pixel addition output operation module; The Backbone network module is used to extract features from the input three-channel image to be detected and obtain coarse-grained saliency features of multiple sizes; The progressive feature fusion module is used to fuse the coarse-grained saliency features obtained from the Backbone network module to obtain fused features; The kernel change edge detection module is used to capture the structured detail features lost by the saliency target during the upsampling process from the coarse-grained saliency features obtained from the Backbone network module; The pixel-by-pixel addition output operation module is used to perform pixel-by-pixel addition fusion on the fused features obtained by the progressive feature fusion module and the structured features obtained by the kernel change edge detection module to obtain a saliency map; The Backbone network module includes a ResNet50 sub-network unit and a VGG16 sub-network unit. A low-resolution feature map is obtained through the ResNet50 sub-network unit, and the low-resolution feature map is subjected to feature extraction through the VGG16 sub-network unit to obtain a coarse-grained saliency feature map; The progressive feature fusion module includes a feature channel reduction operation unit, a deconvolution upsampling operation unit, a simple pixel addition operation unit, a feature fusion operation unit, and a first backpropagation optimization unit; The nuclear change edge detection module includes a structured loss label acquisition unit, an edge detection unit, and a second backpropagation optimization unit. The structured loss label acquisition unit is used to perform downsampling processing on the ground truth label image corresponding to the three-channel image to be detected, and perform upsampling processing on the label image obtained by taking the difference in the downsampling processing in multiple ways, and subtract the label image obtained by the upsampling processing from the ground truth label image corresponding to the three-channel image to be detected pixel by pixel to obtain a structured loss label. The edge detection unit is used to process the coarse-grained saliency feature maps of different sizes output by the Backbone network module by using a convolutional group containing convolutional kernels of different sizes, so that the sizes of the processed feature maps are respectively consistent with the sizes of the respective feature maps output by the Backbone network module, and add the processed feature maps adjacent to each other pixel by pixel to obtain two intermediate edge feature maps, and perform a channel reduction operation on the two obtained intermediate edge feature maps, and perform a pixel-by-pixel addition operation on the two intermediate edge maps obtained after the channel reduction operation to obtain a structured feature map .
2. The edge-enhanced saliency object detection network according to claim 1, wherein The ResNet50 sub-network unit includes a first sampling module, a second sampling module, a third sampling module, and a fourth sampling module. The first sampling module is used to perform 2x downsampling on the three-channel image to be processed, obtaining a low-resolution feature map with a size of 1 / 2 of the three-channel image to be processed. The second sampling module is used to perform 4x downsampling on the three-channel image to be processed, obtaining a low-resolution feature map with a size of 1 / 4 of the three-channel image to be processed. The third sampling module is used to perform 8x downsampling on the three-channel image to be processed, obtaining a low-resolution feature map with a size of 1 / 8 of the three-channel image to be processed. The fourth sampling module is used to perform 16x downsampling on the three-channel image to be processed, obtaining a low-resolution feature map with a size of 1 / 16 of the three-channel image to be processed. The VGG16 sub-network unit includes a first feature extraction module, a second feature extraction module, a third feature extraction module, and a fourth feature extraction module. The first feature extraction module is used to perform feature extraction on the low-resolution feature map output by the first sampling module, obtaining a coarse-grained saliency feature map B_2 with a size of 1 / 2 of the three-channel image to be processed. The second feature extraction module is used to perform feature extraction on the low-resolution feature map output by the second sampling module, obtaining a coarse-grained saliency feature map B_4 with a size of 1 / 4 of the three-channel image to be processed. The third feature extraction module is used to perform feature extraction on the low-resolution feature map output by the third sampling module, obtaining a coarse-grained saliency feature map B_8 with a size of 1 / 8 of the three-channel image to be processed. The fourth feature extraction module is used to perform feature extraction on the low-resolution feature map output by the fourth sampling module, obtaining a coarse-grained saliency feature map B_16 with a size of 1 / 16 of the three-channel image to be processed.
3. The edge-enhanced saliency object detection network according to claim 1, wherein The feature channel reduction operation unit is used to perform a reduction operation on the channels of the saliency feature map B_4 output by the Backbone network module, the channels of the saliency feature map B_8, and the channels processed by the saliency feature map B_16 through a 1 * 1 convolutional kernel, so that the number of channels of the saliency feature map B_4, the number of channels of the saliency feature map B_8, and the number of channels of the saliency feature map B_16 are all the same as the number of channels of the saliency feature map B_2 output by the Backbone network module; the deconvolution upsampling operation unit is used to perform deconvolution upsampling operations on the saliency feature map B_4, the saliency feature map B_8, and the saliency feature map B_16 processed by the channel reduction operation, so that the sizes of the saliency feature map B_4, the saliency feature map B_8, and the saliency feature map B_16 processed by the channel reduction operation are all the same as the size of the saliency feature map B_2 output by the Backbone network module, obtaining the saliency feature map P_1_4, the saliency feature map P_1_8, and the saliency feature map P_1_16; the simple pixel addition operation unit is used to perform pixel-by-pixel addition operations on the saliency feature map B_2 output by the Backbone network module and the saliency feature map P_1_4, the saliency feature map P_1_8, and the saliency feature map P_1_16 output by the deconvolution upsampling operation unit, obtaining three feature maps with the same number of channels and a resolution of 1 / 2 the size of the three-channel image to be processed; the feature fusion operation unit is used to process the three feature maps output by the simple pixel addition operation unit by using a convolutional group containing convolutional kernels of different sizes and fuse the three feature maps into one feature map ; the first backpropagation optimization unit is used to perform loss calculation and backpropagation optimization processing on the feature map output by the feature fusion operation unit through the mask label corresponding to the feature map 。 4. The edge-enhanced saliency object detection network according to claim 1, wherein The second backpropagation optimization unit is used to perform loss calculation and backpropagation optimization processing on the structured feature map by using the structured loss label obtained by the structured loss label acquisition unit. 5. An edge-enhanced saliency object detection algorithm, which performs saliency object detection by applying the edge-enhanced saliency object detection network described in any one of claims 1-4, characterized in that, Comprising: Performing feature extraction on the input three-channel image to be detected through the Backbone network module to obtain coarse-grained saliency features of multiple sizes; Fusing the coarse-grained saliency features obtained from the Backbone network module through the progressive feature fusion module to obtain fused features; Capturing the structured detail features lost by the saliency target during the upsampling process from the coarse-grained saliency features obtained from the Backbone network module through the kernel variation edge detection module; Performing pixel-by-pixel addition fusion on the fused features obtained from the progressive feature fusion module and the structured features obtained from the kernel variation edge detection module through the pixel-by-pixel addition output operation module to obtain a saliency map; Wherein, the Backbone network module includes a ResNet50 sub-network unit and a VGG16 sub-network unit. A low-resolution feature map is obtained through the ResNet50 sub-network unit, and feature extraction is performed on the low-resolution feature map through the VGG16 sub-network unit to obtain a coarse-grained saliency feature map; The progressive feature fusion module includes a feature channel reduction operation unit, a deconvolution upsampling operation unit, a simple pixel addition operation unit, a feature fusion operation unit, and a first backpropagation optimization unit; The nuclear change edge detection module includes a structured loss label acquisition unit, an edge detection unit, and a second backpropagation optimization unit. The structured loss label acquisition unit is used to perform downsampling processing on the true label image corresponding to the three-channel image to be detected, and perform upsampling processing on the label image obtained by taking the difference in the downsampling processing in multiple ways, and subtract the label image obtained by the upsampling processing from the true label image corresponding to the three-channel image to be detected pixel by pixel to obtain a structured loss label. The edge detection unit is used to process the coarse-grained saliency feature maps of different sizes output by the Backbone network module by using a convolutional group containing convolutional kernels of different sizes, so that the sizes of the processed feature maps are respectively consistent with the sizes of the respective feature maps output by the corresponding Backbone network module, and add the processed feature maps adjacent to each other pixel by pixel to obtain two intermediate edge feature maps, and perform channel reduction operation processing on the two obtained intermediate edge feature maps, and perform pixel-by-pixel addition operation on the two intermediate edge maps obtained by the channel reduction operation processing to obtain a structured feature map .
6. The edge enhancement saliency target detection algorithm according to claim 5, characterized in that, The input three-channel image to be detected is subjected to feature extraction by the Backbone network module to obtain coarse-grained saliency features of multiple sizes: The three-channel image to be processed is respectively downsampled by 2 times, 4 times, 8 times, and 16 times through the ResNet50 sub-network unit to obtain low-resolution feature maps of 1 / 2 size, 1 / 4 size, 1 / 8 size, and 1 / 16 size of the three-channel image to be processed; The VGG16 sub-network unit is used to respectively perform feature extraction on the low-resolution feature maps of 1 / 2 size, 1 / 4 size, 1 / 8 size, and 1 / 16 size output by the ResNet50 sub-network unit to obtain the coarse-grained saliency feature map B_2 of 1 / 2 size of the three-channel image to be processed, the coarse-grained saliency feature map B_4 of 1 / 4 size of the three-channel image to be processed, the coarse-grained saliency feature map B_8 of 1 / 8 size of the three-channel image to be processed, and the coarse-grained saliency feature map B_16 of 1 / 16 size of the three-channel image to be processed.
7. The edge-enhanced saliency object detection algorithm according to claim 5, wherein The coarse-grained saliency features obtained from the Backbone network module are fused through the progressive feature fusion module to obtain the fused features, including: The feature channel reduction operation unit uses a 1 * 1 convolutional kernel to perform a reduction operation on the channels of the saliency feature map B_4, the saliency feature map B_8, and the saliency feature map B_16 processed by the Backbone network module; The deconvolution upsampling operation unit performs deconvolution upsampling operations on the saliency feature map B_4, the saliency feature map B_8, and the saliency feature map B_16 processed by the channel reduction operation to obtain the saliency feature map P_1_4, the saliency feature map P_1_8, and the saliency feature map P_1_16; The simple pixel addition operation unit performs pixel-by-pixel addition operations on the saliency feature map B_2 output by the Backbone network module respectively with the saliency feature map P_1_4, the saliency feature map P_1_8, and the saliency feature map P_1_16 output by the deconvolution upsampling operation unit to obtain three feature maps with the same number of channels and a resolution of 1 / 2 size of the three-channel image to be processed; The feature fusion operation unit processes the three feature maps output by the simple pixel addition operation unit by using a convolution group including convolution kernels of different sizes, and fuses the three feature maps into one feature map ; The feature map is optimized by the first backpropagation optimization unit The corresponding mask label is used to perform loss calculation and backpropagation optimization processing on the feature map output by the feature fusion operation unit 8. The edge-enhanced saliency object detection algorithm according to claim 5, wherein Use the structured loss label obtained by the second backpropagation optimization unit through the structured loss label acquisition unit to perform loss calculation and backpropagation optimization processing on the structured feature map and perform loss calculation and backpropagation optimization processing.