Salience Detection and Edge-Guided Optimization Method Based on Multi-Scale Feature Fusion
By adopting multi-scale feature fusion and edge-guided optimization methods in significance detection, the problems of difficulty in detection, blurred edges and non-smoothness in complex backgrounds are solved, and clearer target edges and better segmentation effects are achieved.
Patent Information
- Application Number
- CN202510206192.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-25
AI Technical Summary
The existing significance detection methods have problems such as difficulty in detecting target significant graphs, blurred edges of significant graphs and non-smooth edges in complex contexts.
The significance detection and edge-guided optimization method based on multi-scale feature fusion are adopted. Pre-processing and multi-scale feature fusion are performed through VGG-16 network without full connection layer, edge detection is performed in combination with CFRL edge loss function, and the significance graph is optimized through the ES-fusion module.
The generated significant map is more uniform, the target edge profile is clearer, and the target segmentation between the surrounding environment is better, solving the significance detection problem in complex backgrounds.
Smart Images

Figure CN119723273B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for detecting and processing images, and particularly to a saliency detection and edge-guided optimization method based on multi-scale feature fusion. Background Art
[0002] Identifying salient stimuli in the visual field is a key attentional mechanism in humans. When freely viewing, our eyes tend to fixate on regions in the scene that have unique variations in visual stimuli, such as bright colors, distinct textures, or more complex semantic aspects. If a familiar face or any sudden movement appears, this mechanism will guide our eyes to fixate on the prominent information regions in the scene. In recent years, saliency detection based on this mechanism has played an increasingly important role in applications such as object recognition, image segmentation, and scene understanding. By accurately locating the salient regions that are of interest to the human eye or algorithms in an image, not only can the processing effect of subsequent tasks be optimized, but also the overall computational efficiency can be significantly improved.
[0003] The features extracted by traditional saliency detection methods lack sufficient semantic information and often show certain limitations in complex scenes. Because these methods mainly rely on the contrast calculation and statistical analysis of low-level visual features (such as color, texture, brightness, etc.), although they can effectively handle scenes with simple backgrounds and single targets, the detection accuracy and robustness will significantly decrease when dealing with complex backgrounds, occlusions, and multi-target scenes.
[0004] In recent years, saliency detection methods based on deep learning have gradually become mainstream. By constructing an end-to-end deep neural network model, high-level features with strong representation capabilities can be learned from large-scale data, thus greatly improving the accuracy and robustness of saliency detection. In 2017, researchers proposed a saliency object detection model - UCF based on an encoder-decoder structure. However, while the encoder filters out irrelevant features, it also loses some effective saliency features, and the decoder inevitably introduces some noise information during the upsampling process. In 2018, researchers combined the recurrent computational structure of a recurrent neural network with the encoder-decoder network, added a time series to UCF, and made the basic network be repeatedly trained at different times. The saliency map output by the network at the previous time was regarded as prior knowledge for the network training at the next time, and it was sent into the encoder together with the original image for processing to obtain the saliency map, achieving good results. In order to further improve the boundary detection accuracy of saliency objects, in 2019, researchers made each sub-module between the encoding and decoding parts of the U-Net model maintain a bidirectional feedback connection, and at the same time the model learned to extract boundaries from the saliency predictions output by itself, enhancing the processing ability for boundaries. With the development of deep neural networks, researchers have gradually applied deep feature information to saliency object detection. However, in complex natural scenes, existing saliency detection methods based on deep learning still have problems such as difficult detection of object saliency maps in complex backgrounds, blurred edges of saliency maps, and uneven edges. Summary of the Invention
[0005] The object of the present invention is to solve the technical problems existing in the existing saliency detection methods, such as difficult detection of object saliency maps in complex backgrounds, blurred edges of saliency maps, and uneven edges, and to provide a saliency detection and edge guidance optimization method based on multi-scale feature fusion.
[0006] In order to achieve the above object of the invention, the present invention provides the following technical solutions:
[0007] A saliency detection and edge guidance optimization method based on multi-scale feature fusion, which is characterized in that it includes the following steps:
[0008] S1. Select the VGG-16 network without a fully connected layer as the backbone network, input the original image into it for preprocessing, and then send the preprocessed original image into the multi-scale feature fusion module; in the multi-scale feature fusion module, perform saliency map detection on the preprocessed original image to obtain a predicted saliency map;
[0009] S2. Based on the U-Net network, perform edge detection on the original image using the CFRL edge loss function to obtain an edge map;
[0010] S3. Optimize the predicted saliency map with the edge map through the ES-fusion module to obtain the optimized saliency map, and complete the saliency detection and edge-guided optimization based on multi-scale feature fusion.
[0011] Further, step S1 is specifically as follows:
[0012] S1.1. Select the VGG-16 network without a fully connected layer as the backbone network, input the original image into it for preprocessing, and send the preprocessed original images output by the block2_conv2 layer, block3_conv3 layer, block4_conv3 layer, and block5_conv layer in the VGG-16 network into the multi-scale feature fusion module respectively;
[0013] S1.2. In the multi-scale feature fusion module, use 1×1 and 3×3 convolutional kernels to perform convolution operations on the preprocessed original images output by the block2_conv2 layer, block3_conv3 layer, block4_conv3 layer, and block5_conv layer, and perform batch normalization and relu function activation respectively to obtain the tensors y0 and y1 of each layer;
[0014] S1.3. Use depthwise separable convolution to perform 5×5, 7×7, 9×9, 11×11, 13×13, and 15×15 convolution operations on the preprocessed original images output by the block2_conv2 layer, block3_conv3 layer, block4_conv3 layer, and block5_conv layer respectively, and use the relu function for activation to obtain the tensors y2, y3, y4, y5, y6, and y7 of each layer respectively;
[0015] S1.4. First, add the tensors y2 and y5, y3 and y6, y4 and y7 of each layer to synthesize feature information at different levels, and then splice the results after addition with the tensors y0 and y1 of the corresponding layers together to form feature maps containing multi-feature scales of each layer;
[0016] S1.5. Input the feature maps containing multi-feature scales of each layer into the channel attention module of the CBAM module in the multi-scale feature fusion module, perform global max pooling and global average pooling on the feature maps containing multi-feature scales of each layer in the spatial dimension respectively, compress the spatial dimension to 1 at the same time, and retain the channel information to obtain the feature maps of each layer after global max pooling and global average pooling;
[0017] S1.6. Input the feature maps of each layer after global max pooling and global average pooling into a shared multi-layer perceptron to extract features;
[0018] S1.7. Add the features extracted from each layer, and after activating them through the sigmoid function, obtain the final channel attention weights for each layer respectively;
[0019] S1.8. Multiply the channel attention weights of each layer by the feature maps of the corresponding layers that have undergone global max pooling and global average pooling to obtain the weighted feature maps of each layer processed by the channel attention module, so as to focus on the important information of the channels;
[0020] S1.9. Input the weighted feature maps of each layer into the spatial attention module of the CBAM module, perform max pooling and average pooling on them respectively in the channel dimension, compress the channel dimension to 1, and retain the spatial information to obtain the weighted feature maps of each layer that have undergone max pooling and average pooling;
[0021] S1.10. Perform a concatenation operation on the weighted feature maps of each layer that have undergone max pooling and average pooling along the channel dimension, then the channel dimension will increase accordingly. Extract features through a convolutional layer, then compress the channel dimension to 1, and finally obtain the final spatial attention weights of each layer after activation by the sigmoid function;
[0022] S1.11. Multiply the spatial attention weights of each layer by the weighted feature maps of the corresponding layers that have undergone max pooling and average pooling to obtain the weighted feature maps of each layer processed by the spatial attention module, so as to focus on the spatial positions where the important information is located;
[0023] S1.12. Upsample and fuse the weighted feature maps of each layer processed by the spatial attention module at different scales step by step, and finally generate a high-resolution predicted saliency map.
[0024] Further, in step S2, the CFRL edge loss function includes the cross-entropy loss function , the multi-scale dice loss function and the DAFT loss function ;
[0025] The formula expression of the cross-entropy loss function is:
[0026]
[0027] where: is the total number of original images to be processed, is the th manually marked mask image, is the th predicted edge map;
[0028] The multi-scale dice loss function The formula expression is:
[0029]
[0030] Where: is the weight, is the corresponding scale list, is an element in the corresponding scale list, is the Dice loss at scale;
[0031] The DAFT loss function The formula expression is:
[0032]
[0033] Where: is the weight after introducing the edge enhancement term, is the number of correctly recognized positive samples, is the number of negative samples recognized as positive samples, that is, the number of misreported negative samples; is the number of positive samples recognized as negative samples, that is, the number of missed positive samples; is used to adjust the dynamic weight of the ratio, is used to adjust the dynamic weight of the ratio, is a constant, is a constant;
[0034] The CFRL edge loss function The formula expression is:
[0035] .
[0036] Furthermore, step S3 is specifically as follows:
[0037] S3.1. First, input the predicted saliency map and the edge map into the channel attention module in the ES-fusion module and activate them through the sigmoid function to obtain the channel weights;
[0038] S3.2. Multiply the channel weights by the predicted saliency map to obtain the weighted saliency map, and input the weighted saliency map into the spatial attention module of the ES-fusion module and activate it through the sigmoid function to obtain the saliency map spatial mask;
[0039] S3.3. Multiply the channel weights by the edge map to obtain the weighted edge map;
[0040] S3.4. Add the result of multiplying the weighted edge map by the saliency map spatial mask to the weighted saliency map to obtain an optimized saliency map, thus completing the saliency detection and edge-guided optimization based on multi-scale feature fusion.
[0041] Further, the corresponding scale list = [0.5, 1, 2], which respectively represent halving the size, the original size, and doubling the size;
[0042] The value is ;
[0043] The value is 2.
[0044] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0045] The present invention provides a saliency detection and edge-guided optimization method based on multi-scale feature fusion. Using the VGG-16 network without a fully connected layer as the backbone network, the saliency map detection work is completed by introducing a multi-scale feature fusion module. Considering weak edges and background noise, a CFRL edge loss function is designed. This loss function can not only ensure that the model has the ability to detect edges of different sizes and proportions, but also adjusts the edge weighting method to avoid adding meaningless weights to the background area. Based on this loss function, edge detection is completed through a U-Net network. Considering the problems of blurred boundaries and uneven transitions in the predicted saliency map obtained by the multi-scale feature fusion module, the saliency map is optimized through an ES-fusion module. The predicted saliency map and the edge map are put into the channel attention module of the ES-fusion module and activated by a sigmoid function to obtain channel weights. The channel weights are multiplied by the predicted saliency map to obtain a weighted saliency map. The weighted saliency map is input into the spatial attention module of the ES-fusion module and activated by a sigmoid function to obtain a saliency map spatial mask. Then, the channel weights are multiplied by the edge map to obtain a weighted edge map. Finally, the result of multiplying the weighted edge map by the saliency map spatial mask is added to the weighted saliency map to obtain an optimized saliency map. The generated saliency map is more uniform, the target edge contour is clearer, and the segmentation effect between the target and the surrounding environment is better. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a flowchart of the saliency detection and edge-guided optimization method based on multi-scale feature fusion of the present invention;
[0047] Figure 2 is a schematic diagram of the CBAM module in the embodiment of the saliency detection and edge-guided optimization method based on multi-scale feature fusion of the present invention;
[0048] Figure 3Schematic diagram of the channel attention module in the CBAM module in the embodiment of the saliency detection and edge guidance optimization method based on multi-scale feature fusion of the present invention;
[0049] Figure 4 Schematic diagram of the spatial attention module in the CBAM module in the embodiment of the saliency detection and edge guidance optimization method based on multi-scale feature fusion of the present invention;
[0050] Figure 5 Schematic diagram of the multi-scale feature fusion module in the embodiment of the saliency detection and edge guidance optimization method based on multi-scale feature fusion of the present invention;
[0051] Figure 6 Comparison chart of edge detection effects using different loss functions in the embodiment of the saliency detection and edge guidance optimization method based on multi-scale feature fusion of the present invention, where Figure 6 (a1) and Figure 6 (b1) are the original images, Figure 6 (a2) and Figure 6 (b2) are the edge detection effect diagrams after post-processing using a combination of cross-entropy loss function and multi-scale dice loss function, Figure 6 (a3) and Figure 6 (b3) are the edge detection effect diagrams processed using the CFRL edge loss function;
[0052] Figure 7 Schematic diagram of the ES-fusion module in the embodiment of the saliency detection and edge guidance optimization method based on multi-scale feature fusion of the present invention;
[0053] Figure 8 Comparison chart of the effects of the original image, manually marked mask image, predicted saliency map, and optimized saliency map in the embodiment of the saliency detection and edge guidance optimization method based on multi-scale feature fusion of the present invention, where Figure 8 (a1), Figure 8 (a2) and Figure 8 (a3) are the original images, Figure 8 (b1), Figure 8 (b2) and Figure 8 (b3) are the manually marked mask images, Figure 8 (c1), Figure 8 (c2) and Figure 8 (c3) are the predicted saliency maps, Figure 8 (d1), Figure 8 (d2) and Figure 8 (d3) are the optimized saliency maps. Detailed implementation manners
[0054] The present invention will be further described below with reference to the accompanying drawings and exemplary embodiments.
[0055] Referring to Figure 1 the flowchart shown, the saliency detection and edge-guided optimization method based on multi-scale feature fusion provided by the present invention includes the following steps:
[0056] S1. Select the VGG-16 network without a fully connected layer as the backbone network, input the original image into it for preprocessing, and then send the preprocessed original image into the multi-scale feature fusion module; perform saliency map detection on the preprocessed original image in the multi-scale feature fusion module to obtain a predicted saliency map;
[0057] More specifically, step S1 is as follows:
[0058] S1.1. Select the VGG-16 network without a fully connected layer as the backbone network, input the original image into it for preprocessing, and send the preprocessed original image output from the block2_conv2 layer, block3_conv3 layer, block4_conv3 layer, and block5_conv layer in the VGG-16 network into the multi-scale feature fusion module respectively; the schematic diagram of the multi-scale feature fusion module (MSFF Module) is as Figure 5 shown;
[0059] S1.2. In the multi-scale feature fusion module, use 1×1 and 3×3 convolutional kernels to perform convolutional operations on the preprocessed original image output from the block2_conv2 layer, block3_conv3 layer, block4_conv3 layer, and block5_conv layer, and perform batch normalization and relu function activation respectively to obtain tensors y0 and y1 of each layer;
[0060] S1.3. Use depthwise separable convolution to perform 5×5, 7×7, 9×9, 11×11, 13×13, and 15×15 convolutional operations on the preprocessed original image output from the block2_conv2 layer, block3_conv3 layer, block4_conv3 layer, and block5_conv layer respectively, and use the relu function for activation to obtain tensors y2, y3, y4, y5, y6, and y7 of each layer;
[0061] S1.4. First, add the tensors y2 and y5, y3 and y6, y4 and y7 of each layer to synthesize feature information of different levels, and then splice the added results with the tensors y0 and y1 of the corresponding layers together to form feature maps containing multi-feature scales of each layer;
[0062] S1.5. Input the feature maps of each layer containing multi - feature scales into the channel attention module of the CBAM module in the multi - scale feature fusion module respectively. Perform global max - pooling and global average - pooling on the feature maps of each layer containing multi - feature scales in the spatial dimension respectively, while compressing the spatial dimension to 1 and retaining the channel information, to obtain the feature maps of each layer after global max - pooling and global average - pooling. The CBAM module is as shown in Figure 2 and includes a sequentially connected channel attention module and a spatial attention module. The schematic diagram of the channel attention module is as shown in Figure 3 .
[0063] S1.6. Input the feature maps of each layer after global max - pooling and global average - pooling into a shared multi - layer perceptron to extract features.
[0064] S1.7. Add the features extracted from each layer, and after activation by the sigmoid function, obtain the final channel attention weights of each layer respectively.
[0065] S1.8. Multiply the channel attention weights of each layer by the corresponding feature maps of each layer after global max - pooling and global average - pooling to obtain the weighted feature maps of each layer after being processed by the channel attention module, so as to focus on the important information of the channels.
[0066] S1.9. Input the weighted feature maps of each layer into the spatial attention module of the CBAM module. Perform max - pooling and average - pooling on them respectively in the channel dimension, compress the channel dimension to 1, and retain the spatial information, to obtain the weighted feature maps of each layer after max - pooling and average - pooling. The schematic diagram of the spatial attention module is as shown in Figure 4 .
[0067] S1.10. Perform a concatenation operation on the weighted feature maps of each layer after max - pooling and average - pooling along the channel dimension, then the channel dimension increases accordingly. Then extract features through a convolutional layer, and then compress the channel dimension to 1. Finally, after activation by the sigmoid function, obtain the final spatial attention weights of each layer.
[0068] S1.11. Multiply the spatial attention weights of each layer by the corresponding weighted feature maps of each layer after max - pooling and average - pooling to obtain the weighted feature maps of each layer after being processed by the spatial attention module, so as to focus on the spatial positions where the important information is located.
[0069] S1.12. Perform progressive up - sampling and fusion of the weighted feature maps of each layer after being processed by the spatial attention module at different scales, and finally generate a high - resolution predicted saliency map.
[0070] S2. Based on the U - Net network, perform edge detection on the original image using the CFRL edge loss function to obtain an edge map.
[0071] In step S2, the present invention designs a CFRL (Comprehensive Feature Refinement Loss) edge loss function , and the CFRL edge loss function consists of three parts, including the cross-entropy loss function , the multi-scale dice loss function and the DAFT loss function .
[0072] Among them, the cross-entropy loss function is used to measure the accuracy of pixel-level prediction, and its formula is expressed as:
[0073]
[0074] Where: is the total number of original images to be processed, is the th manually marked mask image, is the th predicted edge map;
[0075] The multi-scale dice loss function is to introduce multi-scale image adjustment, enhance the sensitivity to edges of different scales, and resample (manually marked mask image) and (predicted edge map) at different scales. For each scale , calculate the sum of intersections:
[0076]
[0077] Where: is the resampled image of the manually marked mask image at scale , is the resampled image of the predicted edge map at scale ;
[0078] Then, for each scale calculate (i.e., the Dice loss at scale ):
[0079]
[0080] Where: is a smoothing term, with a value of 1 here, used to avoid the denominator being 0;
[0081] Combined with the weight for Weighted summation is performed to obtain the formula expression of the multi-scale dice loss function:
[0082]
[0083] Where: is the weight, with the default , is the corresponding scale list, and = [0.5, 1, 2], representing reducing the size by half, the original size, and doubling the size respectively. is an element in the corresponding scale list. is the Dice loss at scale
[0084] DAFT loss function The design concept combines the idea of dynamic weight adjustment for special easy and difficult samples and the positive and negative sample adaptive balance mechanism in edge detection, which is suitable for tasks with a serious imbalance in the ratio of positive and negative samples and is very sensitive to fine edge structures.
[0085] First, the basic Tversky Index function is introduced, which is represented by as follows:
[0086]
[0087] Where: is the number of correctly identified positive samples. is the number of negative samples misidentified as positive samples, i.e., the number of false-negative negative samples. is the number of positive samples misidentified as negative samples, i.e., the number of false-positive positive samples. is the weight of is the weight of
[0088] Since the weights and of the basic Tversky Index function are fixed values, here we introduce dynamic weights:
[0089] According to false positives ( ) and false negatives ( ), the dynamic weights used to adjust the ratio and the dynamic weight used to adjust the ratio are calculated:
[0090]
[0091]
[0092] Wherein: and is a constant with a value of 0.5, is a constant with a value of ;
[0093] Introduce dynamic weights based on the Tversky Index function:
[0094]
[0095] Subsequently, introduce an edge enhancement term to assign greater weights to the edge regions:
[0096]
[0097] Wherein: is the weight after introducing the edge enhancement term, is a constant with a value of 2, is the edge map;
[0098] Then, combine with the Focal Loss function to adjust the weights of difficult samples, and the formula expression of the DAFT loss function is:
[0099]
[0100] That is:
[0101]
[0102] Wherein: is the dynamic weight used to adjust the ratio, is the dynamic weight used to adjust the ratio, is a constant with a value of 2;
[0103] Combine the above three loss functions to obtain the CFRL edge loss function designed by the present invention :
[0104] .
[0105] As Figure 6 shown, it is a comparison chart of edge detection effects using different loss functions, wherein Figure 6 (a1) and Figure 6 (b1) are the original images, Figure 6 (a2) and Figure 6 (b2) are the edge detection effect diagrams processed by combining the cross-entropy loss function and the multi-scale dice loss function, Figure 6 (a3) and Figure 6(b3) is the edge detection effect diagram processed by the CFRL edge loss function. It can be seen that the CFRL edge loss function compared with the cross-entropy loss function and the multi-scale dice loss function has a better combined effect, relatively good ability to suppress background noise, and the generated edge map is more coherent.
[0106] S3. Optimize the saliency map by passing the predicted saliency map and the edge map through the ES-fusion module to obtain the optimized saliency map, and complete the saliency detection and edge-guided optimization based on multi-scale feature fusion.
[0107] Step S3 is specifically as follows:
[0108] S3.1. First input the predicted saliency map and the edge map into the channel attention module in the ES-fusion module and activate them through the sigmoid function to obtain the channel weights; the schematic diagram of the ES-fusion module is as Figure 7 shown;
[0109] S3.2. Multiply the channel weights by the predicted saliency map to obtain the weighted saliency map, input the weighted saliency map into the spatial attention module of the ES-fusion module and activate it through the sigmoid function to obtain the saliency map spatial mask;
[0110] S3.3. Multiply the channel weights by the edge map to obtain the weighted edge map;
[0111] S3.4. Add the result of multiplying the weighted edge map by the saliency map spatial mask to the weighted saliency map to obtain the optimized saliency map, and complete the saliency detection and edge-guided optimization based on multi-scale feature fusion.
[0112] As Figure 8 shown, it is the effect comparison diagram of the original image, the manually marked mask image, the predicted saliency map, and the optimized saliency map, where Figure 8 (a1), Figure 8 (a2) and Figure 8 (a3) are the original images, Figure 8 (b1), Figure 8 (b2) and Figure 8 (b3) are the manually marked mask images, Figure 8 (c1), Figure 8 (c2) and Figure 8 (c3) are the predicted saliency maps, Figure 8 (d1), Figure 8 (d2) and Figure 8(d3) is the optimized saliency map. It can be seen that the optimized saliency map obtained after being processed by the saliency detection and edge-guided optimization method based on multi-scale feature fusion of the present invention is more uniform, the target edge contour is clearer, and the segmentation effect with the surrounding environment is better. This method provides a new solution for the accuracy of saliency map detection and the retention of edge details, and is applicable to the target segmentation and edge optimization requirements in complex scenarios.
[0113] The embodiments described above are only descriptions of the specific implementation manners of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A saliency detection and edge guidance optimization method based on multi-scale feature fusion, characterized in that: The following steps are involved: S1. Select the VGG-16 network without a fully connected layer as the backbone network, input the original image into it for preprocessing, and pass the preprocessed original image into the multi-scale feature fusion module; in the multi-scale feature fusion module, perform saliency map detection on the preprocessed original image to obtain a predicted saliency map; specifically: S1.
1. Select the VGG-16 network without fully connected layers as the backbone network, input the original image into it for preprocessing, and send the preprocessed original images output by the block2_conv2 layer, block3_conv3 layer, block4_conv3 layer and block5_conv layer in the VGG-16 network to the multi-scale feature fusion module respectively; S1.
2. In the multi-scale feature fusion module, 1×1 and 3×3 convolution kernels are used to perform convolution operations on the pre-processed original images output by the block2_conv2 layer, block3_conv3 layer, block4_conv3 layer and block5_conv layer, and batch normalization and relu function activation are performed to obtain the tensors y0 and y1 of each layer respectively; S1.3, use separable convolution to perform 5×5, 7×7, 9×9, 11×11, 13×13, and 15×15 convolution operations on the preprocessed original images output by the block2_conv2 layer, block3_conv3 layer, block4_conv3 layer, and block5_conv layer, respectively, and use the relu function to activate, and obtain the tensors y2, y3, y4, y5, y6, and y7 of each layer respectively; S1.4, first add the tensors y2 and y5, y3 and y6, y4 and y7 of each layer to synthesize the feature information of different levels, and then concatenate the added results with the tensors y0 and y1 of each layer to form feature maps of each layer containing multiple feature scales; S1.5, input the feature maps of each layer containing multiple feature scales into the channel attention module of the CBAM module in the multi-scale feature fusion module, perform global maximum pooling and global average pooling on the feature maps of each layer containing multiple feature scales in the spatial dimension, compress the spatial dimension to 1, and retain the channel information, so as to obtain the feature maps of each layer after global maximum pooling and global average pooling; S1.6, input the feature maps of each layer after global maximum pooling and global average pooling into the shared multi-layer perceptron to extract features; S1.7, add the features extracted from each layer, and activate them through the sigmoid function to obtain the final channel attention weights of each layer; S1.8, multiply the channel attention weight of each layer by the feature map of the corresponding layer after global maximum pooling and global average pooling, and obtain the weighted feature map of each layer after being processed by the channel attention module to focus on the important information of the channel; S1.9, input the weighted feature maps of each layer into the spatial attention module of the CBAM module, perform maximum pooling and average pooling on them in the channel dimension, compress the channel dimension to 1, and retain the spatial information, so as to obtain the weighted feature maps of each layer after maximum pooling and average pooling; S1.10, the weighted feature maps of each layer after maximum pooling and average pooling are connected along the channel dimension, and the channel dimension increases accordingly. Then, a convolution layer is used to extract features, and then the channel dimension is compressed to 1. Finally, the final spatial attention weights of each layer are obtained by sigmoid function activation; S1.
11. Multiply the spatial attention weights of each layer by the weighted feature maps of the corresponding layers after maximum pooling and average pooling to obtain the weighted feature maps of each layer processed by the spatial attention module to focus on the spatial location of important information; S1.12, the weighted feature maps of each layer processed by the spatial attention module are gradually upsampled and fused at different scales, and finally a high-resolution predicted saliency map is generated; S2, based on the U-Net network, the CFRL edge loss function is used to perform edge detection on the original image to obtain an edge map; S3. The predicted saliency map and edge map are optimized through the ES-fusion module to obtain the optimized saliency map, completing the saliency detection and edge-guided optimization based on multi-scale feature fusion.
2. The method for saliency detection and edge guidance optimization based on multi-scale feature fusion according to claim 1, characterized in that: In step S2, the CFRL edge loss function L CFRL Including the cross entropy loss function L bce , multi-scale dice loss function L multi-scale-dice And the DAFT loss function L DAFT ; The cross entropy loss function L bce The formula is: Where: N is the total number of original images to be processed, is the i-th manually labeled mask image, is the i-th predicted edge map; The multi-scale dice loss function L multi-scale-dice The formula is: Where: w s is the weight, S is the corresponding scale list, s is an element in the corresponding scale list, Dice s is the Dice loss at scale s; The DAFT loss function L DAFT The formula is expressed as: Where: e is the weight after the edge enhancement term is introduced, TP is the number of positive samples correctly identified, FP is the number of negative samples identified as positive samples, that is, the number of false positive negative samples; FN is the number of positive samples identified as negative samples, that is, the number of false positive positive samples; α t is the dynamic weight used to adjust the FP ratio, β t is the dynamic weight used to adjust the FN ratio, ε is a constant, γ is a constant; The CFRL edge loss function L CFRL The formula is: L CFRL =L bce +L multi-scale-dice +L DAFT .
3. The method for saliency detection and edge guidance optimization based on multi-scale feature fusion according to claim 2, characterized in that: Step S3 is specifically as follows: S3.1, the predicted saliency map and edge map are first input into the channel attention module in the ES-fusion module and activated by the sigmoid function to obtain the channel weight; S3.2, multiply the channel weight by the predicted saliency map to obtain a weighted saliency map, input the weighted saliency map into the spatial attention module of the ES-fusion module and activate it with a sigmoid function to obtain a saliency map spatial mask; S3.3, multiplying the channel weights by the edge map to obtain a weighted edge map; S3.
4. Add the result of multiplying the weighted edge map and the spatial mask of the saliency map to the weighted saliency map to obtain the optimized saliency map, thus completing the saliency detection and edge-guided optimization based on multi-scale feature fusion.
4. The method for saliency detection and edge guidance optimization based on multi-scale feature fusion according to claim 2, characterized in that: The corresponding scale list s=[0.5, 1, 2] represents the size being reduced by half, the original size, and the size being increased by half, respectively; The value of ε is 1×10 -6 ; The value of γ is 2.
Citation Information
Patent Citations
RGB-D-based salient target detection method and storage medium
CN113837223A
Optical remote sensing image saliency target detection method based on attention edge interaction
CN116129289A