A Salient Object Detection Method and System Based on Boundary Enhancement
By adopting multi-level feature fusion, multi-scale information extraction and boundary information extraction methods in the significance object detection technology, the problems of scale changes and pixel blur in the boundary area are solved, and the detection effect of more refined significant boundaries and more consistent significant areas is achieved.
Patent Information
- Application Number
- CN202210467623.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-04-29
AI Technical Summary
The existing significance target detection technology has shortcomings in dealing with scale changes and boundary area pixel blur, and it is difficult to obtain fine significant boundaries and consistent significant areas.
A significance target detection method based on boundary enhancement is adopted, and a feature map containing multi-scale information is generated through multi-level feature fusion and multi-scale information extraction. The boundary information extraction module is used to extract significance boundary information, combined with significance target features for fusion, and finally a mixed loss function is used for model training.
It effectively solves the problems of scale changes and pixel blur in boundary areas, improves the accuracy and consistency of significance target detection, and obtains clearer significance boundaries and more uniform and brighter significance target areas.
Smart Images

Figure CN114821059B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a saliency object detection method and system based on boundary enhancement. Background Art
[0002] The research on saliency object detection has a history of more than twenty years. There are three very important nodes in total. The first wave of development of saliency object detection originated from an article by Itti in 1998. This pioneering work mimics the human attention process and constructs a saliency map from bottom to top using low-level features; the second node is that saliency detection incorporates the concept of an object and becomes a binary segmentation problem of saliency objects, which is more in line with practical applications; the emergence of convolutional neural networks has sparked the third wave of enthusiasm for saliency object detection. Convolutional neural networks have extremely strong feature extraction capabilities, can obtain a larger receptive field, and thus can better detect salient regions in images, which is also the current mainstream method. At present, the field of saliency detection has produced great theoretical and application value, but another value of it is as an auxiliary for many other vision tasks, such as preprocessing for tasks such as object recognition, image editing, and semantic segmentation.
[0003] Scale variation is one of the main challenges in the SOD task. Limited by downsampling operations, it is difficult for CNNs to handle this problem. Different levels of feature layers only have the ability to process specific scales, and the amount of object information contained in features of different resolutions is different. One way is to perform lateral output on each layer of features in the top-down path, upsample them to a unified resolution and then fuse them to obtain an output containing multi-scale information. However, this method only uses separate resolution features in each layer and is not sufficient to handle problems of various scales. There is also a simple strategy of integrating information from feature layers of different resolutions, but this fusion method is prone to information redundancy and noise interference. There is still room for further optimization in the processing methods for scale variation problems.
[0004] During the feature extraction process, detailed information is continuously lost, and the boundary regions of the salient objects obtained by pixel-level saliency methods are often not satisfactory. To obtain a fine salient boundary, there are many innovative methods in addition to the multi-scale feature fusion method. Some methods use a recursive approach to utilize low-level local information to refine high-level features. There are also some methods that use superpixels for preprocessing to extract boundaries before saliency detection or use CRF for postprocessing on the saliency prediction map to maintain object boundaries. Such methods require additional processing and are less efficient. Regarding the choice of loss function, the commonly used training loss function for salient object detection is binary cross-entropy loss. However, the binary cross-entropy loss has a low confidence when judging boundary pixels, resulting in a very blurred boundary and also unable to guarantee the consistency of the salient region. There are many possibilities for improvement in the design of the network structure and loss function for boundary information extraction. Summary of the Invention
[0005] The purpose of the present invention is to provide a saliency object detection method and system based on boundary enhancement to overcome the deficiencies of the prior art.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions:
[0007] A saliency object detection method based on boundary enhancement, comprising the following steps:
[0008] S1. Extract abstract feature maps of different resolutions from the training set images, perform multi-level fusion on the abstract feature maps to obtain a multi-level fusion feature map, and process the multi-level fusion feature map to obtain a feature map containing multi-scale information;
[0009] S2. After converting the information of the obtained feature map containing multi-scale information, perform splicing and fusion to obtain a feature containing boundary information. At the same time, obtain a boundary detection result using the feature after each level of conversion, and then further perform splicing and fusion to obtain a fused boundary detection result;
[0010] S3. After extracting multi-scale information from the feature map containing multi-scale information, splice it with the feature containing boundary information to obtain a saliency object detection result;
[0011] S4. Use the saliency object detection result, each level of boundary detection result, the fused boundary detection result, and the corresponding training set to train the saliency object detection model until the convergence condition is met, and use the trained saliency object detection model for object detection;
[0012] Further, in step S1, the convolutional neural network uses ResNet-50 trained on ImageNet as the backbone of the network, removes the last pooling layer and fully connected layer, and obtains five feature maps of different sizes.
[0013] Furthermore, the network structure formed by using ResNet-50 trained on ImageNet as the backbone of the network includes a multi-level feature aggregation module, a multi-scale information extraction module, and a boundary information extraction module.
[0014] Furthermore, by performing upsampling or pooling to keep the dimensions consistent, adding elements to each other for information supplementation, and then aggregating the features with supplemented information, five multi-level aggregation feature maps of different sizes can be obtained, which can improve the feature expression ability.
[0015] Furthermore, dilated convolutions with different dilation rates are used for sampling, and information of different scales is obtained through different receptive fields, improving the detection ability of the network model for objects with scale changes.
[0016] Furthermore, in the process of gradually fusing features at each level, multi-scale information extraction is performed multiple times on features of different sizes to further fuse multi-scale information.
[0017] Furthermore, using the gradually fused features as input and performing boundary detection on each level of features can extract the boundary information of the object, further refining the saliency object detection result of the network;
[0018] Furthermore, the loss function is used in the training process, and the parameters are adjusted during the backpropagation of the loss.
[0019] Furthermore, the loss function includes a saliency object detection loss and a saliency boundary detection loss. The saliency object detection loss is used to guide the correct classification of saliency object pixel points, and the saliency boundary detection loss is used to guide the correct classification of pixel points in the saliency object boundary region.
[0020] Furthermore, the loss of saliency object detection includes the binary cross-entropy loss BCE for individual pixels and the consistency enhancement loss CEL for the entire image, which can make the detection result more uniformly highlighted.
[0021] Compared with the prior art, the present invention has the following beneficial technical effects:
[0022] The present invention relates to a saliency object detection method based on boundary enhancement, which extracts abstract feature maps of different resolutions from training set images, performs multi-level fusion on the abstract feature maps to obtain multi-level fusion feature maps, and processes the multi-level fusion feature maps to obtain a feature map containing multi-scale information. Taking visual saliency image data as input, a convolutional neural network is used to predict the saliency object region, solving the problems of scale change and pixel blurring in the boundary region in the saliency object detection task. By using feature information of different resolutions to complement each other, the expression ability of single-resolution features is further enhanced. Multi-scale feature extraction is used to extract information of different scales from fixed-resolution features, better solving the problem of object scale change. The boundary is extracted to model the saliency boundary, and after extracting the boundary information, the saliency object feature information is further supplemented, solving the problem of unclear boundary pixels to a certain extent, and obtaining the final saliency object prediction. A hybrid loss function is used to supervise the model training from different levels, highlighting the saliency object region more uniformly and brightly.
[0023] Furthermore, multi-level feature aggregation is adopted to enable features of different scales to aggregate with each other, enhancing the expression ability of fixed-scale features.
[0024] Furthermore, through multi-scale information extraction, multi-scale information is extracted from features of a fixed scale, enhancing the detection ability of the network model for scenarios with large target size changes.
[0025] Furthermore, after extracting the boundary information of the saliency object, the saliency object is supplemented, further improving the quality of model prediction.
[0026] Furthermore, the loss of saliency object detection consists of binary cross-entropy loss and consistency enhancement loss. In particular, the consistency enhancement loss is supervised from the entire image level. On the one hand, it can make the loss function focus more on the foreground, and on the other hand, it can make the loss immune to scale change interference, thereby improving the saliency object detection effect. Description of the Drawings
[0027] Figure 1 is the implementation flowchart of the saliency object detection method with boundary enhancement in the embodiment of the present invention.
[0028] Figure 2 is the network structure diagram of the saliency object detection model with boundary enhancement in the embodiment of the present invention.
[0029] Figure 3 is the internal structure diagram of the multi-level feature aggregation module in the embodiment of the present invention.
[0030] Figure 4 is the internal structure diagram of the multi-scale information extraction module in the embodiment of the present invention.
[0031] Figure 5 It is the internal structure diagram of the boundary information extraction module in the embodiment of the present invention.
[0032] Figure 6 It is the detection effect diagram of the saliency target detection model with boundary enhancement in the embodiment of the present invention. Specific implementation manners
[0033] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0034] The present invention will be further described in detail below with reference to the accompanying drawings:
[0035] As Figure 1 shown, a saliency target detection method based on boundary enhancement includes the following steps;
[0036] S1. Extract abstract feature maps with different resolutions from the training set images, perform multi-level fusion on the abstract feature maps to obtain multi-level fusion feature maps, and process the multi-level fusion feature maps to obtain a feature map containing multi-scale information
[0037] S2. After converting the information of the obtained feature map containing multi-scale information, perform splicing and fusion to obtain a feature containing boundary information. At the same time, use the feature after each level of conversion to obtain a boundary detection result, and then further perform splicing and fusion to obtain a fused boundary detection result;
[0038] S3. Extract multi-scale information from the feature map containing multi-scale information and splice it with the feature containing boundary information to obtain a saliency target detection result;
[0039] S4. Use the saliency target detection result, each level of boundary detection result, the fused boundary detection result, and the corresponding training set to train the saliency target detection model until the convergence condition is met, and use the trained saliency target detection model to perform target detection;
[0040] This application uses publicly available data as the data set, and divides the data set into a training set and a test set.
[0041] The network structure design is as Figure 2As shown in the figure, in the feature extraction stage: The convolutional neural network uses ResNet-50 trained on ImageNet as the backbone of the network, removes the last pooling layer and fully connected layer. The input image is input into the backbone network, and five different levels of abstract feature maps F1 - F5 are extracted after five groups of convolutional operations. The sizes are 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32 of the input respectively, and the number of channels are 64, 256, 512, 1024, 2048 respectively. From F1 to F5, the low-level detailed information decreases continuously, and the high-level semantic information increases gradually.
[0042] The multi-level feature aggregation module, such as Figure 3 shown in the figure, is divided into two stages: complementation and aggregation. From F i-1 to F i+1 the resolution of the features gradually decreases, and the number of channels gradually increases. In the complementation stage S 1 1, the three input features first go through a 1×1 convolution for preprocessing to keep the number of channels consistent. On the one hand, it can reduce the computational amount, and on the other hand, it helps with subsequent element fusion. Then the input feature F i is respectively pooled and upsampled to supplement the information of F i-1 and F i+1 ; F i-1 and F i+1 are respectively pooled and upsampled to supplement the information of F i . The pooling and upsampling operations are to make the complementary features have the same resolution. The feature complementation process is expressed as the following formula:
[0043]
[0044]
[0045]
[0046] In the formula: F′ j represents the feature after reducing the channel dimension of F j ; represents the i-th level feature after supplementation in the complementation stage S 1 1; Conv(·) represents the convolution responsible for changing the channel dimension; ReLU represents the ReLU non-linear activation function; Up(·) represents the upsampling operation; AvgPool(·) represents the average pooling operation. It should be noted that the top-level feature and the bottom-level feature have only one neighbor, so there are only L 2 +L 3 and L 1 +L 2 two channels in the complementation stage.
[0047] The second stage is the feature aggregation stage S2 In this stage, the complementary features from different channels are aggregated to obtain features containing multi-level information as horizontal features and output them to the decoder. The specific formula is as follows:
[0048]
[0049] In the formula: The features are fused in an element-wise addition manner. Similar to the complementary stage, MF 1 and MF 5 only aggregate the features of one neighbor. Additionally, after all element-wise fusion operations in the above two stages, a group of 3×3 convolutions are accompanied. The combination of regularization and the non-linear transformation ReLU is used to further abstract the features.
[0050] Reverse hierarchical feature fusion starts from the top-level MF5. First, multi-scale information extraction is performed on MF5. The feature containing multi-scale information is upsampled to twice the original size, followed by a 1×1 convolution operation to reduce the channel dimension to be consistent with the feature MF4, so that the two features can be added element-wise. Additionally, after the element-wise addition of the feature map elements, a 3×3 convolution operation is attached for further fusion. In this way, it is carried out sequentially. Finally, the multi-level aggregated features M1, M2, M3, M4, and M5 of different scales are hierarchically fused to complete, and four hierarchically fused features h1, h2, h3, and h4 are obtained respectively.
[0051] The multi-scale information extraction module is as Figure 4 shown. Specifically, given an input feature h, in the forward process, the feature first passes through dilated convolutions with dilation rates of 2, 4, and 8 respectively to extract features sh 1 、sh 2 and sh 3 containing different scale information. In the second step, a residual operation is used to add the original feature and these sampled features element-wise, followed by a convolution operation and an activation function to further aggregate the features and improve the non-linear ability, and finally a feature M containing multi-scale information is obtained. As shown in the following formula:
[0052]
[0053] In the formula: Conv 3×3 (·) represents a 3×3 convolution operation; is the fusion of feature layers by element-wise addition.
[0054] The boundary information extraction module is as Figure 5 shown. The four hierarchically fused features h 4 、h 3 、h 2 、h 1As the input, it first passes through an information conversion module to extract the feature eh containing boundary information 4 、eh 3 、eh 2 、eh 1 The information conversion module consists of two convolutional groups of 1×1, 3×3, and 1×1, and a residual connection operation is included in the convolutional group; then the feature containing boundary information is upsampled after reducing the number of channels through a 1×1 convolution to obtain the saliency boundary prediction result e 4 、e 3 、e 2 、e 1 。To transmit the extracted saliency object boundary information to the saliency object prediction branch to make up for the lack of details, the extracted multi-level boundary features eh 4 、eh 3 、eh 2 、eh 1 are upsampled and concatenated along the channels, and then input into the boundary feature aggregation module to obtain the final feature EF containing boundary information. The boundary aggregation module consists of four 3×3 convolutions containing residual operations. The final boundary feature is shown in the following formula:
[0055] EF = EdgeInfo(Concat(Up(eh 1 ), Up(eh 1 ), Up(eh 3 ), Up(eh 4 )))
[0056] In the formula: EdgeInfo(·) represents the boundary feature aggregation module; EF represents the saliency boundary feature that aggregates multi-level information and can be used for fusion with the saliency object feature in the next step.
[0057] During the training process, the backpropagation strategy is used to optimize the parameters of the network, and the loss function is used to assist in training. The loss functions used during the training process are divided into two categories according to different tasks: saliency boundary detection loss and saliency object detection loss. The formula for the total loss function during the training process is shown as follows:
[0058] Loss = L sod + λ 1 L edge
[0059] In the formula: λ 1 is a hyperparameter used to balance the losses of the two tasks, and its value is set to 10 in the experiment.
[0060] Due to the high sparsity of boundary pixels, the number of boundary pixels and non-boundary pixels is highly unbalanced. Therefore, using the balanced binary entropy loss to supervise the salient boundary learning process can solve the problem of pixel imbalance. The formula of the balanced binary entropy loss is expressed as follows:
[0061]
[0062] In the formula: β is the ratio of the number of non-boundary pixels to the total number of all pixels.
[0063] The salient object detection loss function is composed of two loss functions with different focuses, including the binary cross-entropy loss for individual pixel points and the consistency enhancement loss for the foreground region. The total loss is expressed as follows:
[0064] L sod = L bce + L cel
[0065] The binary cross-entropy loss is the most used loss function in the salient object detection task. This loss is a pixel-level loss and converges on all pixels. The formula is shown as follows:
[0066]
[0067] In the formula: P represents the salient object prediction map; p represents a pixel point in P; G represents the ground truth map; g represents a pixel point in G; log(·) is the pixel-level logarithmic operation.
[0068] The consistency enhancement loss is an image-level loss. On the one hand, it can make the loss function focus more on the foreground, and on the other hand, it can make the loss immune to scale changes. The formula of the consistency enhancement loss function is expressed as follows:
[0069]
[0070] In the formula: P represents the salient object prediction map; p represents a pixel point in the prediction map.
[0071] A salient object detection method based on boundary enhancement according to the present invention can solve the problems of large variation in target size, fuzzy prediction of the salient object boundary region, and uneven pixels inside the region in the visual scene for the salient dataset.
[0072] A multi-polymer feature aggregation module is inserted in the network transport layer to enhance the expression ability of the fixed-resolution features by aggregating the feature information of different resolutions in adjacent layers.
[0073] A multi-scale information extraction module is inserted into each level of the network decoder. By extracting multi-scale information from the features of each level, the ability of the network to handle scenarios with large variations in target size is enhanced.
[0074] Based on the gradually fused features, a boundary extraction module is used to detect the boundaries of salient objects, and the boundary features are fused with the salient object features to enhance the detection effect of the network model in the boundary region.
[0075] A hybrid loss function is used for the salient object detection task, and supervision is carried out at both the pixel level and the image level to promote the backpropagation of gradients, strengthen model convergence, and further improve the model training effect;
[0076] This application has achieved competitive Fmax and MAE results on four groups of publicly available salient detection datasets, and its performance is superior to several popular salient object detection methods.
[0077] Embodiment
[0078] A salient object detection method based on boundary enhancement includes the following steps:
[0079] S1, Four groups of publicly available salient datasets are used as the experimental datasets. The specific work process is as follows:
[0080] (1.1), The training set part of the largest dataset is used as the training set of the model, and the test set of the largest dataset and the other three datasets are all used as test sets;
[0081] (1.2), The image dataset is randomly horizontally flipped before being input into network training to achieve data augmentation.
[0082] S2, A feature extraction network is used to extract abstract feature maps with different resolutions and different numbers of channels. The specific work process is as follows:
[0083] (2.1), Remove the last pooling layer and fully connected layer of the ResNet50 network, and only retain the remaining network structure;
[0084] (2.2), The data processed in step (1.2) is input into the ResNet50 feature extraction network in the dimension of (N, C, H, W) to obtain five groups of abstract feature maps with different resolutions and numbers of channels.
[0085] S3, A multi-level feature aggregation module is used to enhance the expression ability of the features extracted by the encoder, as shown in the figure. The specific work process is as follows:
[0086] (3.1), The abstract feature maps extracted in (2.2) are convolutionally processed to change the number of channels, and the number of channels of all features is kept consistent;
[0087] (3.2) The features obtained in (3.1) are complementary to each other. Specifically, between adjacent layers, the low-resolution features are upsampled and then element-wise fused with the high-resolution features, and the high-resolution features are pooled and then element-wise fused with the low-resolution features, so that the features of different resolutions complement each other.
[0088] (3.3) Aggregate the complementary features obtained in (3.2). Specifically, for each level of features, if there are higher-level features, use the higher-level features to be upsampled and then element-wise added to it; if there are lower-level features, pool the lower-level features and then element-wise add them to it. For each level of features in (2.2), a corresponding aggregated feature is obtained, which can enhance the expressive ability of the fixed-resolution features.
[0089] S4. Use the multi-scale information extraction module to extract multi-scale information from the fixed features, enhancing the network's detection ability for targets of different scales. The specific workflow is as follows:
[0090] (4.1) Feed the multi-level aggregated features extracted in (3.3) into three branches in a top-down and parallel manner. The three branches are sampled with convolutions of different dilation rates of 2, 4, and 8 respectively, and use a residual operation to element-wise add the original features and these sampled features.
[0091] (4.2) Upsample the features containing multi-scale information extracted in (4.1), element-wise add them to the corresponding aggregated features in (3.3), and then obtain the features after hierarchical fusion through a 3×3 convolution.
[0092] S5. Use the boundary information extraction module to extract significant boundary information from the hierarchically fused features, further supplementing the significant target information and enhancing the detection effect of the network. The specific workflow is as follows:
[0093] (5.1) Convert the features obtained in (4.2) to obtain a feature map containing boundary information.
[0094] (5.2) Reduce the channel dimension of the features obtained in (5.1) to 1 through a 1×1 convolution, then upsample to obtain the boundary output, and fuse multiple boundary outputs to obtain the fused boundary output.
[0095] (5.3) Upsample the boundary features in (5.1) and then concatenate them along the channel dimension, and further fuse and change the number of channels through the boundary aggregation module to obtain the boundary features.
[0096] S6. Fuse the boundary information and the target information to obtain the final saliency prediction. The specific workflow is as follows:
[0097] (6.1) Re - extract multi - scale information for the last feature in (4.2). Since it is the bottom layer, there is no need to add elements with the horizontal features.
[0098] (6.2) Concatenate the boundary features in (5.3) and the features obtained in (6.1) along the channels, then further perform convolution fusion. After channel transformation and up - sampling, the final saliency prediction map is obtained.
[0099] S7. Use the obtained boundary detection results, object detection results, and the corresponding training set images to train the object detection model: During the training process, binary cross - entropy loss, consistency enhancement loss, and balanced binary cross - entropy loss are used to promote the backpropagation of gradients, strengthen the model convergence, and further improve the training effect.
[0100] S8. For the trained saliency object detection model, use the test image as the input to obtain the results of saliency object detection, as Figure 6 shown. The specific work process is as follows:
[0101] (8.1) For the saliency object detection model described in step S7, use the test set described in step (1.1) as the input to obtain the results of saliency detection.
[0102] (8.2) Compare the detection results of the saliency object detection model described in step (8.1) with the actual saliency object ground - truth map. The saliency object detection model described in step (8.1) has achieved excellent detection effects and performs very well on the Fmax, MAE, and Em metrics on four datasets, as shown in the figure.
Claims
1. A saliency object detection method based on boundary enhancement, characterized in that, it includes the following steps: S1. Extract abstract feature maps of different resolutions from the training set images, perform multi-level fusion on the abstract feature maps to obtain a multi-level fusion feature map, and process the multi-level fusion feature map to obtain a feature map containing multi-scale information; S2. After transforming the information of the obtained feature map containing multi-scale information, splice and fuse it to obtain a feature containing boundary information. At the same time, use the feature after each level of transformation to obtain a boundary detection result, and then further splice and fuse to obtain a fused boundary detection result; Specifically, taking four progressively fused features h 4 , h 3 , h 2 , h 1 as the input, first pass through an information conversion module to extract features eh 4 , eh 3 , eh 2 , eh 1 containing boundary information. The information conversion module consists of two convolutional groups of 1×1, 3×3, and 1×1, and the convolutional group contains a residual connection operation; then, after reducing the number of channels through 1×1 convolution and upsampling the features containing boundary information, the significant boundary prediction results e 4 , e 3 , e 2 , e 1 are obtained; the extracted multi-level boundary features eh 4 , eh 3 , eh 2 , eh 1 are upsampled and concatenated along the channels, and then input into the boundary feature aggregation module to obtain the final feature EF containing boundary information. The boundary aggregation module consists of four 3×3 convolutions containing residual operations; the final boundary feature is shown as follows: EF = EdgeInfo(Concat(Up(eh 1 ), Up(eh 1 ), Up(eh 3 ), Up(eh 4 ))) In the formula: EdgeInfo(·) represents the boundary feature aggregation module; EF represents the saliency boundary feature that aggregates multi-level information and can be used for the next step of fusing with the saliency object feature; S3. Extract multi-scale information from the feature map containing multi-scale information and splice it with the feature containing boundary information to obtain a saliency object detection result; S4. Use the saliency object detection result, the boundary detection result, and the corresponding training set to train the saliency object detection model until the convergence condition is met, and use the trained saliency object detection model for object detection; Use the backpropagation strategy to optimize the parameters of the network, and use the loss function to assist in training. The loss function used in the training process is divided into two categories according to different tasks: saliency boundary detection loss and saliency object detection loss. The formula of the total loss function in the training process is as follows: Loss=L sod +λ 1 L edge where: λ 1 is a hyperparameter used to balance the losses of the two tasks, and its value is set to 10 in the experiment; Use the balanced binary entropy loss to supervise the saliency boundary learning process. The formula of the balanced binary entropy loss is expressed as follows: In the formula: β is the ratio of the number of non-boundary pixels to the total number of all pixels; The saliency object detection loss function is composed of two loss functions with different focuses, including the binary cross-entropy loss for a single pixel point and the consistency enhancement loss for the foreground region. The total loss is expressed as follows: L sod = L bce + L cel The binary cross-entropy loss is the most used loss function in the saliency object detection task. This loss is a pixel-level loss and converges on all pixels. The formula is shown as follows: In the formula: P represents the saliency object prediction map; p represents a pixel point in P; G represents the ground truth map; g represents a pixel point in G; log(·) is a pixel-level logarithmic operation; The consistency enhancement loss is an image-level loss. The formula of the consistency enhancement loss function is expressed as follows: In the formula: P represents the saliency object prediction map; p represents a pixel point in the prediction map.
2. A saliency object detection method based on boundary enhancement according to claim 1, characterized in that, ResNet-50 trained on ImageNet is used as the backbone of the network to extract abstract feature maps of different resolutions from the training set images, and the last pooling layer and fully connected layer are removed to obtain five multi-level fusion feature maps of different sizes.
3. A saliency object detection method based on boundary enhancement according to claim 2, characterized in that, The multi-level fusion feature maps of different scales obtained are fused from top to bottom in a reverse step-by-step manner. Before each fusion, multi-scale information extraction is first performed on the current multi-level fusion feature map, and then upsampling is carried out and fused with the multi-level fusion feature map of the previous layer to obtain a feature map containing multi-scale information.
4. A saliency object detection method based on boundary enhancement according to claim 1, wherein, in step S1, the multi-level fusion of the abstract feature map is to keep the sizes of the feature layer and its adjacent feature layer consistent through upsampling or pooling, add the elements to each other for information supplementation, and then aggregate the features with supplemented information to obtain five multi-level aggregation feature maps of different sizes.
5. A saliency object detection system with boundary enhancement based on the method according to claim 1, wherein, it includes a convolutional feature extraction network module, a multi-level feature aggregation module, a boundary information extraction module, a multi-scale information extraction module and a detection module; the convolutional feature extraction network is used to extract abstract feature maps of different resolutions from the training set images, the multi-level feature aggregation module is used to perform multi-level fusion on the abstract feature maps to obtain multi-level fusion feature maps, and the multi-scale information extraction module is used to perform multi-scale extraction on the multi-level fusion feature maps to obtain feature maps containing multi-scale information; the boundary information extraction module is used to perform information transformation on the obtained feature maps containing multi-scale information and then perform splicing and fusion to obtain features containing boundary information, obtain a boundary detection result using the features after each level of transformation, and then further perform splicing and fusion to obtain a fused boundary detection result, and perform multi-scale information extraction on the feature maps containing multi-scale information and then splice them with the feature information containing boundary information to obtain a saliency object detection result; the detection module is used to train the saliency object detection model according to the detected object detection results and the corresponding training set images until the loss value meets the convergence condition, and use the trained saliency object detection model for object detection.
Citation Information
Patent Citations
Saliency detection method based on boundary enhancement
CN111310767A
Image feature detection method based on Gram matrix and F norm
CN112836708A