Weak supervision salient target detection method based on enhanced multilevel feature aggregation
By combining the improved context anchoring attention mechanism and the method of removing redundant modules, the problem of dynamic attention to context information during feature aggregation in weakly supervised significant object detection is solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202411901717.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-06-13
AI Technical Summary
In weakly supervised significant object detection, how to effectively utilize limited annotation information to improve the accuracy and robustness of the model, especially in the process of feature aggregation, dynamically pay attention to key context information in the image.
A new aggregation module is proposed that combines the improved context anchoring attention (CAA) mechanism, and a removal of redundant module is introduced after the encoder to enhance the fusion capability of low-level, high-level and global context features, dynamically pay attention to key context information and remove irrelevant background interference.
Through the improved aggregation module and the redundant module, the accuracy and robustness of target detection are significantly improved, and the complementarity of each layer's characteristics can be better utilized to reduce the influence of noise and outliers.
Smart Images

Figure BDA0005203365170000031 
Figure BDA0005203365170000091 
Figure FDA0005203365160000021
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and relates to a weakly supervised salient object detection method, in particular to a method for enhancing multi-level feature aggregation based on a convolutional neural network. Background Art
[0002] With the rapid development of deep learning technology, weakly supervised learning has been increasingly widely applied in the field of computer vision. Especially in the task of salient object detection, it has gradually become an important research direction. Traditional salient object detection methods usually rely on a large amount of pixel-level labeled data, and the acquisition cost of these data is high and time-consuming. To solve this problem, weakly supervised salient object detection reduces the dependence on finely labeled data by introducing less labeled information, such as image-level labels, bounding boxes, or even completely unsupervised methods. The rise of this research direction not only reduces the burden of data annotation but also enables salient object detection to be promoted in more practical applications, especially in resource-constrained situations. Meanwhile, with the maturity of deep neural network and self-supervised learning technologies, more and more research attempts to further improve the performance and effect of weakly supervised salient object detection by using context information, global features, and multi-modal data of images.
[0003] In the process of weakly supervised salient object detection, how to effectively utilize limited labeled information to improve the accuracy and robustness of the model is a key issue. If multi-level feature aggregation is enhanced, it can effectively fuse low-level features, high-level features, and global context features, thus providing a richer and more comprehensive representation for the salient object detection task. Low-level features can capture detailed information such as edges and textures; high-level features focus on the semantics and structure of the object; while global context features help the model understand the position and background information of the object in the scene. Through this multi-level feature aggregation, the model can better utilize the complementarity of each layer of features, thereby improving the detection accuracy and robustness.
[0004] The present invention proposes a new aggregation module combining an improved context anchor attention (CAA) mechanism, which combines the improved CAA mechanism with the FIA module. This new aggregation module can dynamically focus on key context information in the image during the feature aggregation process through the improved context anchor attention mechanism, thereby further improving the accuracy of object detection and suppressing the problem of irrelevant background interference. At the same time, a redundant removal module is introduced to further optimize the feature representation in detection. Finally, the effectiveness of the method is verified on the scribble saliency dataset S-DUTS with the improved structure. Summary of the Invention
[0005] The present invention proposes a weakly supervised salient object detection method based on enhanced multi-level feature aggregation. It uses SCWSSOD, which is currently highly effective and accurate, as the network basic framework. By improving the original aggregation module, the fusion ability of low-level, high-level, and global context features is enhanced, and the model's ability to identify targets is improved. At the same time, a redundancy removal module is introduced to optimize the feature representation after the encoder, remove irrelevant information, and improve the detection accuracy and robustness.
[0006] The technical solution adopted by the present invention includes the following steps:
[0007] Step 1, data preparation. Collect the relevant datasets for salient object detection and divide them into two parts: a training set and a test set. The training set is used for model training, while the test set is used to evaluate the model's performance. To improve the model's learning ability under weakly supervised conditions, all images in the training set contain manually marked scribble labels. At the same time, the images in the test set remain in their original state, mainly for subsequent evaluation and verification of the model's performance.
[0008] Step 2, construct the network structure. The present invention replaces and improves the original aggregation module in the SCWSSOD network framework and introduces a redundancy removal module after the encoder. The overall network structure is mainly divided into two parts, as Figure 1 shown.
[0009] The first part is the encoder part. The input image is first preliminarily extracted with features and downsampled by the 7x7 convolutional layer of the ResNet-50 encoder with a stride of 2, and then processed by the 3x3 max pooling layer. Subsequently, multiple residual block groups operate in sequence. The first group of residual blocks follows a 1x1, 3x3, 1x1 convolutional structure and uses skip connections to solve the gradient problem and fuse features. The subsequent residual block groups change in terms of the number of channels, quantity, etc. according to the network design, and finally output a multi-level encoded feature map. Then, a redundancy removal module (Rr module) is introduced after the encoder, as Figure 4 shown. First, multi-scale feature information is obtained through convolutional kernels of different sizes, and the multi-scale input features are added element-wise to obtain the fused feature. The specific expression is shown in Equation (1):
[0010] f' = conv1(f) + conv2(f) (1)
[0011] where, in the redundancy removal module, f represents the input feature, and f' represents the fused feature. conv1 and conv2 are convolutional operations of different sizes, used to extract multi-scale information. Then, global average pooling (GAP) is performed on each channel to generate a global description of the image. Finally, dimensionality reduction is performed through a fully connected layer to improve the calculation efficiency and accurately select the most representative features. The specific description is shown in Equation (2):
[0012] F = relu(norm(Wf')) (2)
[0013] Among them, f' represents the fused feature of the input, F represents the dimensionality-reduced feature, W represents the fully connected layer, relu is the activation function, and norm is used to normalize the elements in the image to [0, 1].
[0014] To obtain the feature information required for object detection, the vector F is mapped back to dimension C through two fully connected layers respectively. Then, a softmax operation is applied to these two vectors in the channel dimension to obtain their respective weight vectors Q j . Then, these weights are multiplied element-wise with the two outputs of the multi-scale features, and finally the output f of the redundant module removal is obtained out , specifically described as shown in equations (3) and (4):
[0015]
[0016]
[0017] Among them, w j represents the j-th fully connected operation, Q i represents the j-th weight vector, represents channel multiplication, conv i (f) represents the feature of the i-th multi-scale
[0018] The second part is the decoder part. The Context Anchor Attention (CAA) mechanism is integrated into the Feature Interleaving Aggregation (FIA) module to form a Context-Aware Enhanced Aggregation module, which is added to the middle layer position of the decoder. Due to the original CAA mechanism, it can only extract the average information of the feature map, which is not fine enough. Therefore, the present invention improves the CAA mechanism, as Figure 3 shown
[0019] First, the input x (the input features of the three branches in the original FIA module) adds max_pool(x) of max pooling on the original basis and sums it with avg_pool(x) of average pooling to obtain the intermediate feature map x', specifically described as shown in equation (5):
[0020] x' = max_pool(x) + avg_pool(x) (5)
[0021] Here, x is the input feature map, and two different pooled feature maps are obtained after max pooling and average pooling operations. After adding them, this operation can fuse the information of the two pooling methods and obtain more context information
[0022] Next, the summed feature map x' is processed using two layers of ordinary convolutions. In this step, depthwise separable convolutions are removed, and only conventional convolution operations are retained. First, the convolution layer conv 1 is used to process x', and then the output result is applied to the convolution layer conv 2 again, as specifically described in Equation (6):
[0023] x” = conv 2 (conv 1 (x')) (6)
[0024] Here, conv 1 and conv 2 are two ordinary convolution layers respectively, including operations such as convolution kernels, strides, and padding. The convolution operation can extract local feature information. x” is the feature map after two layers of convolution, containing more high-order features.
[0025] Finally, after convolution processing, we will use the Sigmoid activation function to perform a non-linear transformation on x” to obtain the attention factor f attn . This factor is used to adjust the importance of features and help the network focus on key regions, as specifically described in Equation (7):
[0026] f attn = σ(x”) (7)
[0027] where σ(·) represents the Sigmoid activation function, and its formula is The output f attn of the Sigmoid activation function is a matrix between 0 and 1, representing the attention intensity at each position.
[0028] Throughout the process, we introduce a residual connection and multiply the input x by the obtained attention factor f attn to obtain the final output f a , as specifically described in Equation (8):
[0029] f a = x · f attn (8)
[0030] Here, the role of the residual connection is to combine the original input x with the attention-adjusted feature f attn to generate the final output. This operation can ensure that the network better retains the original information while dynamically adjusting the weight of each feature through the attention mechanism.
[0031] After the input features of the three branches enter the improved CAA mechanism, the mechanism can dynamically learn the importance of each feature, perform context-aware weighting on the features of different branches from a global perspective, and avoid the fixed weight problem of traditional convolution and pooling methods. Subsequently, convolution operations, batch normalization operations, and activation operations are carried out. These operations further extract and transform the features, and better interweave the features of different branches.
[0032] Step 3, the context-aware enhanced aggregation module (CAEF module) receives three parts of input to fully integrate the features of the three levels, as Figure 2 shown. Among them, the three parts of input are the final output f a obtained after the improved CAA mechanism in Step 2 processes, that is, the low-level features after weighted processing, the high-level features and the global context features
[0033] Specifically, to ensure the consistency of the multiplication operation, first send the underlying feature map into a 1×1 convolutional layer conv 1 to compress it so that it has the same number of channels as the high-level feature . Then perform a 3×3 convolutional layer on the high-level feature , and after upsampling, obtain the semantic mask Furthermore, we multiply the mask by the compressed underlying feature
[0034] In addition, considering that the high-level features will discard some detailed information related to the significant objects, we apply the above fusion strategy in a symmetric manner. Different from the symmetric path mentioned above, a detail mask is generated from the low-level features through a 3×3 convolutional layer Then multiply the mask by the high-level features after upsampling. The symmetric path should add fine-grained detailed information to the predicted saliency map, specifically described as shown in Equations (9) to (12):
[0035]
[0036]
[0037]
[0038]
[0039] In the formula Denote the compressed underlying features, ⊙ denote element-wise multiplication, δ denote the ReLU activation function, upsample is the upsampling operation by bilinear interpolation, and t is the stage index.
[0040] In addition, to model the relationships between different parts of the salient object and mitigate the dilution process of the high-level features, global context features are introduced at each stage. We use the global context features to generate a context mask Then, the mask is multiplied to the compressed underlying features, as specifically described in Equations (13) and (14):
[0041]
[0042]
[0043] Finally, these three levels of features are concatenated and then passed through a 3×3 convolutional layer to obtain the final fused features, as specifically described in Equation (15):
[0044]
[0045] Each of the above convolutional layers is equipped with a batch normalization layer and a ReLU activation function.
[0046] Step 4: Feed the training set in Step 1 into the network of the present invention that has improved the aggregation module and introduced the redundancy removal module for training.
[0047] Step 5: Use the test set images in Step 1 to feed into the network of the present invention trained in Step 4 to obtain the finally predicted saliency map.
[0048] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0049] 1. Traditional feature aggregation methods lack long-range dependence modeling, and the learning of feature importance is relatively single and lacks sufficient adaptability, making it difficult to handle noise and outliers. While the present invention integrates the improved CAA mechanism into the FIA module, designs a context-aware enhanced aggregation module, effectively strengthens the limitations during feature aggregation, and avoids the fixed weight problem of traditional convolution and pooling methods.
[0050] 2. In salient object detection, the top-level features in an image usually carry rich semantic information but may also contain a large amount of redundant or irrelevant background information, which often has a negative impact on the object detection task. In view of this situation, the present invention introduces a redundancy removal module, enabling precise object detection and background suppression to be achieved under the framework of weakly supervised learning, effectively improving the performance of the salient object detection model. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is the overall network structure diagram;
[0052] Figure 2 is the block diagram of the context-aware enhanced aggregation module;
[0053] Figure 3 Flowchart of the improved context-anchored attention mechanism;
[0054] Figure 4 is the block diagram of the redundancy removal module;
[0055] Figure 5a is RGB image 1;
[0056] Figure 5b is the salient prediction map output by SCWSSOD;
[0057] Figure 5c is the salient prediction map of the present invention;
[0058] Figure 6a is RGB image 2;
[0059] Figure 6b is the salient prediction map output by SCWSSOD;
[0060] Figure 6c is the salient prediction map of the present invention;
[0061] Figure 7a is RGB image 3;
[0062] Figure 7b is the salient prediction map output by SCWSSOD;
[0063] Figure 7c is the salient prediction map of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0064] The present invention will be further described in detail below in conjunction with embodiments, but the embodiments of the present invention are not limited thereto.
[0065] This embodiment is implemented under the PyTorch deep learning framework. The computer configuration is as follows: an Intel(R) Xeon(R) Platinum 8352V CPU @ 2.10GHz processor, 48GB of memory, RTX 3080x2 (20GB) graphics cards, and a Linux operating system. As Figures 1 to 4 shown, this embodiment discloses a method for weakly supervised salient object detection based on enhanced multi-level feature aggregation, and the specific situation is as follows:
[0066] (1) Data preparation
[0067] The training set uses the scribble saliency dataset S-DUTS, with a total of 15,572 images for model training. The test set uses 1,000 images from ECSSD, 5,168 images from DUT-OMRON, 850 images from PASCAL-S, 4,447 images from HKU-IS, and 5,019 images from DUTS for evaluating the model performance. All images in the training set contain manually labeled scribble tags. At the same time, the images in the test set remain in their original state, mainly for subsequent evaluation and verification of the model performance.
[0068] (2) Network structure design
[0069] The overall network structure refers to Figure 1 , and the overall network structure includes two parts: an encoder and a decoder. Among them, in this invention, the original aggregation module is improved in the middle layer of the decoder, and a redundancy removal module is introduced after the encoder.
[0070] The context-aware enhanced aggregation module refers to Figure 2 , and the three-branch input features in the module are first processed by an improved CAA mechanism and then aggregated. First, for the low-level feature maps (t = 1, 2, 3), they are first sent into a 1×1 convolutional layer for compression so that their number of channels is the same as that of the high-level features. Then, the high-level features are processed by a 3×3 convolutional layer and upsampled to obtain a semantic mask, and then this mask is multiplied by the compressed low-level features. At the same time, in a mirroring manner, the low-level features are passed through a 3×3 convolutional layer to generate a detail mask, and this mask is multiplied by the upsampled high-level features to add fine-grained information to the predicted saliency map. In addition, global context features are introduced at each stage, and a context mask is generated using them, and this mask is multiplied by the compressed low-level features. Finally, the above three-level features are concatenated so that they converge together, and then processed by a 3×3 convolutional layer equipped with a batch normalization layer and a ReLU activation function to obtain the final fused features.
[0071] The improved CAA mechanism refers to Figure 3, first, an average pooling layer avg_pool is defined with a pooling kernel size of 7, a stride of 1, and a padding of 3. Additionally, a max pooling layer max_pool is introduced with a pooling kernel size of 7, a stride of 1, and a padding of 3. Average pooling and max pooling operations are performed on the input x, and the results are added together to obtain an intermediate feature. Then, two convolutional modules conv1 and conv2 are defined. conv1 changes the number of input channels from channels to channels*2 with a convolutional kernel size of 1, a stride of 1, and a padding of 0, and applies the specified normalization and activation functions; conv2 changes the number of channels back from channels*2 to channels, also using the specified normalization and activation functions. The intermediate feature is sequentially passed through conv1 and conv2 for convolutional operations. Next, a Sigmoid activation function is defined, and the Sigmoid activation function is used to perform a non-linear transformation on the convolutional result to obtain the attention factor f attn . Finally, the residual connection is achieved by multiplying the input x with the attention factor f attn to obtain the final output f a .
[0072] Reference for removing redundant modules Figure 4 , first, conv3x3 with a dilation rate of 1 and conv3x3 with a dilation rate of 2 are used to perform convolutions of different sizes on the input feature f to extract multi-scale information and fuse it into the feature f'. Then, global average pooling (GAP) is performed on each channel of f' to generate a global description. After size reduction through a fully connected layer, a vector F is obtained. Next, F is mapped back to dimension C through two fully connected layers, and then softmax is applied in the channel dimension to obtain a weight vector. Finally, the weight vector is multiplied element-wise with the two outputs of the multi-scale feature to obtain the module output.
[0073] (3) Network training settings
[0074] The Python language is used, and the ResNet-50 deep learning framework is used to build and train a deep neural network model for the network as Figure 1 shown.
[0075] Basic configuration: The path of the dataset, the storage path of the training data is. / data / DUTS. The save path of the trained model is set to escwssod. The batch size batch = 16. The initial learning rate lr = 1e-3. The momentum of the optimizer is 0.9. The weight decay is 5e-4. The total number of training epochs epoch = 40.
[0076] Optimizer: Use the SGD optimizer with Nesterov momentum and adopt different learning rate and momentum adjustment strategies. Among them, a custom learning rate scheduling function get_triangle_lr is used to adjust the learning rate. get_triangle_lr is a triangular learning rate scheduler. As the training progresses, the learning rate gradually increases within a range and then gradually decreases. In each training step, the current learning rate is calculated through the get_triangle_lr function and the optimizer is updated.
[0077] Every 10 batches, record the learning rate and loss information and print the log. When the epoch is greater than 28, save the model parameters every 1 epoch or at the last epoch.
[0078] (4) Selection of evaluation metrics
[0079] The present invention uses the Mean Absolute Error (MAE) and the mean Enhanced-Alignment Measure (mEm) as evaluation metrics. The lower the MAE, the closer the predicted significant region of the model is to the actual one, and the higher the accuracy. The higher the Mean E-measure value, the better the integrity of the region where the significant object is detected and the more accurate the positioning, indicating that the effect of significant object detection is better.
[0080] The comparison of the indicators of SCWSSOD and the present invention on 5 test sets is shown in Table 1. It can be seen that the Mean Absolute Error (MAE) and the mean E-measure of the present invention have been further improved.
[0081] Table 1 Comparison of indicators
[0082]
[0083] Such as Figure 5a Figures 5a, 6a and 7a are RGB images, and Figures 5b, 6b and 7b are the significant prediction maps output by SCWSSOD. Figure 5c Figures 5c, 6c and 7c are the significant prediction maps of the enhanced multi-level feature aggregation of the present invention. It can be seen that the method of the present invention is well reflected in places such as the redundant parts of the human portrait, the discontinuous position in the middle of the train carriage, and the extra background information of the goose. The MAE and mEm values of the two methods are shown in Table 1. It can be seen from the table that the detection effect of the present invention is significantly better than that of the SCWSSOD method.
[0084] The above are only the preferred embodiments of the present invention and are not used to limit the use of the present invention method. Any improvement made within the spirit and principle of the present invention shall be included in the protection scope of the present invention method.
Claims
1. A weakly supervised salient object detection method based on enhanced multi-level feature aggregation, characterized by The following steps are involved: Step 1: Data preparation: collect relevant datasets for salient object detection and divide them into two parts: training set and test set. The training set is used for model training, while the test set is used to evaluate the performance of the model. In order to improve the learning ability of the model under weak supervision, all images in the training set contain manually marked graffiti tags. At the same time, the images in the test set remain in their original state, mainly used for subsequent model performance evaluation and verification. Step 2, constructing a network structure. The present invention replaces and improves the original aggregation module in the SCCSSOD network framework, and introduces a redundancy removal module after the encoder. The overall network structure is mainly divided into two parts; The first part is the encoder part; the input image is first processed by the 7x7 convolution layer of the ResNet-50 encoder with a step size of 2, and the features are initially extracted and downsampled, and then processed by the 3x3 maximum pooling layer. Then multiple residual block groups are operated in sequence. The first group of residual blocks follows the 1x1, 3x3, and 1x1 convolution structures and uses jump connections to solve the gradient problem and fusion features. The subsequent residual block groups change in the number of channels and quantity according to the network design, and finally output a multi-level encoding feature map; then the redundancy removal module (Rr module) is introduced after the encoder; first, multi-scale feature information is obtained through convolution kernels of different sizes, and the multi-scale input features are added element by element to obtain the fusion features. The specific expression is shown in formula (1): f'=conv1(f)+conv2(f) (1) In the redundancy removal module, f represents the input feature and f' represents the fused feature; conv1 and conv2 are convolution operations of different sizes, which are used to extract multi-scale information; then, global average pooling (GAP) is performed on each channel to generate a global description of the image; finally, the size is reduced through the fully connected layer to improve the computational efficiency and accurately select the most representative features. The specific description is shown in formula (2): F = relu(norm(Wf')) (2) Among them, f' represents the input fusion feature, F represents the dimension reduction feature, W represents full connection, relu is the activation function, and norm is used to normalize the elements in the image to [0,1]; In order to obtain the feature information required for target detection, the vector F is mapped back to dimension C through two fully connected layers; then, the softmax operation is applied to the two vectors in the channel dimension to obtain their respective weight vectors Q j ; Then, these weights are multiplied element-wise with the two outputs of the multi-scale features to finally obtain the output f of the redundant removal module out , specifically described as shown in formula (3) and formula (4): Among them, w j represents the jth fully connected operation, Q i represents the jth weight vector, Represents channel multiplication, conv i (f) represents the i-th multi-scale feature; The second part is the decoder part; the contextual anchor attention (CAA) mechanism is integrated into the feature interleaving aggregation (FIA) module to form a context-aware enhanced aggregation module, which is added to the middle layer of the decoder; since the original CAA mechanism can only extract the average information of the feature map, it is not refined enough; therefore, the present invention improves the CAA mechanism; First, the input x (the input features of the three branches in the original FIA module) is added with the maximum pooling max_pool(x) on the original basis and summed with the average pooling avg_pool(x) to obtain the intermediate feature map x'. The specific description is shown in formula (5): x'=max_pool(x)+avg_pool(x) (5) Here, x is the input feature map, and two different pooling feature maps are obtained after the maximum pooling and average pooling operations. After adding them together, the operation can fuse the information of the two pooling methods and obtain more context information; Next, two layers of ordinary convolution are used to process the added feature map x'. In this step, the depth-wise separable convolution is removed, and only the regular convolution operation is retained. First, the convolution layer conv1 is used to process x', and then the convolution layer conv2 is applied to its output again. The specific description is shown in formula (6): x'=conv2(conv1(x')) (6) Here, conv1 and conv2 are two common convolution layers, including convolution kernel, stride, padding and other operations. The convolution operation can extract local feature information; x" is the feature map after two layers of convolution processing, which contains more high-order features; Finally, after the convolution process, we will use the Sigmoid activation function to perform a nonlinear transformation on x' to obtain the attention factor f attn ; This factor is used to adjust the importance of features and help the network focus on key areas. The specific description is shown in formula (7): f attn =σ(x”) (7) Where σ(·) represents the Sigmoid activation function, and its formula is The output f of the Sigmoid activation function attn is a matrix between 0 and 1, indicating the intensity of attention to each position; In the whole process, we introduce residual connection to connect the input x and the obtained attention factor f attn Multiply them to get the final output f a , the specific description is shown in formula (8): f a =x·f attn (8) Here, the role of the residual connection is to combine the original input x with the attention-adjusted feature f attn Combined to generate the final output; this operation ensures that the network better retains the original information, while dynamically adjusting the weight of each feature through the attention mechanism; After the input features of the three branches enter the improved CAA mechanism, the mechanism can dynamically learn the importance of each feature and perform context-aware weighted processing on the features of different branches from a global perspective, avoiding the fixed weight problem of traditional convolution and pooling methods; followed by a series of convolution operations, batch normalization operations, and activation operations; these operations further extract and transform the features to better interweave the features of different branches; Step 3: The context-aware enhanced aggregation module (CAEF module) receives three inputs to fully integrate the features at three levels; the three inputs are the final output f obtained after the improved CAA mechanism in step 2. a , which are the low-level features after weighted processing Advanced Features and global context features Specifically, in order to ensure the consistency of multiplication operations, the underlying feature map is first Send it to the 1×1 convolution layer conv1 to compress it so that it has the same high-level features The same number of channels; then for high-level features Perform a 3×3 convolution layer and upsample to obtain a semantic mask Next, we will mask Multiply by the compressed underlying features In addition, considering that high-level features will discard some detail information related to salient objects, we apply the above fusion strategy in a symmetrical way; different from the symmetrical path mentioned above, a detail mask is generated from low-level features through a 3×3 convolution layer Then the mask Multiplying high-level features After upsampling, the symmetric path should add fine-grained detailed information to the predicted saliency map, as shown in Equations (9) to (12): In the formula represents the compressed underlying features, ⊙ represents element-by-element multiplication, δ represents the ReLU activation function, upsample represents the upsampling operation through bilinear interpolation, and t represents the stage index; In addition, in order to model the relationship between different parts of salient objects and alleviate the dilution process of high-level features, global context features are introduced at each stage. We use global context features To generate the context mask Then, the mask The underlying features multiplied by the compression are specifically described as shown in equations (13) and (14): Finally, the three levels of features are connected in series and then passed through a 3×3 convolutional layer to obtain the final fusion feature, which is specifically described as shown in formula (15): Each of the above convolutional layers is equipped with a batch normalization layer and ReLU activation function; Step 4, transferring the training set of step 1 into the network of the present invention with an improved aggregation module and a redundant removal module for training; Step 5, using the test set image of step 1 to be sent to the network of the present invention trained in step 4 to obtain the final predicted saliency map.