Salient target detection method based on edge perception and attention mechanism
By introducing edge perception and dual channel attention mechanisms in significance object detection, the problem that existing methods are difficult to detect significant targets in complex backgrounds is solved, achieving higher detection accuracy and robustness.
Patent Information
- Application Number
- CN202510390149.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-24
AI Technical Summary
Existing significance object detection methods are difficult to capture fine saliency areas when processing low contrast or blurred images, especially in complex backgrounds and multi-objective environments.
Using a significance object detection method based on edge perception and dual-channel attention mechanism, image boundary information is extracted through edge perception module, and feature weighting is performed in the spatial domain and channel domain respectively, strengthening significant area expression and suppressing background interference.
It improves the accuracy and robustness of object detection, can better adapt to the significant object detection tasks in complex scenarios, and enhances the adaptability and detection accuracy of the model.
Smart Images

Figure CN120198649A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of salient object detection, and in particular, to a salient object detection method based on edge perception and attention mechanism. Background Art
[0002] In the task of object detection, salient object detection aims to accurately separate the eye-catching objects from the complex background. With the rapid development of deep learning technology, most of the existing salient object detection methods rely on convolutional neural networks to extract image features. However, traditional methods often have difficulty in capturing fine salient regions when dealing with low-contrast or blurred images, especially in complex background and multi-object environments.
[0003] In practical applications, although edge perception and attention mechanism have potential, due to the low contrast between objects and the background in the image, or the presence of more interference information in the background, the existing salient object detection methods based on edge and attention mechanism still face some challenges. Especially in complex scenes, how to effectively capture the edge information of the object and combine the global context information to improve the detection accuracy remains a difficult problem to be solved.
[0004] Inspired by the human visual perception mechanism, researchers have begun to explore salient object detection methods based on edge perception and attention mechanism, aiming to enhance the adaptability of the model to complex scenes. As the key structural information in the image, edges not only contain the boundary and shape features of the object, but also provide rich visual cues to help the model accurately distinguish the object from the background. At the same time, the introduction of the attention mechanism enables the model to adaptively focus on the salient regions, reducing background interference while improving the detection accuracy and robustness. Therefore, the detection method that fuses edge perception and attention mechanism has received extensive attention, providing a new research idea for salient object detection.
[0005] The method of edge perception and dual-channel attention mechanism enables the model to more accurately detect the target region by enhancing edge feature expression and adaptive feature screening. The edge perception module is responsible for extracting the boundary information in the image, thus accurately locating the target contour and effectively dealing with the problem of target loss in low-contrast regions. The dual-channel attention mechanism weights the features in the spatial domain and the channel domain respectively to strengthen the expression of the salient region while suppressing background interference, thereby achieving more accurate and robust object detection. It can better adapt to the accurate detection of complex scene images and meet the actual needs of engineering projects.
[0006] The method in the literature "Edge-Aware Multiscale Feature Integration Network for Salient Object Detection in Optical Remote Sensing Images" enhances the salient object based on edge information and realizes the salient object detection in optical remote sensing images by using an edge-aware multiscale feature fusion network. However, this method mainly targets remote sensing data and still has certain limitations in the ability to finely depict the target boundary in complex scenes. The literature "Attentive Feedback Network for Boundary-Aware Salient Object Detection" proposes an attention feedback network that focuses on boundary information and uses a feedback mechanism to enhance the boundary perception ability of salient objects. By gradually optimizing the feature expression of the salient region and reducing background interference at the same time, the boundary of the salient object becomes clearer and more complete. However, the suppression effect of this method on background interference depends on the feedback adjustment of the network structure, and the detection accuracy of the salient object may be affected in the presence of large background noise. Summary of the Invention
[0007] The purpose of the present invention is to provide a salient object detection method based on edge perception and dual-channel attention mechanism. By enhancing edge feature expression and adaptive feature screening, feature weighting is performed in the spatial domain and channel domain respectively to strengthen the expression of the salient region and effectively suppress background interference, thereby improving the detection accuracy and robustness.
[0008] The technical solution of the present invention is as follows: A salient object detection method based on edge perception and attention mechanism, specifically including the following steps:
[0009] The input image is successively passed through each level of convolutional layers to extract features at each level.
[0010] The features at each level except the first convolutional layer are jointly input into the boundary perceptron for image boundary perception.
[0011] The features at each level except the first convolutional layer are respectively input into the corresponding level of the multiscale feature extraction module to obtain features at different scales at the corresponding level.
[0012] Except for the first level and the last level, the edge information of the input image, the features at different scales at the corresponding level, and the reverse channel are jointly input into the dual-channel attention mechanism for feature fusion and alignment to obtain the salient prediction map at the corresponding level; the features at different scales at the last level are obtained through multi-head attention to obtain the last salient prediction map.
[0013] Starting from the last layer, the saliency prediction maps are added layer by layer from bottom to top to obtain the saliency prediction map.
[0014] The convolutional layer adopts five layers.
[0015] The reverse channel is obtained by reverse attention propagation on the previous channel of the same level and the same scale feature.
[0016] The processing process of the boundary perception device is as follows:
[0017] Calculate the gradient information of the input image I by convolution to obtain edge features;
[0018] The specific idea is: for the input image, use the convolution kernel matrix of the Sobel operator to calculate the edge information of the input image; the horizontal direction G of the Sobel operator x and the vertical direction G y are calculated as follows:
[0019]
[0020] where * represents the convolution operation;
[0021] Calculate the gradient magnitude G as the edge response map:
[0022]
[0023] After obtaining the gradient response map G, splice it with the image I2 processed by each convolutional layer:
[0024] F i = Concat(I2, G) (4)
[0025] where Concat(·) represents the channel-level splicing operation, and F i represents the feature output of the i-th layer.
[0026] Furthermore, the multi-scale feature extraction module includes multiple convolutional layers;
[0027] First, perform a series of convolution operations on the image I2 processed by each convolutional layer to extract features of different scales layer by layer; specifically, assume that the resolution of the input image is H×W, and use multiple convolutional layers with different depths to extract features; ensure that the information of the shallow features can be effectively transmitted to the deep layer and enhance the correlation between the multi-scale features; introduce multiple convolutional kernels with different scales in each convolutional layer of the multi-scale feature extraction module to capture features of different scales:
[0028]
[0029] where σ(·) represents the softmax activation function, and Wi represents the weight matrix of the convolutional kernel in the i-th layer, F i-1 represents the feature output in the (i - 1)-th layer, b i is the bias term;
[0030] Perform a concatenation operation on features of different scales to obtain the processed multi-scale features:
[0031] F MDFE = Concat(F i (1×1) , F i (3×3) , F i (5×5) ) (6)
[0032] Concat represents concatenation in the channel dimension;
[0033] Adopt a weighted fusion mechanism to perform weighted fusion on features of different scales:
[0034]
[0035] where α i is the weight coefficient of features of different scales, n represents the number of multi-scale features, F MDFE ' represents the multi-scale features after weighted fusion.
[0036] The dual-channel attention mechanism corresponds to the forward channel and the reverse channel respectively. The forward channel mainly focuses on the global semantic information of the detected target, while the reverse channel combines edge perception information and emphasizes the modeling of local details;
[0037] In the forward channel, channel information modeling is achieved through the channel attention mechanism. First, perform global average pooling on the image feature map I3 composed of F i to extract the global information of each channel; calculate the channel weights through two fully connected networks, and apply the calculated channel weights to the feature map I3 to enhance the overall semantic information of the salient target; at the same time, use Softmax activation to obtain the channel attention score; the channel attention score acts on the image feature map I3, so that important channels are enhanced and unimportant channels are suppressed;
[0038] In the reverse channel, spatial information modeling is achieved through the spatial attention mechanism; first, perform max pooling and average pooling in the channel dimension on the feature map I3 to extract key spatial information; calculate the spatial weights through a 7×7 convolution and use the Softmax activation function to generate the spatial attention map; the spatial attention map is multiplied element-wise with the feature map I3 to strengthen the local edges and details of the salient target, and obtain a doubly enhanced feature map;
[0039] Di = DAM(F i ', I3) (8)
[0040] Among them, D i represents the doubly enhanced feature map of the i-th layer, DAM represents the dual-channel attention mechanism, and F i ' represents the spatial attention map of the i-th layer.
[0041] The multi-head attention mechanism uses Scaled Dot-Product Attention in the Transformer structure for cross-scale information interaction and enhancement:
[0042] M' = MHA(M1, M2, M3,) (9)
[0043] Among them, M' represents the feature map processed by the multi-head attention mechanism, MHA represents the multi-head self-attention mechanism, and M1, M2, and M3 respectively represent the features of three different scales for multi-scale feature processing.
[0044] The feature alignment and fusion adopt a variety of fusion methods combined with the feature alignment mechanism;
[0045] The feature alignment mechanism performs spatial alignment and channel alignment in sequence;
[0046] The doubly enhanced feature map D i is directly weighted with the feature map M' obtained after multi-head attention processing to obtain the feature map D i ' of different scales;
[0047] Spatial alignment: For the features D i ' of different scales, use bilinear interpolation to adjust them to the same spatial size H×W:
[0048] D i aligned = Resize(D i ', H, W) (10)
[0049] where Resize represents the upsampling operation;
[0050] Channel alignment: The number of channels of the features of different scales is different, and 1×1 convolution is used for channel transformation:
[0051] D i aligned = Conv 1×1 (D i aligned ) (11)
[0052] After completing the feature alignment, a variety of fusion strategies are adopted, including addition fusion, multiplication fusion, and concatenation fusion:
[0053] Additive fusion:
[0054]
[0055] Multiplicative fusion:
[0056]
[0057] Cascade fusion:
[0058] R cat = Concat(D1 aligned , D2 aligned ,..., D n aligned ) (14)
[0059] where the Concat operation is performed by concatenating in the channel dimension, followed by channel compression through a 1×1 convolution; finally, the fused feature R is weighted and combined by additive, multiplicative, and cascade methods:
[0060] R = ω1R add + ω2R mul + ω3R cat (15)
[0061] where ω1, ω2, and ω3 are learnable parameters.
[0062] The fused feature R still has a relatively high channel dimension and complex feature expressions. A 1×1 convolutional layer is used to reduce its dimension and extract significant information:
[0063] S' = Conv 1×1 (R) (16)
[0064] The Softmax activation function is used for processing to map the significant prediction to a probability distribution:
[0065]
[0066] where Si′ represents the unnormalized significant score at pixel point i, and Softmax normalizes the entire significant map.
[0067] Advantages of the present invention: By combining edge perception, attention mechanism, and text information fusion, not only the accuracy of significant object detection is improved, but also the robustness and adaptability of the method are enhanced. Especially in complex background and multi-object environments, accurate segmentation of significant objects can be better achieved.
[0068] In view of the problems that most of the previous saliency target detection methods process low-resolution images, high-resolution image samples are scarce, and it is difficult to accurately detect, etc., the present invention proposes a saliency target detection method based on edge perception and dual-channel attention mechanism. By introducing an edge perception module, the boundary information of the target is accurately extracted, and the structural features of the image are enhanced, thereby effectively improving the target detection accuracy under low contrast and complex backgrounds; through the dual-channel attention mechanism, the features are weighted in the spatial domain and the channel domain respectively, adaptively focusing on the salient regions and suppressing background interference, and improving the robustness and accuracy of the detection model. The present invention can effectively adapt to the saliency target detection task in complex scenes and meet the actual engineering requirements. Description of the Drawings
[0069] Figure 1 It is a flowchart of a saliency target detection method based on edge perception and attention mechanism.
[0070] Figure 2 It is a flowchart of the dual-channel attention mechanism.
[0071] Figure 3 It is a schematic diagram of feature fusion and alignment. Detailed Embodiment
[0072] Figure 1 It is the main flowchart of the technical solution of the present invention. As Figure 1 shown, the present invention proposes a saliency target detection method based on edge perception and attention mechanism, including the following steps:
[0073] The input image is successively passed through each convolutional layer to extract features at each level;
[0074] The features at each level except the first convolutional layer are jointly input into the boundary perceptron for image boundary perception;
[0075] The features at each level except the first convolutional layer are respectively input into the corresponding hierarchical multi-scale feature extraction module to obtain features at different scales of the corresponding level;
[0076] Except for the first level and the last level, the edge information of the input image, the features at different scales of the corresponding level, and the reverse channel are jointly input into the dual-channel attention mechanism for feature fusion and alignment to obtain the saliency prediction map of the corresponding level; the features at different scales of the last level are obtained through multi-head attention to obtain the final saliency prediction map;
[0077] The saliency prediction maps start from the last layer and are added layer by layer from bottom to top to obtain the saliency prediction map. The processing process of the boundary perceptron is as follows:
[0078] Calculate the gradient information of the input image I by convolution to obtain edge features;
[0079] The specific idea is as follows: For the input image, use the convolution kernel matrix of the Sobel operator to calculate the edge information of the input image; the horizontal direction G of the Sobel operator x and the vertical direction G y are calculated as follows:
[0080]
[0081] where * represents the convolution operation;
[0082] Calculate the gradient magnitude G as the edge response map:
[0083]
[0084] After obtaining the gradient response map G, concatenate it with the image I2 processed through each convolutional layer:
[0085] F i = Concat(I2, G) (4)
[0086] where Concat(·) represents the channel-level concatenation operation, and F i represents the feature output of the i-th layer. The multi-scale feature extraction module includes multiple convolutional layers;
[0087] First, perform a series of convolution operations on the image I2 processed by each hierarchical convolutional layer to extract features of different scales layer by layer; specifically, assume that the resolution of the input image is H×W, and use multiple convolutional layers with different depths to extract features; ensure that the information of shallow features can be effectively transmitted to deep layers and enhance the correlation between multi-scale features; introduce multiple convolutional kernels with different scales in each convolutional layer of the multi-scale feature extraction module to capture features of different scales:
[0088]
[0089] where σ(·) represents the softmax activation function, W i represents the weight matrix of the i-th convolutional kernel, F i-1 represents the feature output of the (i - 1)-th layer, and b i is the bias term;
[0090] Perform a concatenation operation on the features of different scales to obtain the processed multi-scale features:
[0091] F MDFE = Concat(F i (1×1) , F i (3×3) , F i (5×5)) (6)
[0092] Concat represents concatenation in the channel dimension;
[0093] A weighted fusion mechanism is adopted to perform weighted fusion on features of different scales:
[0094]
[0095] Among them, α i is the weight coefficient of features of different scales, n represents the number of multi-scale features, and F MDFE ' represents the multi-scale features after weighted fusion.
[0096] The dual-channel attention mechanism corresponds to the forward channel and the reverse channel respectively. The forward channel mainly focuses on the global semantic information of the detected target, while the reverse channel combines edge perception information and emphasizes the modeling of local details;
[0097] In the forward channel, channel information modeling is achieved through the channel attention mechanism. First, global average pooling is performed on the image feature map I3 composed of F i to extract the global information of each channel; the channel weights are calculated through a two-layer fully connected network, and the calculated channel weights are applied to the feature map I3 to enhance the overall semantic information of the salient target; at the same time, Softmax activation is used to obtain the channel attention score; the channel attention score is applied to the image feature map I3, so that important channels are enhanced and unimportant channels are suppressed;
[0098] In the reverse channel, spatial information modeling is achieved through the spatial attention mechanism; first, max pooling and average pooling in the channel dimension are performed on the feature map I3 to extract key spatial information; the spatial weights are calculated through a 7×7 convolution, and the Softmax activation function is used to generate the spatial attention map; the spatial attention map is multiplied element-wise with the feature map I3 to strengthen the local edges and detail information of the salient target, and a doubly enhanced feature map is obtained;
[0099] D i = DAM(F i ', I3) (8)
[0100] Among them, D i represents the doubly enhanced feature map of the i-th layer, DAM represents the dual-channel attention mechanism, and F i ' represents the spatial attention map of the i-th layer.
[0101] The multi-head attention mechanism adopts Scaled Dot-Product Attention in the Transformer structure to perform cross-scale information interaction and enhancement:
[0102] M' = MHA(M1, M2, M3,) (9)
[0103] Among them, M' represents the feature map processed by the multi-head attention mechanism, MHA represents the multi-head self-attention mechanism, and M1, M2, and M3 respectively represent the features of three different scales in multi-scale feature processing.
[0104] The feature alignment and fusion adopt a variety of fusion methods combined with a feature alignment mechanism;
[0105] The feature alignment mechanism performs spatial alignment and channel alignment in sequence;
[0106] The doubly enhanced feature map D i is directly weighted with the feature map M' obtained after multi-head attention processing to obtain feature maps D i ';
[0107] Spatial alignment: For features D i ' of different scales, use bilinear interpolation to adjust to the same spatial size H×W:
[0108] D i aligned = Resize(D i ', H, W) (10)
[0109] where Resize represents the upsampling operation;
[0110] Channel alignment: The number of channels of features of different scales is different, and 1×1 convolution is used for channel transformation:
[0111] D i aligned = Conv 1×1 (D i aligned ) (11)
[0112] After completing the feature alignment, a variety of fusion strategies are adopted, including addition fusion, multiplication fusion, and concatenation fusion:
[0113] (1) Addition fusion:
[0114]
[0115] (2) Multiplication fusion:
[0116]
[0117] (3) Concatenation fusion:
[0118] R cat = Concat(D1 aligned , D2 aligned,...,D n aligned ) (14)
[0119] Among them, the Concat operation is performed for splicing in the channel dimension, and then channel compression is performed through 1×1 convolution; finally, the fused feature R is weighted and combined by addition, multiplication, and concatenation:
[0120] R = ω1R add + ω2R mul + ω3R cat (15)
[0121] where ω1, ω2, and ω3 are learnable parameters.
[0122] The fused feature R still has a relatively high channel dimension and complex feature expressions. A 1×1 convolutional layer is used to reduce its dimension and extract significant information:
[0123] S' = Conv 1×1 (R) (16)
[0124] The Softmax activation function is used for processing to map the significant prediction to a probability distribution:
[0125]
[0126] where Si′ represents the unnormalized significant score at pixel point i, and after Softmax normalization, the entire significant map of the detected image is obtained.
[0127] Aiming at the problem that the significant object detection method based on edge perception and attention mechanism has poor performance in detecting significant objects in complex scenes, the present invention proposes a significant object detection method combining edge perception feature extraction and dual-channel attention mechanism. Edge information is introduced through the edge perception module to enhance the contrast between the significant object and the background, thereby improving the accuracy of the object boundary; the dual-channel attention mechanism is used to model channel attention and spatial attention respectively, taking into account global semantic information and local detail features, effectively improving the detection accuracy of the significant region; the multi-head attention mechanism is introduced to fuse features of different scales, realizing full interaction of multi-layer features, and improving the significant object detection ability in complex scenes. The present invention can achieve more refined object prediction in the high-resolution significant object detection task, and has higher robustness and detection accuracy compared with the existing methods.
[0128] To verify the effectiveness of the present invention in salient object detection, the present invention conducts experiments on multiple image datasets by combining edge perception and dual-channel attention mechanism. In the salient object detection task on the DUTS dataset, the method of the present invention achieves a mean absolute error of 0.038, while on the ECSSD dataset, the mean absolute error of salient objects is 0.026, demonstrating strong detection accuracy and robustness.
Claims
1. A salient object detection method based on edge perception and attention mechanism, characterized in that: The specific steps include: The input image is sequentially passed through each level of convolutional layers to extract features at each level; All features except the first convolutional layer are input to the boundary sensor for image boundary perception; The features of each level except the first convolutional layer are respectively input into the multi-scale feature extraction module of the corresponding level to obtain the features of different scales of the corresponding level; Except for the first and last layers, the edge information of the input image, the features of different scales at the corresponding layer, and the back channel are input into the dual-channel attention mechanism for feature fusion and alignment to obtain the saliency prediction map of the corresponding layer; the features of different scales at the last layer are subjected to multi-head attention to obtain the saliency prediction map of the last layer; The saliency prediction map starts from the last layer and is added layer by layer from bottom to top to obtain the saliency prediction map.
2. The method for detecting salient objects based on edge perception and attention mechanism according to claim 1, characterized in that: The convolutional layer adopts five layers.
3. The method for detecting salient objects based on edge perception and attention mechanism according to claim 1, characterized in that: The reverse channel is obtained by reverse attention propagation through the previous channel of the same level and the same scale feature.
4. The method for detecting salient objects based on edge perception and attention mechanism according to claim 1, characterized in that: The processing process of the boundary sensor is as follows: Calculate the gradient information of the input image I by convolution to obtain edge features; The specific idea is: for the input image, use the convolution kernel matrix of the Sobel operator to calculate the edge information of the input image; the horizontal direction G of the Sobel operator x and vertical direction G y The calculation is as follows: Among them, * represents the convolution operation; Calculate the gradient magnitude G as the edge response map: After obtaining the gradient response map G, it is concatenated with the image I2 processed by each convolution layer: F i =Concat(I2,G) (4) Among them, Concat(·) represents the channel-level concatenation operation, F i Represents the feature output of the i-th layer.
5. The method for detecting salient objects based on edge perception and attention mechanism according to claim 1, characterized in that: The multi-scale feature extraction module includes multiple convolutional layers; First, a series of convolution operations are performed on the image I2 processed by each level of convolutional layers to extract features of different scales layer by layer. Specifically, assuming that the resolution of the input image is H×W, multiple convolutional layers of different depths are used to extract features. This ensures that the information of shallow features can be effectively transmitted to the deep layer and enhances the association between multi-scale features. Multiple convolution kernels of different scales are introduced into each convolutional layer of the multi-scale feature extraction module to capture features of different scales: Among them, σ(·) represents the softmax activation function, W i represents the weight matrix of the i-th layer convolution kernel, F i-1 represents the feature output of layer i-1, b i is the bias term; Perform concatenation operations on features of different scales to obtain processed multi-scale features: F MDFE =Concat(F i (1×1) ,F i (3×3) ,F i (5×5) ) (6) Concat means concatenation in the channel dimension; A weighted fusion mechanism is used to perform weighted fusion of features of different scales: Among them, α i is the weight coefficient of different scale features, n represents the number of multi-scale features, F MDFE ' represents the multi-scale features after weighted fusion.
6. The method for detecting salient objects based on edge perception and attention mechanism according to claim 1, characterized in that: The dual-channel attention mechanism corresponds to the forward channel and the backward channel respectively. The forward channel mainly focuses on the global semantic information of the detected target, while the backward channel combines edge perception information and emphasizes the modeling of local details. In the forward channel, channel information modeling is achieved through the channel attention mechanism. First, F i The composed image feature map I3 is globally averaged pooled to extract the global information of each channel; the channel weights are calculated through a two-layer fully connected network, and the calculated channel weights are applied to the feature map I3 to enhance the overall semantic information of the salient target; at the same time, Softmax activation is used to obtain the channel attention score; The channel attention score acts on the image feature map I3, so that important channels are enhanced and unimportant channels are suppressed; In the reverse channel, spatial information modeling is achieved through the spatial attention mechanism. First, the feature map I3 is subjected to maximum pooling and average pooling in the channel dimension to extract key spatial information. The spatial weight is calculated through a 7×7 convolution, and a spatial attention map is generated using a Softmax activation function; the spatial attention map is element-wise multiplied with the feature map I3 to enhance the local edge and detail information of the salient target, thereby obtaining a doubly enhanced feature map; D i =DAM(F i ',I3) (8) Among them, D i represents the doubly enhanced feature map of the i-th layer, DAM represents the dual-channel attention mechanism, and F i ' represents the spatial attention map of the i-th layer.
7. The method for detecting salient objects based on edge perception and attention mechanism according to claim 1, characterized in that: The multi-head attention mechanism uses the Scaled Dot-Product Attention in the Transformer structure to perform cross-scale information interaction and enhancement: M'=MHA(M1,M2,M3,) (9) Among them, M' represents the feature map after processing by the multi-head attention mechanism, MHA represents the multi-head self-attention mechanism, and M1, M2, and M3 represent features of three different scales in multi-scale feature processing.
8. The method for detecting salient objects based on edge perception and attention mechanism according to claim 1, characterized in that: The feature alignment and fusion adopts multiple fusion methods combined with feature alignment mechanism; The feature alignment mechanism sequentially performs spatial alignment and channel alignment; The doubly enhanced feature map D i Directly weight the feature map M' obtained after multi-head attention processing to obtain feature maps D of different scales i '; Spatial alignment: For features D of different scales i ', use bilinear interpolation to adjust to the same spatial size H×W: D i aligned =Resize(D i ',H,W) (10) Where Resize represents the upsampling operation; Channel alignment: The number of feature channels at different scales is different, and 1×1 convolution is used for channel transformation: D i aligned =Conv 1×1 (D i aligned ) (11) After feature alignment, a variety of fusion strategies are used, including additive fusion, multiplicative fusion, and cascade fusion: (1) Additive Fusion: (2) Multiplication Fusion: (3) Cascade fusion: R cat =Concat(D1 aligned ,D2 aligned ,...,D n aligned ) (14) Among them, the Concat operation performs splicing in the channel dimension, and then performs channel compression through 1×1 convolution; the final fusion feature R is a weighted combination of addition, multiplication and cascade: R=ω1R add +ω2R mul +ω3R cat (15) Among them, w1, w2, w3 are learnable parameters.
9. The method for detecting salient objects based on edge perception and attention mechanism according to claim 8, characterized in that: The fused feature R still has a high channel dimension and complex feature expression. It is reduced in dimension through a 1×1 convolution layer to extract significant information: S'=Conv 1×1 (R) (16) The Softmax activation function is used to map the significance prediction into a probability distribution: Among them, Si′ represents the unnormalized saliency score at pixel i, and Softmax normalizes the entire saliency map.
Citation Information
Cited By
Cloud detection method and device, electronic equipment and computer readable storage medium
CN121033367A
Multi-dimensional frequency domain and deformable attention fusion saliency target detection method
CN121190754A
Document-level saliency detection method and related device
CN121542806A
Small target detection method based on combination of target super-resolution, background degradation and attention
CN121746692A