Salient target detection method and system based on edge detection and attention mechanism
By combining edge detection and attention mechanism in the significance object detection method, using parallel branches and multi-scale attention, the accuracy and robustness of target boundary detection in complex scenarios are solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510172353.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
AI Technical Summary
Existing significant object detection methods have problems with accuracy and robustness when dealing with complex scenarios, especially when the lighting is uneven, the background texture is complex, and the target is blocked, it is difficult to accurately detect the target boundaries.
The significance object detection method based on edge detection and attention mechanism is adopted. Through parallel edge branches and significance detection branches, combined with multi-scale attention and edge guidance learning strategies, the interior and boundaries of regional features are improved and the accuracy of object recognition is improved.
Through the combination of edge detection and attention mechanism, the significance target detection method significantly improves the detection accuracy and robustness of target boundaries in complex scenarios, meeting the needs of practical applications.
Smart Images

Figure CN120107615A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer vision technology, and in particular to a method and system for detecting salient targets based on edge detection and attention mechanism. Background Art
[0002] In today's digital age, image and video data are growing explosively. How to quickly and accurately locate salient targets from massive visual information has become one of the key research topics in the field of computer vision, and it plays an indispensable role in many practical application scenarios.
[0003] In the early research on salient object detection, traditional methods mainly rely on manually designed features to identify salient areas. These manual features usually include low-level visual clues such as color, brightness, texture, and contrast. Researchers cleverly design various heuristic algorithms to calculate the saliency of different areas. For example, the saliency model based on the human visual system proposed by Itti et al. simulates the center-periphery mechanism and calculates the feature differences of image regions at different scales to determine salient areas. This method can achieve certain results in simple scenes, but its limitations are exposed when faced with complex real-world scenes. The uneven illumination in complex scenes may cause instability in color and brightness features, and the complex textures and repeated patterns in the background will interfere with the saliency calculation based on texture and contrast. Moreover, when the target is partially occluded, traditional methods often find it difficult to accurately detect the complete target area. These factors seriously restrict the performance of traditional manual feature methods in complex scenes, resulting in low accuracy and robustness of detection results, which is difficult to meet the needs of practical applications.
[0004] With the rapid development of deep learning technology, salient object detection methods based on convolutional neural networks (CNNs) have gradually become mainstream. The powerful automatic feature learning ability of CNN enables it to extract high-level semantic features of images from large-scale data, thus showing great advantages in salient object detection tasks. For example, the fully convolutional network (FCN) architecture is widely used in pixel-level saliency prediction tasks. Through end-to-end training, it can directly output a saliency map of the same size as the input image, and has achieved significant performance improvements over traditional methods on some public datasets. However, although CNN can learn rich semantic information, it still has obvious deficiencies in processing object boundaries. In the forward propagation process of CNN, as the number of network layers increases, the receptive field gradually expands and the resolution of the feature map gradually decreases, which causes the network to lose a lot of detail information in the deep features, resulting in blurry, discontinuous or even wrong detection of object boundaries; especially when the contrast between the target and the background is low or the target boundary is complex, it is difficult for CNN-based methods to accurately locate the target boundary, which is a serious problem for applications that require accurate target segmentation.
[0005] Edge detection is a basic technology in image processing and computer vision. Its main purpose is to identify points in digital images where the brightness changes dramatically (i.e., edge points) so as to outline the contours of objects in the image. Edges usually correspond to the boundaries between the target and the background, the boundaries between different parts of the target, or the boundaries of the texture. Through edge detection, the target in the image can be separated from the background, providing important basic information for subsequent higher-level visual tasks such as target recognition, target segmentation, and shape analysis. Classic edge detection algorithms such as the Canny operator and the Sobel operator determine the edge position by calculating the gradient of the image. However, these traditional edge detection methods are easily affected by noise in complex scenes, and the detection effect of weak edges is poor. They are difficult to be directly applied to the task of salient target detection, resulting in poor overall detection effect of existing salient target detection schemes. Summary of the invention
[0006] In order to solve the above problems, the present invention proposes a salient target detection method and system based on edge detection and attention mechanism. By combining edge detection and attention mechanism, the salient target recognition and positioning as well as the image edge definition are better realized, thereby improving the accuracy of the system's salient target recognition.
[0007] According to some embodiments, the present disclosure adopts the following technical solutions:
[0008] The salient object detection method based on edge detection and attention mechanism includes:
[0009] Acquire an image to be detected;
[0010] Input the image into the trained object detection network to highlight the salient object area and obtain the final salient object detection prediction map;
[0011] Among them, the target detection network constructs a parallel edge branch and a saliency detection branch based on two sets of independent features extracted from the image, respectively emphasizing the selectivity and invariance of features in detecting salient edges and salient regions; the edge branch interactively fuses low-level features with spatial structural details and high-level features with rich semantic knowledge to obtain edge features; the saliency detection branch generates a regional feature containing multi-scale key information through multi-scale attention, and uses the edge-guided learning strategy and the edge features as a guide to improve the interior and boundaries of the regional features to obtain the final salient target region.
[0012] According to some embodiments, the present disclosure adopts the following technical solutions:
[0013] The salient object detection system based on edge detection and attention mechanism includes:
[0014] The image acquisition module is configured to: acquire an image to be detected;
[0015] The object detection module is configured to: input the image into the trained object detection network, highlight the salient object area, and obtain the final salient object detection prediction map;
[0016] Among them, the target detection network constructs a parallel edge branch and a saliency detection branch based on two sets of independent features extracted from the image, respectively emphasizing the selectivity and invariance of features in detecting salient edges and salient regions; the edge branch interactively fuses low-level features with spatial structural details and high-level features with rich semantic knowledge to obtain edge features; the saliency detection branch generates a regional feature containing multi-scale key information through multi-scale attention, and uses the edge-guided learning strategy and the edge features as a guide to improve the interior and boundaries of the regional features to obtain the final salient target region.
[0017] According to some embodiments, the present disclosure adopts the following technical solutions:
[0018] A computer program product includes a computer program, and when the computer program is executed by a processor, the method for detecting a salient object based on edge detection and an attention mechanism is implemented.
[0019] According to some embodiments, the present disclosure adopts the following technical solutions:
[0020] A non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the method for detecting a salient target based on edge detection and attention mechanism is implemented.
[0021] According to some embodiments, the present disclosure adopts the following technical solutions:
[0022] An electronic device comprises: a processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device implements the method for detecting a salient target based on edge detection and an attention mechanism.
[0023] Compared with the prior art, the present invention has the following beneficial effects:
[0024] This application combines edge detection and attention mechanism to better realize the recognition and positioning of significant targets and the definition of image edges; by designing multiple feature aggregation modules and cleverly using multiple convolutions to mine deep semantic information of images, shallow image information is retained as much as possible, and feature images of different scales and different information are fused to improve the accuracy of the system's significant target recognition, including:
[0025] (1) Edge-guided learning and multi-level cross-level fusion: First, considering the different properties of object interior and edge features, two independent sets of features are extracted from the backbone network. Then, they are used to construct two parallel branches to emphasize the selectivity and invariance of features in detecting salient edges and salient regions, respectively. Finally, an edge-guided interaction module (EGI) is further proposed to achieve interaction between the two branches through an edge-guided learning strategy that uses edge information as guidance to simultaneously improve the interior and boundaries of salient objects.
[0026] (2) In addition to utilizing edge information, two specific cross-level fusion modules are designed to fully utilize the semantics and detailed information within the two branches; a high-level interactive fusion module (HIF) is introduced to better provide location guidance of the global context by exploiting the correlation between two adjacent semantic features; and a low-level weighted fusion module (LWF) is designed to supplement the fine information and alleviate the pollution of redundant features by dynamically selecting the input information stream. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings constituting a part of the present disclosure are used to provide a further understanding of the present disclosure. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation on the present disclosure.
[0028] Figure 1This is the edge-aware SOD network structure diagram of Example 1 of the present disclosure.
[0029] Figure 2 This is a diagram of the encoder structure of Embodiment 1 of the present disclosure.
[0030] Figure 3 It is a structural diagram of the atrous spatial convolutional pooling pyramid module (ASPP) of Example 1 of the present disclosure.
[0031] Figure 4 This is a structural diagram of the efficient multi-scale attention module (EMA) of Example 1 of the present disclosure.
[0032] Figure 5 It is a structural diagram of the Advanced Interactive Fusion Module (HIF) of Example 1 of the present disclosure.
[0033] Figure 6 It is a structural diagram of the low-level weighted fusion module (LWF) of Example 1 of the present disclosure.
[0034] Figure 7 It is a redefined structural diagram of Embodiment 1 of the present disclosure.
[0035] Figure 8 This is a structural diagram of the spatial and channel reconstruction convolution (SSCONV) of Example 1 of the present disclosure.
[0036] Fig. 9 It is a structural diagram of the edge guided interaction module (EGI) of Example 1 of the present disclosure.
[0037] Fig.10 It is the GMS flow chart of Example 1 of the present disclosure. DETAILED DESCRIPTION
[0038] The present disclosure is further described below in conjunction with the accompanying drawings and embodiments.
[0039] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present disclosure belongs.
[0040] Example 1
[0041] In one embodiment of the present disclosure, a salient object detection method based on edge detection and attention mechanism is provided, including:
[0042] Step S1: Acquire the image to be detected;
[0043] Step S2: Input the image into the trained object detection network, highlight the salient object area, and obtain the final salient object detection prediction map;
[0044] Among them, the target detection network constructs a parallel edge branch and a saliency detection branch based on two sets of independent features extracted from the image, respectively emphasizing the selectivity and invariance of features in detecting salient edges and salient regions; the edge branch interactively fuses low-level features with spatial structural details and high-level features with rich semantic knowledge to obtain edge features; the saliency detection branch generates a regional feature containing multi-scale key information through multi-scale attention, and uses the edge-guided learning strategy and the edge features as a guide to improve the interior and boundaries of the regional features to obtain the final salient target region.
[0045] As an embodiment, the disclosed salient target detection method based on edge detection and attention mechanism, by combining edge detection and attention mechanism, better realizes salient target recognition and positioning and image edge definition, and improves the accuracy of salient target recognition of the system. The specific implementation process is as follows:
[0046] Step 1: Build an experimental platform and use Python language to build a target detection network, which is called edge-aware SOD network in this embodiment. Call the PyTorch open source framework to implement the construction of convolutional networks. All experiments are performed on the server (NAVIDIA RTX-3080 GPU, video memory 24G).
[0047] Step 2: Prepare training data sets and test data sets. There are a series of classic data sets for salient object recognition. In order to make the experimental results more comprehensive, five data sets are selected to conduct multi-faceted tests on the edge-aware SOD network. The five classic data sets are ECSSD, PASCAL-S, DUT-OMRON, HKU-IS, and DUTS-E.
[0048] Step 3: Preprocess the data set, crop the photos in the data set to make them uniform in size, and subtract irrelevant parts of the photos; then use Gaussian filtering and other filtering techniques to denoise the photos. Enhance the data by randomly and horizontally flipping the images in advance and randomly adjust the image size to [224, 256, 288, 320, 352].
[0049] Step 4: Pre-train the backbone network of the edge-sensing SOD network. The backbone network is mainly used to extract and encode image information, corresponding to the five light blue stereo blocks in the encoder, to facilitate follow-up processing by subsequent modules. This embodiment selects RESNET-50 as the backbone network, and pre-trains it, and saves the training path and training parameter weights.
[0050] When pre-training and initializing the parameters of the backbone network, the data training set ImageNet is selected for unified training operations; for other parameters in the training model, a series of parameters are randomly generated using the random package, and then the optimal parameters are updated and adjusted through continuous training.
[0051] In the training phase, stochastic gradient descent (SGD) is used to optimize the network, with momentum set to 0.9 and weight decay set to 0.0005; the initial learning rate is set to 0.005 and the others are set to 0.0005; the training process requires a total of 66 epochs with a mini-batch size of 6; in the inference phase, the image size is randomly resized to [224, 256, 288, 320, 352] to obtain predictions without any other post-processing techniques.
[0052] Step 5: Train the overall model. Set the model parameters the same as in step 4 and train the model in the DUT-TR dataset.
[0053] Step 6. Test on the five data sets in step 2. In order to better reflect the effect of the edge-aware SOD network, three common indicators, F value, MAE value, and S value, are used to intuitively display the results of the edge-aware SOD network, where the F value parameter is set to 0.3.
[0054] The edge-aware SOD network is described in detail below.
[0055] Recently, deep convolutional neural networks (CNNs) have demonstrated their powerful feature representation capabilities and have been successfully applied to SOD with satisfactory performance; although significant improvements have been made, there are still two challenges that need to be addressed in order to make further progress:
[0056] The first challenge is how to effectively extract and utilize edge information to assist saliency reasoning. Inspired by the observation that boundary regions are proper subsets of corresponding salient regions, some algorithms mine edge cues from salient object information under the supervision of edge labels. They often fuse edge and salient features by simple concatenation or point-wise summation. These methods achieve impressive detection performance. However, they ignore the fact that the edge and internal regions of salient objects in the network have different characteristics. Specifically, internal features should remain unchanged from appearance changes to ensure the consistency of intra-class features, while boundary features should be selective enough to separate salient objects from background regions. Therefore, the above methods may encounter the problem of insufficient coordination between internal and edge features. As a result, even with clear boundaries, salient objects are easily disturbed by background edges or have obvious internal defects.
[0057] The second challenge is how to make full use of multi-level features. Generally speaking, high-level features with rich semantic knowledge can provide strong guidance for the coarse localization of salient objects, while low-level features retain spatial structural details suitable for depicting object boundaries. Based on the above observations, current SOD methods usually combine semantics and detailed information in the decoder by utilizing a series of identical progressive merging processes. Although these methods have achieved milestone results, they ignore the gap difference between the two input features of each cross-level fusion process. Therefore, some methods may find it difficult to suppress background noise or capture complete salient objects in complex scenes.
[0058] In order to solve the above problems, this embodiment proposes a novel edge-aware SOD network, such as Figure 1 As shown in the figure, it includes encoder, atrous spatial convolutional pooling pyramid module (ASPP), efficient multi-scale attention module (EMA), high-level interactive fusion module (HIF), low-level weighted fusion module (LWF), edge-guided interaction module (EGI), gated multi-scale predictor (GMS), etc., which are described below:
[0059] 1. Encoder
[0060] In the deep learning architecture, the encoder is a neural network module that is mainly used to extract features and learn representations of input data (such as text, images, speech, etc.); it gradually converts the input data into a low-dimensional, abstract feature representation. This process is usually achieved through a series of convolutional layers, pooling layers (when processing data such as images) or fully connected layers (when processing general vector data).
[0061] Taking image data as an example, in the encoder part of the convolutional neural network (CNN), the initial convolution layer extracts low-level features of the image with spatial structural details, such as edges, textures, etc.
[0062] As the number of network layers increases, the spatial dimension of the data is reduced through pooling operations (such as maximum pooling and average pooling), while the convolutional layer continues to extract more advanced high-level features with rich semantic knowledge, such as the shape of the object, category-related features, etc.
[0063] Ultimately, the encoder converts the original high-dimensional image data into a low-dimensional feature vector, which contains an abstract representation of the input image and can be used for subsequent classification, retrieval, generation and other tasks.
[0064] In this embodiment, Figure 2 As shown in the figure, based on RESNET-50, the encoder is built using 1x1 convolution and dilated spatial convolution pooling pyramid module ASPP to realize image feature mining.
[0065] Like most salient object detection models, the network structure proposed in this example is also an encoder-decoder structure. The encoder part is specifically:
[0066] Based on the commonly used RESNET-50 architecture, pre-training on ImageNet was performed, the last global pooling layer and fully connected layer were removed from the original RESNET-50 network, and only five sequentially stacked cubic blocks (i.e. Figure 2 The five blue blocks in ) serve as the backbone feature extractor.
[0067] Given an input image, first, the original side features of the corresponding level are obtained through these five residual blocks using the improved RESNET-50. Then, they are processed through 1x1 convolution and the atrous spatial convolution pooling pyramid module ASPP to generate two independent sets of features, namely, Figure 1 Medium E 1 、E 2 、E 3 is a set of features, F 1 、F 2 、F 3 、F 4 、F 5 as another set of features and pass them to two parallel branches, the decoder for saliency detection and edge detection.
[0068] Among them, the 1x1 convolution and atrous spatial convolution pooling pyramid module ASPP, specifically, applies two groups of 1×1 convolution layers to reduce the channels of three lower-level common features, and the features of the last two layers (i.e., the outputs of the residual blocks) are respectively input into the corresponding ASPP modules to enhance the semantic information, where the expansion rate of the module is set to {2, 4, 6}.
[0069] The specific process is as follows:
[0070] (1) The input image is resized to 224x224 pixels with 3 channels through cropping and other scaling transformations.
[0071] (2) Build the initial convolutional block, i.e. the first blue block, which consists of a convolutional layer, a batch normalization layer, a ReLU activation layer, and a maximum pooling layer.
[0072] Convolution layer: Use a 7×7 convolution kernel, 64 filters, and a stride of 2 to perform preliminary feature extraction on the input image.
[0073] Batch Normalization Layer: A batch normalization layer is added after the convolution layer to normalize the data and improve the stability of training.
[0074] ReLU activation layer: Introduce the nonlinear activation function ReLU to enhance the expressive power of the network.
[0075] Maximum pooling layer: A maximum pooling layer with a pooling kernel of 3×3 and a stride of 2 is used to further reduce the size of the feature map and reduce the amount of computation.
[0076] (3) Build four residual modules, namely:
[0077] The first residual module: contains 3 residual blocks, each residual block uses 64 filters, and the convolution kernel sizes in its structure are mainly 1×1 and 3×3, which are used to extract and process low-level features of the image.
[0078] The second residual module: includes 4 residual blocks, each with 128 filters. In this stage, the residual block with Bottleneck structure is used. The number of channels is first reduced to 64 through 1×1 convolution, and then the main features are processed through 3×3 convolution. Finally, another 1×1 convolution is used to restore the number of channels to 128. This structure helps to reduce the amount of calculation and the number of parameters, while effectively extracting more advanced semantic features.
[0079] The third residual module: includes 6 residual blocks, each block has 256 filters, and also adopts the Bottleneck structure to further deepen the network and extract more complex features.
[0080] The fourth residual module: includes 3 residual blocks, each with 512 filters, and continues to use the Bottleneck structure to learn higher-level semantic information of the image.
[0081] The residual block is the core building block in the ResNet (Residual Network) architecture, and is designed to solve the gradient vanishing and degradation problems that occur during the training of deep neural networks. As the number of network layers continues to increase, traditional neural networks may face the situation where the gradient gradually approaches zero during the back propagation process, making the network difficult to train; the residual block introduces a "short-cut connection" to make it easier for the network to learn the identity mapping, so that it can be effectively trained even if the network is deepened.
[0082] The main components of the residual block are:
[0083] Convolution Layers: The residual block contains multiple convolution layers, which are used to extract features of the input data. The parameters of these convolution layers (such as convolution kernel size, step size, padding method, etc.) are set according to the specific network architecture and task requirements; for example, in ResNet-50, some convolution layers use 1×1 and 3×3 convolution kernels to adjust the number of channels and extract spatial features.
[0084] Batch Normalization Layers: Batch Normalization Layers are usually added after each convolutional layer. It normalizes the data of each batch to make the data distribution more stable, accelerate the network training process, and alleviate the gradient vanishing problem to a certain extent.
[0085] Activation Functions: Common activation functions such as ReLU (Rectified Linear Unit) are used to introduce nonlinear factors and enhance the network's expressiveness. The ReLU function sets input values less than zero to zero and values greater than zero to remain unchanged, which allows the network to learn more complex functional relationships.
[0086] Short-cut Connection: This is the key innovation of the residual block. The shortcut connection directly skips one or more convolutional layers and then adds the input data to the output after the convolutional layer and other operations. For example, assuming that the input of the residual block is x, and the output after a series of convolution, normalization and activation operations is F(x), then the final output of the residual block is F(x)+x. This structure enables the network to learn the residual function F(x)=yx. If some layers of the network do not learn effective feature transformations, then it can directly learn the identity mapping (i.e., F(x)=0), thereby ensuring that the network performance does not decrease with the increase in the number of layers.
[0087] Bottleneck Residual Block: It is used for deeper networks, such as ResNet-50 and deeper network architectures. Its structure is to first reduce the number of channels through a 1×1 convolution layer (to reduce the dimension), then perform the main feature extraction through a 3×3 convolution layer, and finally restore the number of channels through a 1×1 convolution layer. This design can greatly reduce the amount of calculation and the number of parameters while maintaining the performance of the network, and also includes shortcut connections to achieve residual learning.
[0088] Among them, the atrous spatial convolutional pooling pyramid module (ASPP), such as Figure 3As shown in the figure, it consists of dilated convolution, pooling layer, pyramid module and other parts; the dilated convolution is 3x3 convolution, and three 3x3 dilated convolutions form a pyramid module; the pooling layer is 1x1 pooling, and the 1x1 convolution is a common convolution block. Specifically:
[0089] Atrous Convolution: Also called dilated convolution, it is a variant of the convolution operation. Unlike ordinary convolution, it inserts some "holes" between the convolution kernel elements and controls the size of the receptive field by setting different dilation rates. For example, when the dilation rate is 1, it is an ordinary convolution. When the dilation rate is greater than 1, the convolution kernel performs convolution operations on the input feature map in a jumping manner. In this way, the receptive field can be increased without increasing the size of the convolution kernel and the amount of calculation, thereby obtaining a wider range of contextual information.
[0090] Pooling layer: It is a downsampling operation, which mainly includes maximum pooling and average pooling. Maximum pooling selects the maximum value in the local area as the output, while average pooling calculates the average value of the local area. Pooling operation can reduce the dimension of data and the amount of calculation, while retaining important feature information and improving the robustness and anti-interference ability of the model.
[0091] Pyramid Module: Generally refers to building a multi-level structure, similar to the shape of a pyramid; in computer vision, pyramid modules are usually used to process information of different scales. For example, the Feature Pyramid Network (FPN) constructs a multi-scale feature pyramid to fuse the high-resolution, low-semantic information of the shallow network with the low-resolution, high-semantic information of the deep network to meet the detection and segmentation requirements of targets of different sizes in tasks such as target detection and semantic segmentation.
[0092] Through this module, the following functions are realized:
[0093] Multi-scale feature extraction and fusion: Atrous spatial convolution can extract multi-scale spatial features at different dilation rates, and the pooling operation further adjusts the scale of the feature map. By constructing a pyramid structure, these features of different scales can be stored and fused in layers. For example, in the lower-level pyramid, atrous convolution with a smaller dilation rate and a larger pooling size can be used to obtain the detailed features of small targets. At higher levels, atrous convolution with a larger dilation rate and a smaller pooling size or no pooling is used to obtain the global semantic features of large targets.
[0094] Contextual information enhancement: The use of dilated convolution enables the module to obtain a wider range of contextual information without losing too much spatial resolution. At different levels of the pyramid, by properly setting the dilation rate and pooling strategy of the dilated convolution, the contextual semantics of the features can be gradually enriched, from the context of local details (such as the texture context of the edge of the target) to the global context (such as the relationship between the target and the surrounding environment), providing more comprehensive information support for subsequent visual tasks.
[0095] Adapting to targets of different sizes and scene complexities: In tasks such as target detection or segmentation, the size of the target and the complexity of the scene vary. This pyramid module can automatically adjust the feature extraction method according to the target size and scene complexity. For small targets and complex background scenes, it pays more attention to the fine features at the bottom of the pyramid. For large targets and simple scenes, it can use the abstract semantic features of the high-level pyramid to improve the performance of the model in various situations.
[0096] This embodiment provides an example of ASPP, specifically:
[0097] (1) Accepts a feature map as input, which has shape (C, H, W), where C represents the number of input channels, H represents the input height, and W represents the input width.
[0098] (2) Multi-scale dilated spatial convolution stage: The module sets up multiple parallel branches, each of which is used to generate feature maps of different scales. Then, dilated convolution calculations are performed in these branches to generate multi-scale features. The dilated convolution calculation formula is as follows:
[0099]
[0100] Where X is the input feature map, c is the output channel index, h and w are the height and width indices of the output feature map, c′ is the input channel index, h′ and w′ are the position indices within the convolution kernel, and K i is the convolution kernel of the i-th branch.
[0101] (3) Perform an average pooling operation on the output results in step (2).
[0102] (4) Pyramid construction: Multi-scale feature maps are constructed into a pyramid structure according to their scales. Generally speaking, feature maps with smaller scales (higher resolution) are placed at the bottom of the pyramid, and feature maps with larger scales (lower resolution) are placed at the top of the pyramid. For example, if there is a high-resolution feature map obtained by a small dilation rate atrous convolution and a low-resolution feature map obtained by a large dilation rate atrous convolution and pooling operation, the former is placed at the bottom and the latter is placed at the top.
[0103] (5) Output the results.
[0104] 2. Efficient Multi-Scale Attention Module (EMA)
[0105] Attention mechanism is of great significance in salient target recognition. It can focus on key information and highlight the target area. The attention mechanism enables the model to automatically focus on the salient target area in the image and ignore irrelevant information such as the background. For example, the channel attention mechanism can assign weights to feature maps of different channels, enhance channel features related to salient targets, and suppress unimportant channel features, so that the model can more effectively use key feature information to identify targets, improve the accuracy and effectiveness of feature expression, and thus improve the recognition accuracy.
[0106] In the task of salient object recognition, the attention mechanism helps the model locate and recognize objects more accurately, especially for complex situations such as multiple objects, small objects, and low contrast between objects and backgrounds, which can significantly improve recognition accuracy. For example, in medical images, accurate identification of tiny lesion areas provides a more accurate basis for disease diagnosis.
[0107] The attention mechanism can guide the model to learn effective feature representations more quickly, reduce the learning and processing of irrelevant information, thereby accelerating the convergence of the model, improving training efficiency, and enabling the model to achieve better performance in a shorter time. Human vision can quickly and accurately identify salient targets in complex and changing environments. Models based on the attention mechanism also have similar capabilities and can better cope with salient target recognition tasks under complex conditions such as different lighting, angles, and occlusions, with stronger robustness and generalization capabilities.
[0108] In salient object recognition, information at different scales is important for accurate object recognition. The attention mechanism can effectively integrate multi-scale information, capturing both the global features of the object and the local details of the object, thereby more comprehensively describing the salient object and improving the accuracy and completeness of recognition. For example, when recognizing large objects, attention is paid to both the overall outline and the local key parts of the object to more accurately determine the target category.
[0109] EMA is a module used to process multi-scale information in computer vision or other related fields. Its core purpose is to efficiently focus on the key information of input data at different scales through the attention mechanism. This module aims to solve the problem that traditional methods may not be able to effectively utilize multi-scale features when dealing with situations involving multiple target sizes, complex scene structures, etc.
[0110] The module, such as Figure 4 As shown in the figure, it consists of a multi-scale feature extraction part, an attention mechanism part, and a feature fusion part, which are:
[0111] (1) Multi-scale feature extraction: usually some structures that can generate multi-scale features are used, such as convolutional layers with different step lengths, pooling layer combinations, or dilation convolutions with different expansion rates. These operations can extract feature maps of different resolutions from the original input data, representing information at different scales. For example, feature maps with lower resolution may contain more abstract semantic information and are suitable for identifying large targets or overall scene structures. Feature maps with higher resolution retain more detailed information and are important for detecting small targets or fine parts of targets.
[0112] Specifically, the input feature map is fed into multiple parallel convolution branches, each of which uses different convolution kernel size, step size and other parameters to extract features of different scales.
[0113] (2) Attention mechanism part: including spatial attention submodule and channel attention submodule:
[0114] Spatial attention submodule: It focuses on the importance distribution of feature maps in the spatial dimension. It highlights the importance of different positions for the task by calculating the weight of each spatial position. For example, in target detection tasks, spatial attention can enable the model to focus on areas where targets may appear and reduce attention to background areas. There are many ways to calculate spatial attention weights. A common way is to perform global average pooling and global maximum pooling on the feature map, and then pass the pooled results through a small convolutional neural network or a fully connected network to generate a spatial attention weight map.
[0115] Global pooling: For each scale of feature maps, global average pooling and global maximum pooling operations are performed respectively to compress each feature map into two different vectors, which respectively summarize the overall average information and maximum activation information of the feature map in the spatial dimension.
[0116] Weight generation network: After concatenating the two vectors obtained above, input them into a small convolutional neural network or multi-layer perceptron composed of a 1x1 convolutional layer and an activation function. After calculation, it outputs a spatial attention weight map with the same spatial size as the input feature map. Each element in the map represents the importance of the corresponding position to the task in the spatial dimension.
[0117] Feature weighting: The spatial attention weight map is multiplied element-by-element with the original feature maps of each scale, so that the model can focus on more important areas in the spatial dimension, enhance the feature representation of key areas, and suppress the information of unimportant areas.
[0118] Channel attention submodule: Attention is allocated to the channel dimension of the feature map. Each channel usually represents a specific feature type in a convolutional neural network. The channel attention mechanism can evaluate the relevance of each channel to the current task; for example, in an image classification task, some channels may contain features that are highly relevant to the target category (such as texture, color, etc.); the calculation of channel attention weights can be achieved by pooling the channel features and then passing them through a small network containing an activation function.
[0119] Channel pooling: For the channel dimension of each scale feature map, global average pooling and global maximum pooling are performed respectively to obtain two statistical features of each channel, and then these statistical features are concatenated into a new vector.
[0120] Channel weight calculation: The concatenated vector is input into a small network containing a fully connected layer and an activation function (such as Sigmoid) to calculate the attention weight of each channel. These weights reflect the relevance of the features contained in each channel to the task.
[0121] Channel feature adjustment: Use the calculated channel attention weights to perform weighted summation on the original channel features, enhance the features of important channels, suppress less important channels, and make the model pay more attention to channel information related to the task.
[0122] (3) Feature fusion: The weighted feature maps of different scales obtained through the attention mechanism will be fused here. The fusion method can be a simple weighted summation or a more complex feature concatenation followed by integration through a convolutional layer. Through this fusion, the model can comprehensively utilize the key information filtered by attention at different scales to generate a feature representation containing important information at multiple scales for subsequent task processing, such as classification, detection, or segmentation.
[0123] Weighted fusion: Feature maps of different scales processed by the attention mechanism are fused, usually by weighted summation. Each scale feature map is assigned a weight, which can be adaptively adjusted according to the task and data, or can be a pre-set fixed value. The feature maps of each scale are then added together according to their weights to obtain a fused feature map that integrates important information at different scales.
[0124] Convolutional integration: After weighted summation, a convolutional layer may be used to integrate the fused features. Through one or more 1x1 convolutional layers, the number of channels and spatial structure of the fused feature map can be further adjusted to make the features smoother and more compact, and better suited to subsequent task processing, such as input to a classifier for image classification, used to generate a saliency map for target detection, etc.
[0125] 3. High-level interactive fusion module (HIF) and low-level weighted fusion module (LWF)
[0126] How to effectively aggregate features at different levels is a key issue in salient object detection. As mentioned earlier, most SOD methods usually adopt a unified progressive fusion process to exploit high-level semantic information and low-level detailed information; however, due to the differential fusion of the gaps between cross-layer input features, this unified integration may interfere with the transmission of global structured cues or introduce ambiguous features.
[0127] To this end, this embodiment designs two specific modules (HIF and LWF) for different cross-level features, effectively realizing the complementary advantages of multi-level features.
[0128] 1. Advanced interactive fusion module:
[0129] The difference between two consecutive high-level features is small compared to other high-level features, and has a strong ability to capture the global background. Therefore, a high-level interactive fusion module (HIF) is proposed to provide the location information of salient objects by exploiting the cross-layer correlation of these two deeper features, such as Figure 5 As shown, the specific steps are:
[0130] (1) Input F into the module 5 and F 4 , and perform three 1x1 convolutions on each of them to obtain
[0131] (2) Matrix multiplication, The matrix is multiplied, and the two results are processed by the Softmax function respectively.
[0132] (3) Perform matrix multiplication on the results of step (1) and step (2) as shown in the figure to obtain And give them different weights.
[0133] (4) Add the four results obtained in pairs according to their weights and perform concat processing.
[0134] (5) The result of step (4) is processed by 3x3 convolution to obtain the output result.
[0135] 2. Low-level weighted fusion module (LWF)
[0136] In cross-layer fusion, there are large differences between low-level features; in the process of fusion with high-level features, low-level features contain detailed clues but also have more background noise, while high-level information has rich semantic information but lacks fine information; for this reason, a novel low-level weighted fusion module is proposed, which selectively integrates the above two information streams (i.e., low-level features with spatial structural details and high-level features with rich semantic knowledge) to suppress redundant information and supplement important features, such as Figure 6 As shown, the specific steps are:
[0137] (1) Input f into the module i+1 and f i , and both are processed by spatial and channel reconstruction convolution (SSCONV) to obtain and The two are multiplied at the pixel level, and the result of the multiplication is then compared with Perform concat splicing to get
[0138] (2) The result obtained in step (1) is processed by activation function and 3x3 convolution, and then respectively and Multiply and add the results to get
[0139] (3) Respectively and Redefine it to get and The redefinition here is Figure 7 As shown, After convolution processing, Perform pixel multiplication and addition alternately; finally, subtract the two results at the pixel level, and finally obtain the subtraction result through convolution. Same reason.
[0140] (4) The two results obtained in step (3) are concatenated and then processed by 3x3 convolution to obtain the final output feature map.
[0141] The following is an explanation of the spatial and channel reconstruction convolution (SSCONV), as follows: Figure 8 As shown, SSCONV includes a spatial reconstruction unit, a channel reconstruction volume unit, and a 1x1 convolution:
[0142] Spatial reconstruction unit: The main focus is on reconstructing the information of the feature map in the spatial dimension (usually the height and width of a two-dimensional image); traditional convolution operations will change the spatial structure of the feature map when extracting features, and spatial reconstruction convolution aims to restore, adjust or enhance this spatial structure information in a certain way; for example, in some cases, due to multiple convolution and pooling operations, the spatial resolution is reduced and the details are lost. The spatial reconstruction convolution can try to restore some of the lost details or reorganize the spatial features to better represent the shape, position and other information of the target.
[0143] Channel reconstruction unit: focuses on reconstructing the channel dimension of the feature map; in convolutional neural networks, each channel can be regarded as a specific feature representation; the purpose of channel reconstruction convolution is to adjust the relationship between channels, such as reducing redundant channels, enhancing important channels, or generating new discriminative channels; through this reconstruction, the feature representation can be made more efficient and targeted to meet specific task requirements, such as classification, detection or segmentation.
[0144] The specific processing steps of spatial and channel reconstruction convolution (SSCONV) are:
[0145] (1) Receive a feature map whose shape is (C in , H in , W in ), and then upsample it via transposed convolution;
[0146] (2) Spatial reconstruction: Perform global average pooling and maximum pooling operations on the upsampled feature map and concatenate the two. Perform 1x1 convolution and sigmoid activation on the concatenated result to obtain a spatial attention weight map. Then, multiply the spatial attention weight map with the upsampled feature map to highlight important spatial areas.
[0147] (3) Channel reconstruction: Use convolution to adjust the channels of the input feature map without changing the spatial size; perform average pooling and maximum pooling on the feature map after adjusting the channel; concatenate the pooling results, and obtain the channel attention weight through convolution and activation; multiply the channel attention weight with the feature map after adjusting the channel to highlight the important channels.
[0148] (4) Add the results of spatial reconstruction and channel reconstruction to obtain the final output feature map.
[0149] 4. Edge Guided Interaction Module (EGI)
[0150] After capturing the salient features (i.e., regional features) and edge features, there is still a key issue, that is, how to use edge information to improve the representation ability of segmentation features; the common practice is to integrate these two features by cascading or element-by-element addition, which can easily lead to interference from the edges, noisy edge features or redundant information affecting the final detection accuracy of the model.
[0151] In order to alleviate this problem, this embodiment proposes an edge-guided interaction module, which improves and utilizes edge information through an edge-guided learning strategy. Based on the selectivity of edge features and the invariance of internal features, the two-stream structure (i.e., the edge branch and the saliency branch) learns boundary details and salient areas respectively, and introduces EGI to enhance salient features under the guidance of edge cues; based on the edge-guided learning strategy, the edge information is improved and utilized to process the cross-branch interaction of edge cues and saliency cues. It does not directly fuse, but regards the edge features as the guidance of lower-level cues, so that salient objects form clear boundaries, such as Fig. 9 As shown, the edge guided interaction module (EGI) is specifically:
[0152] (1) Input simultaneously to the module and First of all A 1x1 convolution operation is performed, and then the two results are multiplied by matrix multiplication and processed by the Softmax function to obtain an output result.
[0153] (2) A new output result is obtained through 1x1 convolution processing, and this result is multiplied with the output result obtained in step (1) to obtain a new result.
[0154] (3) Perform pixel-wise multiplication with the result obtained in step (2).
[0155] (4) Add the result in step (3) at pixel level to obtain the final output of the module.
[0156] 5. Gated Multi-Scale Predictor (GMS)
[0157] Fig.10 It is the flow chart of GMS, such as Fig.10 As shown, the specific process is as follows:
[0158] Model training phase:
[0159] (1) The output (network features) of the previous link is transmitted to this module and enters two branches at the same time. One branch does not perform any processing; the other branch enters the gate for processing.
[0160] (2) The gate can process the transmitted feature map into modules of five scales [56, 64, 72, 80, 88] (i.e., the blue three-dimensional figure in the flowchart), but only one of them is used during the training process, which is determined by the dataset size of the dataset itself.
[0161] (3) Both the unprocessed branches and the gated processed results are processed by a classifier, where a multi-layer perceptron classifier is used.
[0162] (4) Upsample the results processed by the classifier, add the branches that have been gated in the early stage, and then add the results of the branches that have not been gated to output a feature map. (The existing loss and GMS loss directly output by this module are not used here)
[0163] Model testing phase:
[0164] (1) The output (network features) of the previous link is transmitted to this module and enters two branches at the same time. One branch does not perform any processing; the other branch enters the gate for processing.
[0165] (2) Gating can process the transmitted feature map into modules with five scales of [56, 64, 72, 80, 88] (i.e., the blue three-dimensional figure in the flowchart). Here, the feature map is processed into five scales of [56, 64, 72, 80, 88] through gating.
[0166] (3) Both the unprocessed branches and the gated processed results are processed by a classifier, where a multi-layer perceptron classifier is used.
[0167] (4) Upsample the results processed by the classifier, add the branches that have been previously gated, and then add them to the results of the branches that have not been gated, and output the final result.
[0168] In the training process of the above edge-aware SOD network, the loss function used consists of two parts, namely the edge branch loss function and the significance branch loss function. After calculating the two separately, the final loss function is obtained by adding the two together.
[0169] The edge branch loss function is generated by comparing the edge results Bn output in different links with the standard edge result graph. Here, the category-balanced binary cross entropy loss function is used for calculation, and the five results are added together to obtain the final loss function of the branch.
[0170] The saliency branch loss function is generated by comparing the edge results Pn output in different links with the saliency recognition result graph. Here, the IoU loss function is used for calculation, and the three results are added to obtain the final loss function of the branch.
[0171] The final loss function is the sum of the edge branch loss function and the significance branch loss function.
[0172] Example 2
[0173] In one embodiment of the present disclosure, a salient object detection system based on edge detection and attention mechanism is provided, including:
[0174] The image acquisition module is configured to: acquire an image to be detected;
[0175] The object detection module is configured to: input the image into the trained object detection network, highlight the salient object area, and obtain the final salient object detection prediction map;
[0176] Among them, the target detection network constructs a parallel edge branch and a saliency detection branch based on two sets of independent features extracted from the image, respectively emphasizing the selectivity and invariance of features in detecting salient edges and salient regions; the edge branch interactively fuses low-level features with spatial structural details and high-level features with rich semantic knowledge to obtain edge features; the saliency detection branch generates a regional feature containing multi-scale key information through multi-scale attention, and uses the edge-guided learning strategy and the edge features as a guide to improve the interior and boundaries of the regional features to obtain the final salient target region.
[0177] Example 3
[0178] In one embodiment of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the method for detecting a salient object based on edge detection and an attention mechanism is implemented.
[0179] Example 4
[0180] In one embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the method for detecting a salient target based on edge detection and an attention mechanism is implemented.
[0181] Example 5
[0182] In one embodiment of the present disclosure, an electronic device is provided, including: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes the method for salient target detection based on edge detection and attention mechanism.
Claims
1. A salient object detection method based on edge detection and attention mechanism, characterized in that: include: Acquire an image to be detected; Input the image into the trained object detection network to highlight the salient object area and obtain the final salient object detection prediction map; Among them, the target detection network constructs a parallel edge branch and a saliency detection branch based on two sets of independent features extracted from the image, respectively emphasizing the selectivity and invariance of features in detecting salient edges and salient regions; the edge branch interactively fuses low-level features with spatial structural details and high-level features with rich semantic knowledge to obtain edge features; the saliency detection branch generates a regional feature containing multi-scale key information through multi-scale attention, and uses the edge-guided learning strategy and the edge features as a guide to improve the interior and boundaries of the regional features to obtain the final salient target region.
2. The method for salient object detection based on edge detection and attention mechanism as claimed in claim 1, characterized in that: The two sets of independent features are extracted using an encoder built on the basis of resnet-50 using 1x1 convolution and dilated spatial convolution pooling pyramid modules, specifically: Remove the last global pooling layer and fully connected layer from the original resnet-50 network, and use only five sequentially stacked residual blocks as the backbone feature extractor to extract the original side features of the corresponding level; The original side features are processed to generate low-level features with spatial structure details and high-level features with rich semantic knowledge, which are passed to two parallel branches.
3. The method for salient object detection based on edge detection and attention mechanism as claimed in claim 1, characterized in that: The edge branch mines edge features from low-level features and high-level features under the supervision of edge labels, and the saliency detection branch mines regional features from low-level features and high-level features under the supervision of regional labels.
4. The method for detecting salient objects based on edge detection and attention mechanism according to claim 1, wherein: The edge branch and the saliency detection branch adopt a cross-level fusion method, including a high-level interactive fusion module and a low-level weighted fusion module; The high-level interactive fusion module exploits the cross-layer correlation between two adjacent semantic features to provide the location information of salient objects; The low-level weighted fusion module selectively integrates low-level features and high-level features. To suppress redundant information and supplement important features.
5. The method for detecting salient objects based on edge detection and attention mechanism according to claim 1, wherein: The saliency detection branch also includes an efficient multi-scale attention module, which consists of a multi-scale feature extraction part, an attention mechanism part, and a feature fusion part. The attention mechanism is used to efficiently focus on key information of input data at different scales.
6. The method for salient object detection based on edge detection and attention mechanism as claimed in claim 1, characterized in that: The edge-guided learning strategy proposes an edge-guided interaction module to process cross-branch interactions of edge features and regional features based on convolution, matrix multiplication, pixel-level multiplication and pixel-level addition.
7. A salient object detection system based on edge detection and attention mechanism, characterized in that: include: The image acquisition module is configured to: acquire an image to be detected; The object detection module is configured to: input the image into the trained object detection network, highlight the salient object area, and obtain the final salient object detection prediction map; Among them, the target detection network constructs a parallel edge branch and a saliency detection branch based on two sets of independent features extracted from the image, respectively emphasizing the selectivity and invariance of features in detecting salient edges and salient regions; the edge branch interactively fuses low-level features with spatial structural details and high-level features with rich semantic knowledge to obtain edge features; the saliency detection branch generates a regional feature containing multi-scale key information through multi-scale attention, and uses the edge-guided learning strategy and the edge features as a guide to improve the interior and boundaries of the regional features to obtain the final salient target region.
8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for detecting a salient object based on edge detection and attention mechanism described in any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by the processor, the salient target detection method based on edge detection and attention mechanism as described in any one of claims 1-6 is implemented.
10. An electronic device, characterized in that: include: A processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes the salient target detection method based on edge detection and attention mechanism as described in any one of claims 1-6.
Citation Information
Cited By
Target detection method based on edge extraction and accurate positioning
CN120510406A
Automatic welding method of reinforcement cage welding robot
CN120551726A
Real-time target detection method and system for intelligent image processing
CN120765917A