A Salient Object Detection Method Based on Brain-Inspired Dual-Process CNN-Transformer Network
By adopting a brain-inspired dual-process CNN-Transformer network in significance object detection, combined with the pyramid attention U-Net and RECA-Transformer modules, the problem of insufficient feature redundancy and fine edge feature capture in the prior art is solved, and higher detection accuracy and performance are achieved.
Patent Information
- Application Number
- CN202411699171.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-11-26
AI Technical Summary
The prior art has feature redundancy in the detection of significance targets, making it difficult to accurately segment out significant targets or regions, ignores visual attribute information of channel dimensions, and the Transformer structure is insufficient in capturing fine edge features.
The dual-process CNN-Transformer network based on brain-inspired is adopted, combining the pyramid attention U-Net and RECA-Transformer modules, spatial features are extracted through the dual-branch pyramid U-Net module, and global and channel-level feature information is captured through the RECA-Transformer module, and boundary prediction is optimized using attention mechanism.
It significantly improves model performance and prediction accuracy, can better balance global information capture and boundary refinement processing, and generate more accurate significance goals.
Smart Images

Figure CN119741469B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and image processing in computer science, and particularly relates to a saliency object detection method based on a brain-inspired dual-process CNN-Transformer network. Background Art
[0002] Saliency object detection mimics the human eye visual system to perform object localization and segmentation in a visual scene to quickly find the most important object or region. It has been widely applied to various tasks in computer vision, such as object detection, semantic segmentation, AR / VR, graphic captioning, etc.
[0003] In existing saliency object detection methods, most methods still rely on the multi-scale feature extraction architecture of convolutional neural networks (CNNs). These methods capture the global structure and local details of an image through a multi-scale feature fusion network, thereby improving the accuracy of saliency prediction. However, the feature information obtained in this way often has redundancy, making it difficult to accurately segment the most salient object or region in subsequent saliency object detection. In addition, existing methods usually focus on the spatial features of an image and ignore the importance of visual attribute information in the channel dimension. In a deep network, channel-level features imply different visual attributes, and these attributes can provide a new information reference dimension to better support the saliency object detection task. Recently, the encoder-decoder architecture based on Transformer has attracted wide attention as an alternative to traditional methods. This type of architecture generates a query-driven global representation through attention to explain the relationship between an object and its surrounding image context.
[0004] More and more research shows that the hierarchical operation of the human visual perception system is similar to the structure of the ventral visual stream (VVS), and this theory has inspired the design of hybrid CNN-Transformer architectures. More and more evidence shows that brain recognition is based on the Dual-Process Detection Theory, which has had a profound impact on psychology and cognitive neuroscience. As an important framework, the Dual-Process Detection Theory aims to explain the cognitive mechanism in the human recognition and memory processes. In human visual saliency detection, when an object enters the visual scene, the brain will immediately initiate a series of complex processing steps. This process can be explained in detail by the Dual-Process Detection Theory. Specifically, as Figure 10As shown, the visual processing of the brain is divided into two stages: familiarity and recollection. When the primary visual cortex receives light information from an object, it quickly and unconsciously begins to process the basic features of the image, such as edges, orientations, and basic shapes. This initial processing is similar to the function of a convolutional neural network, which quickly extracts and identifies local features of an image through a hierarchical structure. Then, the brain forms a preliminary sense of recognition of the object, a process called familiarity. If more detailed information is needed, the brain will undergo a slower and more conscious recollection process, which involves the hippocampus and related memory regions. This stage requires integrating information from different sources to confirm the identity of the object and its importance in a specific context. This more complex process is similar to the way Transformers work, capturing long-range dependencies and comprehensively processing information through self-attention mechanisms and global feature modeling.
[0005] Salient object detection not only requires identifying salient objects in an image but also ensuring that the predicted objects have precise and clear boundaries. Simply relying on multi-scale feature fusion of CNNs can improve feature extraction to some extent but often leads to the loss of fine-grained features and cannot retain complete information. On the other hand, the Transformer structure can efficiently process and model global information and multi-scale features of images but still has deficiencies in capturing fine-grained edge features. Therefore, a method that combines the Transformer structure with multi-scale feature fusion of CNNs can better balance global information capture and boundary refinement, thus achieving better results in salient object detection. Summary of the Invention
[0006] The object of the present invention is to provide a salient object detection method based on a brain-inspired dual-process CNN-Transformer network, which solves the problems existing in existing CNN methods in salient object detection tasks, such as feature redundancy, difficulty in accurately segmenting salient objects or regions, ignoring visual attribute information in the channel dimension, and the existing Transformer structure, although capable of capturing global information, still has deficiencies in capturing fine-grained edge features.
[0007] The technical solution adopted by the present invention is as follows:
[0008] A salient object detection method based on a brain-inspired dual-process CNN-Transformer network, comprising the following steps:
[0009] S1: Construct a dataset for salient object detection and divide the dataset into a training set, a test set, and a validation set;
[0010] S2: Construct a network model based on Pyramid Attention U-Net and RECA-Transformer. The network model based on Pyramid Attention U-Net and RECA-Transformer includes a dual-branch Pyramid U-Net module and a Residual Efficient Channel Attention Transformer module (RECA-Transformer module). Among them, the dual-branch Pyramid U-Net module is an encoder-decoder architecture that includes an encoder part and a decoder part.
[0011] The encoder part of the dual-branch Pyramid U-Net module includes a pre-trained VGG16 network and a spatial attention module (SS). The pre-trained VGG16 network is used to gradually extract the feature information of the input image. At the same time, combined with the spatial attention module (SS), a pyramid structure is formed to more effectively extract the spatial feature representation X and visual feature V of the input image, and the spatial feature representation X is enhanced by combining the visual feature V.
[0012] The decoder part of the dual-branch Pyramid U-Net module uses transposed upsampling (TransposedUpsampling), bilinear upsampling (Bilinear Upsampling), a spatial context pooling module (Sup), and a residual coordinate attention module (Residual Coordinate Attention Model, i.e., RCAM). The transposed upsampling, bilinear upsampling, and the spatial context pooling module are used crosswise to retain the spatial information Y of the feature map. The residual coordinate attention module is used to generate the spatial attention feature map Y. s ;
[0013] The Residual Efficient Channel Attention Transformer (RECA-Transformer) includes a RECA transformation block (RECA Transformer Block). The RECA transformation block (RECA Transformer Block) performs feature extraction and captures global characteristics through a multi-layer attention mechanism and weight learning.
[0014] S3: Use the training set to train the network model based on Pyramid Attention U-Net and RECA-Transformer, use the validation set to adjust the model parameters, determine whether there is overfitting, and adopt the combination of the basic loss function BaseLoss and the edge attention loss function EdgeLoss as the attention boundary optimization loss function TotalLoss of the model to complete the training of the network model based on Pyramid Attention U-Net and RECA-Transformer; use the test set to test the network performance of the network model based on Pyramid Attention U-Net and RECA-Transformer;
[0015] S4: Adopt the trained network model based on Pyramid Attention U-Net and RECA-Transformer for salient object detection.
[0016] Further, the encoder part of the dual-branch pyramid U-Net module includes five stages, namely the first stage composed of the 2nd to 5th layers of the VGG16 network and the spatial attention module, the second stage composed of the 5th to 10th layers of the VGG16 network and the spatial attention module, the third stage composed of the 10th to 17th layers of the VGG16 network and the spatial attention module, the fourth stage composed of the 17th to 24th layers of the VGG16 network and the spatial attention module, and the fifth stage composed of the 24th to 31st layers of the VGG16 network and the spatial attention module; among them, the VGG16 network uses the pre-trained convolutional layer and pooling layer to extract the feature information of the image. In each convolutional stage, the size of the feature map is successively reduced to 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image; the spatial attention module consists of a convolutional layer and an effective channel attention module. The convolutional layer in the spatial attention module is used to extract the spatial feature representation X of the feature map, and the effective channel attention module is used to enhance the spatial feature representation.
[0017] Further, the expression of the convolutional layer in the spatial attention module is as follows:
[0018] X = (Conv2D 3×3 (ReLU(Conv2D 1×1 (x c ))))
[0019] where X represents the spatial feature representation of the feature map, Conv2D 3×3 (·) represents two-dimensional convolution with a kernel size of 3×3, ReLU(·) represents the ReLU activation function, Conv2D 1×1 (·) represents two-dimensional convolution with a kernel size of 1×1, and x cRepresents the input feature map of the spatial attention module;
[0020] The effective channel attention module includes a globally average pooling layer GAP, a one-dimensional convolutional layer, and a Sigmoid activation function stacked in sequence. The output of the Sigmoid activation function is multiplied by the input of the globally average pooling layer along the channels and used as the output of the effective channel attention module. The expression of the effective channel attention module is as follows:
[0021] X′ = ECA(X) = X · σ(Conv1D 1×1 (AvgPool(X))),
[0022] where X′ represents the enhanced feature, X represents the spatial feature representation of the feature map, ECA(·) represents the processing function of the effective channel attention module, σ(·) refers to applying the Sigmoid activation to the result of the one-dimensional convolution to obtain the channel attention weight, Conv1D 1×1 (·) represents one-dimensional convolution with a kernel size of 1×1, and AvgPool(·) represents the globally average pooling operation;
[0023] That is, the expression of the spatial attention module is as follows:
[0024] X′ = ECA(Conv2D 3×3 (ReLU(Conv2D 1×1 (x c ))))。
[0025] Furthermore, the decoder part of the dual-branch pyramid U-Net module contains six stages. The first stage of the decoder part includes a residual coordinate attention module (RCAM) and a pixel shuffle module (PixelShuffle); the second stage of the decoder part includes bilinear upsampling, a spatial context pooling module (Sup), a residual coordinate attention module (RCAM), and slicing and concatenation operations; the third, fourth, and fifth stages of the decoder part all include transposed upsampling, bilinear upsampling, a spatial context pooling module (Sup), a residual coordinate attention module (RCAM), and slicing and concatenation operations; the sixth stage of the decoder part includes transposed upsampling and slicing and concatenation operations.
[0026] Furthermore, the spatial context pooling module (Sup) first uses a 1×1 convolution operation to map the input feature map x of the spatial context pooling module into a single-channel feature map G. The dimension of the obtained feature map G is (N, 1, H, W). Subsequently, the H channel and the W channel of the feature map G are combined together, and the H and W dimensions of the feature map G are merged to obtain a feature map G′. The dimension of the feature map G′ is (N, 1, H×W); then, a Softmax operation is performed to calculate the weight of each position of the feature map G′, and the normalized feature map G′ after channel merging is applied to the input feature map x, and an enhanced feature map is obtained by element-wise multiplication.
[0027] Furthermore, the residual coordinate attention module (RCAM) includes a parallel horizontal global average pooling layer (XAvg Pool) and a vertical global average pooling layer (Y Avg Pool), a concatenation operation (Conact), a 3×3 convolution layer, a batch normalization layer (Batch Norm), a non-linear activation layer (Non-linear), followed by two parallel 1×1 convolution layers and two parallel Sigmoid activation functions. The outputs of the two Sigmoid activation functions are extended to the size of the input feature map, and then the extended feature map is combined with the input feature map of the residual coordinate attention module by element-wise multiplication, and the obtained output is added to the input feature map of the residual coordinate attention module through a residual connection to complete the channel expansion of the input feature map. The residual coordinate attention module enhances the representation ability of the feature map by combining spatial and channel attention mechanisms; its principle is to generate attention weights for each channel through global average pooling, 1×1 convolution, and Sigmoid activation, and apply them to the input feature map, so that the model can better capture local and global feature information.
[0028] The residual coordinate attention module is used to perform channel expansion on the input feature map. First, the input feature map of the residual coordinate attention module is globally averaged in the horizontal and vertical directions through a parallel horizontal global average pooling layer (X Avg Pool) and a vertical global average pooling layer (Y Avg Pool) to extract global average information and retain spatial information; subsequently, the horizontally and vertically pooled feature maps are concatenated in the second dimension to generate a new feature map for further processing and analysis. The specific workflow is as follows:
[0029] y = concat(x h , x w , dim = 2);
[0030] where x h and x wrespectively represent the pooling results in the vertical and horizontal directions, concat(·) represents the concatenation operation, and y represents the concatenated feature map;
[0031] Then, a 3×3 convolutional layer and a batch normalization layer are used to reduce the dimension of the concatenated feature map, and at the same time, non-linearity is introduced through the ReLU activation function:
[0032] y′ = ReLU(BN(Conv2D 3×3 )(y));
[0033] where y′ represents the feature map after dimensionality reduction, ReLU(·) represents the ReLU activation function, BN(·) represents the normalization process, and Conv2D 3×3 (·) represents a two-dimensional convolution with a kernel size of 3×3;
[0034] Next, the dimensionality-reduced feature map y′ is separated into a horizontal feature map x w1 and a vertical feature map x h1 , and attention weights are generated through a 1×1 convolutional layer and a Sigmoid activation function respectively:
[0035] x h1 ,x w1 = split(y′, [h, w], dim = 2);
[0036] x h2 = Sigmoid(Conv2D 1×1 (x h1 ));
[0037] x w2 = Sigmoid(Conv2D 1×1 (x w1 ));
[0038] where split(·) represents splitting the feature map in the second dimension, Sigmoid(·) represents the sigmoid activation function, x h2 represents the vertical feature map of the attention weight, and x w2 represents the horizontal feature map of the attention weight;
[0039] Then, the generated vertical feature map x h2 of the attention weight and the horizontal feature map x w2 of the attention weight are expanded to the size of the input feature map and combined with the input feature map through element-wise multiplication:
[0040] x h3 = x h2 ·expand(h, w);
[0041] x w3 = x w2 · expand(h, w);
[0042] f = identity · x h3 · x w3 ;
[0043] Among them, f represents the output result combined by element-wise multiplication, expand(h, w) represents expanding the height and width of the attention weight feature map to h and w, identity represents the initial input feature map of the residual coordinate attention module (RCAM), x h3 represents the expanded vertical feature map, and x w3 represents the expanded horizontal feature map; finally, the input and output features are combined through a residual connection.
[0044] Furthermore, the residual efficient channel attention transformer module (RECA-Transformer module) first performs preliminary processing on the input data through a feature patch partition module (Patch Partition). The input data is divided into small patches, and each small piece of input data is converted into a feature vector representation through a linear embedding module (Linear Embedding); the converted feature vector representations are successively processed through multiple consecutive RECA transformation blocks (RECA Transformer Block×2) and the linear embedding module; then, the spatial resolution of the feature map is reduced through a feature patch merging module (Patch Merging), and feature extraction processing is performed through two single RECA transformation blocks (RECA Transformer Block×1) at the top of the model; the extracted data features are adjusted in resolution through a patch expanding module (Patch Expanding) to gradually increase the spatial resolution of the feature map to achieve upsampling, thereby generating a high-resolution output feature; it is successively processed through multiple consecutive RECA transformation blocks (RECA Transformer Block×2) and the linear embedding module again, where the consecutive RECA transformation blocks at this stage will receive the feature information passed across layers from the consecutive RECA transformation blocks at the front end; the extracted data features are successively processed again and then the spatial resolution of the feature map is gradually reduced through the feature patch merging module (Patch Merging), while increasing the channel dimension to achieve efficient feature extraction and computational optimization; finally, the channel dimension of the feature map is adjusted using linear projection (Linear Projection) as the output to obtain the final feature map.
[0045] Among them, the RECA Transformer Block introduces the Effective Channel Attention (ECA) module and the Residual MLP (Re-MLP) module, aiming to enhance useful features and suppress useless features through adaptive weighting in the channel dimension. Specifically, the ECA module calculates the attention weights for each channel and weights the channel features according to these weights, enabling the model to focus more on key features and thus capture and utilize important information more accurately in complex visual tasks.
[0046] Furthermore, the consecutive RECA Transformer Block (RECA Transformer Block×2) contains two connected RECA Transformer Blocks. The RECA Transformer Block includes, in sequence, Layer Normalization (LN), Window-based Multi-Head Self-Attention (W-MSA), Layer Normalization (LN), the Residual MLP (Re-MLP) module, and the ECA module. Among them, the input of the first Layer Normalization (LN) is added to the output of the Window-based Multi-Head Self-Attention (W-MSA) as the input of the second Layer Normalization (LN), and the input of the second Layer Normalization (LN) is added to the output of the ECA module as the output of the RECA Transformer Block; in the consecutive RECA Transformer Block (RECA Transformer Block×2), the output of the first RECA Transformer Block is used as the input of the second RECA Transformer Block.
[0047] In addition, a residual connection is introduced in the Residual MLP (Re-MLP) module. The Residual MLP module includes, in sequence, the ECA module, the fully connected layer fc1, the fully connected layer fc2, and the fully connected layer fc3. The residual connection acts on the fully connected layers fc1 and fc2, and the output of the fully connected layer fc1 is added to the output of the fully connected layer fc2 as the input of the fully connected layer fc3; the residual connection not only strengthens the roles of the two fully connected layers fc1 and fc2 but also effectively alleviates the problem of gradient disappearance; by introducing the residual connection, the gradients in the model can flow smoothly during the backpropagation process, promoting the effective propagation of gradients.
[0048] Furthermore, the calculation formula of the attention boundary optimization loss function TotalLoss, which is composed of the basic loss function BaseLoss and the boundary attention loss function EdgeLoss, is as follows:
[0049] TotalLoss = BaseLoss + β·EdgeLoss;
[0050]
[0051] AttentionWeights = 1 + α·Edges;
[0052]
[0053] Among them, BaseLoss is the basic loss, which calculates the cross-entropy loss between the predicted output image of the model and the true label (targets). The edge map (edges) is generated by calculating the gradient components of the true mask map in the vertical and horizontal directions, so as to obtain a high-quality edge map. β is the weight of the edge loss, EdgeLoss is the edge optimization loss, and CrossEntropy(p i ,t i ) represents the cross-entropy loss function, p i represents the predicted value, t i represents the actual value, N represents a total of N points, AttentionWeights represents the attention weight, α represents the proportion of the edge weight, and Edges represents the edge.
[0054] The beneficial effects of the present invention are as follows:
[0055] (1) Based on the dual-process detection theory, the present invention collaboratively applies CNN and Transformer to salient object detection, forming a hybrid architecture based on the pyramid attention U-Net and the residual efficient channel attention transformer model. The dual-branch pyramid attention U-Net module (CNN module) is responsible for extracting extensive and general shallow features, and the residual efficient channel attention transformer module (Transformer module) generates high-level representations for specific tasks and captures global relationships through the attention mechanism. By combining the advantages of both, the present invention significantly improves the model performance and prediction accuracy.
[0056] (2) The present invention combines visual features through the spatial attention module and the residual coordinate attention module, enhancing the spatial feature representation and balancing the model's attention to regions of different sizes. In addition, the present invention constructs local region features that incorporate visual attribute information in the channel dimension, which helps the model retain more local image context information.
[0057] (3) The present invention constructs a residual efficient channel attention transformer module (RECA-Transformer module) and an attention boundary optimization loss function, which are respectively used to capture the global context information of the image and perform the image boundary prediction task. The method of the present invention can effectively solve the problem that it is difficult to achieve the integrity of fine edge features in the prior art during saliency target prediction, and generate more accurate saliency targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0059] Figure 1 It is a structural diagram of the network model based on the pyramid attention U-Net and RECA-Transformer in the embodiment of the present invention;
[0060] Figure 2 It is a schematic structural diagram of the residual coordinate attention module in the embodiment of the present invention;
[0061] Figure 3 It is a schematic structural diagram of the effective channel attention module (ECA) in the embodiment of the present invention;
[0062] Figure 4 It is a schematic structural diagram of the Re-MLP module in the embodiment of the present invention;
[0063] Figure 5 It is a schematic structural diagram of the continuous RECA conversion block (RECA-Transformer Block×2) in the embodiment of the present invention;
[0064] Figure 6 It is a Precision-Recall evaluation index diagram in the embodiment of the present invention;
[0065] Figure 7 It is a comparison visualization result diagram of the method of the present invention and ten other existing saliency target detection methods;
[0066] Figure 8 It is a detection result diagram of the module ablation visualization of the method of the present invention;
[0067] Figure 9 It is a simplified flowchart of the network framework built based on brain inspiration in the present invention.
[0068] Figure 10 It is a schematic structural diagram of the ventral visual stream (VVS) of the dual process of brain recognition. DETAILED DESCRIPTION
[0069] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention can also be implemented in other ways different from those described herein, and therefore, the present invention is not limited to the limitations of the specific embodiments disclosed below.
[0070] Unless otherwise defined, the technical terms or scientific terms used herein should be understood by people with ordinary skills in the field described in this application. The "first", "second" and similar words used in this application do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, similar words such as "one" or "one" do not indicate quantity restrictions, but indicate the existence of at least one. Similar words such as "connect" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship also changes accordingly.
[0071] like Figure 1 As shown in FIG. 1 , a salient object detection method based on a brain-inspired dual-process CNN-Transformer network comprises the following steps:
[0072] Step S1: construct a data set for salient object detection, and divide the data set into a training set, a test set, and a validation set. In the embodiment of the present invention, five data sets widely used in the field of salient object detection are used, namely, the ECSSD data set, the PASCAL-S data set, the HKUIS data set, the DUT-OMRON data set, and the DUTS data set, and the data in the five data sets are divided into a training set, a test set, and a validation set according to a ratio of 7:2:1.
[0073] The ECSSD dataset contains 1,000 high-resolution images with complex backgrounds and diverse scenes. The salient objects in each image are of various types and have prominent locations. Compared with other datasets, the ECSSD dataset has precise annotations and is suitable for evaluating the performance of salient object detection algorithms when processing complex scenes.
[0074] The PASCAL-S dataset contains 850 images with intricate scenes, diverse backgrounds, and salient objects covering multiple categories. Each image is carefully annotated to ensure accurate representation of the objects. Compared with the ECSSD dataset, the PASCAL-S dataset has slightly fewer images, but it poses higher requirements for the adaptability of algorithms in different backgrounds.
[0075] The HKUIS dataset contains 4447 images, which are challenging due to the presence of multiple unconnected salient objects, overlapping boundaries, and low color contrast. These characteristics make it particularly valuable for testing the robustness of saliency object detection models. Compared with the ECSSD dataset and the PASCAL-S dataset, the HKUIS dataset is larger in quantity and focuses more on the detection effect in difficult scenarios.
[0076] The DUT-OMRON dataset contains 5168 images, and each image shows one or two salient objects with rich visual features. Compared with the HKUIS dataset, the DUT-OMRON dataset has higher object richness and diversity, making it suitable for comprehensively evaluating the performance of detection methods under complex object conditions.
[0077] The DUTS dataset contains 15572 images, which is the largest publicly available dataset in the field of saliency object detection. This dataset is divided into a training set (DUTS-TR, 10553 images) and a test set (DUTS-TE, 5019 images), covering a variety of scenes and object categories. Compared with the ECSSD dataset, the PASCAL-S dataset, the HKUIS dataset, and the DUT-OMRON dataset, the DUTS dataset has obvious advantages in terms of data scale and scene diversity, and can comprehensively test the generalization ability of the model.
[0078] Step S2: Construct a network model based on the Pyramid Attention U-Net and RECA-Transformer. The network model based on the Pyramid Attention U-Net and RECA-Transformer includes a dual-branch pyramid U-Net module and a Residual Efficient Channel Attention Transformer module (RECA-Transformer module). Among them, the dual-branch pyramid U-Net module is an encoder-decoder architecture that includes an encoder part and a decoder part.
[0079] The encoder part of the dual-branch pyramid U-Net module uses the pre-trained VGG16 network as the backbone network and combines the spatial attention module (SS) to form a pyramid structure; the convolutional layer and pooling layer in the pre-trained VGG16 network are used to gradually extract the feature information of the input image. At the same time, the pre-trained VGG16 network combines the spatial attention module (SS) to form a pyramid structure, so as to more effectively extract the visual feature V of the input image. Among them, the spatial attention module (SS) consists of a convolutional layer and an efficient channel attention module (ECA), which is used to enhance the spatial feature representation X by combining the visual feature V.
[0080] The decoder part of the dual-branch pyramid U-Net module contains six stages, in which transposed upsampling, bilinear upsampling, spatial context pooling module (Sup) and residual coordinate attention module (Residual Coordinate Attention Model, namely RCAM) are used; the transposed upsampling, bilinear upsampling and spatial context pooling module are used crosswise to retain the spatial information y of the feature map; the residual coordinate attention module is used to generate the spatial attention feature map Y. s 。
[0081] The residual efficient channel attention transformer module (RECA-Transformer module) introduces the efficient channel attention module (Efficient Channel Attention, namely ECA) and the residual MLP module (Residual-MLP, namely Re-MLP module) on the basis of the Swin-Transformer backbone network.
[0082] In the encoder part of the dual-branch pyramid U-Net module used in the embodiments of the present invention, a pre-trained VGG16 network is adopted as the backbone network, which is divided into five stages. The first stage consists of the second to fifth layers of the VGG16 network and a spatial attention module. The second stage consists of the fifth to tenth layers of the VGG16 network and a spatial attention module. The third stage consists of the tenth to seventeenth layers of the VGG16 network and a spatial attention module. The fourth stage consists of the seventeenth to twenty-fourth layers of the VGG16 network and a spatial attention module. The fifth stage consists of the twenty-fourth layer to the last layer (i.e., the thirty-first layer) of the VGG16 network and a spatial attention module. This VGG16 network uses pre-trained convolutional layers and pooling layers to gradually extract the feature information of the image. In each convolutional stage, the size of the feature map decreases sequentially, gradually decreasing from the original image to 1 / 2 (stage 1), 1 / 4 (stage 2), 1 / 8 (stage 3), 1 / 16 (stage 4), and 1 / 32 (stage 5), forming a pyramid structure. As Figure 1 shown, the input feature map first undergoes feature extraction by the second to fifth layers of the VGG16 network and the spatial attention module in the first stage of the encoder part respectively, and the extracted feature information is fused to obtain the feature information g1, and then it is input to the second stage of the encoder part for processing; the fifth to tenth layers of the VGG16 network and the spatial attention module in the second stage of the encoder part respectively perform feature extraction on the feature information g1, and the extracted feature information is fused to obtain the feature information g2, and then it is input to the third stage of the encoder part for processing; the tenth to seventeenth layers of the VGG16 network and the spatial attention module in the third stage of the encoder part respectively perform feature extraction on the feature information g2, and the extracted feature information is fused to obtain the feature information g3, and then it is input to the fourth stage of the encoder part for processing; the seventeenth to twenty-fourth layers of the VGG16 network and the spatial attention module in the fourth stage of the encoder part respectively perform feature extraction on the feature information g3, and the extracted feature information is fused to obtain the feature information g4, and then it is input to the fifth stage of the encoder part for processing; the twenty-fourth layer to the last layer (i.e., the thirty-first layer) of the VGG16 network and the spatial attention module in the fifth stage of the encoder part respectively perform feature extraction on the feature information g4, and the extracted feature information is fused to obtain the feature information g5; that is, the five stages of the encoder part sequentially output the feature information g1, feature information g2, feature information g3, feature information g4, and feature information g5 with gradually decreasing sizes.
[0083] The spatial attention module (SS) consists of a convolutional layer and an efficient channel attention module (ECA). The convolutional layer of the spatial attention module is used to extract the spatial feature representation X of the feature map, and the efficient channel attention module aims to enhance the spatial feature representation X. As Figure 3 shown, the efficient channel attention module (ECA) includes a global average pooling layer GAP, a one-dimensional convolutional layer, and a Sigmoid activation function stacked in sequence, and the output of the Sigmoid activation function is multiplied by the input of the global average pooling layer along the channels as the output of the efficient channel attention module;
[0084] The expression of the convolutional layer of the spatial attention module is as follows:
[0085] X = (Conv2D 3×3 (ReLU(Conv2D 1×1 (x c ))))
[0086] where X represents the spatial feature representation of the feature map, Conv2D 3×3 (·) represents two-dimensional convolution with a kernel size of 3×3, ReLU(·) represents the ReLU activation function, Conv2D 1×1 (·) represents two-dimensional convolution with a kernel size of 1×1, and x c represents the input feature map of the spatial attention module;
[0087] The expression of the efficient channel attention module is as follows:
[0088] X' = ECA(X) = X·σ(Conv1D 1×1 (AvgPool(X)))
[0089] where X' represents the feature after enhancement processing, X represents the spatial feature representation of the feature map, ECA(·) represents the processing function of the efficient channel attention module, σ(·) refers to performing Sigmoid activation on the result of the one-dimensional convolution to obtain the channel attention weight, Conv1D 1×1 (·) represents one-dimensional convolution with a kernel size of 1×1, and AvgPool(·) represents the global average pooling operation.
[0090] That is, the expression of the spatial attention module of the present invention is as follows:
[0091] X' = ECA(Conv2D 3×3 (ReLU(Conv2D 1×1 (x c ))))
[0092] In the embodiment of the present invention, the decoder part of the dual-branch pyramid U-Net module includes six stages. As Figure 1 shown, the structure of the decoder part is introduced in detail as follows. In the first stage of the decoder part, in view of the lack of corresponding high-level features, the feature information g5 output by the fifth stage (Stage5) of the encoder part is processed first. The feature information g5 obtained from the fifth stage of the encoder part of the dual-branch pyramid U-Net module is input into the first stage of the decoder part. The structure of the first stage of the decoder part is as Figure 1 shown in (a) of 5 . The first stage of the decoder part is processed by a Residual Coordinate Attention Model (RCAM) and a PixelShuffle module to obtain the feature information S
[0093] . Among them, the basic principle of the Residual Coordinate Attention Model (RCAM) is to improve the representation ability of the feature map by coordinating the spatial and channel attention mechanisms; the Residual Coordinate Attention Model uses global average pooling, 1×1 convolution, and the Sigmoid activation function to generate attention weights for each channel and apply them to the input feature map, so as to effectively capture local and global features.
[0093] The obtained feature information S 5 is input into the second stage of the decoder part for processing. The structure of the second stage of the decoder part is as Figure 1 shown in (b) of 5 . The second stage of the decoder part first processes the feature information S 5 through Bilinear Upsampling and a Spatial Context Pooling module (Sup), and then sums it with the feature information g4 output by the fourth stage (Stage4) of the encoder part to obtain the feature information R 21 . Then, the feature information R 21 is input into the Residual Coordinate Attention Model (RCAM) for processing to obtain the feature information R 22 . And then, the feature information S 5 input into the first stage of the decoder part and the feature information R 22 are sliced and concatenated (Slicing and Concatenation) to obtain the feature information R 23 . Finally, the feature information R 23 and the feature information S 5 are summed to output the feature information S 4 processed by the second stage of the decoder part.
[0094] The feature information S 4 enters the third stage of the decoder part for processing. The structure of the third stage of the decoder part is as Figure 1As shown in (c), first perform transposed upsampling on the feature information S 4 to obtain the feature information R 31 . At the same time, perform bilinear upsampling and spatial context pooling module (Sup) processing on the feature information g3 transmitted from the third stage (Stage3) of the encoder part to obtain the feature information R 32 . Then, sum the feature information R 32 with the feature information g4 transmitted from the fourth stage (Stage4) of the encoder part to obtain the fused feature information R 33 . Then, use the residual coordinate attention module (RCAM) to perform spatial enhancement on the fused feature information R 33 to obtain the feature information R 34 . Then, slice and concatenate the feature information R 34 with the feature information R 31 to obtain the feature information R 35 . This processing method enables the model to fully consider high-level and low-level feature information at different scales, thereby ensuring that the finally generated feature map has high resolution and high quality. Finally, the output feature information R 35 is summed with the feature information R 31 again to output the feature information S after the third stage of the decoder part 3 .
[0095] Feature information S 3 enters the fourth stage of the decoder part for processing. The structure of the fourth stage of the decoder part is as shown in (c) of Figure 1 . The processing process is the same as that of the third stage of the decoder part. First, perform transposed upsampling on the feature information S 3 to obtain the feature information R 41 . At the same time, perform bilinear upsampling and spatial context pooling module (Sup) processing on the feature information g2 transmitted from the second stage (Stage2) of the encoder part to obtain the feature information R 42 . Then, sum the feature information R 42 with the feature information g3 transmitted from the third stage (Stage3) of the encoder part to obtain the fused feature information R 43 . Then, perform spatial enhancement through the residual coordinate attention module (RCAM) to obtain the feature information R 44 . Then, the feature information R 44 is combined with the feature information R 41Perform slicing and concatenation to obtain the feature information R 45 The finally output feature information R 45 Once again, sum with the feature information R 41 to output the feature information S after the fourth stage of processing in the decoder part 2 .
[0096] Feature information S 2 Enters the fifth stage of the decoder part. The structure of the fifth stage of the decoder part is as shown in Figure 1 (c). Similarly, first perform transposed upsampling on the feature information S 2 to obtain the feature information R 51 . At the same time, perform bilinear upsampling and spatial context pooling module (Sup) processing on the feature information g1 transmitted from the first stage (Stage1) of the encoder part to obtain the feature information R 52 . Then sum the feature information R 52 with the feature information g2 transmitted from the second stage (Stage2) of the encoder part to obtain the feature information R 53 . Then perform spatial enhancement through the residual coordinate attention module (RCAM) to obtain the feature information R 54 . Then sum the feature information R 54 with the feature information R 51 to perform slicing and concatenation (Slicing and Concatenation) to obtain the feature information R 55 . The finally output feature information R 55 Once again, sum with the feature information R 51 to output the feature information S after the fifth stage of processing 1 .
[0097] The obtained feature information S 1 is then input to the sixth stage of the decoder part for processing, as shown in Figure 1 (d). The sixth stage of the decoder part first performs transposed upsampling on the feature information S 1 to obtain the feature information R 61 . Then slice and concatenate (Slicing and Concatenation) the feature information R 61 with the feature information g1 transmitted from the first stage (Stage1) of the encoder part to obtain the feature information after the encoder-decoder architecture processing of the double-branch pyramid U-Net module
[0098] Among them, the working principle of the spatial context pooling module (Sup) is as follows: First, apply a 1×1 convolution operation to the input feature map x of the spatial context pooling module to map it to a single-channel feature map G, and the dimension of the output feature map G is (N, 1, H, W). Subsequently, combine the H channel and the W channel together, merge the H and W dimensions of the feature map G, that is, (N, 1, H×W), to obtain the feature map G'. Then perform the Softmax operation to calculate the weight at each position of the feature map G' as described below:
[0099] G′ = C reshape = reshape(G, (N, 1, H×W));
[0100] A = Softmax(C reshape , dim = 2);
[0101] Among them, reshape(·) represents resizing the feature map, C reshape , dim = 2 means processing in the second dimension, Softmax(·) represents performing the Softmax operation, and A represents the feature map obtained after normalization by the Softmax operation.
[0102] The purpose of the Softmax operation is to normalize the weights so that the sum of the weights at all positions on the feature map is 1, and the weight corresponding to each position represents the importance of the corresponding position. Specifically, the Softmax operation normalizes the feature map G' after channel merging and applies it to the original input feature map x (i.e., the input feature map of the spatial context pooling module), and obtains the enhanced feature map by element-wise multiplication. This process can be summarized as follows:
[0103] A = Softmax(C reshape , dim = 2) = torch.matmul(x.unsqueeze(1), G);
[0104] In the above formula, A represents the feature map obtained after normalization by the Softmax operation, torch.matmul(·) represents matrix multiplication, unsqueeze(1) represents removing the first dimension, and x represents the input feature map of the spatial context pooling module.
[0105] In the embodiment of the present invention, the Residual Coordinate Attention Module (RCAM) uses a cyclic convolutional attention mechanism. The Residual Coordinate Attention Module (RCAM) can effectively model the long-range dependencies in the feature map, improve the model's context understanding ability, and enhance the confidence of the predicted image. By capturing these dependencies, the Residual Coordinate Attention Module (RCAM) not only maintains the fine boundary structure but also strengthens the network's ability to aggregate and interpret the spatial and context information of each region of the input image, thus achieving stronger and more reliable predictions. As Figure 2 shown, the Residual Coordinate Attention Module includes a horizontal global average pooling layer (X Avg Pool) and a vertical global average pooling layer (Y Avg Pool) in parallel, a concatenation operation (Conact), a 3×3 convolutional layer, a batch normalization layer (Batch Normalization, i.e., Batch Norm), a non-linear activation layer (Non-linear), followed by two 1×1 convolutional layers in parallel and two Sigmoid activation functions in parallel. The outputs of the two Sigmoid activation functions are extended to the size of the input feature map, and then the extended feature map is combined with the input feature map of the Residual Coordinate Attention Module through element-wise multiplication, and the obtained output is added to the input feature map of the Residual Coordinate Attention Module through a residual connection to complete the channel expansion of the input feature map. Using a 3×3 depth convolutional layer to perform spatial feature extraction on each channel can effectively extract spatial features and maintain computational efficiency, thereby enhancing the spatial expression ability of the feature map.
[0106] Specifically, first, the input feature map of the Residual Coordinate Attention Module is subjected to global average pooling in the horizontal and vertical directions respectively through the horizontal global average pooling layer (X Avg Pool) and the vertical global average pooling layer (Y Avg Pool) in parallel to extract global average information and retain spatial information. Subsequently, the horizontally and vertically pooled feature maps are concatenated in the second dimension to generate a new feature map for further processing and analysis. The specific workflow is as follows:
[0107] y = concat(x h , x w , dim = 2);
[0108] where x h represents the pooling result in the vertical direction, x w represents the pooling result in the horizontal direction, concat(·) represents the concatenation operation, and y represents the concatenated feature map. Then, a 3×3 convolutional layer and a batch normalization layer are used to reduce the dimension of the concatenated feature map, and at the same time, a ReLU activation function is used to introduce non-linearity:
[0109] y' = ReLU(BN(Conv2D 3×3 )(y));
[0110] where y' represents the feature map after dimensionality reduction, ReLU(·) represents the ReLU activation function, BN(·) represents batch normalization, and Conv2D 3×3 (·) represents two-dimensional convolution with a kernel size of 3×3.
[0111] Next, the feature map y' after dimensionality reduction is split into a horizontal feature map x w1 and a vertical feature map x h1 , and attention weights are generated through a 1×1 convolutional layer and the Sigmoid activation function respectively:
[0112] x h1 , x w1 = split(y', [h, w], dim = 2);
[0113] x h2 = Sigmoid(Conv2D 1×1 (x h1 ));
[0114] x w2 = Sigmoid(Conv2D 1×1 (x w1 ));
[0115] where split(·) represents splitting the feature map in the second dimension, Sigmoid(·) represents the sigmoid activation function, x h2 represents the vertical feature map of the attention weight, and x w2 represents the horizontal feature map of the attention weight.
[0116] Then, the generated vertical feature map x h2 of the attention weight and the horizontal feature map x w2 of the attention weight are expanded to the size of the input feature map and combined with the input feature map through element-wise multiplication:
[0117] x h3 = x h2 ·expand(h, w);
[0118] x w3 = x w2 ·expand(h, w);
[0119] f = identity · x h3 · x w3 ;
[0120] Among them, f represents the output result combined by element-wise multiplication, expand(h, w) represents expanding the height and width of the attention weight feature map to h and w, identity represents the original input feature map of the Residual Coordinate Attention Module (RCAM), and x h3 represents the expanded vertical feature map, and x w3 represents the expanded horizontal feature map; finally, the input and output features are combined through a residual connection.
[0121] The structure of the Residual Efficient Channel Attention Transformer Module (RECA-Transformer Module) in the embodiments of the present invention is as Figure 1 shown. First, the input data is preliminarily processed by a Feature Patch Partition Module (Patch Partition). The input data is segmented into small patches, and each small piece of input data is converted into a feature vector representation through a Linear Embedding Module (Linear Embedding). The converted feature vector representations are successively processed through multiple consecutive RECA Transformation Blocks (RECA Transformer Block×2) and Linear Embedding Modules level by level. Then, the spatial resolution of the feature map is reduced through a specific Feature Patch Merging Module (Patch Merging), and it is processed through two single RECA Transformation Blocks (RECA Transformer Block×1) at the top of the model. Among them, the RECA Transformation Block (RECA Transformer Block) is used to perform feature extraction on the feature vector representation, and capture global characteristics through a multi-layer attention mechanism and weight learning. Multiple consecutive RECA Transformation Blocks and Linear Embedding Modules gradually further optimize the features, ensuring that important information can be gradually extracted through multi-level processing. The extracted data features are adjusted in resolution through a specific Patch Expanding Module (Patch Expanding), gradually increasing the spatial resolution of the feature map to achieve upsampling, so as to generate high-resolution output features. It is processed through multiple consecutive RECA Transformation Blocks (RECA Transformer Block×2) and Linear Embedding Modules level by level again. Among them, the consecutive RECA Transformation Blocks at this stage will receive the feature information passed across layers from the consecutive RECA Transformation Blocks in the front end, enabling the feature information extracted in the early stage to be effectively accessed by subsequent modules, improving the overall modeling ability. The data features extracted through level-by-level processing are gradually reduced in the spatial resolution of the feature map through a specific Feature Patch Merging Module (Patch Merging), while increasing the channel dimension, realizing efficient feature extraction and computational optimization. Finally, the channel dimension of the feature map is adjusted using Linear Projection (LinearProjection) as the output to obtain the final feature map.
[0122] Among them, the RECA Transformer Block includes an Efficient Channel Attention (ECA) module and a Residual MLP (Re-MLP) module, which enhance useful features and suppress useless features through adaptive weighting in the channel dimension. Specifically, the ECA module calculates the attention weights for each channel and weights the channel features according to these weights, enabling the model to focus more on important features and thus more accurately capture and utilize key information when dealing with complex vision tasks. The RECA Transformer Block×2 module in this embodiment represents consecutive RECA Transformer Blocks, containing two connected RECA Transformer Blocks; RECA Transformer Block×1 represents a single RECA Transformer Block. As Figure 5 shown, the RECA Transformer Block includes, connected in sequence, Layer Normalization (LN), Window-based Multi-Head Self-Attention (W-MSA), Layer Normalization (LN), a Residual MLP (Re-MLP) module, and an Efficient Channel Attention (ECA) module. Among them, the input of the first Layer Normalization (LN) is added to the output of the Window-based Multi-Head Self-Attention (W-MSA) and used as the input of the second Layer Normalization (LN), and the input of the second Layer Normalization (LN) is added to the output of the Efficient Channel Attention (ECA) module and used as the output of the RECA Transformer Block. At the same time, in the consecutive RECA Transformer Block×2, the output of the first RECA Transformer Block is used as the input of the second RECA Transformer Block. In addition, a residual connection is introduced in the Residual MLP (Re-MLP) module. As Figure 4 shown, the Residual MLP module includes, connected in sequence, an Efficient Channel Attention (ECA) module, three fully connected layers fc1, fc2, and fc3. The residual connection acts on the fully connected layers fc1 and fc2, and the output of the fully connected layer fc1 is added to the output of the fully connected layer fc2 and used as the input of the fully connected layer fc3. This not only enhances the roles of the two fully connected layers fc1 and fc2 but also effectively alleviates the problem of gradient disappearance, promotes the smooth flow of gradients during backpropagation, and improves the training effect of the model.
[0123] Step S3: Use the training set to train the network model based on the Pyramid Attention U-Net and RECA-Transformer constructed in Step S2. Use the validation set to adjust the model parameters, determine whether there is overfitting, and adopt the combination of the basic loss function BaseLoss and the edge attention loss function EdgeLoss as the attention boundary optimization loss function TotalLoss of the model to complete the training of the network model based on the Pyramid Attention U-Net and RECA-Transformer; use the test set to test the network performance of the network model based on the Pyramid Attention U-Net and RECA-Transformer; the calculation formula of the attention boundary optimization loss function TotalLoss composed of the basic loss function BaseLoss and the edge attention loss function EdgeLoss is as follows:
[0124] TotalLoss = BaseLoss + β·EdgeLoss;
[0125]
[0126] AttentionWeights = 1 + α·Edges;
[0127]
[0128] Among them, BaseLoss is the basic loss function, which calculates the cross-entropy loss between the predicted output image of the model and the true label (targets). The edge map (edges) is generated by calculating the gradient components of the true mask map in the vertical and horizontal directions, so as to obtain a high-quality edge map. β is the weight of the edge loss, and EdgeLoss is the edge attention loss function. CrossEntropy(p i ,t i ) represents the cross-entropy loss function, p i represents the predicted value, t i represents the actual value, N represents a total of N points, AttentionWeights represents the attention weight, α represents the proportion of the edge weight, and Edges represents the edge.
[0129] Step S4: Use the trained network model based on the Pyramid Attention U-Net and RECA-Transformer to perform saliency target detection on the image to be detected.
[0130] The following combines specific experimental data to illustrate the technical effects of the embodiments of the present invention:
[0131] In the training process of the embodiments of the present invention, the PyTorch 2.3 framework is used, and the hardware configuration includes an i9-13900K CPU and an RTX 3090 GPU (24GB). The optimizer selects Stochastic Gradient Descent (SGD), the initial learning rate is set to 0.005, and the momentum is 0.9. The batch size is 15. The training, test, and validation data are divided according to the ratio of 7:2:1. During the training process, the image size is adjusted to 352×352. Note that the parameter settings of the boundary optimization loss function TotalLoss are β = 2 and α = 1.5.
[0132] In the experiment, five commonly used saliency object detection evaluation metrics are used to evaluate the performance of the model, namely F-measure, Enhanced-alignment Measure (E-measure), Mean Absolute Error (MAE), Precision-Recall, and Structure-measure (S-measure). The detailed descriptions of the evaluation metrics are as follows:
[0133] (1) F-measure is a comprehensive metric for evaluating the performance of the model, which combines Precision and Recall to balance the trade-off between them. The calculation formula is as follows:
[0134]
[0135] Among them, F β represents the result value of the F-measure, and β 1 is the trade-off parameter, usually set to 1, indicating the balance between Precision and Recall. The higher the result value of the F-measure, the better the overall performance of the model in detecting saliency objects.
[0136] (2) Enhanced-alignment Measure (E-measure) aims to evaluate the structural similarity between the predicted saliency map and the ground truth saliency map. It comprehensively considers the pixel-level similarity and the overall structural information. The calculation formula is as follows:
[0137]
[0138] Among them, E φ represents the result value of the Enhanced-alignment Measure, and P φ and R φ represent the adjusted Precision and Recall respectively. E-measure pays more attention to global consistency in saliency detection evaluation, and Eφ The larger the value, the higher the structural similarity between the detection result and the true target.
[0139] (3) Mean Absolute Error (MAE) is used to measure the pixel difference between the predicted saliency map and the true saliency map. Its calculation formula is as follows:
[0140]
[0141] Where P(i, j) represents the pixel value of the predicted saliency map, G(i, j) represents the pixel value of the true saliency map, and H×W represents the total number of pixels in the image. The lower the value of MAE, the closer the predicted saliency map by the model is to the true saliency map.
[0142] (4) Precision-Recall curve is used to evaluate the performance of the model at different thresholds. Precision represents the proportion of pixels that are actually significant targets among the pixels predicted as significant regions by the model, and Recall represents the proportion of all actual significant targets that are correctly detected. By plotting the Precision-Recall curve at different thresholds, the performance of the model can be visually evaluated.
[0143] (5) Structure-measure (S-measure) is used to measure the structural similarity between the saliency map and the true saliency map. S-measure combines Region-aware similarity and Object-aware similarity, and its calculation formula is as follows:
[0144] S α =α 1 ×S o +(1 - α 1 )×S r ;
[0145] Where S α represents the result value of the structure similarity measure, S o represents the object similarity, S r represents the region similarity, and α 1 is the weight coefficient, usually set to 0.5. The larger the result value of S-measure, the more consistent the predicted saliency map by the model is with the true saliency map in terms of structure.
[0146] To verify the effectiveness of the method, the proposed saliency object detection method of the present invention was tested on five datasets (i.e., ECSSD dataset, PASCAL-S dataset, HKU-IS dataset, DUT-OMRON dataset, DUTS dataset), and compared with 10 recently published neural network-based saliency object detection methods based on five evaluation metrics (i.e., F-measure, E-measure, MAE, Precision-Recall, S-measure). The existing comparison methods include F3Net, LDF, MSFNet, PAKRN, SelfReformer, EDN, BBRF, MENet, CSD and A2S-v3. The saliency images of all comparison methods were provided by the authors of each method or generated by running their publicly available code. As shown in Table 1, in terms of the four evaluation metrics of F-measure, MAE, E-measure and S-measure (where the higher the F-measure and S-measure values, the better; the lower the MAE value, the better; the higher the E-measure value, the better), the method of the present invention is generally superior to all comparison methods, especially outstanding on the ECSSD dataset, PASCAL-S dataset and DUT-OMRON dataset; on the HKU-IS dataset and DUTS dataset, the method of the present invention is not much different from the method with the best performance in other evaluation metrics. Precision-Recall is also a commonly used evaluation metric in saliency object detection, which is used to measure the detection accuracy and balance ability of the saliency object detection model at different thresholds, and is particularly suitable for scenarios dealing with unbalanced datasets. As Figure 6 shown, the method of the present invention is superior to all the methods being compared in terms of the Precision-Recall evaluation metric.
[0147] Table 1 Quantitative evaluation of our method and other ten techniques on five benchmark datasets (the bold data represent the best evaluation experimental data of all methods and the experimental evaluation data of our method in the last row):
[0148]
[0149] The experimental results in Table 1 show that the overall experimental evaluation results of our method in the last row are significantly better than the previous methods, demonstrating the superior performance of our method in salient object detection. For example, when compared with ten other methods, our method achieved the best results in terms of F-measure, MAE, and S-measure on the ECSSD dataset, the best results in terms of E-measure and S-measure on the PASCAL-S dataset, and the best performance in terms of F-measure, MAE, and E-measure on the DUT-OMRON dataset. Additionally, the evaluation metrics on the HKU-IS and DUTS datasets are also relatively high. To verify the performance of the method proposed in the present invention, as Figure 7 shown, the visual superiority of the inventive method is demonstrated.
[0150] To verify the effectiveness of each module proposed in the present invention, ablation comparison experiments were conducted. The F-measure, mean absolute error (MAE), enhanced alignment measure (E-measure), and structural similarity measure (S-measure) of the double-branch pyramid attention U-Net module alone, the network model based on pyramid attention U-Net and RECA-Transformer using the general boundary loss function module (cross-entropy loss and Dice loss), the residual efficient channel attention transformer module alone, and the complete network model of the present invention based on pyramid attention U-Net and RECA-Transformer were analyzed on the DUTS dataset, PASCAL-S dataset, and ECSSD dataset. The detection results are shown in Table 2.
[0151] Table 2:
[0152]
[0153] As can be seen from Table 2, ablation experiments on the double-branch pyramid attention U-Net module (CNN module), the residual efficient channel attention transformer module (RECA-Transformer module), and the network model based on pyramid attention U-Net and RECA-Transformer using the general boundary loss function module can show that these modules play an important role in the inventive method, verifying the contribution of the inventive method to improving the performance of the salient object detection task. Figure 8 The visual superiority of the inventive method is demonstrated.
[0154] Figure 9This is a simplified flowchart of the network model built based on brain inspiration in the present invention. Specifically, the upper-layer network's dual-branch pyramid attention U-Net module (CNN module) is used to extract extensive and general shallow features, while the lower-layer network's residual efficient channel attention transformer module (RECA-Transformer module) focuses on generating task-specific high-level representations and capturing global relationships through the attention mechanism; by combining the advantages of both, the accuracy of prediction can be enhanced.
[0155] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A salient object detection method based on a brain-inspired dual-process CNN-Transformer network, characterized in that: The following steps are involved: S1: construct a dataset for salient object detection, and divide the dataset into a training set, a test set, and a validation set; S2: constructing a network model based on pyramid attention U-Net and RECA-Transformer, wherein the network model based on pyramid attention U-Net and RECA-Transformer includes a dual-branch pyramid U-Net module and a residual efficient channel attention converter module; wherein the dual-branch pyramid U-Net module is an encoder-decoder architecture including an encoder part and a decoder part; The encoder part of the dual-branch pyramid U-Net module includes a pre-trained VGG16 network and a spatial attention module; the pre-trained VGG16 network is used to gradually extract feature information of the input image, and at the same time, the spatial attention module is combined to form a pyramid structure, so as to more effectively extract the spatial feature representation X and visual feature V of the input image, and the spatial feature representation X is enhanced in combination with the visual feature V; The decoder part of the dual-branch pyramid U-Net module uses transposed upsampling, linear interpolation upsampling, spatial context pooling module and residual coordinate attention module; the transposed upsampling, linear interpolation upsampling and spatial context pooling modules are used interchangeably to retain the spatial information of the feature map; the residual coordinate attention module is used to generate a spatial attention feature map; The residual efficient channel attention converter includes a RECA conversion block for feature extraction, which captures global characteristics through multi-layer attention mechanism and weight learning; S3: Use the training set to train the network model based on the pyramid attention U-Net and RECA-Transformer, use the validation set to adjust the model parameters, determine whether it is overfitting, use the combination of the basic loss function BaseLoss and the boundary attention loss function EdgeLoss as the model's attention boundary optimization loss function TotalLoss, and complete the training of the network model based on the pyramid attention U-Net and RECA-Transformer; use the test set to test the network performance of the network model based on the pyramid attention U-Net and RECA-Transformer; S4: Use the trained pyramid attention U-Net and RECA-Transformer based network models for salient object detection.
2. The method for salient object detection based on the brain-inspired dual-process CNN-Transformer network according to claim 1, characterized in that: The encoder part of the dual-branch pyramid U-Net module includes five stages, namely, the first stage consisting of the 2nd to 5th layers of the VGG16 network and the spatial attention module, the second stage consisting of the 5th to 10th layers of the VGG16 network and the spatial attention module, the third stage consisting of the 10th to 17th layers of the VGG16 network and the spatial attention module, the fourth stage consisting of the 17th to 24th layers of the VGG16 network and the spatial attention module, and the fifth stage consisting of the 24th to 31st layers of the VGG16 network and the spatial attention module; wherein the VGG16 network uses pre-trained convolutional layers and pooling layers to extract feature information of the image, and in each convolution stage, the size of the feature map is successively reduced to 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image; the spatial attention module consists of a convolutional layer and an effective channel attention module, the convolutional layer in the spatial attention module is used to extract the spatial feature representation X of the feature map, and the effective channel attention module is used to enhance the spatial feature representation.
3. The method for salient object detection based on the brain-inspired dual-process CNN-Transformer network according to claim 2, characterized in that: The expression of the convolutional layer in the spatial attention module is as follows: X=(Conv2D 3×3 (ReLU(Conv2D 1×1 (x c )))), Among them, X represents the spatial feature representation of the feature map, Conv2D 3×3 (·) represents a two-dimensional convolution, the convolution kernel size is 3×3, ReLU (·) represents the ReLU activation function, Conv2D 1×1 (·) represents a two-dimensional convolution, the convolution kernel size is 1×1, x c Represents the input feature map of the spatial attention module; The effective channel attention module includes a global average pooling layer GAP, a one-dimensional convolutional layer and a Sigmoid activation function stacked in sequence, and the output of the Sigmoid activation function is multiplied by the input of the global average pooling layer along the channel as the output of the effective channel attention module; the expression of the effective channel attention module is as follows: X′=ECA(X)=X·σ(Conv1D 1×1 (AvgPool(X))), Among them, X′ represents the features after enhancement processing, ECA(·) represents the processing function of the effective channel attention module, σ(·) refers to the Sigmoid activation of the result of one-dimensional convolution to obtain the channel attention weight, Conv1D 1×1 (·) represents one-dimensional convolution with a kernel size of 1×1, and AvgPool(·) represents a global average pooling operation; That is, the expression of the spatial attention module is as follows: X′=ECA(Conv2D 3×3 (ReLU(Conv2D 1×1 (x c ))))。 4. The method for salient object detection based on a brain-inspired dual-process CNN-Transformer network according to claim 1, characterized in that: The decoder part of the dual-branch pyramid U-Net module includes six stages. The first stage of the decoder part includes a residual coordinate attention module and a pixel reconstruction module; the second stage of the decoder part includes linear interpolation upsampling, spatial context pooling module, residual coordinate attention module, slicing and splicing operations; the third, fourth and fifth stages of the decoder part all include transposition upsampling, linear interpolation upsampling, spatial context pooling module, residual coordinate attention module, slicing and splicing operations; the sixth stage of the decoder part includes transposition upsampling, slicing and splicing operations.
5. The method for salient object detection based on a brain-inspired dual-process CNN-Transformer network according to claim 1, characterized in that: The spatial context pooling module first uses a 1×1 convolution operation to map the input feature map x of the spatial context pooling module to a single-channel feature map G, and the dimension of the obtained feature map G is (N, 1, H, W). Subsequently, the H channel and the W channel of the feature map G are combined together, and the H and W dimensions of the feature map G are merged to obtain a feature map G′, and the dimension of the feature map G′ is (N, 1, H×W). Then, a Softmax operation is performed to calculate the weight of each position of the feature map G′, and the feature map G′ after channel merging is normalized and applied to the input feature map x, and an enhanced feature map is obtained by element-wise multiplication.
6. The method for salient object detection based on the brain-inspired dual-process CNN-Transformer network according to claim 1, characterized in that: The residual coordinate attention module includes a parallel horizontal global average pooling layer and a vertical global average pooling layer, a splicing operation, a 3×3 convolution layer, a batch normalization layer, and a nonlinear activation layer, followed by two parallel 1×1 convolution layers and two parallel Sigmoid activation functions. The outputs of the two Sigmoid activation functions are expanded to the size of the input feature map, and then the expanded feature map is combined with the input feature map of the residual coordinate attention module through element-by-element multiplication, and the obtained output is added to the input feature map of the residual coordinate attention module through a residual connection to complete the channel expansion of the input feature map; The process of the residual coordinate attention module for channel expansion of the input feature map is as follows: first, the input feature map of the residual coordinate attention module is subjected to global average pooling in the horizontal and vertical directions respectively through the parallel horizontal global average pooling layer and the vertical global average pooling layer to extract the global average information and retain the spatial information; then, the horizontal and vertical pooled feature maps are concatenated in the second dimension to generate a new feature map for further processing and analysis; the expression is as follows: y=concat(x h ,x w ,dim=2); Among them, x h and x w They represent the pooling results in the vertical and horizontal directions respectively, concat(·) represents the concatenation operation, and y represents the concatenated feature map; Then, a 3×3 convolutional layer and a batch normalization layer are used to reduce the dimension of the concatenated feature map, and the nonlinear characteristics are introduced through the ReLU activation function. The expression is as follows: and′=ReLU(BN(Conv2D 3×3 (and))); Among them, y′ represents the feature map after dimensionality reduction, ReLU(·) represents the ReLU activation function, BN(·) represents normalization, Conv2D 3×3 (·) represents a two-dimensional convolution with a kernel size of 3×3; Next, the reduced feature map y′ is separated into horizontal feature maps x w1 and vertical feature map x h1 , and generate attention weights through 1×1 convolution layer and Sigmoid activation function respectively, the expression is as follows: x h1 ,x w1 =split(y′,[h,w],dim=2); x h2 =Sigmoid(Conv2D 1×1 (x h1 )); x w2 =Sigmoid(Conv2D 1×1 (x w1 )); Among them, split(·) means splitting the feature map in the second dimension, Sigmoid(·) means the sigmoid activation function, x h2 represents the attention weight vertical feature map, x w2 Represents the attention weight level feature map; Then, the generated attention weight vertical feature map x h2 and the attention weight level feature map x w2 Expand to the size of the input feature map and combine them with the input feature map by element-wise multiplication, as follows: x h3 =x h2 ·expand(h,w); x w3 =x w2 ·expand(h,w); f=identity·x h3 ·x w3 ; Among them, f represents the output result after element-by-element multiplication, expand(h,w) represents the length and width of the attention weight feature map expanded to h and w, identity represents the initial input feature map of the residual coordinate attention module, and x h3 represents the expanded vertical feature map, x w3 represents the expanded horizontal feature map; finally, the input and output features are combined through the residual connection.
7. The method for salient object detection based on a brain-inspired dual-process CNN-Transformer network according to claim 1, characterized in that: The residual efficient channel attention converter module first performs preliminary processing on the input data through the feature block partitioning module. The input data is divided into small blocks, and each small block of input data is converted into a feature vector representation through a linear embedding module; the converted feature vector representation is sequentially processed by multiple continuous RECA conversion blocks and linear embedding modules; then, the spatial resolution of the feature map is reduced through the feature block merging module, and feature extraction is performed through two single RECA conversion blocks at the top of the model; The extracted data features are adjusted in resolution through the block expansion module, gradually increasing the spatial resolution of the feature map to achieve upsampling, thereby generating high-resolution output features; It is processed step by step again through multiple continuous RECA conversion blocks and linear embedding modules. The continuous RECA conversion block at this stage will receive the feature information transmitted across layers by the continuous RECA conversion block at the front end. The extracted data features are processed step by step again and then the spatial resolution of the feature map is gradually reduced through the feature block merging module, while the channel dimension is increased to achieve efficient feature extraction and calculation optimization. Finally, linear projection is used to adjust the channel dimension of the feature map as output to obtain the final feature map.
8. The method for salient object detection based on a brain-inspired dual-process CNN-Transformer network according to claim 1, characterized in that: The continuous RECA conversion block includes two connected RECA conversion blocks, wherein the RECA conversion block includes layer normalization, window-based multi-head self-attention mechanism, layer normalization, residual MLP module and effective channel attention module connected in sequence, wherein the input of the first layer normalization is added to the output of the window-based multi-head self-attention mechanism as the input of the second layer normalization, and the input of the second layer normalization is added to the output of the effective channel attention module as the output of the RECA conversion block; in the continuous RECA conversion block, the output of the first RECA conversion block is used as the input of the second RECA conversion block.
9. The method for salient object detection based on the brain-inspired dual-process CNN-Transformer network according to claim 8, characterized in that: The residual MLP module includes an effective channel attention module, a fully connected layer fc1, a fully connected layer fc2, and a fully connected layer fc3 which are connected in sequence, and the output of the fully connected layer fc1 is added to the output of the fully connected layer fc2 through a residual connection and used as the input of the fully connected layer fc3.
10. The method for salient object detection based on a brain-inspired dual-process CNN-Transformer network according to claim 1, characterized in that: The calculation formula of the attention boundary optimization loss function TotalLoss composed of the basic loss function BaseLoss and the boundary attention loss function EdgeLoss is as follows: TotalLoss=BaseLoss+β·EdgeLoss; AttentionWeights=1+α·Edges; Among them, BaseLoss is the basic loss, which calculates the cross entropy loss between the output image predicted by the model and the true label. The edge map edges is generated by calculating the gradient components of the true mask map in the vertical and horizontal directions to obtain a high-quality edge map. β is the weight of the edge loss, EdgeLoss is the edge optimization loss, and CrossEntropy(p i ,t i ) represents the cross entropy loss function, p i represents the predicted value, t i Represents the actual value, N represents a total of N points, AttentionWeights represents the attention weight, α represents the proportion of edge weight, and Edges represents the edge.
Citation Information
Patent Citations
Remote sensing image saliency target detection method
CN118015332A
Saliency target detection method based on double-flow coding-decoding structure network
CN118865037A