Salient target detection method based on Mama network bidirectional guidance model

The salient object detection method of the Mamba network bidirectional guidance model combines high-level semantic information and low-level edge information, which solves the problems of insufficient detection accuracy and efficiency in existing technologies and achieves higher accuracy and more efficient salient object detection.

CN120747447APending Publication Date: 2025-10-03NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510613304.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing salient object detection methods have shortcomings in utilizing high-level semantic information and low-level edge information, resulting in poor detection accuracy and computational efficiency, especially limited generalization ability in complex scenarios.

Method used

A salient object detection method based on the Mamba network bidirectional guidance model is adopted. Through the combination of encoder branch, edge branch, salient branch and decoder branch, high-level semantic information and low-level edge information are used for saliency prediction, including feature extraction of encoder branch, expansion of receptive field and spatial enhancement of edge branch, two-stage feature fusion of salient branch and feature refinement of decoder branch.

Benefits of technology

The accuracy of saliency detection is improved, the computational efficiency is enhanced, and it is easy to deploy in actual application systems. The network learning efficiency is optimized through the joint training method of saliency supervision and edge supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747447A_ABST
    Figure CN120747447A_ABST
Patent Text Reader

Abstract

The invention discloses a saliency target detection method based on a Mama network two-way guidance model. The method comprises the steps that a two-way model framework based on the Mama network is composed of an encoder branch, an edge branch, a saliency branch and a decoder branch; the encoder branch performs block division on the obtained original image to be detected, inputs the image to the Mama feature extractor and performs down-sampling to obtain five-layer features, the first three-layer low-layer features are respectively subjected to convolution processing and then are used as edge features to be input into the edge branch, the first four-layer features are used as significant features to be input into the significant branch, and the significant branch is used as edge features to be input into the edge branch; the fifth-layer high-level features are subjected to two-stage mixed attention processing, global features are obtained, and positioning guidance is provided for subsequent feature fusion; the edge branch performs receptive field expansion and spatial enhancement on the input edge features to obtain edge fusion features. According to the invention, the precision and efficiency of saliency target detection can be well improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a salient target detection method based on a Mamba network bidirectional guidance model, and belongs to the technical field of image salient target detection. Background Art

[0002] Salient object detection aims to simulate the human visual attention mechanism and automatically identify the most eye-catching areas or objects in an image or video. Its purpose is to quickly and accurately locate and extract important objects from complex backgrounds, providing a foundation for subsequent analysis tasks. This technology is of great significance in a variety of fields, such as intelligent surveillance, autonomous driving, and content-aware image editing. It can improve processing efficiency and accuracy, enhance the user experience, and enable automated and intelligent operations. By focusing on key information, salient object detection helps reduce computing resource consumption and improve the effectiveness of various computer vision applications.

[0003] Salient object detection methods can be broadly categorized into two main categories: traditional feature engineering-based methods and deep learning-based methods. Traditional feature engineering-based methods primarily rely on manually designed features, such as low-level visual features like color, texture, and edges, and calculate the saliency value of each pixel or region by analyzing the distribution of these features in the image. Deep learning-based methods can automatically learn high-level feature representations from data, often with stronger expressive power and higher accuracy. However, the latter typically require a large amount of labeled data for training, and the model complexity is high. Data acquisition is costly and time-consuming, and they also suffer from problems such as insufficient adaptability to complex scenarios and limited generalization capabilities. This is primarily due to insufficient understanding and utilization of high-level semantic information or edge information. Some methods attempt to design more complex network structures to improve detection accuracy, but this results in an exponential increase in computational effort, thereby reducing detection efficiency.

[0004] To this end, it is necessary to design a suitable salient object detection model with high detection accuracy and computational efficiency so that it can be better applied to downstream tasks.

[0005] The information disclosed in this background section is only intended to enhance understanding of the overall background of the invention and should not be considered as an admission or any form of suggestion that the information constitutes the prior art already known to a person of ordinary skill in the art. Summary of the Invention

[0006] The present invention aims to overcome the shortcomings of the existing technology by providing a salient object detection method based on the Mamba network bidirectional guidance model. This method, based on the Mamba network bidirectional guidance model saliency prediction method, better utilizes high-level semantic information and low-level edge information, thereby improving saliency detection accuracy. The feature extraction method based on the Mamba backbone is significantly superior in computational efficiency and is easy to deploy in practical application systems.

[0007] To achieve the above object, the present invention is implemented by adopting the following technical solutions:

[0008] The present invention discloses a salient object detection method based on a Mamba network bidirectional guidance model, comprising:

[0009] Obtain the original image to be detected;

[0010] Input the original image into the trained Mamba network-based bidirectional guidance model to obtain the saliency prediction result corresponding to the original image;

[0011] The bidirectional guidance model based on the Mamba network includes an encoder branch, an edge branch, a significant branch, and a decoder branch. The specific processing method is as follows:

[0012] The encoder branch divides the original image to be detected into blocks and inputs it into the Mamba feature extractor for downsampling to obtain five layers of features. The first three low-level features are convolved and input into the edge branch as edge features. The first four layers of features are input into the salient branch as salient features. The fifth layer of high-level features undergoes two-stage mixed attention processing to obtain global features.

[0013] The edge branch expands the receptive field and spatially enhances the input edge features to obtain edge fusion features;

[0014] The salient branch fuses the input global features, salient features, and edge fusion features in two stages. The first stage is to guide the fusion of the global features with the salient features of the current layer to obtain semantic guidance features. The second stage is to embed the semantic guidance features into the edge information of the edge fusion features to obtain aggregate features.

[0015] The decoder branch transfers and refines the aggregated features of each level obtained by the saliency branch and the global features to obtain the final saliency prediction result.

[0016] Furthermore, the encoder branch includes an image segmentation module, four Mamba feature extraction modules, three downsampling modules, and a high-level feature two-stage hybrid attention global module; the edge branch includes two edge feature expansion receptive field space enhancement modules; the saliency branch includes four two-stage two-directional feature weighted enhancement modules; and the decoder branch includes four decoder modules.

[0017] Furthermore, the encoder branch divides the acquired original image to be detected into blocks, inputs it into the Mamba feature extractor and performs downsampling, extracts five layers of features as salient features, and the first three layers of low-level features are respectively subjected to convolution processing as edge features. The expression of the extracted features is as follows:

[0018] {F n |n=1, 2, 3, 4, 5}

[0019] {E n =Conv_1(F n )|n=1,2,3}

[0020] Where n represents the sequence of features; F n represents the nth layer of salient features extracted by the Mamba backbone network; E n Represents the edge features of the nth layer; Conv_1 represents 1×1 convolution.

[0021] F1 and E1 represent the results of the original image after being processed by the image division module. The corresponding feature resolution is H×W×C, where H and W represent the height and width of the image respectively, and C is 32, which is the number of channels of the features F1 and E1; F2 and E2 represent the results of the upper-level features after being processed by the Mamba feature extraction module, and the feature resolution is the same as F1 and E1; F n (n=3, 4, 5), E3 represents the result of the upper-level feature after being processed by the downsampling module and the Mamba feature extraction module, and the corresponding feature resolution is

[0022] Furthermore, the Mamba feature extraction module adopts the Mamba network structure, and the feature extraction process and expression are as follows:

[0023] First, after the initial processing of normalization, linear transformation, depth convolution, and SiLU activation function, the expression of the process is as follows:

[0024]

[0025] Where x represents the upper-layer features extracted by the upper-layer Mamba backbone network; LN represents Layer Norm normalization; Linear represents linear transformation; DWConv represents depthwise convolution; SiLU represents SiLU activation function; Indicates the result of the above processing.

[0026] Then, we use the Multi-Scanning 2D Module (MS-SS2D), normalization, and linear transformation; and then perform a residual connection. The result of the above process is added to x pixel by pixel to preserve the original information and enhance the training stability of the model. The expression of this process is as follows:

[0027]

[0028] In the formula, MS-SS2D represents the multi-directional 2D scanning method. Performs scans in eight different directions, flattens the 2D features into 1D sequences, obtains the long-range dependencies of each sequence, merges the eight sequences using a summation operation, and reshapes the result into the same 2D structure as the input features, thereby generating features with stronger global expressiveness and better robustness. LN stands for Layer Norm normalization; Linear stands for linear transformation. represents pixel-by-pixel addition; x′ represents the result of residual connection.

[0029] Finally, normalization and a feed-forward network (FFN) are used to process the residual connection result x′ to enhance the feature expression capability, and then the residual connection result is connected with x′ using pixel-by-pixel addition to obtain the output result y. The expression of the output result is as follows:

[0030]

[0031] Where LN stands for Layer Norm normalization; FFN stands for feedforward network, which consists of input layer, hidden layer and output layer. The input data is processed by linear combination and nonlinear activation function of each layer in sequence through the forward propagation algorithm. represents pixel-by-pixel addition, and y represents the features extracted by the Mamba backbone network at this layer.

[0032] Furthermore, the high-level feature two-stage hybrid attention global module is used to perform two-stage hybrid attention processing on the high-level feature F5 extracted by the fourth Mamba feature extraction module to obtain the global feature expressing the semantic target. Provides semantic guidance and positioning for subsequent modules. The processing flow and expression are as follows:

[0033] The module is divided into a spatial attention mechanism part and a channel attention mechanism part. The input features are enhanced by the former and the latter, and the output of the former serves as the input of the latter.

[0034] Spatial attention mechanism part: Five parallel branches are set up to process the input features. Among them, four pooling layer branches with similar structures are used to capture rich global context at different spatial scales and improve the ability to obtain global information; the visual state space processing branch is used to alleviate the severe and uneven degradation of input features in spatial areas.

[0035] The pooling layer branch first applies a series of pooling layers with pooling window sizes of 1×1, 2×2, 3×3 and 6×6 to the input high-level feature F5 to obtain the pooled feature i represents the feature order; then 1×1 convolution is used to compress the number of channels of these pooled features to Get compression features Then use bilinear interpolation upsampling to compress the features Upsample to the same size as the high-level feature F5 to obtain the upsampled features The expression of the process is as follows:

[0036]

[0037] Where Conv_1 represents 1×1 convolution; Upsample represents bilinear interpolation upsampling.

[0038] Upsample the features By pixel-by-pixel addition fusion, we can obtain multi-scale features F at different spatial scales. 5′ , the expression of the process is as follows:

[0039]

[0040] Where, Represents pixel-by-pixel addition.

[0041] The visual state space processing branch is processed in sequence through normalization, linear transformation, depth convolution, SiLU activation function, MS-SS2D scanning, and normalization to obtain the feature The expression of the process is as follows:

[0042]

[0043] Where LN represents Layer Norm normalization; Linear represents linear transformation; DWConv represents depthwise convolution; SiLU represents SiLU activation function; MS-SS2D represents multi-directional 2D scanning.

[0044] Then, after multiplication branches and residual connections, the features contain rich global information and the enhanced features F are integrated. 5″ , the expression of the process is as follows:

[0045]

[0046] Where LN represents Layer Norm normalization; Linear represents linear transformation; SiLU represents SiLU activation function; ⊙ represents pixel-by-pixel multiplication; Represents pixel-by-pixel addition.

[0047] Next, we introduce two dynamic trainable weight parameters α and β to control the multi-scale feature F 5′ , Enhanced feature F 5″ The contribution of the information integration with the original high-level feature F5 is expressed as follows: As part of the spatial attention mechanism output. The weighted features The expression is as follows:

[0048]

[0049] Where α and β represent dynamic trainable weight parameters; Concat represents channel dimension concatenation.

[0050] Channel attention mechanism part: weighted features After a 1×1 convolution, the width and height dimensions are expanded into a one-dimensional vector and transposed to obtain a two-dimensional matrix HW represents the number of height and width multiplication, C represents the number of channels; then the two-dimensional matrix After three parallel branches, Q, K, and V are obtained to represent query, key, and value respectively. The expressions of Q, K, and V are as follows:

[0051]

[0052] Where W Q 、W K and W V Represent the weight matrices of query, key, and value respectively.

[0053] Transpose K and multiply it with Q to get the correlation matrix A, which is expressed as follows:

[0054]

[0055] Where, represents matrix multiplication; K T represents the transpose of K; A ij It represents the inner product of the i-th column in K and the j-th column in Q, that is, the correlation between two different channel vectors.

[0056] Each column of the correlation matrix A is normalized using the Softmax function and constrained to (0,1); then matrix multiplication of V and A is performed to obtain the global features after channel saliency enhancement. Provide semantic guidance and positioning for subsequent modules. The expression is as follows:

[0057]

[0058] Where, Represents matrix multiplication.

[0059] Furthermore, the edge branch includes two edge feature expansion receptive field spatial enhancement modules, which are used for edge features E1, E2, and E3 to expand the receptive field and perform spatial enhancement to obtain edge fusion features E″ to guide the salient features to obtain a fine salient map with clear boundaries.

[0060] The two edge feature expansion receptive field space enhancement modules, the first module input is edge features E1, E2, feature resolution size is H × W × 32, the output is feature E'; the second module input is the feature E' and edge feature E3 output of the previous module, feature resolution size is H × W × 32 and The subsequent process description takes the second module as an example. Its internal process and expression are as follows:

[0061] First, the edge feature E3 is subjected to bilinear interpolation upsampling processing to make it have the same feature resolution as the feature E′, and the upsampled feature is obtained. Then the feature E′ is combined with the upsampled feature Perform pixel-by-pixel addition fusion to obtain the initial fusion feature E c The process is expressed as follows:

[0062]

[0063] Where Ypsample represents bilinear interpolation upsampling; Represents pixel-by-pixel addition.

[0064] The expanded receptive field part is used to initialize the fusion feature E c Expand the receptive field to ensure that the edge details are not lost while giving it more spatial details and obtaining the expanded feature E R Its internal process is:

[0065] First, the initial fusion feature E c Through four parallel branches i(i∈{1, 2, 3, 4}), i represents the i-th parallel branch, l1, l2, l3 respectively contain a 3x3 convolution, and set different expansion rates to 5, 8, 11, l4 contains a 1x1 convolution; then, the outputs of the four branches are spliced ​​in the channel dimension and then subjected to a 1x1 convolution to obtain the expanded feature E R The expression is as follows:

[0066] E R =Conv_1(Concat(Conv_3_5(E c ), Conv_3_8(E c ), Conv_3_11(E c ), Conv_1(E c )))

[0067] Where Conv_3_5, Conv_3_8, and Conv_3_11 represent 3x3 convolutions with dilation rates of 5, 8, and 11, respectively; Conv_1 represents 1x1 convolution; and Concat represents channel dimension concatenation.

[0068] The spatial attention part is used to expand the feature E R To enhance spatial attention, the internal process is as follows:

[0069] First, expand the feature E R The width and height dimensions are expanded into one-dimensional vectors and transposed to obtain a two-dimensional matrix E R ∈R HW ×C , HW represents the number of height and width multiplied, C represents the number of channels; then, the two-dimensional matrix E R ∈R HW×C After three parallel branches, the channel is reduced in dimension to obtain three matrices Q, K, and V, which represent query, key, and value respectively. The expressions of Q, K, and V are as follows:

[0070] Q=E R W Q 、K=E R W K 、V=E R W v

[0071] Where W Q 、W K and W V Represent the weight matrices of query, key, and value respectively.

[0072] Next, the correlation matrix B is obtained by matrix multiplication of the transpose of Q and K, which is expressed as follows:

[0073]

[0074] Where, represents matrix multiplication; K T represents the transpose of K; B ij It represents the inner product of the i-th column in Q and the j-th column in K, that is, the correlation between vectors at two different spatial positions.

[0075] Each column of the correlation matrix B is normalized using the Softmax function and constrained to (0,1); then matrix multiplication of B and V is performed, and the channel dimension is restored to obtain the feature E after spatial significance enhancement S The feature E after spatial saliency enhancement S The expression is as follows:

[0076]

[0077] Where, Represents matrix multiplication.

[0078] The feature E S With the characteristic E′, Multiply each pixel separately to obtain processed features at different levels i represents the feature sequence, which is connected in the channel dimension and then fused through 3x3 convolution to generate an edge fusion feature E″ with a clear boundary as the output. The expression of the process is as follows:

[0079]

[0080] Where ⊙ represents pixel-by-pixel multiplication; Concat represents channel dimension concatenation; Conv_3 represents 3x3 convolution.

[0081] Furthermore, the significant branch includes four two-stage two-directional feature weighted enhancement modules for global features. Distinctive feature F i (i∈{4, 3, 2, 1}), edge fusion feature E″, i represents the significant feature level, the input features are divided into two stages for fusion to obtain the aggregated feature F i′ (i∈{4, 3, 2, 1}), which not only embeds global information and edge information but also reduces noise, which is conducive to generating accurate salient regions.

[0082] The four two-stage two-directional feature weighted enhancement modules, the input of the first module is the global feature The current layer’s significant feature F4 and edge fusion feature E″ are output as aggregated feature F 4′ The input of the second, third, and fourth modules is the aggregated feature F output by the previous module i+1′(i∈{3, 2, 1}), the current layer salient features F i (i∈{3, 2, 1}), edge fusion feature E″, output is the aggregate feature F i′ (i∈{3, 2, 1}). The subsequent process description takes the first module as an example, and its process and expression are as follows:

[0083] The two-stage feature fusion, the first stage is the global feature The guidance fusion of the current layer's significant feature F4 is used to obtain the semantic guidance feature The second stage is semantic guidance features Embed the edge information of the edge fusion feature E″ to obtain the aggregated feature F 4′ .

[0084] The first stage guides integration, and its internal process is:

[0085] First, the global features After two parallel 1x1 convolutions, the matrix dimension transformation operation is projected to and C′ takes Then, matrix multiplication is used to calculate the affinity between pixel pairs, and the attention weight W is generated by the Softmax function. G ∈R HW×HW The process is expressed as follows:

[0086]

[0087] Where, Softmax represents the Softmax function; Represents matrix multiplication; Indicates the process value calculated in the first stage.

[0088] Focus on the weight W G The feature map f is obtained by matrix multiplication with the salient feature F4 after 1x1 convolution. G ∈R C′×H×W ; Then map the feature f G By adding pixel by pixel with the salient feature F4, it is added back to the salient feature to achieve feature enhancement and obtain the semantic guidance feature The expression of the process is as follows:

[0089]

[0090] Where Conv_1 represents 1×1 convolution; Represents matrix multiplication; Represents pixel-by-pixel addition.

[0091] The second stage embeds edge information, and its internal process is as follows:

[0092] First, the edge fusion feature E″ is projected onto the and C′ takes Then, matrix multiplication is used to calculate the affinity between pixel pairs, and the attention weight W is generated through the Softmax function. E ∈R HW×HW The process is expressed as follows:

[0093]

[0094] Where, Softmax represents the Softmax function; Represents matrix multiplication; Indicates the process value calculated in the second stage.

[0095] Focus on the weight W E And the semantic guidance features after 1x1 convolution Through matrix multiplication, we get a feature map f that focuses on sharp boundaries. E ∈R C′×H×W ; Then map the feature f E and semantic guidance features By adding it back to the semantic guidance feature through pixel-by-pixel addition, edge feature enhancement is achieved to obtain the aggregated feature F 4′ , as the output of the first two-stage two-directional feature weighted enhancement module. The expression of the process is as follows:

[0096]

[0097] Where Conv_1 represents 1x1 convolution; Represents matrix multiplication; Represents pixel-by-pixel addition.

[0098] Furthermore, the decoder branch includes four decoder modules for global features and aggregate features F i′ (i∈{4, 3, 2, 1}), i represents the level of the aggregated feature. Through multiple stages of feature fusion and nonlinear transformation, feature information is transferred and refined between different levels, gradually enhancing the representative ability and detail information of the feature, and obtaining the decoded feature. Provide high-quality feature representation for subsequent tasks to better capture the boundaries and internal structures of the target and improve detection accuracy.

[0099] The four decoder modules, the input of the first module is the global feature Aggregate feature F 4′ , the output is the decoded feature The input of the second, third, and fourth modules is the decoding features output by the previous module and the aggregate feature F i′ (i∈{3, 2, 1}), the output is the decoding feature

[0100] The first, second, and third decoder modules have two input feature resolutions of different sizes. After input, the low-resolution features are first upsampled to the same size as the high-resolution features using bilinear interpolation. The fourth decoder module has two input feature resolutions of the same size. The subsequent process description uses the fourth decoder module as an example. Its process and expression are as follows:

[0101] Decoding features and aggregate features F 1′ A 3×3 convolution is used to extract local features and reduce the number of channels. Then, normalization and SiLU activation functions are used for nonlinear and standardization processing. The two processed features are then combined with the original aggregated feature F 1′ , the three are fused by pixel-by-pixel addition to obtain refined features The expression of the process is as follows:

[0102]

[0103] Where Conv_3 represents 3×3 convolution; LN represents Layer Norm normalization; SiLU represents SiLU activation function; Represents pixel-by-pixel addition.

[0104] Refine features After 3×3 convolution, normalization, and SiLU activation, 1×1 convolution is used to reduce dimensionality and computational complexity. Finally, the fourth decoder module adds a final projection layer at the end, ensuring that the decoder branch's final output image maintains the same spatial dimensions as the encoder module's initial input image, resulting in the final saliency prediction image.

[0105] Furthermore, the training method based on the Mamba network bidirectional guidance model is as follows:

[0106] Get the training image and the corresponding true value image;

[0107] Inputting the training image into a bidirectional guidance model based on a Mamba network to obtain a saliency prediction map;

[0108] According to the true value image, the saliency prediction map is supervised and executed on two branches respectively, and the overall loss function of the feedback saliency prediction model is constructed, and finally a trained bidirectional guidance saliency prediction model based on the Mamba network is obtained.

[0109] The supervision performed on the two branches includes: saliency supervision is used to constrain saliency prediction, using binary cross entropy (BCE) loss and intersection over union (IoU) loss for joint supervision; edge supervision is used to improve edge features, using a class-balanced binary cross entropy loss function as a supervision signal.

[0110] Furthermore, the saliency supervision loss function consists of two parts: binary cross entropy (BCE) loss function and intersection-over-union (IoU) loss function, and the expression is as follows:

[0111]

[0112]

[0113] Where, P represents the significant predicted value; G s represents the true value of saliency; H and W represent the height and width of the image respectively; (i, j) represents the pixel position of the image; Sum(i, j) and Mul(i, j) represent P and G s The sum and product at pixel position (i, j); l B represents the binary cross entropy (BCE) loss; l I represents the Intersection over Union (IoU) loss.

[0114] Furthermore, the expression of the edge supervision loss function is as follows:

[0115]

[0116] Where B represents the edge prediction value; G e Represents the true edge value; H and W represent the height and width of the image respectively; (i, j) represents the pixel position of the image; Used to balance the contribution of salient pixels and non-salient pixels, where v+, v- and v represent G e Positive, negative and total pixels; l b represents the marginal supervision loss.

[0117] Furthermore, the expression of the overall loss function is as follows:

[0118]

[0119] Where G s represents the true value of significance; P n Represents the significant predicted value, where n represents its subscript; Ge Indicates the marginal truth value; B n Represents the edge prediction value, where n represents its subscript; l B represents the binary cross entropy (BCE) loss; l I represents the Intersection over Union (IoU) loss; l b represents the marginal supervision loss; Indicates overall loss.

[0120] Beneficial effects:

[0121] 1. The present invention's salient object detection method based on the Mamba network bidirectional guidance model, compared to conventional detection methods based on convolutional networks, better utilizes both high-level semantic information and low-level edge information, thereby improving saliency detection accuracy. The feature extraction method based on the Mamba backbone offers superior computational efficiency and facilitates deployment in practical application systems.

[0122] 2. The present invention adopts a combination of saliency supervision and edge supervision during training, improves edge features through edge supervision, optimizes the learning efficiency of the overall network, and makes the overall model easier to train. BRIEF DESCRIPTION OF THE DRAWINGS

[0123] Figure 1 This is a flowchart of a salient target detection method based on a Mamba network bidirectional guidance model provided in an embodiment.

[0124] Figure 2 Schematic diagram of the structure of the Mamba feature extractor provided in the embodiment.

[0125] Figure 3 It is a structural diagram of the high-level feature two-stage hybrid attention global module provided in the embodiment.

[0126] Figure 4 3 is a schematic diagram of the structure of the edge feature expansion receptive field space enhancement module provided in the embodiment.

[0127] Figure 5 3 is a schematic structural diagram of the spatial enhancement part of the edge feature expansion receptive field spatial enhancement module provided in the embodiment.

[0128] Figure 6 3 is a schematic structural diagram of a two-stage two-directional feature weighted enhancement module provided in an embodiment.

[0129] Figure 7 3 is a schematic structural diagram of a decoder module provided in an embodiment. DETAILED DESCRIPTION

[0130] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.

[0131] This embodiment provides a salient object detection method based on the Mamba network bidirectional guidance model framework, including the following steps:

[0132] Obtain the original image to be detected;

[0133] The original image is input into the trained Mamba network-based bidirectional guided saliency prediction model to obtain the saliency prediction result corresponding to the original image.

[0134] Among them, the Mamba network bidirectional guidance saliency prediction model processes the acquired original image to be detected through four branches. The specific processing method is as follows:

[0135] The encoder branch divides the original image to be detected into blocks and inputs it into the Mamba feature extractor for downsampling to obtain five layers of features. The first three low-level features are convolved and input into the edge branch as edge features. The first four layers of features are input into the salient branch as salient features. The fifth layer of high-level features undergoes two-stage mixed attention processing to obtain global features.

[0136] The edge branch expands the receptive field and spatially enhances the input edge features to obtain edge fusion features;

[0137] The salient branch fuses the input global features, salient features, and edge fusion features in two stages. The first stage is to guide the fusion of the global features with the salient features of the current layer to obtain semantic guidance features. The second stage is to embed the semantic guidance features into the edge information of the edge fusion features to obtain aggregate features.

[0138] The decoder branch transfers and refines the aggregated features of each level obtained by the saliency branch and the global features to obtain the final saliency prediction result.

[0139] The technical concept of this invention is that, compared with conventional detection methods based on convolutional network structures, the proposed saliency prediction method based on the Mamba network bidirectional guidance model better utilizes both high-level semantic information and low-level edge information, thereby improving saliency detection accuracy. The feature extraction method based on the Mamba backbone offers superior computational efficiency and is easy to deploy in practical application systems.

[0140] like Figure 1As shown in the figure, the bidirectional guided saliency prediction model based on the Mamba network consists of an encoder branch, an edge branch, a saliency branch, and a decoder branch. The encoder branch includes an image segmentation module, four Mamba feature extraction modules, three downsampling modules, and a high-level feature two-stage hybrid attention global module; the edge branch includes two edge feature expansion receptive field spatial enhancement modules; the saliency branch includes four two-stage two-directional feature weighted enhancement modules; and the decoder branch includes four decoder modules.

[0141] The encoder branch divides the original image to be detected into blocks, inputs it into the Mamba feature extractor and performs downsampling, extracts five layers of features as salient features, and uses the first three layers of low-level features as edge features after convolution processing. The expression of the extracted features is as follows:

[0142] {F n |n=1, 2, 3, 4, 5}

[0143] {E n =Conv_1(F n )|n=1,2,3}

[0144] Where n represents the sequence of features; F n represents the nth layer of salient features extracted by the Mamba backbone network; E n Represents the edge features of the nth layer; Conv_1 represents 1×1 convolution.

[0145] like Figure 2 As shown in the figure, the Mamba feature extraction module adopts the Mamba network structure. The feature extraction process and expression are as follows:

[0146] First, after the initial processing of normalization, linear transformation, depth convolution, and SiLU activation function, it is expressed as follows:

[0147]

[0148] Where x represents the upper-layer features extracted by the upper-layer Mamba backbone network; LN represents Layer Norm normalization; Linear represents linear transformation; DWConv represents depthwise convolution; SiLU represents SiLU activation function; Indicates the result of the above processing;

[0149] Then, we use Multi-Scanning 2D Module (MS-SS2D), normalization, and linear transformation; then we perform residual connection and add the result of the above process to x pixel by pixel. This can be expressed as:

[0150]

[0151] Wherein, MS-SS2D represents multi-directional 2D scanning mode; LN represents Layer Norm normalization; Linear represents linear transformation; represents pixel-by-pixel addition; x′ represents the result of residual connection;

[0152] Finally, normalization and a feed-forward network (FFN) are used to process x′, which is then concatenated with x′ using pixel-by-pixel addition to obtain the output y, which can be expressed as:

[0153]

[0154] Where, LN represents Layer Norm normalization; FFN represents feedforward network; represents pixel-by-pixel addition, and y represents the features extracted by the Mamba backbone network at this layer.

[0155] like Figure 3 As shown in the figure, the high-level feature two-stage hybrid attention global module is located at the end of the encoder branch, which is used to process the high-level feature F5 extracted by the fourth Mamba feature extraction module in two stages of hybrid attention to obtain the global feature that expresses the semantic target. The specific method is:

[0156] The input features are enhanced by the spatial attention mechanism and the channel attention mechanism. The former outputs weighted features. As the latter input, the global feature is finally obtained

[0157] Among them, the specific process of the spatial attention mechanism is: setting five branches to process the input features, including four pooling layer branches with similar structures and one visual state space processing branch.

[0158] The specific process of the pooling layer branch is as follows: first, a series of pooling layers with pooling window sizes of 1×1, 2×2, 3×3, and 6×6 are applied to the input high-level feature F5 to obtain the pooled feature i represents the feature order; then 1×1 convolution is used to compress the number of channels of these pooled features to Get compression features Then compress the feature Use bilinear interpolation upsampling to obtain upsampled features It can be expressed as:

[0159]

[0160] Where Conv_1 represents 1×1 convolution; Upsample represents bilinear interpolation upsampling.

[0161] Upsample the features By pixel-by-pixel addition fusion, we can obtain the multi-scale feature F 5′ , which can be expressed as:

[0162]

[0163] Where, Represents pixel-by-pixel addition.

[0164] The specific process of the visual state space processing branch is as follows: first, the input high-level feature F5 is subjected to normalization, linear transformation, depth convolution, SiLU activation function, MS-SS2D scanning, and normalization in sequence to obtain the feature It can be expressed as:

[0165]

[0166] Where LN represents Layer Norm normalization; Linear represents linear transformation; DWConv represents depthwise convolution; SiLU represents SiLU activation function; MS-SS2D represents multi-directional 2D scanning.

[0167] Then, after multiplication branches and residual connections, the enhanced features F are integrated 5″ , which can be expressed as:

[0168]

[0169] Where LN represents Layer Norm normalization; Linear represents linear transformation; SiLU represents SiLU activation function; ⊙ represents pixel-by-pixel multiplication; Represents pixel-by-pixel addition.

[0170] Next, we introduce two dynamic trainable weight parameters α and β to control the multi-scale feature F 5′ , Enhanced feature F 5″ Contribution in information integration with original high-level features F5, integrated weighted features As part of the output of the spatial attention mechanism. Expressed by formula:

[0171]

[0172] Where α and β represent dynamic trainable weight parameters; Concat represents channel dimension concatenation.

[0173] Among them, the specific process of the channel attention mechanism is: weighted features After a 1×1 convolution, the width and height dimensions are expanded into a one-dimensional vector and transposed to obtain a two-dimensional matrix HW represents the number of height and width multiplication, C represents the number of channels; then the two-dimensional matrix After three parallel branches, Q, K, and V are obtained, which represent query, key, and value respectively. The formulas of Q, K, and V are as follows:

[0174]

[0175] Where W Q 、W K and W V Represent the weight matrices of query, key, and value respectively.

[0176] Transpose K and multiply it with Q to get the correlation matrix A, which can be expressed as:

[0177]

[0178] Where, represents matrix multiplication; K T represents the transpose of K.

[0179] Each column of the correlation matrix A is normalized using the Softmax function and constrained to (0,1); then matrix multiplication of V and A is performed to obtain the global feature It can be expressed as:

[0180]

[0181] Where, Represents matrix multiplication.

[0182] like Figure 1 and Figure 4 、 Figure 5 As shown in Figure 1, there are two edge feature expansion and spatial enhancement modules, located in the edge branch, for edge features E1, E2, and E3. The specific method is: the input features are expanded and spatially enhanced to obtain edge fusion features E″.

[0183] The input and output of the two modules are as follows: the input of the first module is edge features E1 and E2, and the output is feature E′; the input of the second module is feature E′ and edge feature E3 output by the previous module, and the output is edge fusion feature E″. The subsequent process description takes the second module as an example.

[0184] First, the edge feature E3 is subjected to bilinear interpolation upsampling to obtain the upsampled feature Then the feature E′ is combined with the upsampled feature Perform pixel-by-pixel addition fusion to obtain the initial fusion feature E c . It can be expressed as:

[0185]

[0186] Where Upsample represents bilinear interpolation upsampling; Represents pixel-by-pixel addition.

[0187] Among them, the expanded receptive field part is the initial fusion feature E c Perform expansion receptive field processing to obtain expansion feature E R The specific process is as follows:

[0188] First, the initial fusion feature E c Through four parallel branches i (i∈{1, 2, 3, 4}), i represents the i-th parallel branch, l1, l2, l3 each contain a 3x3 convolution, and set different expansion rates represented by d, with values ​​of 5, 8, and 11, l4 contains a 1x1 convolution; then, the outputs of the four branches are spliced ​​in the channel dimension and then subjected to a 1x1 convolution to obtain the expanded feature E R . It can be expressed as:

[0189] E R =Conv_1(Concat(Conv_3_5(E c ), Conv_3_8(E c ), Conv_3_11(E c ), Conv_1(E c )))

[0190] Where Conv_3_5, Conv_3_8, and Conv_3_11 represent 3x3 convolutions with dilation rates of 5, 8, and 11, respectively; Conv_1 represents 1x1 convolution; and Concat represents channel dimension concatenation.

[0191] Among them, the spatial attention part is for the expansion feature E R To enhance spatial attention, the specific process is as follows:

[0192] First, expand the feature E R The width and height dimensions are expanded into one-dimensional vectors and transposed to obtain a two-dimensional matrix E R ∈R HW ×C , HW represents the number of height and width multiplied, C represents the number of channels; then, the two-dimensional matrix E R ∈R HW×CAfter three parallel branches, the channel is reduced in dimension to obtain three matrices Q, K, and V, which represent query, key, and value respectively. The Q, K, and V are expressed as follows:

[0193] Q=E R W Q 、K=E R W K 、V=E R W V

[0194] Where W Q 、W K and W V Represent the weight matrices of query, key, and value respectively.

[0195] Next, the correlation matrix B is obtained by matrix multiplication of the transpose of Q and K, which is expressed as follows:

[0196]

[0197] Where, represents matrix multiplication; K T represents the transpose of K.

[0198] Each column of the correlation matrix B is normalized using the Softmax function and constrained to (0,1); then matrix multiplication of B and V is performed, and the channel dimension is restored to obtain the feature E after spatial significance enhancement S . It can be expressed as:

[0199]

[0200] Where, Represents matrix multiplication.

[0201] The feature E S With the characteristic E′, Multiply each pixel separately to obtain the processed features i represents the feature sequence, and it is connected in the channel dimension, and then fused through 3x3 convolution to obtain the edge fusion feature E″ as the output. It can be expressed as:

[0202]

[0203] Where ⊙ represents pixel-by-pixel multiplication; Concat represents channel dimension concatenation; Conv_3 represents 3x3 convolution.

[0204] like Figure 1 and Figure 6 As shown, there are four two-stage two-directional feature weighted enhancement modules, located in the saliency branch, for global features. Distinctive feature Fi (i∈{4, 3, 2, 1}), edge fusion feature E″, i represents the significant feature level, the input features are divided into two stages for fusion to obtain the aggregated feature F i′ (i∈{4, 3, 2, 1}).

[0205] The specific method of two-stage feature fusion is: the first stage is global feature The guidance fusion of the current layer's significant feature F4 is used to obtain the semantic guidance feature The second stage is semantic guidance features Embed the edge information of the edge fusion feature E″ to obtain the aggregated feature F 4′ .

[0206] Among them, the input and output of the four modules are: the input of the first module is the global feature The current layer’s significant feature F4 and edge fusion feature E″ are output as aggregated feature F 4′ The input of the second, third, and fourth modules is the aggregated feature F output by the previous module i+1′ (i∈{3, 2, 1}), the current layer salient features F i (i∈{3, 2, 1}), edge fusion feature E″, output is the aggregate feature F i′ (i∈{3, 2, 1}). The following process description takes the first module as an example.

[0207] Among them, the first stage guides integration, and its internal process is:

[0208] First, the global features After two parallel 1x1 convolutions, the matrix dimension transformation operation is projected to and C′ takes Then, matrix multiplication is used to calculate the affinity between pixel pairs, and the attention weight W is generated by the Softmax function. G ∈R HW×HW . It can be expressed as:

[0209]

[0210] Where, Softmax represents the Softmax function; Represents matrix multiplication; Indicates the process value calculated in the first stage.

[0211] Focus on the weight W G The feature map f is obtained by matrix multiplication with the salient feature F4 after 1x1 convolution. G ∈R C′×H×W ; Then map the feature fG By adding pixel by pixel with the salient feature F4, the semantic guidance feature is obtained It can be expressed as:

[0212]

[0213] Where Conv_1 represents 1×1 convolution; Represents matrix multiplication; Represents pixel-by-pixel addition.

[0214] Among them, the second stage embeds edge information, and its internal process is:

[0215] First, the edge fusion feature E″ is projected onto the and C′ takes Then, matrix multiplication is used to calculate the affinity between pixel pairs, and the attention weight W is generated through the Softmax function. E ∈R HW×HW . It can be expressed as:

[0216]

[0217] Where, Softmax represents the Softmax function; Represents matrix multiplication; Indicates the process value calculated in the second stage.

[0218] Focus on the weight W E And the semantic guidance features after 1x1 convolution Through matrix multiplication, we get a feature map f that focuses on sharp boundaries. E ∈R C′×H×W ; Then map the feature f E and semantic guidance features The aggregate feature F is obtained by pixel-by-pixel addition. 4′ , as the module output. It can be expressed as:

[0219]

[0220] Where Conv_1 represents 1x1 convolution; Represents matrix multiplication; Represents pixel-by-pixel addition.

[0221] like Figure 1 and Figure 7 As shown, there are four decoder modules in total, located in the decoder branch, for global features and aggregate features F i′(i∈{4, 3, 2, 1}), i represents the level of the aggregated feature. Through multiple stages of feature fusion and nonlinear transformation, the decoding feature is obtained

[0222] Among them, the input and output of the four decoder modules are: the input of the first decoder module is the global feature Aggregate feature F 4′ , the output is the decoded feature The input of the second, third, and fourth decoder modules is the decoded features output by the previous decoder module and the aggregate feature F i′ (i∈{3, 2, 1}), the output is the decoding feature The subsequent process description takes the fourth module as an example.

[0223] Decoding features and aggregate features F 1′ After 3×3 convolution, normalization, and SiLU activation function, the original aggregate feature F 1′ By pixel-by-pixel addition fusion, refined features are obtained It can be expressed as:

[0224]

[0225] Where Conv_3 represents 3×3 convolution; LN represents Layer Norm normalization; SiLU represents SiLU activation function; Represents pixel-by-pixel addition.

[0226] Refine features After 3×3 convolution, normalization, SiLU activation function, and 1×1 convolution, a final projection layer is added at the end to obtain the final saliency prediction result image.

[0227] like Figure 1 As shown in Figure 2, the training method based on the Mamba network bidirectional guidance model is as follows:

[0228] Get the training image and the corresponding true value image;

[0229] Inputting the training image into a bidirectional guidance model based on a Mamba network to obtain a saliency prediction map;

[0230] According to the true value image, the saliency prediction map is supervised and executed on two branches respectively, and the overall loss function of the feedback saliency prediction model is constructed, and finally a trained bidirectional guidance saliency prediction model based on the Mamba network is obtained.

[0231] Among them, the supervision performed on the two branches includes: saliency supervision adopts binary cross entropy (BCE) loss and intersection over union (IoU) loss for joint supervision; edge supervision adopts category-balanced binary cross entropy loss function as the supervision signal.

[0232] Furthermore, the saliency supervision loss function consists of two parts: binary cross entropy (BCE) loss function and intersection-over-union (IoU) loss function, and the expression is as follows:

[0233]

[0234]

[0235] Where, P represents the significant predicted value; G s represents the true value of saliency; H and W represent the height and width of the image respectively; (i, j) represents the pixel position of the image; Sun(i, j) and Mul(i, j) represent P and G s The sum and product at pixel position (i, j); l B represents the binary cross entropy (BCE) loss; l I represents the Intersection over Union (IoU) loss.

[0236] Furthermore, the expression of the edge supervision loss function is as follows:

[0237]

[0238] Where B represents the edge prediction value; G e Represents the true edge value; H and W represent the height and width of the image respectively; (i, j) represents the pixel position of the image; Used to balance the contribution of salient pixels and non-salient pixels, where v+, v- and v represent G e Positive, negative and total pixels; l b represents the marginal supervision loss.

[0239] Furthermore, the expression of the overall loss function is as follows:

[0240]

[0241] Where G s represents the true value of significance; P n Represents the significant predicted value, where n represents its subscript; G e Indicates the marginal truth value; B n Represents the edge prediction value, where n represents its subscript; l B represents the binary cross entropy (BCE) loss; l I represents the Intersection over Union (IoU) loss; l b represents the marginal supervision loss; Indicates overall loss.

[0242] In summary, in view of the fact that the salient target detection model based on deep learning in the prior art does not fully understand and utilize high-level semantic information or edge information, resulting in low image detection accuracy, high model complexity, high data acquisition cost and time-consuming, which brings difficulties to actual deployment and application, this application proposes a salient target detection method based on the Mamba network bidirectional guidance model, which uses the Mamba network and the bidirectional guidance model to solve the above problems. By using the salient target detection method based on the Mamba network bidirectional guidance model, compared with the general detection method based on the convolutional network structure, the saliency prediction method based on the Mamba network bidirectional guidance model better utilizes high-level semantic information and low-level edge information, thereby improving the accuracy of saliency detection. The feature extraction method based on the Mamba backbone is superior in computational efficiency and is easy to deploy in actual application systems. During training, a combination of saliency supervision and edge supervision is adopted. Edge features are improved through edge supervision, the learning efficiency of the overall network is optimized, and the overall model is easier to train.

[0243] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0244] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0245] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0246] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The above is only a preferred embodiment of the present invention. It should be noted that those skilled in the art can make several improvements and modifications without departing from the principles of the present invention. These improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A salient object detection method based on the Mamba network bidirectional guidance model, characterized by include: The acquired original image to be detected is processed through four branches in the bidirectional guidance model based on the Mamba network; The four branches include an encoder branch, an edge branch, a significant branch, and a decoder branch. The specific processing method includes: The encoder branch divides the original image to be detected into blocks and inputs it into the Mamba feature extractor for downsampling to obtain five layers of features. The first three low-level features are convolved and input into the edge branch as edge features. The first four layers of features are input into the salient branch as salient features. The fifth layer of high-level features undergoes two-stage mixed attention processing to obtain global features. The edge branch expands the receptive field and spatially enhances the input edge features to obtain edge fusion features; The salient branch fuses the input global features, salient features, and edge fusion features in two stages. The first stage is to guide the fusion of the global features with the salient features of the current layer to obtain semantic guidance features. The second stage is to embed the semantic guidance features into the edge information of the edge fusion features to obtain aggregate features. The decoder branch transfers and refines the aggregated features of each level obtained by the saliency branch and the global features to obtain the final saliency prediction result.

2. The method for salient object detection based on the Mamba network bidirectional guidance model according to claim 1, characterized in that: The encoder branch includes an image segmentation module, four Mamba feature extraction modules, three downsampling modules, and a high-level feature two-stage hybrid attention global module. The edge branch includes two edge feature expansion receptive field space enhancement modules. The saliency branch includes four two-stage two-directional feature weighted enhancement modules. The decoder branch includes four decoder modules.

3. The method for detecting salient objects based on the Mamba network bidirectional guidance model according to claim 1, wherein: The encoder branch divides the acquired original image to be detected into blocks, inputs it into the Mamba feature extractor and performs downsampling, extracts five layers of features as salient features, and the first three layers of low-level features are respectively subjected to convolution processing as edge features. The expression of the extracted features is as follows: {F n |n=1,2,3,4,5} {E n =Conv_1(F n )|n=1,2,3} Where n represents the sequence of features; F n represents the nth layer of salient features extracted by the Mamba backbone network; E n Represents the edge features of the nth layer; Conv_1 represents 1×1 convolution; F1 and E1 represent the results of the original image after being processed by the image division module. The corresponding feature resolution is H×W×C, where H and W represent the height and width of the image respectively, and C is 32, which is the number of channels of the features F1 and E1; F2 and E2 represent the results of the upper-level features after being processed by the Mamba feature extraction module, and the feature resolution is the same as F1 and E1; F n (n=3, 4, 5), E3 represents the result of the upper-level feature after being processed by the downsampling module and the Mamba feature extraction module, and the corresponding feature resolution is 4. The method for detecting salient objects based on the Mamba network bidirectional guidance model according to claim 1, wherein: The Mamba feature extraction module adopts the Mamba network structure, and the feature extraction process and expression include: First, after the initial processing of normalization, linear transformation, depth convolution, and SiLU activation function, the expression of the process is as follows: Where x represents the upper-layer features extracted by the upper-layer Mamba backbone network; LN represents Layer Norm normalization; Linear represents linear transformation; DWConv represents depthwise convolution; SiLU represents SiLU activation function; Indicates the result of the above processing; Then, a multi-directional 2D scanning (Multi-Scanning 2D Module, MS-SS2D) is used, normalized, and linearly transformed. A residual connection is then performed, and the result of the above processing is added to x pixel by pixel to preserve the original information and enhance the training stability of the model. The expression of the process is: In the formula, MS-SS2D represents the multi-directional 2D scanning method. Performs scans in eight different directions, flattens the 2D features into 1D sequences, obtains the long-range dependencies of each sequence, merges the eight sequences using a summation operation, and reshapes the result into the same 2D structure as the input features, thereby generating features with stronger global expressiveness and better robustness. LN stands for Layer Norm normalization; Linear stands for linear transformation. represents pixel-by-pixel addition; x′ represents the result of residual connection; Finally, normalization and a feed-forward network (FFN) are used to process the residual connection result x′ to enhance the feature expression capability, and then the residual connection result x′ is connected with x′ using pixel-by-pixel addition to obtain the output result y. The expression of the output result is: Where LN stands for Layer Norm normalization; FFN stands for feedforward network, which consists of input layer, hidden layer and output layer. The input data is processed by linear combination and nonlinear activation function of each layer in sequence through the forward propagation algorithm. represents pixel-by-pixel addition, and y represents the features extracted by the Mamba backbone network at this layer.

5. The method for detecting salient objects based on the Mamba network bidirectional guidance model according to claim 1, wherein: The high-level feature two-stage hybrid attention global module is used to process the high-level feature F5 extracted by the fourth Mamba feature extraction module in two-stage hybrid attention to obtain the global feature that expresses the semantic target. Provide semantic guidance and positioning for subsequent modules. The processing flow and expressions include: The module is divided into a spatial attention mechanism part and a channel attention mechanism part. The input features are enhanced by the former and the latter, and the output of the former serves as the input of the latter. Spatial attention mechanism: Five parallel branches are set up to process input features. Among them, four pooling layer branches with similar structures are used to capture rich global context at different spatial scales and improve the ability to obtain global information. The visual state space processing branch is used to alleviate the severe and uneven degradation of input features in spatial regions. The pooling layer branch first applies a series of pooling layers with pooling window sizes of 1×1, 2×2, 3×3 and 6×6 to the input high-level feature F5 to obtain the pooled feature i represents the feature order; then 1×1 convolution is used to compress the number of channels of these pooled features to Get compression features Then use bilinear interpolation upsampling to compress the features Upsample to the same size as the high-level feature F5 to obtain the upsampled features The expression of the process is: Where Conv_1 represents 1×1 convolution; Upsample represents bilinear interpolation upsampling; Upsample the features By pixel-by-pixel addition fusion, we can obtain multi-scale features F at different spatial scales. 5′ , the expression of the process is: Where, represents pixel-by-pixel addition; The visual state space processing branch is processed in sequence through normalization, linear transformation, depth convolution, SiLU activation function, MS-SS2D scanning, and normalization to obtain the feature The expression of the process is as follows: Where LN represents Layer Norm normalization; Linear represents linear transformation; DWConv represents depthwise convolution; SiLU represents SiLU activation function; MS-SS2D represents multi-directional 2D scanning; Then, after multiplication branches and residual connections, the features contain rich global information and the enhanced features F are integrated. 5″ , the expression of the process is as follows: Where LN represents Layer Norm normalization; Linear represents linear transformation; SiLU represents SiLU activation function; ⊙ represents pixel-by-pixel multiplication; represents pixel-by-pixel addition; Introduce two dynamic trainable weight parameters α and β to control the multi-scale feature F 5′ , Enhanced feature F 5″ The contribution of the information integration with the original high-level feature F5 is expressed as follows: As part of the output of the spatial attention mechanism, the weighted features The expression is as follows: Where α and β represent dynamic trainable weight parameters; Concat represents channel dimension concatenation; Channel attention mechanism part: weighted features After a 1×1 convolution, the width and height dimensions are expanded into a one-dimensional vector and transposed to obtain a two-dimensional matrix HW represents the number of height and width multiplication, C represents the number of channels; then the two-dimensional matrix After three parallel branches, Q, K, and V are obtained to represent query, key, and value respectively. The expressions of Q, K, and V are: Where W Q 、W K and W V The weight matrices representing query, key, and value respectively; Transpose K and multiply it with Q to get the correlation matrix A, which is expressed as follows: Where, represents matrix multiplication; K T represents the transpose of K; A ij represents the inner product of the i-th column in K and the j-th column in Q, that is, the correlation between two different channel vectors; Each column of the correlation matrix A is normalized using the Softmax function and constrained to (0,1); then matrix multiplication of V and A is performed to obtain the global features after channel saliency enhancement. Provide semantic guidance and positioning for subsequent modules. The global features The expression is: Where, Represents matrix multiplication.

6. The method for detecting salient objects based on the Mamba network bidirectional guidance model according to claim 1, wherein: The edge branch includes two edge feature expansion receptive field spatial enhancement modules, which are used for edge features E1, E2, and E3 to expand the receptive field and perform spatial enhancement to obtain edge fusion features E″, so as to guide the salient features to obtain a fine saliency map with clear boundaries; The two edge feature expansion receptive field space enhancement modules, the first module input is edge features E1, E2, feature resolution size is H × W × 32, the output is feature E'; the second module input is the feature E' and edge feature E3 output of the previous module, feature resolution size is H × W × 32 and The subsequent process description takes the second module as an example. Its internal process and expression are as follows: First, the edge feature E3 is subjected to bilinear interpolation upsampling processing to make it have the same feature resolution as the feature E′, and the upsampled feature is obtained. Then the feature E′ is combined with the upsampled feature Perform pixel-by-pixel addition fusion to obtain the initial fusion feature E c , the expression of the process is: Where Upsample represents bilinear interpolation upsampling; represents pixel-by-pixel addition; The expanded receptive field part is used to initialize the fusion feature E c Expand the receptive field to ensure that the edge details are not lost while giving it more spatial details and obtaining the expanded feature E R , its internal process is: First, the initial fusion feature E c Through four parallel branches i (i∈{1, 2, 3, 4}), i represents the i-th parallel branch, l1, l2, l3 respectively contain a 3x3 convolution, and set different expansion rates to 5, 8, 11, l4 contains a 1x1 convolution; then, the outputs of the four branches are spliced ​​in the channel dimension and then subjected to a 1x1 convolution to obtain the expanded feature E R , the expression is: AND R =Conv_1(Concat(Conv_3_5(E c ),Conv_3_8(And c ),Conv_3_11(E c ),Conv)1(E c ))) Where Conv_3_5, Conv_3_8, and Conv_3_11 represent 3x3 convolutions with dilation rates of 5, 8, and 11, respectively; Conv_1 represents 1x1 convolution; and Concat represents channel dimension concatenation. The spatial attention part is used to expand the feature E R To enhance spatial attention, the internal process is as follows: First, expand the feature E R The width and height dimensions are expanded into one-dimensional vectors and transposed to obtain a two-dimensional matrix E R ∈R HW×C , HW represents the number of height and width multiplied, C represents the number of channels; then, the two-dimensional matrix E R ∈R HW×C After three parallel branches, the channel is reduced in dimension to obtain three matrices Q, K, and V, which represent query, key, and value respectively. The expressions of Q, K, and V include: Q=E R W Q 、K=E R W K 、V=E R W V Where W Q 、W K and W V The weight matrices representing query, key, and value respectively; Next, the correlation matrix B is obtained by matrix multiplication of the transpose of Q and K, which is expressed as follows: Where, represents matrix multiplication; K T represents the transpose of K; B ij represents the inner product of the i-th column in Q and the j-th column in K, that is, the correlation between vectors at two different spatial positions; Each column of the correlation matrix B is normalized using the Softmax function and constrained to (0,1); then matrix multiplication of B and V is performed, and the channel dimension is restored to obtain the feature E after spatial significance enhancement S , the feature E after spatial saliency enhancement S The expression is as follows: Where, Represents matrix multiplication; The feature E S With the characteristic E′, Multiply each pixel separately to obtain processed features at different levels i represents the feature sequence, which is connected in the channel dimension and then fused through 3x3 convolution to generate an edge fusion feature E″ with a clear boundary as the output. The expression of the process is: Where ⊙ represents pixel-by-pixel multiplication; Concat represents channel dimension concatenation; Conv_3 represents 3x3 convolution.

7. The method for detecting salient objects based on the Mamba network bidirectional guidance model according to claim 1, wherein: The significant branch includes four two-stage two-directional feature weighted enhancement modules for global features Distinctive feature F i (i∈{4, 3, 2, 1}), edge fusion feature E″, i represents the significant feature level, the input features are divided into two stages for fusion to obtain the aggregated feature F i′ (i∈{4, 3, 2, 1}), which not only embeds global information and edge information but also reduces noise, which is conducive to generating accurate salient regions; The four two-stage two-directional feature weighted enhancement modules, the input of the first module is the global feature The current layer’s significant features F4 and edge fusion features E” are output as aggregated features F 4′ The input of the second, third, and fourth modules is the aggregated feature F output by the previous module i+1′ (i∈{3,2,1}), the current layer’s salient features F i (i∈{3,2,1}), edge fusion feature E", the output is the aggregate feature F i′ (i∈{3,2,1}); The two-stage feature fusion, the first stage is the global feature The guidance fusion of the current layer's significant feature F4 is used to obtain the semantic guidance feature The second stage is semantic guidance features Embed the edge information of the edge fusion feature E" to obtain the aggregate feature F 4′ ; The first stage guides integration, and its internal process is: First, the global features After two parallel 1x1 convolutions, the matrix dimension transformation operation is projected to and C′ takes Then, matrix multiplication is used to calculate the affinity between pixel pairs, and the attention weight W is generated by the Softmax function. G ∈R HW×HW , the expression of the process is as follows: Where, Softmax represents the Softmax function; Represents matrix multiplication; Indicates the process value calculated in the first stage; Focus on the weight W G The feature map f is obtained by matrix multiplication with the salient feature F4 after 1x1 convolution. G ∈R C′H×W ; Then map the feature f G By adding pixel by pixel with the salient feature F4, it is added back to the salient feature to achieve feature enhancement and obtain the semantic guidance feature The expression of the process is: Where Conv_1 represents 1×1 convolution; Represents matrix multiplication; represents pixel-by-pixel addition; The second stage embeds edge information, and its internal process is as follows: First, the edge fusion feature E″ is projected onto the and C′ takes Then, matrix multiplication is used to calculate the affinity between pixel pairs, and the attention weight W is generated through the Softmax function. E ∈R HW×HW , the expression of the process is as follows: Where, Softmax represents the Softmax function; Represents matrix multiplication; Indicates the process value calculated in the second stage; Focus on the weight W E And the semantic guidance features after 1x1 convolution Through matrix multiplication, we get a feature map f that focuses on sharp boundaries. E ∈R C′×H×W ; Then map the feature f E and semantic guidance features By adding it back to the semantic guidance feature through pixel-by-pixel addition, edge feature enhancement is achieved to obtain the aggregated feature F 4′ , as the output of the first two-stage two-directional feature weighted enhancement module, the expression of the process is: Where Conv_1 represents 1x1 convolution; Represents matrix multiplication; Represents pixel-by-pixel addition.

8. The method for salient object detection based on the Mamba network bidirectional guidance model according to claim 1, characterized in that: The decoder branch includes four decoder modules for global features and aggregate features F i′ (i∈{4, 3, 2, 1}), i represents the level of the aggregated feature. Through multiple stages of feature fusion and nonlinear transformation, the feature information is transferred and refined between different levels, gradually enhancing the representative ability and detail information of the feature to obtain the decoding feature. Provide high-quality feature representation for subsequent tasks to better capture the boundaries and internal structures of the target and improve detection accuracy; The four decoder modules, the input of the first module is the global feature Aggregate feature F 4′ , the output is the decoded feature The input of the second, third, and fourth modules is the decoding features output by the previous module and the aggregate feature F i′ (i∈{3, 2, 1}), the output is the decoding feature The two input features of the first, second, and third decoder modules have different resolutions. After input, bilinear interpolation is first used to upsample the low-resolution features to the same size as the high-resolution features. The two input features of the fourth decoder module have the same resolution. The subsequent process description takes the fourth decoder module as an example. Its process and expression are as follows: Decoding features and aggregate features F 1′ A 3×3 convolution is used to extract local features and reduce the number of channels. Then, normalization and SiLU activation functions are used for nonlinear and standardization processing. The two processed features are then combined with the original aggregated feature F 1′ , the three are fused by pixel-by-pixel addition to obtain refined features The expression of the process is: Where Conv_3 represents 3×3 convolution; LN represents Layer Norm normalization; SiLU represents SiLU activation function; represents pixel-by-pixel addition; Refine features After 3×3 convolution, normalization, and SiLU activation function, 1×1 convolution is used to reduce the dimension and reduce the computational effort. Finally, the fourth decoder module adds a final projection layer at the end to ensure that the final output image of the decoder branch is consistent with the spatial size of the original input image of the encoder module, resulting in the final saliency prediction result image.

9. The method for detecting salient objects based on the Mamba network bidirectional guidance model according to claim 1, wherein: The training method based on the Mamba network bidirectional guidance model is as follows: Get the training image and the corresponding true value image; Inputting the training image into a bidirectional guidance model based on a Mamba network to obtain a saliency prediction map; According to the true value image, the saliency prediction map is supervised, performed on two branches respectively, and an overall loss function of the feedback saliency prediction model is constructed, and finally a trained bidirectional guidance saliency prediction model based on the Mamba network is obtained; The supervision performed on the two branches includes: saliency supervision is used to constrain saliency prediction, using binary cross entropy (BCE) loss and intersection over union (IoU) loss for joint supervision; edge supervision is used to improve edge features, using a class-balanced binary cross entropy loss function as the supervision signal. The expression of the edge supervision loss function is as follows: Where B represents the edge prediction value; G e Represents the true edge value; H and W represent the height and width of the image respectively; (i, j) represents the pixel position of the image; Used to balance the contribution of salient pixels and non-salient pixels, where v+, v- and v represent G e Positive, negative and total pixels; represents the marginal supervision loss; The expression of the overall loss function is as follows: Where G s represents the true value of significance; P n Represents the significant predicted value, where n represents its subscript; G e Indicates the marginal truth value; B n Represents the edge prediction value, where n represents its subscript; represents the binary cross entropy (BCE) loss; represents the intersection-over-union (IoU) loss; represents the marginal supervision loss; Indicates overall loss.

10. The method for salient object detection based on the Mamba network bidirectional guidance model according to claim 9, characterized in that: The loss function of the saliency supervision consists of two parts: binary cross entropy (BCE) loss function and intersection over union (IoU) loss function, and the expression is: Where, P represents the significant predicted value; G s represents the true value of saliency; H and W represent the height and width of the image respectively; (i, j) represents the pixel position of the image; Sum(i, j) and Mul(i, j) represent the P and G s The sum and product at pixel location (i, j); represents the binary cross entropy (BCE) loss; represents the Intersection over Union (IoU) loss.

Citation Information

Cited By

  • Target detection method and device

    CN121685938A