Significant target detection method based on Mama classifier

Through the Mamba filter's significant object detection method, the PVTv2-B3 network and selective state space model are used to solve the shortcomings in detection accuracy and efficiency of VSCode and TCGNet, and achieve higher accuracy and robust significant object detection.

CN120259607APending Publication Date: 2025-07-04CHANGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510290205.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing significant object detection methods such as VSCode and TCGNet have shortcomings in detection accuracy and computing efficiency, especially in complex scenarios, with limited performance improvement.

Method used

The significant object detection method of the Mamba screener is adopted to extract deep features through the PVTv2-B3 network, and the cross-scale feature aggregation is performed using channel splitting and integrated screener. Significant information is integrated through selective state space models, and the weighted binary cross entropy and cross-mix loss function is trained.

Benefits of technology

It improves the accuracy and robustness of significant object detection, can better identify significant objects in complex scenarios, reduce irrelevant information, and enhances the model's detection ability of targets of different sizes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259607A_ABST
    Figure CN120259607A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a salient target detection method based on a Mama classifier, and the method comprises the steps: collecting a to-be-detected target image, and carrying out the depth feature extraction and cross-scale feature aggregation; dividing the input features of different scales into a plurality of primary and secondary channel split groups by using a channel split classifier, and transmitting significant channel information by using a selective state space model; and combining the secondary split groups by using a channel integration classifier, and integrating context saliency information through a residual projection selective state space model and a cascade connection selective state space model to complete construction of a saliency object detection network. According to the method, the problem that the detection precision is insufficient when the VSCode network and the TCGNet network process the salient target is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a salient object detection method based on a Mamba sieve. Background Art

[0002] The task of salient object detection aims to mimic the human eye's visual attention mechanism and automatically identify the regions or objects in an image scene that are most attractive to the human eye.

[0003] Luo Ziyang et al.'s VSCode: General Visual Salient and Camouflaged Object Detection with 2D Prompt Learning has achieved great improvements in salient object detection by using the attention mechanism and text prompts based on the Transformer backbone network; however, since the attention mechanism usually calculates the attention of each token (feature) separately, it calculates the association between one token and other tokens as its attention context; however, this method usually calculates the attention context for each token separately without accumulation, which will lead to irrelevant noise in the attention information; hindering the exploration of significant semantics and affecting the detection accuracy;

[0004] Liu Yi et al.'s TCGNet: Type-correlation guidance for salient object detection is based on the technology of capsule networks; although TCGNet improves the performance of salient object detection by integrating contrast clues and part-whole relationship clues, its high computational complexity and dependence on the capsule network structure limit the efficiency and flexibility in practical applications; at the same time, the improvement of accuracy in extremely complex scenarios is limited. Summary of the Invention

[0005] Aiming at the deficiencies of the existing methods, the present invention solves the problem of insufficient detection accuracy when the VSCode network and the TCGNet network process salient objects.

[0006] The technical solution adopted by the present invention is that the salient object detection method based on a Mamba sieve includes the following steps:

[0007] Step 1, collect the target image to be detected, and perform deep feature extraction and cross-scale feature aggregation;

[0008] As a preferred embodiment of the present invention, the PVTv2-B3 network is used for deep features of different image sizes and channels.

[0009] As a preferred embodiment of the present invention, cross-scale feature aggregation is to perform upsampling convolution batch normalization on adjacent depth features, then perform multiplication, addition, and concatenation, and then perform convolution activation operation to obtain backbone integrated features.

[0010] Step 2: Use a channel splitting sieve to divide input features of different scales into several primary and secondary channel splitting groups, and use a selective state space model to transfer significant channel information.

[0011] As a preferred embodiment of the present invention, the division of the primary channel splitting group includes:

[0012] Backbone integrated feature F i Split into S primary splitting groups along the channel dimension The formula is:

[0013]

[0014] Among them, Conv1(·) represents a 1×1 convolution operation; split(·,p) represents splitting the input into p primary splits along the channel dimension.

[0015] As a preferred embodiment of the present invention, the division of the secondary channel splitting group includes:

[0016] Divide the first primary splitting group s1 into three secondary splitting groups, and the formula is:

[0017]

[0018] After cascading the j-1th secondary splitting group and the jth primary splitting group and inputting them into the SSM, split them into three secondary splitting groups again, and the formula is:

[0019]

[0020] Among them, SSM(·) is a selective scanning mechanism.

[0021] Step 3: Merge the secondary splitting groups through a channel integration sieve, and integrate the context significant information through a residual projection selective state space model and a cascaded connection selective state space model to complete the construction of a significant object detection network.

[0022] As a preferred embodiment of the present invention, the formula for merging the secondary splitting groups through a channel integration sieve is:

[0023]

[0024] Among them, DWCov(·) and Linear(·) represent depth convolution and linear mapping layers respectively, Is the original screened feature.

[0025] As a preferred embodiment of the present invention, the integration of context significant information in the residual projection selective state space model includes:

[0026] Perform a linear mapping on to obtain system matrices {B1, C1, Δ1} and {B2, C2, Δ2};

[0027] Construct system matrices and The formula is:

[0028]

[0029] Construct a shared matrix C based on the Gaussian distributions of C1 and C2;

[0030] Based on C, C1, and C2, and the outputs F and of the selective state space model of rp1 and F rp2 ;

[0031]

[0032]

[0033] Wherein, is the hidden state in the scanning mechanism.

[0034] As a preferred embodiment of the present invention, the integration of context significant information in the cascade-connected selective state space model includes:

[0035] Cascade the features F rp1 and F rp2 Cascade

[0036] Linearly map F cat to obtain the matrix {B cat , C cat};

[0037] Construct the matrix

[0038] The integrated output of the cascade state space model is F con , that is:

[0039]

[0040] Wherein, D cat is the identity matrix;

[0041] For F conPerform channel attention and residual connection to obtain the fused feature M i .

[0042] As a preferred embodiment of the present invention, perform upsampling on the fused feature M i and generate prediction maps of different scales through convolution and concatenation.

[0043] As a preferred embodiment of the present invention, the salient object detection network adopts a weighted binary cross-entropy loss function and an intersection over union loss function.

[0044] The present invention first extracts basic depth features from the input image and aggregates the cross-scale features of the backbone network to obtain depth features with a richer receptive field; inputs the depth features after different-scale interactions into the Mamba-driven channel splitting sieve; inputs the channel segmentation groups of different scales into the Mamba-driven channel integration sieve, thereby parsing the long-range context salient information based on the Mamba sieve; and then jointly trains the salient object detection network by predicting the salient map and borrowing the loss function.

[0045] Advantages of the present invention:

[0046] 1. The screening network (MSNet) driven by Mamba in the present invention introduces the selective state space model into salient object detection, and successfully dynamically filters out irrelevant information by using the implicit recursive mechanism of the selective scanning mechanism, while emphasizing information-rich features and identifying salient semantics for detection; that is, integrating the correlation characteristics of Mamba in the modeling of long-range dependence relationships into salient object detection to screen out salient semantics from a noisy and complex environment, and improving the inference accuracy of the salient object method based on the Mamba sieve;

[0047] 2. The Mamba-driven channel splitting sieve and the Mamba-driven channel integration sieve are designed to discover the complementarity between visual semantic channels, which helps to detect salient objects in complex scenarios with different sizes, etc.;

[0048] 3. Simulations show that the present invention is more robust to salient objects of different sizes when detecting salient objects in complex scenarios. Description of the Drawings

[0049] Figure 1 is the flow chart of the salient object detection method based on the Mamba sieve of the present invention;

[0050] Figure 2 is the block diagram of the Mamba-driven channel splitting sieve of the present invention;

[0051] Figure 3 is the Mamba-driven channel integration sieve of the present invention;

[0052] Figure 4 It is the experimental result simulation diagram of the existing technologies VSCode and TCGNet under the significant object detection dataset. Detailed implementation manners

[0053] The present invention will be further described below in conjunction with the accompanying drawings and embodiments. This figure is a simplified schematic diagram, which only illustrates the basic structure of the present invention in a schematic manner. Therefore, it only shows the components related to the present invention.

[0054] As Figure 1 shown, the significant object detection method based on the Mamba sieve includes the following steps:

[0055] Step 1: Collect the target image to be detected and perform backbone feature integration on the image;

[0056] Specifically, it includes:

[0057] Step 11: Extract deep features from the target image;

[0058] Select PVTv2-B3 as the basic deep feature extraction network, crop the image to 384×384, and extract multi-level deep features of different scales through PVTv2-B3. The deep features are characterized as: {E i}(i = 1, 2, 3, 4). The sizes of these deep features are 96×96, 48×48, 24×24, and 12×12 respectively, and the feature channels are 64, 128, 320, and 512 respectively;

[0059] Step 12: Cross-scale feature aggregation;

[0060] In order to fuse adjacent-scale features to aggregate richer semantic information; the two adjacent-scale inputs E i and E i+1 are subjected to a series of operations such as upsampling, convolution, and batch normalization for feature activation to further obtain deep features t i and t i+1 . The formula is:

[0061]

[0062] where CBG(·) represents the combined operation of 3×3 convolution, batch normalization, and GeLU; Up(·) is the 2-fold bicubic interpolation operation.

[0063] Finally, the results of multiplying and adding t i and t i+1 with the two respectively are cascaded, and a series of operations such as convolution and activation are performed to aggregate adjacent scales to obtain the backbone integrated feature F i , that is:

[0064]

[0065] Among them, CBR(·) represents the combined operation of 3×3 convolution, batch normalization, and ReLU; [·] represents the concatenation operation along the channel dimension, and ⊙ represents the element-wise multiplication.

[0066] Specifically, in order to endow the model with a richer global receptive field, F1 is obtained from the cross-scale features E1 and E4, that is

[0067]

[0068] Among them, Up4(·) is the 4-fold bicubic interpolation upsampling operation; finally, the cross-scale feature F1 is obtained, that is:

[0069]

[0070] Such as Figure 2 , Step 2: Construct a Mamba-driven channel-split sifter (MCSS, Mamba-driven channel-split sifter);

[0071] Step 21: Obtain the primary channel split groups;

[0072] First, in the backbone integrated features {F i}(i = 2, 3, 4), the Mamba-driven channel split filter aims to filter out irrelevant information while identifying significant semantics across channels; specifically, the backbone integrated feature F i is first split into S primary split groups along the channel dimension That is:

[0073]

[0074] Among them, Conv1(·) represents the convolution operation using a 1×1 convolution kernel to enrich the channels of the feature map F i ; split(·, p) represents the operation of splitting the input into p primary splits along the channel dimension.

[0075] Step 22: Perform secondary channel split group iteration;

[0076] First, the first primary split group s1 is divided into three secondary split groups, that is:

[0077]

[0078] Subsequently, the third secondary split group is concatenated with the second primary split group and then input into the selective state space model SSM and split into three secondary split groups again, that is:

[0079]

[0080] Similarly, the iterative process of the j-th primary splitting group can be expressed as:

[0081]

[0082] It should be noted that for the final primary splitting group s S , two rather than three secondary splitting groups are obtained, namely:

[0083]

[0084] where SSM(·) is the selective scanning mechanism;

[0085] Based on this, the Mamba-driven channel splitting sifter explores the interdependence across feature channels. In each iteration, the selective scanning mechanism in Mamba is adopted to filter out irrelevant information while emphasizing the significant semantics crucial to the target.

[0086] Such as Figure 3 , Step 3: Construct a Mamba-driven channel-merging sifter (MCMS);

[0087] Step 31: Integrate the secondary channel groups;

[0088] Each primary splitting group will obtain a dual splitting group, namely and The dual splitting groups are respectively cascaded to form the original screening features and That is:

[0089]

[0090] Subsequently, the two respectively pass through a linear layer and a depth convolution to enrich the features, that is:

[0091]

[0092] where DWCov(·) and Linear(·) represent depth convolution and linear mapping layer respectively.

[0093] Step 32: Use the residual-projection selective state space model to interactively screen features (RPSSM; residual-projection selective SSM);

[0094] The residual-projection selective state space model uses the selective scanning mechanism to interactively screen features from an interactive perspective and Two screening features are fed into the residual projection state space model, and the system matrices {B1, C1, Δ1} and {B2, C2, Δ2} of both are obtained through linear mapping; these system matrices endow the model with context awareness ability; among them, the Δ i (i = 1, 2) parameter is used to adjust the hidden state transition in the state space model to make it input-dependent, thereby enhancing the model's adaptability to different input patterns; specifically, Δ i By controlling the speed or intensity of the hidden state update, the model can process sequence data more flexibly; D i (i = 1, 2) matrix represents the mapping from input to output directly, that is, without the direct influence of the hidden state, and is usually set to the identity matrix; in addition, A i (i = 1, 2) matrix is one of the key components of the model, which is used to define the dynamic change of the system hidden state and is initialized by the HIPPO theory, and the default initialization method is S4D-Real; that is, for the A i matrix in the real number case, its nth element is defined as Here, i' is the imaginary unit; based on this, two other groups of system matrices can be obtained and for the calculation of the state space model, that is:

[0095]

[0096] Among these parameters, matrices C1 and C2 are responsible for decoding the information from the hidden state to calculate the output; in order to interact the two screening features and a shared matrix C is used to calculate the output of the selective state space model to promote the interaction of relevant information between the two; to achieve this, the shared matrix C is modeled based on the Gaussian distributions of C1 and C2, that is:

[0097]

[0098] where μ1, μ2 and σ1, σ2 represent the mean and variance respectively.

[0099] Based on the shared matrix C and matrices C1 and C2, the output F and of the selective state space model for interacting the screening features can be realized rp1 and F rp2 , that is:

[0100]

[0101] where, is the hidden state in the scanning mechanism and helps generate the output The hidden state is updated over time, reflecting how the information in the input sequence accumulates and evolves over time;

[0102] Step 33: Integrate the screening features using the cascaded state space model (concatenation selective SSM, ConSSM);

[0103] The cascaded state space model performs selective scanning on the two enhanced interactive merged screening features from a cascaded perspective and implement the selective scanning mechanism;

[0104] First, cascade the output features F rp1 and F rp2 That is:

[0105]

[0106] Similarly, the set of system matrices {B cat , C cat} is obtained by linear mapping from F cat That is:

[0107] B cat , C cat = Linear(F cat ), (20)

[0108] Most importantly, Δ cat is cascaded from Δ rp1 and Δ rp2 That is:

[0109] Δ rp1 = Linear(Linear(F rp1 )), Δ rp2 = Linear(Linear(F rp2 ))), (21)

[0110] Δ cat = [Δ rp1 , Δ rp2 , (22)

[0111] Based on this, additional system matrices can be obtained That is:

[0112]

[0113] Among them, the method for obtaining A cat is the same as that for A i ;

[0114] Finally, the integrated output of the cascaded state space model is F con , that is:

[0115]

[0116] where D cat is the same as D i is the identity matrix;

[0117] To retain the input information, channel attention and residual connection are performed on the merged feature enhancement to obtain the final integrated result M i , that is:

[0118]

[0119]

[0120] where LN(·) and CA(·) represent layer normalization and channel attention operations respectively.

[0121] Step Four: Prediction map generation;

[0122] Fuse M at different scales in a deep-to-shallow manner i , and generate prediction maps p at different scales through upsampling, convolution, and cascading i , and apply bilinear interpolation to restore it to the original size P i , and use the final fused result P1 as the final image saliency map;

[0123] p i = CBR([Up(M i ), p i+1 ), i = 1, 2, (29)

[0124] where p3 is obtained from M3 and M4 through upsampling and cascading operations, that is:

[0125] p3 = CBR[Up(M4), M3]), (30)

[0126] Finally, p i is restored to the original size P by applying bilinear interpolation upsampling operation i , that is:

[0127] P i = BiCBR(p i ), i = 1, 2, 3, (31)

[0128] where Bi(·) represents the use of bilinear interpolation upsampling operation.

[0129] Step Five: Network training;

[0130] Under the supervision of the ground truth G, the weighted binary cross-entropy loss function and the intersection over union (IoU) loss function are used to jointly train the salient object detection network.

[0131] The cross-entropy loss function is:

[0132]

[0133] where G is the ground truth, P is the prediction result, and n is the pixel index.

[0134] The intersection over union (IoU) loss function is:

[0135]

[0136] where n is the pixel index.

[0137] The weighted loss function is:

[0138] L W = λ1 * L B + λ2 * L I , (34

[0139] where λ1 and λ2 are hyperparameters, and λ1 = 2, λ2 = 1.

[0140] The combined error function is:

[0141]

[0142] where c i ∈ {3, 2, 1}.

[0143] Simulation experiments:

[0144] Simulation conditions: All experiments are implemented on the Pytorch framework. The encoder is initialized with pre-trained weights from PVT-V2-B3, while the other parts are randomly initialized. During training, the DUTS-TR dataset (containing 10,553 images) is used to train the proposed model. DUTS-TR is a widely used training dataset for salient object detection tasks. The input image I and the ground truth map G are resized to 384×384 using bilinear interpolation. Data augmentation techniques such as random horizontal flipping and color jittering are adopted. The SGD optimizer with a momentum of 0.9 and a weight decay of 5×10 -4 is used to optimize the model.

[0145] Simulation content and result analysis:

[0146] Simulation 1: The proposed method is compared with existing efficient salient object detection methods on the public image salient object dataset for salient object detection experiments, and some results are visually compared; such as Figure 4As shown, where the original image represents the image used for experimental input in the database, and the ground truth map represents the binary map manually calibrated. Subsequently, the prediction results of the MSNet of the present invention and the prior art VSCode and TCGNet are given respectively;

[0147] From Figure 4 It can be seen that compared with the prior art, the present invention has better integrity in detecting salient objects, more obvious background suppression effect, and more accurate effect in detecting salient objects in complex scenes.

[0148] Simulation 2: The present invention and the existing salient object detection methods VSCode and TCGNet are used to conduct salient object detection experiments in the public SOD image database, and the well-recognized evaluation indicators (i.e., F-measure value, MAE value, S-measure value, and E-measure value) are used to objectively evaluate the experimental results. The evaluation simulation results are shown in Table 1;

[0149] Table 1

[0150]

[0151] It can be seen from Table 1 that compared with the prior art VSCode and TCGNet, the present invention has higher F-measure value, S-measure value, E-measure value, and lower MAE value, thus indicating that the present invention has better integrity and consistency in detecting salient objects, and fully demonstrating the effectiveness and superiority of the method of the present invention.

[0152] Taking the ideal embodiments of the present invention as the inspiration, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of the present invention. The technical scope of the present invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.

Claims

1. A significant object detection method based on a Mamba sieve, characterized in that, It includes the following steps: Step 1: Collect the target image to be detected, and perform depth feature extraction and cross-scale feature aggregation; Step 2: Use the channel splitting sieve to split the input features of different scales into several primary and secondary channel splitting groups, and use the selective state space model to transfer significant channel information; Step 3: Use the channel integration sieve to merge the secondary splitting groups, and integrate the context significant information through the residual projection selective state space model and the cascaded connection selective state space model to complete the construction of the significant object detection network.

2. The significant object detection method based on the Mamba sieve according to claim 1, characterized in that The primary channel splitting group splitting includes: Backbone integrated feature F i Split into S primary split groups along the channel dimension The formula is: Among them, Conv1(·) represents the 1×1 convolution operation; split(·,p) represents splitting the input into p primary splits along the channel dimension.

3. The method for significant object detection based on a Mamba sieve according to claim 2, wherein The secondary channel splitting group splitting includes: Divide the first primary split group s1 into three secondary split groups, and the formula is: After cascading the (j-1)th secondary split group and the jth primary split group and inputting them into the SSM, split them into three secondary split groups again, and the formula is: Among them, SSM() is the selective scanning mechanism.

4. The significant object detection method based on the Mamba sieve according to claim 1, characterized in that, The formula for merging the secondary split groups through the channel integration sieve is: Among them, DWCov(·) and Linear(·) represent depth convolution and linear mapping layers respectively, which are the original screened features.

5. The significant object detection method based on the Mamba sieve according to claim 4, characterized in that, The residual projection selective state space model integrating the context significant information includes: Pair is linearly mapped to obtain system matrices {B1, C1, Δ1} and {B2, C2, Δ2}; Construct the system matrix and The formula is: Construct a shared matrix C based on the Gaussian distributions of C1 and C2; Based on C, C1, and C2, and and the output F of the selective state space model rp1 and F rp2 ; Among them, is the hidden state in the scanning mechanism.

6. The significant object detection method based on the Mamba sieve according to claim 5, characterized in that, The cascaded connection selective state space model integrating the context significant information includes: Feature F rp1 and F rp2 are cascaded to obtain Linearly map F cat to obtain matrix {B cat , C cat}; Construct a matrix The integrated output of the cascaded state-space model is F con , that is: Among them, D cat is the identity matrix; For F con Perform channel attention and residual connection to obtain the fused feature M i .

7. The significant object detection method based on the Mamba sieve according to claim 6, characterized in that, Perform upsampling, convolution, and concatenation on the fused feature M i to generate prediction maps of different scales.

8. The significant object detection method based on the Mamba sieve according to claim 1, characterized in that The significant object detection network adopts the weighted binary cross-entropy loss function and the intersection over union loss function.

9. The significant object detection method based on a Mamba sieve according to claim 1, characterized in that, Use the PVTv2-B3 network for depth features of different image sizes and channels.

10. The significant object detection method based on the Mamba sieve according to claim 1, wherein The cross-scale feature aggregation is to perform upsampling convolution batch normalization on adjacent depth features, then multiply and add them and cascade them, and then perform convolution activation operations to obtain the backbone integrated features.