Significance target detection method based on global-local attention interaction mechanism
By recoding the backbone network and combining RGB and depth map information, the global-local attention interaction mechanism and a new decoder mechanism are adopted to solve the problems of low efficiency and inaccurate positioning in the existing technology, and more efficient and accurate significance object detection is achieved.
Patent Information
- Application Number
- CN202510223350.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-27
Smart Images

Figure CN120047675A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and in particular relates to a method for detecting salient targets based on a global-local attention interaction mechanism. Background Art
[0002] Salient Object Detection (SOD) is a method of simulating the human visual perception system to locate the most attractive objects in a scene. The main task is to accurately locate and segment salient areas in a given set of images by imitating the visual attention mechanism in the human perception system. At present, there are two mainstream solutions to the problem of salient object detection (SOD) based on deep learning. One is to use multi-layer perceptrons (MLPs) for salient object detection, while the other is to use fully convolutional neural networks (FCNs) for salient object detection. Unlike the fully convolutional neural network method, although the first type of model uses CNN to extract high-level features, due to the use of MLP, the spatial information in the features extracted by CNN cannot be preserved. The second type of model based on fully convolutional neural networks is used to solve the problem of semantic segmentation. Since salient object detection is essentially a segmentation task, many researchers adopt FCN-based architectures because of their ability to preserve spatial information.
[0003] Although the model has made significant improvements in detection accuracy in recent years, there are still problems such as model size expansion, low reasoning efficiency, inaccurate positioning of salient target areas, and blurred edge pixels. Summary of the invention
[0004] The purpose of the present invention is to solve the problems existing in the prior art and provide a method for salient target detection based on a global-local attention interaction mechanism. By re-encoding the backbone network, the network can pay more attention to more representative features and reduce attention to unimportant information. The RGB image information and the depth image information are combined by means of a proxy branch pre-fusion. Then, through the global-local attention mechanism, the model can have both the ability of the global attention mechanism to obtain global context information and the ability of the local attention mechanism to enrich the semantic information of salient target details. Finally, a new decoder mechanism is used to improve the accuracy of the model through multi-scale aggregation processing and an adaptive strength loss function, so that the salient target area can be more accurately located. This method can improve the efficiency of salient target detection without losing accuracy as much as possible.
[0005] To achieve the above object, the technical solution of the present invention is: a method for saliency object detection based on a global-local attention interaction mechanism, re-encoding the backbone network, combining RGB image information and depth image information through the proxy branch pre-fusion method, and then enabling the model to have both the ability to obtain global context information of the global attention mechanism and the ability to enrich the detailed semantic information of the saliency object of the local attention mechanism through the global-local attention mechanism. Finally, a new decoder mechanism is used to accurately locate the saliency object region through multi-scale aggregation processing and an adaptive intensity loss function.
[0006] In an embodiment of the present invention, the method includes the following steps:
[0007] Step S1: Take different target images from different backgrounds through a camera, and separately organize the photos into an RGB data set and a Depth data set;
[0008] Step S2: Input the RGB data set and Depth data set images into the backbone network for convolution processing to extract feature maps, then perform two-dimensional batch normalization on the feature pairs, and then pass through an activation function, and finally perform downsampling on the feature maps through a max-pooling operation;
[0009] Step S3: Pass the processed RGB feature map and Depth feature map through the proxy branch processing mechanism to form a proxy feature map X that fuses the information of both; separately pass the RGB feature map and the proxy feature map X Merge ; through the global-local attention mechanism, and also pass the Depth feature map and the proxy feature map X Merge through the global-local attention mechanism to generate new RGB feature maps and Depth feature maps; respectively pass the output results through the proxy branch processing mechanism, the global-local attention mechanism, and the proxy branch processing mechanism again, so that the model can obtain global context information through the global attention mechanism, accurately locate the saliency object region, and at the same time perform local attention processing of large kernel convolution on the saliency region through the local attention mechanism to enrich the detailed semantic information of the saliency object; Merge Step S4: Output the result after being processed by the global-local attention mechanism and merged, and process it through a three-modal multi-layer decoder; the three-modal multi-layer decoder performs three-modal hybrid multi-scale aggregation processing on the feature map to enable the three-modal multi-layer decoder to capture multi-level semantic information with complementary information; at the same time, the three-modal multi-layer decoder adopts a new adaptive pixel intensity loss function to refine the target edge pixels through multi-composite function weighting, improve the model accuracy, and make it more accurately locate the saliency object region, and finally output the saliency object detection result.
[0010] Step S4: Output the result after being processed by the global-local attention mechanism and merged, and process it through a three-modal multi-layer decoder; the three-modal multi-layer decoder performs three-modal hybrid multi-scale aggregation processing on the feature map to enable the three-modal multi-layer decoder to capture multi-level semantic information with complementary information; at the same time, the three-modal multi-layer decoder adopts a new adaptive pixel intensity loss function to refine the target edge pixels through multi-composite function weighting, improve the model accuracy, and make it more accurately locate the saliency object region, and finally output the saliency object detection result.
[0011] In an embodiment of the present invention, step S1 is specifically as follows:
[0012] Step S11: Use a camera to find and obtain images of different target objects under different backgrounds;
[0013] Step S12: Shoot different target images in different pre-screened scenarios and aggregate them into the dataset of the current product;
[0014] Step S13: Use prior knowledge and a model to obtain the depth map Depth of all the pictures in the dataset, infer the depth information of each pixel point, and separately organize them into an RGB dataset and a Depth dataset.
[0015] In an embodiment of the present invention, step S2 is specifically as follows:
[0016] Step S21: Input the RGB dataset and the Depth dataset Data RGB and Data Depth into the re-encoding backbone network ViT respectively to obtain an RGB feature map and a Depth feature map, which are called R RGB and R Depth ;
[0017] Step S22: Perform normalization processing on the obtained RGB feature map R RGB and Depth feature map R Depth respectively, where R RGB-B and R Depth-B are the results after normalization of R RGB and R Depth respectively, and Bn is the normalization operation; the operation is as follows:
[0018] R RGB-B = Bn(R RGB )
[0019] R Depth-B = Bn(R Depth )
[0020] Step S23: Pass the obtained R RGB-B and R Depth-B through an activation function respectively to enable the model to learn complex features; where R RGB-B-R and R Depth-B-R are the results after passing through the activation function, and relu is the operation on the activation function; the operation process is as follows:
[0021] R RGB-B-R = relu(R RGB-B )
[0022] R Depth-B-R = relu(RDepth-B )
[0023] Step S24. Take the obtained R RGB-B-R and R Depth-B-R Perform a max pooling operation to downsample the feature map, extract more representative features, and reduce the computational complexity, where X RGB and X Depth are the extracted feature maps after passing through the activation function, and Maxpool is the max pooling operation; the operation process is as follows:
[0024] X RGB = Maxpool(R RGB-B-R )
[0025] X Depth = Maxpool(R Depth-B-R ).
[0026] In an embodiment of the present invention, the backbone network is modified and re-encoded, that is, a coding network guided by the recruitment of a part of feature information is added to the backbone network structure, so that the network can pay more attention to more representative features and reduce the attention to unimportant information at the same time. The specific operation of the recruitment of feature information is as follows:
[0027]
[0028] where conv 7 represents a large kernel convolution operation of 7*7. Through the large kernel convolution operation, more global information in the image is captured with a larger receptive field. conv 1 / 4 represents a 1*1 convolution operation and changes the output channels to 1 / 4 times the input channels. conv 4 represents a 1*1 convolution operation and changes the output channels to 4 times the input channels. By changing the channels twice, more important information in the network is recruited to obtain more representative target features and reduce the acquisition of unimportant information; the Sigmoid operation represents the activation function added after the convolution operation; finally, the original feature map is added back, and a residual connection is used to improve the network training efficiency, simplify the learning objective, and improve the convergence speed.
[0029] In an embodiment of the present invention, the specific content of step S3 is as follows:
[0030] Step S31. Perform a confluence operation on the X RGB and X Depth output in step S2 to obtain a proxy fusion branch X Merge that fuses RGB information and depth map information;
[0031] Step S32. Perform operations on X RGB and XDepth Respectively fuse with the proxy to obtain the tributary X Merge Perform global-local attention operation to obtain the attention-featured map X RGB-AT and X Depth-AT , where the Attention operation represents the global-local attention operation; the operation is as follows:
[0032] X RGB-AT = Attention(X RGB , X Merge )
[0033] X Depth-AT = Attention(X Depth , X Merge )
[0034] Step S33: Perform triple-branch proxy pre-fusion on the feature maps X RGB-AT and X Depth-AT with X Merge again to obtain the proxy pre-fused X RGB-AT and the tributary X Depth-AT of X Merge-1 , and finally output X RGB-AT and X Depth-AT and the proxy pre-fused tributary X Merge-1 , and input the three feature maps into the next step;
[0035] Step S34: Perform global-local attention operations on X RGB-AT and X Depth-AT respectively with the proxy pre-fused X Merge-1 to obtain the attention-featured maps X RGB-AT-2 and X Depth-AT-2 ; the operation is as follows:
[0036] X RGB-AT-2 = Attention(X RGB-AT , X Merge-1 )
[0037] X Depth-AT-2 = Attention(X Depth-AT , X Merge-1 )
[0038] Step S35: Perform triple-branch proxy pre-fusion on the feature maps X RGB-AT-2 and X Depth-AT-2 with X Merge-1 again to obtain the proxy pre-fused X RGB-AT and the tributary X Depth-AT of X Merge-2 , and finally output X RGB-AT-2 and X Depth-AT-2 and the proxy pre-fused tributary XMerge-2 , input the three feature maps into the next step.
[0039] In an embodiment of the present invention, the proxy fusion branch method is as follows:
[0040] For X output in step S2 RGB and X Depth feature maps, unify the number of channels, and then, for the result X RGB of adding X Depth and X add as well as the result X RGB of multiplying X Depth and X mul , then perform a concatenation operation on the two to obtain the proxy fusion branch X Merge , where cat is the concatenation operation and conv merge is a convolution operation whose function is to change the number of channels, and through convolution, the number of channels of the concatenated result is made the same as before; the above proxy fusion branch method can enable the model to retain the integrity of the independent feature information of a single modality and also obtain the relevant information among the three modalities, further enhancing the model's feature expression ability. The specific operations of the overall process are as follows:
[0041] X add =(X RGB +X Depth )
[0042]
[0043] X Merge =conv merge (cat(X add , X mul )).
[0044] In an embodiment of the present invention, the global-local attention method is as follows:
[0045] For X output in step S31 RGB and X Depth as well as X Merge , divide them into two groups for the global-local attention mechanism, where X RGB and X Merge perform a global-local attention mechanism once, and X Depth and X Merge perform a global-local attention mechanism once; for the first group, X RGB performs the global attention mechanism and X Merge performs the local attention mechanism, then concatenate the results of the two attention mechanisms, and adjust the number of channels through a convolution once; for the second group, X Depth performs the global attention mechanism and X MergePerform local attention mechanism; then concatenate the results of the two attention mechanisms, and adjust the number of channels through a convolution operation; where Attention-g represents the global attention mechanism, and X RGB-AT-G is the feature map after the attention mechanism, Attention-p represents the local attention mechanism, and X Merge-p-rbg is the feature map after the local attention mechanism of the RBG map, and X Merge-p-d is the feature map after the local attention mechanism of the depth map, and conv att represents the convolution operation after the attention mechanism, and cat represents the concatenation operation; the overall process is as follows:
[0046] X RGB-AT-G = Attention-g(X RGB )
[0047] X Merge-p-rbg = Attention-p(X Merge )
[0048] X RGB-AT = conv att (cat(X RGB-AT-G , X Merge-p ))
[0049] X Depth-AT-G = Attention-g(X Depth )
[0050] X Merge-p-d = Attention-p(X Merge )
[0051] X Depth-AT = conv att (cat(X Depth-AT-G , X Merge-p ))。
[0052] In an embodiment of the present invention, the step S4 is specifically
[0053] Step S41: Take the obtained X RGB-AT and X RGB-AT-2 as a group and input them into the decoder; take the obtained X Depth-AT and X Depth-AT-2 as a group, and take the obtained X Merge and X Merge-1 as a group and input them into the decoder;
[0054] Step S42: Perform multi-scale aggregation processing on the three groups of feature maps respectively, and then perform concatenation and convolution operations on them to obtain the results X RGB-out , X Depth-out , X Merge-out;
[0055] Step S43: Input X RGB-out , X Depth-out , X Merge-out into a brand-new adaptive pixel intensity loss function to calculate the loss, and finally output the training model to obtain the significant detection target result.
[0056] In an embodiment of the present invention, the specific method of the multi-scale aggregation processing is as follows:
[0057] For the input X RGB-AT and X RGB-AT-2 which are RGB groups, the input X Depth-AT and X Depth-AT-2 which are Depth groups, and the input X Merge and X Merge-1 which are XMerge groups; the multi-scale aggregation processing is carried out in three groups; among them, for the input of the RGB group, first perform a splicing operation on X RGB-AT and X RGB-AT-2 , then change its number of channels to obtain X RGB-AT-CAT , then respectively add X RGB-AT and X RGB-AT-CAT to obtain X RGB-AT-CAT-ADD1 , add X RGB-AT-2 and X RGB-AT-CAT to obtain X RGB-AT-CAT-ADD2 , then splice X RGB-AT-CAT-ADD1 and X RGB-AT-CAT-ADD2 again, and change its number of channels through a convolution to obtain the result X RGB-out ; among them, for the input of the Depth group, first perform a splicing operation on X Depth-AT and X Depth-AT-2 , then change its number of channels to obtain X Depth-AT-CAT , then respectively add X Depth-AT and X Depth-AT-CAT to obtain X Depth-AT-CAT-ADD1 , add X Depth-AT-2 and X Depth-AT-CAT to obtain X Depth-AT-CAT-ADD2 , then splice X Depth-AT-CAT-ADD1 and X Depth-AT-CAT-ADD2 again, and change its number of channels through a convolution to obtain the result X Depth-out ; among them, for the input of the Merge group, first perform a splicing operation on X Merge and X Merge-1 , then change its number of channels to obtain X Merge-AT-CAT , then respectively add X Merge and X Merge-AT-CAT to obtain X Merge-AT-CAT-ADD1 , add X Merge-1 and X Depth-AT-CAT to obtain XMerge-AT-CAT-ADD2 , then splice X again Merge-AT-CAT-ADD1 and X Merge-AT-CAT-ADD2 , and change its number of channels through a single convolution to obtain the result X Merge-out ; The overall operation process is as follows:
[0058] X RGB-AT-CAT = conv RGB (cat(X RGB-AT , X RGB-AT-2 ))
[0059] X RGB-AT-CAT-ADD1 = X RGB-AT-CAT + X RGB-AT
[0060] X RGB-AT-CAT-ADD2 = X RGB-AT-CAT + X RGB-AT-2
[0061] X RGB-out = conv RGB-2 (cat(X RGB-AT-CAT-ADD1 , X RGB-AT-CAT-ADD2 ));
[0062] X Depth-AT-CAT = conv Depth (cat(X Depth-AT , X Depth-AT-2 ))
[0063] X Depth-AT-CAT-ADD1 = X Depth-AT-CAT + X Depth-AT
[0064] X Depth-AT-CAT-ADD2 = X Depth-AT-CAT + X Depth-AT-2
[0065] X Depth-out = conv Depth-2 (cat(X Depth-AT-CAT-ADD1 , X Depth-AT-CAT-ADD2 ));
[0066] X Merge-AT-CAT = conv Merge (cat(X Merge , X Merge-1 ))
[0067] X Merge-AT-CAT-ADD1 = X Merge-AT-CAT + X Merge
[0068] X Merge-AT-CAT-ADD2 = X Merge-AT-CAT + X Merge-1
[0069] X Depth-out =conv Merge-2 (cat(X Merge-AT-CAT-ADD1 ,X Merge-AT-CAT-ADD2 ))
[0070] where conv RGB and conv RGB-2 The convolution operation of the RGB group makes the number of channels of the concatenated feature map the same as before. Depth and conv Depth-2 is the convolution operation of the Depth group, so that the number of channels of the concatenated feature map is the same as before, conv Merge and conv Merge-2 It is the convolution operation of the Merge group, so that the number of channels of the concatenated feature map is the same as before.
[0071] Compared with the prior art, the present invention has the following beneficial effects:
[0072] 1. Compared with the existing salient target methods, this invention combines depth map information and RBG map information for salient target detection tasks while maintaining accuracy as much as possible, so that the two modal information complement each other. RGB information provides rich texture and color features, while depth map information provides the spatial geometric structure and depth relationship between objects. This combination can effectively improve the accuracy of the model in complex scenes. ;
[0073] 2. The present invention innovates the proxy tributary and global-local attention mechanism methods, which enable the model to obtain global context information through the global attention mechanism and accurately locate the salient target area. At the same time, the local attention mechanism performs local attention processing of large kernel convolution on the salient area to enrich the semantic information of the salient target details;
[0074] 3. The present invention provides a new decoder, which performs trimodal hybrid multi-scale aggregation processing on feature maps, so that the decoder captures multi-level semantic information with complementary information. At the same time, the decoder adopts a new adaptive pixel intensity loss function, which refines the target edge pixels through multiple composite function weighting, improves the model accuracy, and makes it more accurately locate the significant target area, thus realizing significant target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 The figure is a flow chart of the method of the present invention. DETAILED DESCRIPTION
[0076] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings.
[0077] The present invention provides a method for saliency object detection based on a global-local attention interaction mechanism, re-encodes the backbone network, combines RGB image information and depth image information through the proxy branch pre-fusion method, and then enables the model to have both the ability to obtain global context information of the global attention mechanism and the ability to enrich the detailed semantic information of the saliency object of the local attention mechanism through the global-local attention mechanism. Finally, a new decoder mechanism is used to accurately locate the saliency object region through multi-scale aggregation processing and an adaptive intensity loss function.
[0078] The following is the specific implementation process of the present invention.
[0079] Please refer to Figure 1 , the present invention provides a saliency object detection method for saliency object detection based on a global-local attention interaction mechanism, including the following steps:
[0080] Step S1: Use an industrial camera installed on the production line in cooperation with a fixed diffused light source to take the top view / left view or right view of qualified products, and organize the photos into an RBG dataset and a Depth dataset;
[0081] Step S2: Input the RGB dataset and Depth dataset images into the backbone network for image processing, and perform normalization operations and activation function operations on the processed images;
[0082] Step S3: Perform proxy fusion operations on the processed RGB feature map and Depth feature map to obtain a proxy fusion feature map Merge, and perform a global-local attention mechanism on the RGB feature map, Depth feature map, and Merge feature map Merge;
[0083] Step S4: Put the feature maps after the attention mechanism into a new decoder respectively, perform three-modal hybrid multi-scale aggregation processing on them, and strengthen the edge information through a new adaptive pixel intensity loss function to obtain the saliency object.
[0084] In this embodiment, step S1 specifically includes the following steps:
[0085] Step S11: Use the camera to find and obtain images of different target objects under different backgrounds;
[0086] Step S12: Take different target images in different scenarios pre-screened by humans and aggregate them into the dataset of the current product;
[0087] Step S13: Use prior knowledge and the model to obtain the depth map Depth of all the pictures in the dataset, infer the depth information of each pixel point, and organize them into an RGB dataset and a Depth dataset respectively.
[0088] In this embodiment, step S2 specifically includes the following steps:
[0089] Step S21: Organize two groups of pictures into an RGB data set and a Depth data set Data RGB and Data Depth Input them into the re-encoded backbone network (Vision Transformer) (perform re-encoding on the Vision Transformer), and obtain their RGB data set feature maps and Depth data set feature maps respectively, which are called R RGB and R Depth .
[0090] Step S22: Perform normalization processing on the obtained RGB data set feature map R RGB and Depth data set feature map R Depth respectively to reduce the instability in model training, where R RGB-B and R Depth-B are the results after normalization of R RGB and R Depth respectively, and Bn is the normalization operation. The operation is as follows:
[0091] R RGB-B = Bn(R RGB )
[0092] R Depth-B = Bn(R Depth )
[0093] Step S23: Pass the obtained R RGB-B and R Depth-B through the activation function respectively to enable the model to learn complex features. Where R RGB-B-R and R Depth-B-R are the results after passing through the activation function, and relu is the operation on the activation function. The operation process is as follows:
[0094] R RGB-B-R = relu(R RGB-B )
[0095] R Depth-B-R = relu(R Depth-B )
[0096] Step S24: Pass the obtained R RGB-B-R and R Depth-B-R through the max pooling operation to perform downsampling on the feature maps, extract more representative features and reduce the computational complexity, where X RGB and X Depth are the extracted feature maps after passing through the activation function, and perform the max pooling operation. The operation process is as follows:
[0097] X RGB = Maxpool(R RGB-B-R )
[0098] X Depth = Maxpool(R Depth-B-R )
[0099] In this embodiment, step S3 specifically includes the following steps:
[0100] Step S31: Perform a confluence operation on the X RGB and X Depth feature maps to obtain a proxy fusion branch X that combines RGB information and depth map information Merge .
[0101] Step S32: Perform global-local attention operations on X RGB and X Depth respectively with the proxy fusion branch X Merge to obtain attention-based feature maps X RGB-AT and X Depth-AT ; The operation is as follows:
[0102] X RGB-AT = Attention(X RGB , X Merge )
[0103] X Depth-AT = Attention(X Depth , X Merge )
[0104] Step S33: Perform three-branch proxy pre-fusion on the above feature maps X RGB-AT and X Depth-AT with X Merge again to obtain the proxy pre-fusion X RGB-AT and the branch X Depth-AT of X Merge-1 , and finally output X RGB-AT and X Depth-AT and the proxy pre-fusion branch X Merge-1 , and input the three feature maps into the next step.
[0105] Step S34: Perform global-local attention operations on X RGB-AT and X Depth-AT respectively with the proxy pre-fusion X Merge-1 to obtain attention-based feature maps X RGB-AT-2 and X Depth-AT-2 ; The operation is as follows:
[0106] X RGB-AT-2= Attention(X RGB-AT , X Merge-1 )
[0107] X Depth-AT-2 = Attention(X Depth-AT , X Merge-1 )
[0108] Step S35: Perform three-branch proxy pre-fusion on the above feature maps X RGB-AT-2 and X Depth-AT-2 and X Merge-1 again to obtain the proxy pre-fusion X RGB-AT and the tributary X Depth-AT of X Merge-2 , and finally output X RGB-AT-2 and X Depth-AT-2 and the proxy pre-fusion tributary X Merge-2 , and input the three feature maps into the next step.
[0109] In this embodiment, step S4 specifically includes the following steps
[0110] Step S41: Take the obtained X RGB-AT and X RGB-AT-2 as a group and input them into the decoder; take the obtained X Depth-AT and X Depth-AT-2 as a group, and take the obtained X Merge and X Merge-1 as a group and input them into the decoder;
[0111] Step S42: Perform multi-scale aggregation processing on the three groups of feature maps respectively, and then perform splicing and convolution operations to obtain the results X RGB-out , X Depth-out , X Merge-out ;
[0112] Step S43: Input X RGB-out , X Depth-out , X Merge-out into a new adaptive pixel intensity loss function to calculate the loss, and finally output the training model to obtain the significant detection target result.
[0113] The above are the preferred embodiments of the present invention. All changes made according to the technical solutions of the present invention, when the functions and effects produced do not exceed the scope of the technical solutions of the present invention, shall fall within the protection scope of the present invention.
Claims
1. A method for salient target detection based on global-local attention interaction mechanism, characterized in that: The backbone network is re-encoded, and the RGB image information and depth image information are combined through the pre-fusion of proxy tributaries. Then, the global-local attention mechanism is used to enable the model to have both the ability of the global attention mechanism to obtain global contextual information and the ability of the local attention mechanism to enrich the semantic information of salient target details. Finally, a new decoder mechanism is used to accurately locate the salient target area through multi-scale aggregation processing and adaptive strength loss function.
2. The method for salient target detection based on global-local attention interaction mechanism according to claim 1, characterized in that: The steps include: Step S1, take different target images from different backgrounds through a camera, and organize the photos into RGB data sets and Depth data sets respectively; Step S2: Input the RGB dataset and the Depth dataset images into the backbone network for convolution processing to extract feature maps, then perform two-dimensional batch normalization on the feature pairs, and then use the activation function and finally perform downsampling operation on the feature maps through the maximum pooling operation; Step S3: The processed RGB feature map and Depth feature map are processed through the proxy branch processing mechanism to form a proxy feature map X that integrates the information of the two. Merge ; The RGB feature map and the proxy feature map X are respectively Merge Through the global-local attention mechanism, the Depth feature map and the proxy feature map X are also Merge Through the global-local attention mechanism, new RGB feature maps and Depth feature maps are generated; the output results are processed again through the proxy tributary processing mechanism, the global-local attention mechanism and the proxy tributary processing mechanism, so that the model can obtain global context information through the global attention mechanism and accurately locate the salient target area. At the same time, the local attention mechanism is used to perform local attention processing of large kernel convolution on the salient area to enrich the semantic information of the salient target details; Step S4, output the result after being processed and merged by the global-local attention mechanism, and process it through a trimodal multi-layer decoder; the trimodal multi-layer decoder performs trimodal mixed multi-scale aggregation processing on the feature map, so that the trimodal multi-layer decoder captures multi-level semantic information with complementary information; at the same time, the trimodal multi-layer decoder adopts a new adaptive pixel intensity loss function, and refines the target edge pixels through multiple composite function weighting, thereby improving the accuracy of the model, making it more accurately locate the significant target area, and finally outputting the significant target detection result.
3. The method for salient target detection based on global-local attention interaction mechanism according to claim 2, characterized in that: The step S1 is specifically as follows: Step S11, searching and acquiring different target object images under different backgrounds through a camera; Step S12: photograph different target images in different scenes that have been pre-screened, and combine them into a data set for the current product; Step S13: Obtain the depth map Depth of all images in the data set through prior knowledge and models, infer the depth information of each pixel, and organize them into RGB data set and Depth data set respectively.
4. The method for salient target detection based on global-local attention interaction mechanism according to claim 2, characterized in that: The step S2 is specifically as follows: Step S21: RGB dataset and Depth dataset Data RGB and Data Depth Input the re-encoding backbone network ViT to obtain the RGB feature map and the Depth feature map, respectively, which are called R RGB and R Depth ; Step S22: respectively obtain the RGB feature map R RGB And Depth feature map R Depth Normalization is performed, where R RGB-B and R Depth-B R RGB and R Depth The result after normalization, Bn is the normalization operation; its operation is as follows: R RGB-B =Bn(R RGB ) R Depth-B =Bn(R Depth Step S23: RGB-B and R Depth-B After activation functions, the model can learn complex features; R RGB-B-R and R Depth-B-R is the result after the activation function, and relu is the operation on the activation function; the operation process is as follows: R RGB-B-R =relu(R RGB-B ) R Depth-B-R =relu(R Depth-B Step S24: RGB-B-R and R Depth-B-R After the maximum pooling operation, the feature map is downsampled to extract more representative features and reduce the computational complexity, where X RGB and X Depth It is the extracted feature map after the activation function, and Maxpool is the maximum pooling operation; its operation process is as follows: X RGB =Maxpool(R RGB-B-R ) X Depth =Maxpool(R Depth-B-R )。 5. The method for salient target detection based on global-local attention interaction mechanism according to claim 4, characterized in that: The backbone network is modified and re-encoded, that is, a coding network for feature information collection and guidance is added to the backbone network structure, so that the network can pay more attention to more representative features and reduce attention to unimportant information. The specific operations of feature information collection and guidance are as follows: Conv7 represents a 7*7 large kernel convolution operation. The convolution operation is performed through the large kernel to capture more global information in the image with a larger receptive field. 1 / 4 It represents a 1*1 convolution operation, and converts the output channel into 1 / 4 times the input channel. Conv4 represents a 1*1 convolution operation, and converts the output channel into 4 times the input channel. Through two channel conversions, more important information in the network is recruited, more representative target features are obtained, and non-important information is obtained. The Sigmoid operation represents the activation function added to it after the convolution operation. Finally, the original feature map is added back through a residual connection to improve the network training efficiency, simplify the learning objectives, and increase the convergence speed.
6. The method for salient target detection based on global-local attention interaction mechanism according to claim 4, characterized in that: The step S3 is specifically as follows: Step S31: X output from step S2 RGB and X Depth Perform a merging operation to obtain a proxy fusion branch X that fuses RGB information and depth map information Merge ; Step S32: X RGB and X Depth Respectively merge tributaries X with the agent Merge Perform global-local attention operation to obtain the feature map X with attention RGB-AT and X Depth-AT , the Attention operation represents the global-local attention operation; its operation is as follows: X RGB-AT =Attention(X RGB ,X Merge ) X Depth-AT =Attention(X Depth ,X Merge ) Step S33: feature map X RGB-AT and X Depth-AT With X Merge Perform three-branch proxy pre-fusion again to obtain proxy pre-fusion X RGB-AT and X Depth-AT Tributaries X Merge-1 , and finally output X RGB-AT and X Depth-AT and proxy pre-fusion tributary X Merge-1 , the three feature maps are input into the next step; Step S34: RGB-AT and X Depth-AT Pre-integrate with the agent X Merge-1 Perform global-local attention operation to obtain the feature map X with attention RGB-AT-2 and X Depth-AT-2 ; Its operation is as follows: X RGB-AT-2 =Attention(X RGB-AT ,X Merge-1 ) X Depth-AT-2 =Attention(X Depth-AT ,X Merge-1 ) Step S35: feature map X RGB-AT-2 and X Depth-AT-2 With X Merge-1 Perform three-branch proxy pre-fusion again to obtain proxy pre-fusion X RGB-AT and X Depth-AT Tributaries X Merge-2 , and finally output X RGB-AT-2 and X Depth-AT-2 and proxy pre-fusion tributary X Merge-2 , and input the three feature maps into the next step.
7. The method for salient target detection based on global-local attention interaction mechanism according to claim 6, characterized in that: The proxy fusion tributary method is: For the X output from step S2 RGB and X Depth The feature map is unified in terms of the number of channels, and then X RGB and X Depth The result of adding the two is X add and X RGB and X Depth The result of multiplying the two is X mul Then, the two are spliced to obtain the proxy fusion tributary X Merge , where cat is the concatenation operation, conv merge It is a convolution operation, which changes the number of channels. Through convolution, the number of channels of the spliced result is consistent with that before. The above proxy fusion tributary method can enable the model to retain the integrity of the independent feature information of a single modality, and can also obtain the relevant information between the three modalities, further improving the model's feature expression ability. The specific operations of the overall process are as follows: X add =(X RGB +X Depth ) X Merge =conv merge (cat(X add ,X mul ))。 8. The method for salient target detection based on global-local attention interaction mechanism according to claim 6, characterized in that: The global-local attention method is: The output of step S31 is X RGB and X Depth and X Merge Divide into two groups for global-local attention mechanism, where X RGB and X Merge Perform a global-local attention mechanism, X Depth and X Merge Perform a global-local attention mechanism; the first group consists of X RGB Make a global attention mechanism, by X Merge Perform a local attention mechanism, then concatenate the results of the two attention mechanisms, and then adjust the number of channels through a convolution; the second group consists of X Depth Make a global attention mechanism, by X Merge Perform a local attention mechanism; then concatenate the results of the two attention mechanisms, and adjust the number of channels through a convolution; Attention-g represents the global attention mechanism, X RGB-AT-G is the feature map after the attention mechanism, Attention-p represents the local attention mechanism, X Merge-p-rbg is the feature map after the local attention mechanism of the RBG map, X Merge-p-d is the feature map after the local attention mechanism of the depth map, conv att represents the convolution operation after the attention mechanism, and cat represents the concatenation operation; the overall process is as follows: X RGB-AT-G =Attention-g(X RGB ) X Merge-p-rbg =Attention-p(X Merge X RGB-AT =conv att (cat(X RGB-AT-G ,X Merge-p )) X Depth-AT-G =Attention-g(X Depth ) X Merge-p-d =Attention-p(X Merge ) X Depth-AT =conv att (cat(X Depth-AT-G ,X Merge-p ))。 9. The method for salient target detection based on global-local attention interaction mechanism according to claim 6, characterized in that: The step S4 is specifically as follows: Step S41: Get the X RGB-AT and X RGB-AT-2 As a group, input the decoder; the obtained X Depth-AT and X Depth-AT-2 As a group, the obtained X Merge and X Merge-1 As a group, input decoder; Step S42: perform multi-scale aggregation processing on the three group feature maps respectively, and then perform splicing and convolution operations on them to obtain the result X RGB-out , X Depth-out , X Merge-out ; Step S43: X RGB-out , X Depth-out , X Merge-out Input the adaptive pixel intensity loss function to calculate the loss, and finally output the training model to obtain the saliency detection target result.
10. The method for salient target detection based on global-local attention interaction mechanism according to claim 9, characterized in that: The specific method of the multi-scale aggregation processing is: For the input X RGB-AT and X RGB-AT-2 For the RGB group, enter X Depth-AT and X Depth-AT-2 For the Depth group, enter X Merge and X Merge-1 Merge group; multi-scale aggregation processing is carried out in three groups; for the input of RGB group, first X RGB-AT and X RGB-AT-2 Perform a splicing operation, then change the number of channels to obtain X RGB-AT-CAT , and then use X RGB-AT With X RGB-AT-CAT Add to get X RGB-AT-CAT-ADD1 , use X RGB-AT-2 With X RGB-AT-CAT Add to get X RGB-AT-CAT-ADD2 , and then splice X again RGB-AT-CAT-ADD1 and X RGB-AT-CAT-ADD2 , and obtain the result X by changing the number of channels through a convolution RGB-out , where for the input of the Depth group, first Depth-AT and X Depth-AT-2 Perform a splicing operation, then change the number of channels to obtain X Depth-AT-CAT , and then use X Depth-AT With X Depth-AT-CAT Add to get X Depth-AT-CAT-ADD1 , use X Depth-AT-2 With X Depth-AT-CAT Add to get X Depth-AT-CAT-ADD2 , and then splice X again Depth-AT-CAT-ADD1 and X Depth-AT-CAT-ADD2 , change the number of channels through a convolution, and get the result X Depth-out ; For the input of the Merge group, first Merge and X Merge-1 Perform a splicing operation, then change the number of channels to obtain X Merge-AT-CAT , and then use X Merge With X Merge-AT-CAT Add to get X Merge-AT-CAT-ADD1 , use X Merge-1 With X Depth-AT-CAT Add to get X Merge-AT-CAT-ADD2 , and then splice X again Merge-AT-CAT-ADD1 and X Merge-AT-CAT-ADD2 , change the number of channels through a convolution, and get the result X Merge-out ; The overall operation process is as follows: X RGB-AT-CAT =conv RGB (cat(X RGB-AT ,X RGB-AT-2 )) X RGB-AT-CAT-ADD1 =X RGB-AT-CAT +X RGB-AT X RGB-AT-CAT-ADD2 =X RGB-AT-CAT +X RGB-AT-2 X RGB-out =conv RGB-2 (cat(X RGB-AT-CAT-ADD1 ,X RGB-AT-CAT-ADD2 )); X Depth-AT-CAT =conv Depth (cat(X Depth-AT ,X Depth-AT-2 )) X Depth-AT-CAT-ADD1 =X Depth-AT-CAT +X Depth-AT X Depth-AT-CAT-ADD2 =X Depth-AT-CAT +X Depth-AT-2 X Depth-out =conv Depth-2 (cat(X Depth-AT-CAT-ADD1 ,X Depth-AT-CAT-ADD2 )); X Merge-AT-CAT =conv Merge (cat(X Merge ,X Merge-1 )) X Merge-AT-CAT-ADD1 =X Merge-AT-CAT +X Merge X Merge-AT-CAT-ADD2 =X Merge-AT-CAT +X Merge-1 X Depth-out =conv Merge-2 (cat(X Merge-AT-CAT-ADD1 ,X Merge-AT-CAT-ADD2 )) where conv RGB and conv RGB-2 The convolution operation of the RGB group makes the number of channels of the concatenated feature map the same as before. Depth and conv Depth-2 is the convolution operation of the Depth group, so that the number of channels of the concatenated feature map is the same as before, conv Merge and conv Merge-2 It is the convolution operation of the Merge group, so that the number of channels of the concatenated feature map is the same as before.