Salient Object Detection Method and Device Based on Cross-Modal Feature Fusion and Progressive Decoding

Through the dual-stream Swin Transformer encoder and cross-modal attention fusion module combined with the progressive fusion decoder, the problems of feature redundancy and waste of computing resources in RGB-D significance object detection are solved, and efficient significance object detection is achieved.

CN115908789BActive Publication Date: 2025-07-22DALIAN NATIONALITIES UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211576796.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-09
Publication Date
2025-07-22
Estimated Expiration
2042-12-09

AI Technical Summary

Technical Problem

The existing RGB-D significance object detection methods require the state-of-the-art effects through the addition of additional feature enhancement or edge generation modules, resulting in feature redundancy and waste of computing resources, while limiting further development of model design.

Method used

The dual-stream Swin Transformer encoder is used to extract multi-level and multi-scale RGB features and depth features, feature fusion is performed through a cross-modal attention fusion module, and a progressive fusion decoder is used to fuse low-level features step by step to avoid the use of additional modules.

Benefits of technology

Effectively integrate information from different modes and levels, improve the model's ability to adapt to goals at different scales, achieve accurate and significant predictions, and reduce waste of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115908789B_ABST
    Figure CN115908789B_ABST
Patent Text Reader

Abstract

The present invention discloses a saliency object detection method and device based on cross-modal feature fusion and progressive decoding. The present invention extracts multi-level and multi-scale RGB features and depth features of the image to be detected through a two-stream Swin Transformer encoder; fuses the multi-level and multi-scale RGB features and depth features through a cross-modal attention fusion module to obtain fused features; decodes the high-level fused features in the fused features through a progressive fusion decoder, and progressively fuses low-level features during the decoding process. The present invention solves the problem that the prior art needs to add additional feature enhancement or edge generation modules to achieve the state-of-the-art effect, which inevitably causes feature redundancy and waste of computing resources, and also limits the further development of the design of saliency object detection models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and particularly to a saliency object detection method and device for cross-modal feature fusion and progressive decoding. Background Art

[0002] Salient Object Detection (SOD) aims to simulate the human visual perception system to detect the most attractive regions in an image and accurately segment them, and has a wide range of applications in the field of computer vision, such as object recognition, content-based image retrieval, object segmentation, image editing, video analysis, and visual tracking.

[0003] In recent years, Convolutional Neural Networks (CNNs) have been widely applied in this field and achieved great success, breaking through the performance bottleneck of traditional methods. However, new challenges also arise. For example, the detection effect in complex scenarios (such as cluttered backgrounds, multiple objects, different illuminations, transparent objects, etc.) is often not ideal. With the increasing popularity of depth cameras such as Kinect and RealSense, RGB-D salient object detection that incorporates depth information has become an attractive research direction, and a large number of related studies have emerged. The large amount of spatial structure, 3D layout, and object boundary information contained in depth maps greatly improves the detection effect in complex scenarios.

[0004] Since there are significant differences in the information contained between RGB images and depth images, how to effectively integrate the complementary information between different modalities has become a key issue in RGB-D salient object detection. Some studies directly integrate depth maps with RGB maps as four-channel inputs, but this method does not fully consider the distribution differences between the two modalities, so it cannot effectively integrate cross-modal information. Other researchers regard depth features as auxiliary information and directly extract or enhance them using independent networks and then fuse them into RGB features. For example, Zhu et al. use an independent sub-network to extract depth features and then directly merge these features into the RGB network. Fan et al. use channel and spatial attention to mine depth information clues and then fuse the depth information into RGB features in an auxiliary manner.

[0005] During the feature extraction process, some detailed information will inevitably be lost, which in turn leads to the phenomenon of blurred boundaries in salient predictions. To address this issue, most existing algorithms obtain edge information by designing additional modules and matching corresponding objective functions for them. For example, Liu et al. designed an edge-aware module to obtain structural information from low-level depth features to generate edge features and used them to guide the decoding process. Ji et al. extracted boundary information from low-level RGB features by designing an edge collaborator and imposed additional supervision on it to emphasize the object boundary.

[0006] On the other hand, due to the significant differences in the sizes of salient objects, multi-scale context feature aggregation becomes the key to accurately locating salient objects. To address this issue, existing algorithms often utilize feature enhancement modules based on the attention mechanism or ASPP to extract multi-scale information from the highest-level features. For example, Zhao et al. designed a PAFE module based on ASPP, which treats different spatial positions unequally when aggregating multi-scale features to enhance the representation ability of salient regions. Similarly, Zhao et al. proposed a FoldASPP module to capture context information and locate salient objects at different scales.

[0007] Although the above-mentioned several mechanisms can improve the performance of salient object detection in various aspects, most algorithms often need to add additional feature enhancement or edge generation modules to achieve state-of-the-art results. This inevitably causes feature redundancy and waste of computing resources, and also limits the further development of the design of salient object detection models. Therefore, it is necessary to propose a method for salient object detection with cross-modal feature fusion and progressive decoding to solve the above problems. Summary of the Invention

[0008] The purpose of the present invention is to provide a method for salient object detection with cross-modal feature fusion and progressive decoding to solve the problems in the prior art that additional feature enhancement or edge generation modules are required to achieve state-of-the-art results, which inevitably causes feature redundancy and waste of computing resources, and also limits the further development of the design of salient object detection models.

[0009] The present invention provides a method for salient object detection with cross-modal feature fusion and progressive decoding, including:

[0010] Obtain the image to be detected;

[0011] Extract multi-level and multi-scale RGB features and depth features from the image to be detected through a dual-stream Swin Transformer encoder;

[0012] Fuse the multi-level and multi-scale RGB features and depth features through a cross-modal attention fusion module to obtain fused features;

[0013] Decode the high-level fused features in the fused features through a progressive fusion decoder, and gradually fuse low-level features during the decoding process.

[0014] Further, extracting multi-level and multi-scale RGB features and depth features from the image to be detected through a dual-stream Swin Transformer encoder includes:

[0015] Duplicate the depth image into 3 channels;

[0016] The to-be-detected image is segmented into non-overlapping blocks through a fragment segmentation operation;

[0017] Four-stage different-scale features are respectively obtained from the RGB image and the depth image, where the RGB feature is denoted as f i r (i = 1, 2, 3, 4), and the depth feature is denoted as f i d (i = 1, 2, 3, 4). Each stage consists of a fragment fusion layer and multiple stacked Swin Transformer blocks, where the fragment fusion layer in the first stage is replaced by a linear embedding layer.

[0018] Furthermore, the multi-level and multi-scale RGB features and depth features are fused through a cross-modal attention fusion module to obtain fused features, including:

[0019] The high-level adjacent layer features of the input feature f i r / d are scaled, and the top layer is replaced by the current layer to maintain alignment; the spatial resolution is adjusted to the same as the current level through an upsampling operation; the two input features are concatenated and the number of channels is aligned with f through a convolutional layer to obtain F i r / d ; F i r / d is concatenated with F i r to obtain multi-scale feature F i d i ;

[0020]

[0021] where, UP(·) represents the bilinear interpolation upsampling operation, Cat(·) represents the concatenation operation, and Conv(·) represents the 3*3 convolution operation;

[0022] Two one-dimensional average pooling operations are used to embed direction information into the multi-scale feature F i ; they are concatenated and input into a conversion layer to compress the channels; the feature map embedded with direction information is separated along the x and y directions, and then encoded attention maps are generated in their respective directions through an encoded attention layer and multiplied with F i to achieve channel attention perception;

[0023] Spatial attention perception is obtained through a spatial attention module, and the output is multiplied with F i to obtain the final fused feature F i fuse : ​

[0024] F i fuse = F i × SA(F i × CA x (ConvBS(p x (F i ), p y (F i ))) × CA y (ConvBS(p x (F i ), p y (F i ))))(2)

[0025] Among them, p x and p y represent average pooling operations in the horizontal and vertical directions; ConvBS(·) represents a conversion layer composed of a convolutional layer, a BN layer, and a Sigmoid layer; CA x (·) and CA y (·) represent the generation of encoded attention in the x and y directions, which is implemented through a convolutional layer containing a Sigmoid layer, and SA(·) represents a spatial attention layer.

[0026] Furthermore, the high-level fusion features in the fusion features are decoded by a progressive fusion decoder, and low-level features are gradually fused during the decoding process, including:

[0027] After obtaining the fusion feature F i fuse using the cross-modal attention fusion module, the high-level fusion feature f4 fuse is input into the progressive fusion decoder for decoding, and low-level features are gradually fused during the decoding process; three residual convolutional modules with different dimensions are used instead of a single convolutional layer for decoding, and the specific process is as follows:

[0028] F final = RCM1(Cat(RCM2((Cat(RCM3(Cat(F4 fuse , F3 fuse ))), F2 fuse ))), F1 fuse )) (3)

[0029] Among them, RCM i (·) represents a residual convolutional module, Cat(·) represents a concatenation operation, and F final (·) represents the final feature.

[0030] Furthermore, the decoding method of the residual convolutional module includes:

[0031] The input features are passed through a depthwise separable convolutional layer and an LN layer; the number of channels is adjusted through two pointwise convolutional layers; the input features and the output features are added together, and the feature size is adjusted through an upsampling layer. The specific process is as follows:

[0032] RCM(f) = UP(f + PW2(σ(PW1(LN(DW(f)))))) (4)

[0033] Among them, σ(·) is the GELU activation function, UP(·) represents the upsampling layer, f represents the input features, DW(·) represents the depthwise separable convolutional layer, PW(·) represents the pointwise convolutional layer, and LN(·) represents the normalization layer.

[0034] Furthermore, the method further includes:

[0035] From the high-level feature f4 fuse and each level of residual convolutional module respectively generate a saliency prediction map P i (i = 1, 2, 3), and a hybrid loss composed of BCE loss and IoU loss is used to supervise it.

[0036] Furthermore, the BCE loss L BCE is defined as:

[0037]

[0038] Among them, W and H respectively represent the width and height of the image, P(x, y) represents the predicted coordinates, and G(x, y) represents the ground truth coordinates.

[0039] Furthermore, the IoU loss L IoU is defined as:

[0040]

[0041] Among them, W and H respectively represent the width and height of the image, P(x, y) represents the predicted coordinates, and G(x, y) represents the ground truth coordinates.

[0042] Furthermore, the overall loss L of the model is defined as:

[0043]

[0044] Among them, P i is the generated saliency prediction map, and G is the ground truth map.

[0045] The present invention also provides a saliency object detection device for cross-modal feature fusion and progressive decoding, including:

[0046] An image acquisition module, configured to acquire an image to be detected;

[0047] A two-stream Swin Transformer encoder for extracting multi-level and multi-scale RGB features and depth features from the image to be detected;

[0048] A cross-modal attention fusion module for fusing the multi-level and multi-scale RGB features and depth features to obtain fused features;

[0049] A progressive fusion decoder for decoding the high-level fused features in the fused features and gradually fusing low-level features during the decoding process.

[0050] The present invention has the following beneficial effects: A saliency object detection method and device for cross-modal feature fusion and progressive decoding provided by the present invention obtains an image to be detected; extracts multi-level and multi-scale RGB features and depth features from the image to be detected through a two-stream Swin Transformer encoder; fuses the multi-level and multi-scale RGB features and depth features through a cross-modal attention fusion module to obtain fused features; decodes the high-level fused features in the fused features through a progressive fusion decoder and gradually fuses low-level features during the decoding process; the present invention proposes a new framework with a relatively simple structure for the RGB-D saliency object detection task: encoding - feature fusion - decoding, that is, composed of a two-stream Swin Transformer encoder, a cross-modal attention fusion module, and a progressive fusion decoder. The cross-modal attention fusion module combines encoding attention and spatial attention, efficiently aggregates multi-scale information between different modalities and different levels of features, improves the adaptability of the model to targets of different scales, and effectively integrates complementary information between different modalities from multiple dimensions. The progressive fusion decoder fuses low-level features in a progressive manner and refines the features using residual convolutional blocks, and can retain the detailed information in the low-level features without an additional boundary awareness module or loss function, achieving accurate saliency prediction. Description of the Drawings

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0052] Figure 1 is a flowchart of the saliency object detection method for cross-modal feature fusion and progressive decoding provided by the present invention;

[0053] Figure 2 is an overall framework diagram of the model;

[0054] Figure 3 It is a diagram of the cross-modal attention fusion module;

[0055] Figure 4 It is a diagram of the residual convolution module;

[0056] Figure 5 It is a qualitative comparison diagram of the present invention and the state-of-the-art RGB-D saliency models. Detailed implementation manners

[0057] Please refer to Figure 1 , the embodiments of the present invention provide a saliency object detection method for cross-modal feature fusion and progressive decoding, including:

[0058] S101, obtaining the image to be detected.

[0059] S102, extracting multi-level and multi-scale RGB features and depth features from the image to be detected through a two-stream Swin Transformer encoder.

[0060] Feature extraction is a crucial step in the saliency object detection task. Most of the previous saliency object detection models adopt a CNN-based backbone network for feature extraction. However, due to the inherent limitations of the receptive field of the convolutional kernel, the network shows deficiencies in extracting global features. To address this issue, Swin Transformer uses a sliding window operation based on self-attention to achieve global information modeling and reduces the square computational complexity to linear computational complexity, greatly reducing the computational cost. Therefore, the present invention uses two Swin Transformers as the backbone networks to extract multi-scale features from the RGB image and the depth image respectively. Considering complexity and efficiency, the Swin-B version is adopted in the present invention.

[0061] As Figure 2 shown, first, the depth image is copied to 3 channels to be consistent with the RGB image. Next, the image to be detected is segmented into non-overlapping patches through a patch segmentation operation; then, features of different scales in 4 stages are obtained from the RGB image and the depth image respectively, where the RGB features are denoted as f i r (i = 1, 2, 3, 4), and the depth features are denoted as f i d (i = 1, 2, 3, 4). Each stage consists of a patch fusion layer and multiple stacked Swin Transformer blocks, where the patch fusion layer in the first stage is replaced by a linear embedding layer.

[0062] S103, fusing the multi-level and multi-scale RGB features and depth features through a cross-modal attention fusion module to obtain fused features.

[0063] In the RGB-D salient object detection task, RGB features contain a large amount of texture information, while depth features focus on spatial position information. How to effectively utilize RGB features and depth features and fully exploit the complementary information between features to achieve cross-modal feature fusion is an important issue in the RGB-D salient object detection task. To address this problem, the present invention designs a Cross-Modal Attention Fusion Module (CAM), which combines one-dimensional encoded attention and spatial attention to obtain a larger range of attention information without increasing the computational burden, thereby effectively realizing the cross-modal fusion of RGB features and depth features. At the same time, when most salient object detection methods deal with small salient objects, the detection effect is often not ideal. This is because the multi-level feature scales are fixed and the interaction of context information is not sufficient enough to cope with the changes in the scales of salient objects. Therefore, the present invention scales and fuses the high-level adjacent layer features of the current layer features to obtain the guidance of high-level semantic information and rich multi-scale context information, thereby improving the detection ability for objects of different scales.

[0064] As Figure 3 shown, first, the high-level adjacent layer features of the input feature f i r / d are scaled, and the top layer is replaced by the current layer to maintain alignment; the spatial resolution is adjusted to the same as that of the current layer through an upsampling operation; the two input features are concatenated and the number of channels is aligned with f through a convolutional layer to obtain F i r / d ; F i r / d is concatenated with F i r to obtain the multi-scale feature F i d i i .

[0065]

[0066] Among them, UP(·) represents the bilinear interpolation upsampling operation, Cat(·) represents the concatenation operation, and Conv(·) represents the 3*3 convolutional operation.

[0067] Two one-dimensional average pooling operations are used to embed the direction information into the multi-scale feature F i ; they are concatenated and input into a conversion layer to compress the channels; the feature map embedded with the direction information is separated along the x and y directions, and then encoded attention maps are generated in their respective directions through an encoded attention layer, and are combined with F iMultiply to achieve channel attention perception. Obtain spatial attention perception through the spatial attention module, and multiply the output with F i to get the final fused feature f i fuse , and this process can be described as:

[0068] f i fuse = F i × SA(F i × CA x (ConvBS(p x (F i ), p y (F i ))) × CA y (ConvBS(p x (F i ), p y (F i ))))(2)

[0069] Among them, p x and p y represent average pooling operations in the horizontal and vertical directions; ConvBS(·) represents a conversion layer composed of a convolutional layer, a BN layer, and a Sigmoid layer; CA x (·) and CA y (·) represent the generation of encoded attention along the x and y directions, which is achieved through a convolutional layer containing a Sigmoid layer, and SA(·) represents the spatial attention layer.

[0070] Through the cross-modal attention fusion module designed by the present invention, the depth feature and the RGB feature are fully combined to enhance the feature representation of the target of interest, and through scaling and fusion operations, more multi-scale context information is introduced, improving the adaptability to targets of different scales.

[0071] S104, Decode the high-level fusion feature in the fusion feature through a progressive fusion decoder, and gradually fuse low-level features during the decoding process.

[0072] During the process of extracting high-level features, some edge information is often lost. On the other hand, the upsampling operation during the decoding process will also introduce certain noise information. To address these problems, the present invention designs a progressive fusion decoder, gradually integrating low-level features during the decoding process to supplement the edge contour information, and reducing the impact of noise through residual convolutional blocks.

[0073] Please refer to Figure 2 , when using the cross-modal attention fusion module to obtain the fusion feature F i fuseAfter that, the high-level fusion feature F4 fuse is input into the progressive fusion decoder for decoding, and the low-level features are gradually fused during the decoding process; three residual convolutional modules with different dimensions are adopted to replace the single convolutional layer for decoding, and the specific process is as follows:

[0074] F final = RCM1(Cat(RCM2((Cat(RCM3(Cat(F4 fuse , F3 fuse ))), F2 fuse ))), F1 fuse )) (3)

[0075] where RCM i (·) represents the residual convolutional module, Cat(·) represents the concatenation operation, and F final (·) represents the final feature.

[0076] The RCM structure is as Figure 3 shown. The decoding method of the residual convolutional module includes: First, the input feature passes through a depthwise separable convolution (Depth-Wise, DW) layer and an LN layer; the number of channels is adjusted through two pointwise convolution (Point-Wise, PW) layers; the input feature is added to the output feature, and the feature size is adjusted through an upsample (Upsample, UP) layer, and the specific process is as follows:

[0077] RCM(f) = UP(f + PW2(σ(PW1(LN(DW(f)))))) (4)

[0078] where σ(·) is the GELU activation function, UP(·) represents the upsample layer, f represents the input feature, DW(·) represents the depthwise separable convolution layer, PW(·) represents the pointwise convolution layer, and LN(·) represents the normalization layer.

[0079] In this embodiment, the method further includes: generating significant prediction maps P fuse (i = 1, 2, 3, 4) from the high-level feature F4 i and each level of the residual convolutional module respectively, and supervising them with a hybrid loss composed of the BCE loss and the IoU loss.

[0080] The BCE loss L BCE is defined as:

[0081]

[0082] where W and H respectively represent the width and height of the image, P(x, y) represents the predicted coordinates, and G(x, y) represents the ground truth coordinates.

[0083] IoU Loss L IoU is defined as:

[0084]

[0085] where W and H represent the width and height of the image respectively, P(x, y) represents the predicted coordinates, and G(x, y) represents the ground truth coordinates.

[0086] The overall loss L of the model is defined as:

[0087]

[0088] where P i is the generated saliency prediction map, and G is the ground truth map.

[0089] The method proposed in the present invention was evaluated on six challenging RGB-D saliency object detection datasets. They are representative datasets in saliency object detection and play a very important role in model training.

[0090] Dut contains 1200 images captured by Lytro cameras in real-life scenarios. NLPR includes 1000 images with single or multiple salient objects. NJU2K includes 2003 stereo images with different resolutions. DES contains 135 indoor images collected by Microsoft Kinect. SIP contains 1000 high-resolution images highlighting people. The LFSD dataset mainly includes 100 images taken by Lytro cameras, with multiple small targets and complex backgrounds.

[0091] For fair comparison, the same training datasets as in [reference] were adopted, including 1,485 images from the NJU2K dataset, 700 images from the NLPR dataset, and 800 images from the DUT, totaling 2985 samples to train the algorithm of the present invention. The remaining images of the NJU2K, NLPR, and DUT datasets, as well as the entire datasets of SIP, DES, and LFSD, were used for testing.

[0092] The present invention adopted four widely used evaluation metrics to evaluate this model, namely E-measure (E ξ ), S-measure (S m ), F-measure (F β ), and mean absolute error (MAE). Specifically, E-measure (E ξ ) is used to measure the local pixel-level error and global image-level error. S-measure (S m ) evaluates the regional perception and object perception spatial structure similarity of the saliency map. F-measure (F β) is the weighted harmonic mean of precision and recall, which can be used to evaluate the overall performance of the system. MAE measures the average of the absolute differences per pixel between the saliency map and the ground truth map. In the experiment, the values of the E metric and the F metric are adaptive.

[0093] During the training and testing phases, the sizes of the input RGB images and depth images are adjusted to 384×384. At the same time, the depth image is duplicated into 3 channels to be consistent with the RGB image. During the training process, augmentation strategies such as random flipping, rotation, and boundary cropping are adopted for the training images to prevent overfitting. The Swin-B pre-trained model is used to initialize the parameters of the backbone network, and the remaining parameters are initialized to the PyTorch default settings. The Adam optimizer is used to train the network, with the BatchSize set to 8, the initial learning rate to 5e-5, and the learning rate divided by 10 every 100 epochs. The model of the present invention is trained on a machine with a single NVIDIA GTX3090 GPU. The model converges in about 150 epochs, and the training time is about 12 hours.

[0094] Comparison with the state-of-the-art methods: The model is compared with 12 latest RGB-D salient object detection models, namely CoNet, AILNet, DCF, TriTransNet, EBFSP, HAINet, JL-DCF, SwinNet, BPGNet, SPSN, C2DFNet, and CIRNet. To ensure the fairness of the comparison results, the evaluated saliency maps are provided by the authors or generated by running the source code.

[0095] The quantitative results on 6 widely used datasets are shown in Table 1. According to the results of the 4 evaluation metrics, it can be seen that the algorithm proposed in the present invention has achieved excellent results on the 6 datasets. Among them, the best results are obtained in all metrics on DUT, DES, and LFSD, thus verifying the effectiveness and generalization of the algorithm of the present invention. It is worth noting that the improvement effect of the algorithm of the present invention on the LFSD dataset is significant, and each metric is improved by about 1% compared with the second-best result. This dataset contains multiple small targets and complex backgrounds, which shows that the algorithm of the present invention has strong robustness in difficult scenarios.

[0096] Table 1 Quantitative metrics of the advanced algorithms and the algorithm proposed in the present invention on six RGB-D datasets

[0097]

[0098] To qualitatively evaluate the performance of this algorithm, the results of the algorithm of the present invention are visually compared with those of some representative latest algorithms, including some representative difficult scenarios, such as the cases where the foreground and background are similar (lines 1-2), complex scenarios (lines 3-4), low-quality depth maps (lines 5-6), multi-targets (lines 7-8) and small targets (lines 9-10). The results are as Figure 5 shown. It can be seen from the results that the model of the present invention can more accurately locate and segment salient targets, and can still ensure excellent detection performance in difficult scenarios, verifying the effectiveness and robustness of the model.

[0099] Ablation experiment: Effectiveness of the cross-modal attention fusion module

[0100] 1) Verify the effectiveness of the fusion strategy. The following experiments were carried out: (a) Fusing the current-level features of the RGB image and the depth image as the baseline model. (b) Fusing the low-level features and the current-level features of the RGB image and the depth image. (c) Fusing the high-level features, low-level features and current-level features of the RGB image and the depth image. (d) Fusing the high-level features of the RGB image and the current-level features, and the current-level features of the depth image. (e) The multi-scale context feature aggregation adopted in the present invention, that is, fusing the high-level features and the current-level features of the RGB image and the depth image.

[0101] The experimental results are shown in Table 2. It can be seen from the results that compared with the baseline model, experiments (b), (c), (d) and (e) all have a certain degree of performance improvement, verifying the effectiveness of multi-scale features. On the other hand, comparing experiments (b), (c) and (e), it can be seen that the results of experiments (c) and (e) are better than those of experiment (b), which shows that compared with the supplementary details brought by low-level features, the guiding role of high-level features is more crucial. And the result of experiment (e) is better than that of experiment (c). On the one hand, it may be because some background noise information is introduced in the process of fusing low-level features; on the other hand, the participation of low-level features and high-level features in the fusion will also cause a certain degree of feature redundancy, thus affecting the final result. Finally, by comparing experiments (d) and (e), it can be seen that the multi-scale features of the depth image and the RGB image can both make certain contributions.

[0102] Table 2 Ablation experiment results of multi-scale context feature aggregation. The red result is the best, and the blue is the second best

[0103]

[0104] 2) Verify the effectiveness of the fusion module. Four experiments were carried out: (a) Cross-modal fusion of RGB features and depth features using the channel-spatial attention module proposed by CBAM. (b) On the basis of the CBAM module, a strategy of scaling and fusing adjacent layer features was added. (c) Using the cross-modal fusion module (CM) proposed in JL-DCF, the same strategy of scaling and fusing adjacent layer features was also adopted. (d) Using the cross-modal attention fusion module proposed in the present invention.

[0105] The experimental results are shown in Table 3. It can be observed from experiments (a) and (b) that, compared with the baseline model, the strategy of scaling and fusing adjacent layer features can significantly improve the detection results, and there are obvious improvements in the results of the four evaluation indexes on the three data sets. Moreover, comparing experiments (b), (c), and (d), it can be seen that, compared with the other two feature fusion modules, the CAM of the present invention has obtained the best results. This shows that the cross-modal attention fusion module proposed in the present invention, by means of the one-dimensional encoded attention mechanism, can obtain longer-distance attention information on each dimension and then combine them, so as to effectively realize the cross-modal fusion of RGB features and depth features.

[0106] Table 3 Ablation experiment results of the cross-modal attention fusion module

[0107]

[0108] Verify the effectiveness of the progressive fusion decoder: Replace the residual convolution block in the progressive fusion decoder with a single-layer convolution to form a decoder as the baseline model, and compare the performance gap between the progressive fusion decoder (PFD) and the progressive decoder (PFD') without fusing low-level features. The experimental results are shown in Table 4. Comparing experiments (a) and (b), it can be seen that, compared with the single-layer convolution decoder, the progressive fusion decoder can further extract and retain effective significant information by means of the residual convolution block, while reducing the introduction of noise in the process of fusing low-level features. Comparing experiments (b) and (c), it can be seen that fusing low-level features can significantly improve the detection effect. This is because edge detail information is often lost in the process of extracting high-level features, and the effective supplement can be obtained by fusing low-level features to achieve the accurate segmentation of significant objects.

[0109] Table 4 Ablation experiment results of the progressive fusion decoder

[0110]

[0111] Verify the effectiveness of the loss function: A series of ablation experiments were carried out on the hybrid loss function adopted by the present invention to verify its effectiveness, including (a) BCE loss function. (b) IoU loss function. (c) Hybrid loss function composed of BCE and IoU. (d) The loss function of the present invention, that is, applying a deep supervision strategy to multi-level features on the basis of the hybrid loss function. The experimental results are shown in Table 5. The BCE loss function focuses on supervising all pixels while the IoU loss function mainly focuses on the foreground. By combining the two, their advantages can be taken into account. Experiment (c) is improved compared with (a) and (b) in most metrics, which verifies this conjecture. In addition, introducing the deep supervision strategy to generate predictions and conduct supervision from multi-level features helps to further correct the prediction results by using multi-scale information.

[0112] Table 5 Ablation experiment results of different loss functions, the red results are the best, and the blue are the second best

[0113]

[0114] Verify the redundancy problem of the additional module

[0115] In order to further verify the redundancy problem of the additional module, based on the model of the present invention, an experimental analysis was carried out on the necessity of using additional additional modules for feature enhancement and edge generation in the RGB-D salient object detection task. The model of the present invention was used as the baseline model, and a feature enhancement module and an edge generation module were added respectively.

[0116] The experimental results are shown in Table 6. Comparing experiments (a) and (b), it can be seen that after adding the feature enhancement module, there is a certain degree of reduction in all metrics of the three datasets. This is because the cross-modal attention fusion module of the present invention has obtained a large amount of multi-scale information by using the strategy of scaling and fusing adjacent layer features, and the semantic information in the high-level features is already rich enough. Therefore, adding an additional feature enhancement module will cause overfitting and affect the final result. Comparing experiments (a) and (c), it can be seen that adding a separate edge generation module does not significantly affect the final detection result. This is because the progressive fusion decoder designed by the present invention incorporates low-level features during the decoding process, obtains a large amount of edge information, and filters out the noise through residual convolutional blocks. Therefore, excellent results can be obtained without relying on an additional edge generation module.

[0117] Table 6 Ablation experiment results of additional modules

[0118]

[0119] In view of the problems of feature redundancy and low efficiency caused by the additional modules in RGB-D salient object detection for achieving precise boundary prediction and feature enhancement, a concise RGB-D salient object detection framework is designed from the perspective of module necessity. A cross-modal attention fusion module is used to achieve complementary fusion of depth features and RGB features, and multi-scale information is mined by integrating context features. In addition, a progressive fusion decoder is designed to fuse and extract detailed information in low-level features during the decoding process to achieve precise saliency prediction. Experimental results on six datasets show that this method improves the performance to a new level compared with other state-of-the-art algorithms.

[0120] The present invention also provides a salient object detection device for cross-modal feature fusion and progressive decoding, including:

[0121] An image acquisition module for acquiring an image to be detected;

[0122] A two-stream Swin Transformer encoder for extracting multi-level and multi-scale RGB features and depth features from the image to be detected;

[0123] A cross-modal attention fusion module for fusing the multi-level and multi-scale RGB features and depth features to obtain fused features;

[0124] A progressive fusion decoder for decoding the high-level fused features in the fused features and gradually fusing low-level features during the decoding process.

[0125] The embodiment of the present invention also provides a storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, it implements some or all of the steps in the embodiments of the salient object detection method for cross-modal feature fusion and progressive decoding provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM for short), a random access memory (RAM for short), etc.

[0126] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the parts that contribute to the prior art can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present invention.

[0127] For the same and similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the embodiments of the saliency object detection device for cross-modal feature fusion and progressive decoding, since it is basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method embodiments.

[0128] The description in the method embodiments can be referred to.

[0129] The above-described embodiments of the present invention do not constitute a limitation to the protection scope of the present invention.

Claims

1. A saliency object detection method based on cross-modal feature fusion and progressive decoding, characterized in that Including: Obtain the image to be detected; Extract multi-level and multi-scale RGB features and depth features from the image to be detected through a two-stream Swin Transformer encoder; Fuse the multi-level and multi-scale RGB features and depth features through a cross-modal attention fusion module to obtain fused features, specifically including: for the input features of the high-level adjacent layer features perform scaling, with the top layer replaced by the current layer to maintain alignment; adjust the spatial resolution to the same as the current level through upsampling operations; concatenate the two input features and align the number of channels with to obtain Concatenate with to obtain the multi-scale feature F i ; Among them, UP(·) represents the bilinear interpolation upsampling operation, Cat(·) represents the concatenation operation, and Conv(·) represents the 3*3 convolution operation; Use two one-dimensional average pooling operations to process the multi-scale feature F i Embed the direction information; concatenate them and input them into a conversion layer to compress the channels; separate the feature map embedded with direction information along the x and y directions, then generate encoded attention maps in their respective directions through an encoded attention layer, and multiply them with F i to achieve channel attention perception; Obtain spatial attention perception through the spatial attention module, and multiply the output by F i to obtain the final fused feature Among them, p x and p y represent average pooling operations in the horizontal and vertical directions; ConvBS(·) represents a conversion layer composed of a convolutional layer, a BN layer, and a Sigmoid layer; CA x (·) and CA y (·) represent the generation of channel attention along the x and y directions, which is implemented by a convolutional layer containing a Sigmoid layer, and SA(·) represents the spatial attention layer; Decode the high-level fusion features in the fusion features through a progressive fusion decoder, and gradually fuse the low-level features during the decoding process. Specifically, it includes: Decode the high-level fusion features in the fusion features through a progressive fusion decoder, and gradually fuse the low-level features during the decoding process, including: After obtaining the fusion features using the cross-modal attention fusion module the high-level fusion features are input into the progressive fusion decoder for decoding, and low-level features are gradually fused during the decoding process; three residual convolutional modules with different dimensions are adopted to replace the individual convolutional layers for decoding, and the specific process is as follows: Among them, RCM i (·) represents the residual convolution module, Cat(·) represents the concatenation operation, and F final (·) represents the final feature.

2. The saliency object detection method for cross-modal feature fusion and progressive decoding according to claim 1, wherein Extract multi-level and multi-scale RGB features and depth features from the image to be detected through a two-stream Swin Transformer encoder, including: Copy the depth image into 3 channels; Segment the image to be detected into non-overlapping blocks through a fragment segmentation operation; Extract features of different scales at four stages from the RGB image and the depth image respectively, where the RGB features are represented as and the depth features are represented as Each stage consists of a fragment fusion layer and multiple stacked Swin Transformer blocks, where the fragment fusion layer of the first stage is replaced by a linear embedding layer.

3. The saliency object detection method for cross-modal feature fusion and progressive decoding according to claim 1, wherein The decoding method of the residual convolution module includes: Pass the input feature through a depthwise separable convolution layer and an LN layer; Adjust the number of channels through two pointwise convolution layers; Add the input feature and the output feature, and adjust the feature size through an upsampling layer. The specific process is as follows: RCM(f) = UP(f + PW2(σ(PW1(LN(DW(f))))) (4) Among them, σ(·) is the GELU activation function, UP(·) represents the upsampling layer, f represents the input feature, DW(·) represents the depthwise separable convolution layer, PW(·) represents the pointwise convolution layer, and LN(·) represents the regularization layer.

4. The saliency object detection method for cross-modal feature fusion and progressive decoding according to claim 3, wherein The method further includes: From the high-level features and each level of residual convolution module respectively generate the significant prediction map P i (i = 1, 2, 3, 4), and use the hybrid loss composed of BCE loss and IoU loss to supervise it.

5. The saliency object detection method for cross-modal feature fusion and progressive decoding according to claim 4, wherein BCE loss L BCE is defined as: (5) Among them, W and H respectively represent the width and height of the image, P(x, y) represents the predicted coordinates, and G(x, y) represents the ground truth coordinates.

6. The saliency object detection method for cross-modal feature fusion and progressive decoding according to claim 5, wherein IoU loss L IoU is defined as: Among them, W and H respectively represent the width and height of the image, P(x, y) represents the predicted coordinates, and G(x, y) represents the ground truth coordinates.

7. The saliency object detection method for cross-modal feature fusion and progressive decoding according to claim 6, wherein The overall loss L of the model is defined as: Among which P i is the generated significant prediction map, and G is the ground truth map.

8. A saliency object detection device for cross-modal feature fusion and progressive decoding, characterized in that, Including: An image acquisition module for obtaining the image to be detected; A two-stream SwinTransformer encoder for extracting multi-level and multi-scale RGB features and depth features from the image to be detected; Cross-modal attention fusion module, which is used to fuse the multi-level and multi-scale RGB features and depth features to obtain fused features, specifically including: for the input features of the high-level adjacent layer features scale them, and the highest layer is replaced by the current layer to maintain alignment; adjust the spatial resolution to the same as the current level through upsampling operation; concatenate the two input features and align the number of channels with through a convolutional layer to obtain Concatenate with to obtain the multi-scale feature F i ; Among them, UP(·) represents the bilinear interpolation upsampling operation, Cat(·) represents the concatenation operation, and Conv(·) represents the 3*3 convolution operation; Using two one-dimensional average pooling operations to process the multi-scale feature F i Embed the direction information; concatenate them and input to the conversion layer to compress the channels; separate the feature map embedded with direction information along the x and y directions, then generate encoded attention maps in their respective directions through the encoded attention layer, and multiply with F i to achieve channel attention perception; Obtain spatial attention perception through the spatial attention module, and multiply the output by F i to obtain the final fused feature Among them, p x and p y represent average pooling operations in the horizontal and vertical directions; ConvBS(·) represents a conversion layer composed of a convolutional layer, a BN layer, and a Sigmoid layer; CA x (·) and CA y (·) represent the generation of channel attention along the x and y directions, which is achieved through a convolutional layer containing a Sigmoid layer, and SA(·) represents the spatial attention layer; A progressive fusion decoder for decoding the high-level fusion features in the fusion features and gradually fusing the low-level features during the decoding process. Specifically, it includes: Decode the high-level fusion features in the fusion features through a progressive fusion decoder, and gradually fuse the low-level features during the decoding process, including: After obtaining the fusion features using the cross-modal attention fusion module the high-level fusion features are input into the progressive fusion decoder for decoding, and low-level features are gradually fused during the decoding process; three residual convolutional modules with different dimensions are adopted to replace the individual convolutional layers for decoding, and the specific process is as follows: Among them, RCM i (·) represents the residual convolution module, and Cat(·) represents the concatenation operation, F final (·) represents the final feature.