A salient object detection method based on multi-modal feature refinement and fusion

CN117975216BActive Publication Date: 2026-09-29NANCHANG HANGKONG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410192199.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-21
Publication Date
2026-09-29
Estimated Expiration
2044-02-21

AI Technical Summary

Technical Problem

[0005]本发明的目的在于,提供一种基于多模态特征细化和融合的显著性物体检测方法,缓解了现有RGB-D显著性检测方法中存在的目标边缘检测模糊及内部检测不完整的问题,并且能够有效对特征进行增强并将增强后的特征进行充分融合,缓解了现有模型在特征提取阶段所提取到特征模态单一以及包含过多非显著特征的问题

Benefits of technology

(1)本发明通过跨层级交叉注意力融合模块ICFM和深度特征聚合模块DFAM分别对RGB特征和深度特征进行细化增强;跨层级交叉注意力融合模块ICFM利用来自上层的语义来指导低级特征的生成,使用交叉注意力操作来融合多分辨率特征,并沿着自上而下的路径逐步传播更高层次的语义信息,从而可以有效地抑制背景中的冗余特征并强调显著区域;深度特征聚合模块DFAM通过膨胀卷积和分组卷积以及残差网络的思想,对传入的特征进行通道维度的分割并将分割之后的特征进行卷积,在卷积过程中将相邻的结果进行融合获得包含多尺度信息的特征,最后将包含多尺度信息的特征与膨胀卷积特征以及原特征进行融合,可以有效的获取多尺度特征以及得到更大的感受野,达到强化显著特征的目的;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117975216B_ABST
    Figure CN117975216B_ABST
Patent Text Reader

Abstract

The application relates to a salient object detection method based on multi-modal feature refinement and fusion, which comprises the following steps: constructing a salient object detection model, proposing RGB features and depth features through a Swin-Transformer backbone network, refining and enhancing the features of the two modes through a cross-layer cross-attention fusion module and a depth feature aggregation module, complementarily fusing the features through a multi-level attention mechanism module, generating cross-modal features through a grouping attention unit, decoding through a decoder, fusing the decoding result with edge feature extraction of a CNN refinement unit to generate a predicted salient map. The application alleviates the problems of fuzzy target edge detection and incomplete internal detection in the existing salient detection method, can effectively enhance and sufficiently fuse the features, and alleviates the problems of single feature mode and excessive non-salient features in the feature extraction stage of the existing model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, specifically to a salient object detection method based on multimodal feature refinement and fusion. Background Technology

[0002] In recent years, with the deepening of deep learning research, salient object detection (SOD) has been significantly improved. However, identifying salient objects in complex scenes still faces challenges, such as when the target and background colors are similar, when the scene is too bright or too dark due to lighting changes, or when encountering complex foregrounds like transparent objects. Relying solely on RGB images for accurate detection remains a significant challenge. Inspired by the fact that the human eye can perceive not only appearance information such as color, shape, and texture in a scene, but also capture depth information through a binocular vision system, forming a sense of depth, depth maps, as an intuitive representation of depth information, have been introduced into RGB salient object detection to supplement this information.

[0003] Early work on RGB-D salient object detection focused on handcrafted depth features with prior knowledge, such as the idea that closer objects make objects more salient. These methods typically used previously employed contrast-based network frameworks for predictions at various scales. While handcrafted feature-based RGB-D salient object detection has achieved considerable success, it still faces significant challenges in dealing with salient objects in complex backgrounds and with noisy depth information.

[0004] Due to the limitations of handcrafted feature-based methods and the superiority of deep learning techniques in extracting complex features, more and more RGB-D detection work has adopted deep neural networks, especially convolutional neural networks, in recent years. These include single-stream and two-stream RGB-D salient object detection models. Therefore, some researchers have designed single-stream networks, mimicking RGB salient object detection, and achieved some success. In their 2019 paper, "Contrast prior and fluid pyramid integration for RGBD salient object detection," Zhao JX et al. proposed a fluid pyramid module to combine features at different levels through numerous short connections to enhance the information representation capabilities of RGB-D. This cleverly integrates deep features into the proposed single-stream network through a feature enhancement module. However, the structure of single-stream networks limits their integration; the information contained in depth maps and RGB images has different emphases and cannot be easily integrated. Two-stream networks, on the other hand, can effectively utilize information, resulting in more comprehensive feature extraction. Therefore, two-stream networks are more widely used. This means that each network branch will have a corresponding CNN and a corresponding node to fuse them to obtain the final saliency prediction. Based on the depth of the fusion point, two-stream networks can be categorized into three types: early fusion, late fusion, and multi-level fusion. For example, Qu L, He S, and Zhang J et al., in their 2017 paper "RGBD salient object detection via deep fusion," proposed fusing low-level features from RGB and deep information streams as a shared CNN for training. However, this approach suffers from insufficient feature extraction and an inability to distinguish features from different modalities. In contrast to the early fusion mechanism, Han J, Chen H, and Liu N et al., in their 2017 paper "CNNs-based RGB-D saliency detection via cross-view transfer and multiview fusion," proposed combining features from the two-stream network to obtain robust global contextual information. This method achieves better object localization than early fusion, but due to the lack of shallow fusion information, the resulting prediction map is relatively coarse. To explore multi-level cross-modal features, recent works have shifted their focus to designing multi-level cross-modal fusion frameworks. Summary of the Invention

[0005] The purpose of this invention is to provide a salient object detection method based on multimodal feature refinement and fusion, which alleviates the problems of blurred target edge detection and incomplete internal detection in existing RGB-D salient object detection methods. Furthermore, it can effectively enhance features and fully fuse the enhanced features, thereby alleviating the problems of existing models extracting features with single modalities and containing too many non-salient features during the feature extraction stage.

[0006] The technical solution adopted in this invention is: a salient object detection method based on multimodal feature refinement and fusion, comprising the following steps: S1: Construct a salient object detection model, which includes a Swin-Transformer backbone network, a cross-level attention fusion module ICFM, a deep feature aggregation module DFAM, a multi-level attention mechanism module MAF, a group attention unit GA, and a decoder; The Swin-Transformer backbone network includes an RGB branch and a Depth branch. Both the RGB branch and the Depth branch include a patch embedding module and n stages. The RGB branch is used to extract RGB features from different stages of the input image. The Depth branch is used to extract depth features at different stages of the input image. ; The cross-level cross-attention fusion module (ICFM) has n-1 stages, and the RGB features extracted from the first n-1 stages of the RGB branch are from different stages. The features are refined by inputting them into n-1 cross-level cross-attention fusion modules (ICFM) to generate multi-resolution features. ; The Deep Feature Aggregation (DFAM) module has n-1 stages, and the depth features extracted from the first n-1 stages of the Depth branch are from different stages. The enhanced features are generated by inputting each of the n-1 deep feature aggregation modules (DFAM). ; The Multi-Level Attention (MAF) module is used to extract RGB features from the nth stage of the RGB branch. and the depth features extracted in the nth stage of the Depth branch Complementary fusion is performed to generate attention-enhanced features h, which enhances the expressive power of high-level semantics of the features; The group attention unit (GA) is configured with n units, and the GA is used to fuse the RGB features of the corresponding stage. and depth features The system generates cross-modal features, which are processed by the decoder to generate the decoded features of the current stage. These features are then input into the Group Attention Unit (GA) of the next stage to guide the interactive fusion of RGB and depth features. The input to the GA of the nth stage is the output of the Multi-Level Attention (MAF) module, while the inputs to the GAs of the remaining n-1 stages are the multi-resolution features of the corresponding stage. and enhanced features And the decoding features of the previous stage output by the decoder; S2: Train the salient object detection model, use the trained salient object detection model to detect the input image, and output the decoding features of each stage by the decoder. S3: Use the CNN thinning unit to extract the edge features of the RGB image in the input image, and fuse them with the decoding features of the nth stage through attention. Regress the fused features into a salient object image to generate a predicted saliency map.

[0007] Furthermore, the RGB features of the three consecutive stages are... , and The coupling is divided into two groups, where the first group consists of the RGB features of the i-th stage. and the RGB features of the (i+1)th stage The second group represents the RGB features of the (i+1)th stage. and the RGB features of the (i+2)th stage Let i = 1, 2, ..., n, and input the two pairs of paired RGB features into the i-th cross-level attention fusion module ICFM. For each pair of paired RGB features, the higher-stage RGB features are copied twice and embedded as key vector K and value vector V, respectively. The lower-stage RGB features are copied once and embedded as query vector Q. Standard-scale dot product attention operations and linear projections are performed on each pair of key vector K, query vector Q, and value vector V to generate attention features ca. The two attention features ca are transformed in terms of channel and spatial resolution and then added to the lower-stage RGB features to generate fused features z. An improved ResNet module is used to enhance the fused features z, generating multi-resolution features. The improved ResNet module includes layer normalization, two consecutive linear MLP layers, and residual connections.

[0008] Furthermore, the specific expression for the feature processing of the i-th cross-level attention fusion module ICFM is as follows: ; ; ; ; ; ; in, Represents the RGB features of the i-th stage The corresponding query vector after copying This represents the weight matrix corresponding to the query vector Q. Represents the RGB features of the (i+1)th stage The corresponding key vector after copying, This represents the weight matrix corresponding to the key vector K. Represents the RGB features of the (i+1)th stage The corresponding value vector after copying This represents the weight matrix corresponding to the value vector V. Represents the query vector Key vector Sum value vector The attention features obtained from the calculation, Represents the query vector Key vector Sum value vector Dimensions Represents the key vector transpose, Indicates a linear operation. This represents the activation function operation. This represents the fusion feature generated by the i-th cross-level attention fusion module ICFM. This indicates an upsampling operation with a multiplier of 2. This indicates an upsampling operation with a multiplier of 4. This represents the multi-resolution feature generated by the i-th cross-level attention fusion module ICFM. This represents two consecutive linear layer operations. Presentation layer normalization operation, This indicates an addition operation.

[0009] Furthermore, the deep feature aggregation module DFAM expands the receptive field of the corresponding stage's deep features through dilated convolutional layers to generate dilated convolutional features X. dconv Simultaneously, the depth features of the corresponding stage are convolved and then segmented into N grouped features X along the channel dimension. These grouped features X are then convolved, and adjacent results are fused during the convolution process to generate N convolutionally fused features Y. These N convolutionally fused features Y are then convolved to generate feature Y'. Finally, the features X are dilated and convolved...dconv The deep features and feature Y' of the corresponding stage are fused to generate enhanced feature Y. d .

[0010] Furthermore, the specific expression for the feature processing of the i-th deep feature aggregation module DFAM is as follows: ; ; ; ; ; ; in, Let C represent the depth features of the i-th stage, and let C denote a 1×1 convolution operation. Represents the depth features of the i-th stage The features obtained after convolution are denoted by Split, which represents the segmentation of the features along the channel dimension, and N represents the number of segments. This represents the feature of the j-th group. Represents the j-th convolutional fusion feature. This represents a 3×3 convolution operation. Represents N convolutional fused features Features are generated after convolution. concat represents the concatenation operation of channel dimensions. Represents the depth features of the i-th stage The dilated convolution feature obtained after the dilated convolution operation, Dconv represents the dilated convolution operation. This indicates an addition operation.

[0011] Furthermore, the Multi-Level Attention Mechanism (MAF) module includes two attention fusion modules (AFM) and one multimodal fusion module (MFM). The AFM module comprises a spatial attention mechanism, a channel attention mechanism, and a cross-level attention fusion module (ICFM). The two AFM modules respectively handle the RGB features of the nth stage. and the depth features of the nth stage Enhancement is performed to generate RGB attention fusion features. Features fused with deep attention RGB attention fusion features Features fused with deep attention Adding them together yields the attention fusion features. and RGB attention fusion features Deep attention fusion features and attention fusion features Input the multimodal fusion module (MFM); The multimodal fusion module MFM is based on RGB attention fusion features. Deep attention fusion features and attention fusion features Generate a confusion matrix HY, and then fuse the confusion matrix HY with the RGB attention features. Features fused with deep attention Perform calculations to generate RGB attention-enhanced features. and deep attention enhancement features .

[0012] Furthermore, the specific expression for the feature processing of the multi-level attention mechanism module MAF is as follows: ; ; ; ; ; ; in, This represents the RGB attention fusion feature. This represents the deep attention fusion feature. Indicates attention fusion features, This indicates the operation of the spatial attention mechanism. This indicates the channel attention mechanism operation. This represents the convolution operation. Represents the confusion matrix. Indicates shape reshaping. This represents a 1×1 convolution operation. This represents RGB attention enhancement features. This indicates a deep attention enhancement feature.

[0013] Furthermore, a loss function is introduced as a supervision signal during the decoding process of the decoder and the generation of the predicted saliency map by the CNN refinement unit. The binary cross-entropy loss is used as the loss function to measure the SSIM loss and IOU loss for structural similarity. The specific calculation formula is as follows: ; ; in, Indicates the overall loss. This represents the combination of SSIM loss and IOU loss. This represents the final predicted saliency map. This indicates a correct label that was manually marked. This represents the saliency map for the prediction at the m-th stage. This represents the correct label after downsampling in the m-th stage. , Indicates the loss of bce. Indicates SSIM loss, This indicates Iou's loss.

[0014] Furthermore, the CNN thinning unit uses the first two layers of the VGG network backbone to extract edge features of the RGB image in the input image.

[0015] The beneficial effects of this invention are as follows: (1) In this invention, the cross-level attention fusion module ICFM and the deep feature aggregation module DFAM refine and enhance the RGB features and deep features respectively. The cross-level attention fusion module ICFM uses semantics from the upper layer to guide the generation of low-level features, uses cross-attention operation to fuse multi-resolution features, and gradually propagates higher-level semantic information along the top-down path, thereby effectively suppressing redundant features in the background and emphasizing salient regions. The deep feature aggregation module DFAM uses the ideas of dilated convolution, grouped convolution and residual network to segment the input features by channel dimension and convolve the segmented features. During the convolution process, adjacent results are fused to obtain features containing multi-scale information. Finally, the features containing multi-scale information are fused with the dilated convolution features and the original features, which can effectively obtain multi-scale features and obtain a larger receptive field, thereby achieving the purpose of strengthening salient features. (2) The present invention uses the multi-level attention mechanism module MAF to perform complementary fusion of RGB features and deep features to generate attention-enhanced features, which enhances the expressive power of high-level semantics of features, can more effectively transmit encoder features to the decoder level, and refines input features and suppresses noise in the process. (3) The present invention uses group attention unit (GA) to provide global guidance for the interaction process at each position. The mask inside the group attention unit (GA) suppresses the negative impact when features from different scales and different modalities interact, so that the feature information of the two modalities can fully interact under the guidance of the saliency guidance vector, and generate a more accurate prediction saliency map. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart of an embodiment of the present invention; Figure 2 This is a structural diagram of the cross-level attention fusion module ICFM in an embodiment of the present invention; Figure 3 This is a structural diagram of the Deep Feature Aggregation Module (DFAM) in an embodiment of the present invention; Figure 4 The diagram shows the structure of the Multi-Level Attention Mechanism Module (MAF) in this embodiment of the invention, where (a) is a schematic diagram of the network structure of the MAF, (b) is a schematic diagram of the network structure of the Attention Fusion Module (AFM), and (c) is a schematic diagram of the network structure of the Multi-Modal Fusion Module (MFM). Figure 5 This is a visual comparison of the output results of the present invention with existing RGB-D saliency object detection methods on the NLPR dataset, LFSD dataset, NJU2K dataset, DUT dataset, STERE1000 dataset, and COME15K dataset. Detailed Implementation

[0018] To better understand the above-described objects, features, and advantages of the present invention, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Many specific details are set forth in the following description to provide a thorough understanding of the invention; however, the invention may be practiced in other ways different from those described herein, and therefore, the invention is not limited to the specific embodiments disclosed below.

[0019] like Figure 1 As shown in the figure, this invention proposes a salient object detection method based on multimodal feature refinement and fusion. The method uses a set of corresponding RGB images and depth images as input images for salient object detection, and includes the following steps: S1: Construct a salient object detection model, which includes a Swin-Transformer backbone network, a cross-level attention fusion module ICFM, a deep feature aggregation module DFAM, a multi-level attention mechanism module MAF, a group attention unit GA, and a decoder.

[0020] The Swin-Transformer backbone network includes an RGB branch and a Depth branch. Both the RGB branch and the Depth branch include a patch embedding module and n stages. The RGB branch is used to extract RGB features from different stages of the input image. The Depth branch is used to extract depth features at different stages of the input image. In this embodiment of the invention, n is set to 4. Since the features extracted by the Swin-Transformer backbone network still retain a lot of noise information, this embodiment of the invention refines the features of different modalities separately, enhancing salient features in the feature maps and suppressing non-obvious features.

[0021] For the features of the RGB modality, a cross-level cross-attention fusion module (ICFM) is used for feature refinement. The ICFM module has n-1 stages, and the RGB features extracted from the first n-1 stages of the RGB branch represent the features from different stages. The features are refined by inputting them into n-1 cross-level cross-attention fusion modules (ICFM) to generate multi-resolution features. .

[0022] Specifically, such as Figure 2 As shown, this embodiment of the invention uses RGB features from three consecutive stages. , and The coupling is divided into two groups, where the first group consists of the RGB features of the i-th stage. and the RGB features of the (i+1)th stage The second group represents the RGB features of the (i+1)th stage. and the RGB features of the (i+2)th stage Let i = 1, 2, ..., n. Two pairs of paired RGB features are input into the i-th cross-level attention fusion module (ICFM). The higher-stage RGB features in each pair are copied twice, embedded as key vector K and value vector V respectively, while the lower-stage RGB features are copied once, embedded as query vector Q. Key vector K, value vector V, and query vector Q have the same dimension, d. A standard-proportion dot product attention operation and linear projection are performed on each pair of key vector K, query vector Q, and value vector V to generate attention feature ca, ensuring that the number of channels in attention feature ca is the same as that of the RGB features in the i-th stage. The two attention features ca are transformed in terms of both channel and spatial resolution, and then added to the lower-stage RGB features to generate a fused feature z. An improved ResNet module is used to enhance the fused feature z, generating a multi-resolution feature. The improved ResNet module includes layer normalization, two consecutive linear MLP layers, and residual connections.

[0023] For the i-th cross-level attention fusion module ICFM, the specific expression for its feature processing is as follows: ; ; ; ; ; ; in, Represents the RGB features of the i-th stage The corresponding query vector after copying This represents the weight matrix corresponding to the query vector Q. Represents the RGB features of the (i+1)th stage The corresponding key vector after copying, This represents the weight matrix corresponding to the key vector K. Represents the RGB features of the (i+1)th stage The corresponding value vector after copying This represents the weight matrix corresponding to the value vector V. Represents the query vector Key vector Sum value vector The attention features obtained from the calculation, Represents the query vector Key vector Sum value vector Dimensions Represents the key vector transpose, Indicates a linear operation. This represents the activation function operation. This represents the fusion feature generated by the i-th cross-level attention fusion module ICFM. This indicates an upsampling operation with a multiplier of 2. This indicates an upsampling operation with a multiplier of 4. This represents the multi-resolution feature generated by the i-th cross-level attention fusion module ICFM. This represents two consecutive linear layer operations. Presentation layer normalization operation, This indicates an addition operation.

[0024] Since the value of n in this embodiment of the invention is 4, the fusion feature generated by the i-th cross-level cross-attention fusion module ICFM is... The multi-resolution features generated by the i-th cross-level cross-attention fusion module ICFM The expression can be represented as: ; .

[0025] The cross-level attention fusion module ICFM uses semantics from higher levels to guide the generation of low-level features, uses cross-attention operations to fuse multi-resolution features, and gradually propagates higher-level semantic information along a top-down path, thereby effectively suppressing redundant features in the background and emphasizing salient regions.

[0026] For the features of the depth modality, a Deep Feature Aggregation Module (DFAM) is used for feature refinement. The DFAM consists of n-1 modules, representing the depth features extracted from the first n-1 stages of the Depth branch at different stages. The enhanced features are generated by inputting each of the n-1 deep feature aggregation modules (DFAM). .

[0027] Specifically, such as Figure 3 As shown, the Deep Feature Aggregation Module (DFAM) expands the receptive field of the corresponding stage's deep features through dilated convolutional layers to generate dilated convolutional features X. dconv Simultaneously, the depth features of the corresponding stage are convolved and then segmented into N grouped features X along the channel dimension. These grouped features X are then convolved, and adjacent results are fused during the convolution process to generate N convolutionally fused features Y. These N convolutionally fused features Y are then convolved to generate feature Y'. Finally, the features X are dilated and convolved... dconv The deep features and feature Y' of the corresponding stage are fused to generate enhanced feature Y. d .

[0028] The specific expression for the feature processing of the i-th deep feature aggregation module DFAM is as follows: ; ; ; ; ; ; in, Let C represent the depth features of the i-th stage, and let C denote a 1×1 convolution operation. Represents the depth features of the i-th stage The features obtained after convolution are denoted by Split, which represents the segmentation of the features along the channel dimension, and N represents the number of segments. This represents the feature of the j-th group. Represents the j-th convolutional fusion feature. This represents a 3×3 convolution operation. Represents N convolutional fused features Features are generated after convolution. concat represents the concatenation operation of channel dimensions. Represents the depth features of the i-th stage The dilated convolution feature obtained after the dilated convolution operation, Dconv represents the dilated convolution operation. This indicates an addition operation.

[0029] In this embodiment of the invention, the kernel of the dilated convolutional layer is 3×3, and the dilation rate R is 2. Since N is also 4, the j-th convolutional fusion feature... The expression can be represented as: .

[0030] The Multi-Level Attention (MAF) module is used to extract RGB features from the nth stage of the RGB branch. and the depth features extracted in the nth stage of the Depth branch Complementary fusion is performed to generate attention-enhanced features h, which enhance the expressive power of high-level semantics. The multi-level attention mechanism module MAF can more effectively transfer encoder features to the decoder level, and refine the input features and suppress noise in the process.

[0031] like Figure 4 As shown in (a), the Multi-Level Attention Mechanism (MAF) module includes two attention fusion modules (AFM) and one multimodal fusion module (MFM); the attention fusion module (AFM) includes a spatial attention mechanism, a channel attention mechanism, and a cross-level attention fusion module (ICFM).

[0032] like Figure 4 As shown in (b), the two attention fusion modules (AFM) respectively process the RGB features of the nth stage. and the depth features of the nth stage Enhancement is performed to generate RGB attention fusion features. Features fused with deep attention RGB attention fusion features Features fused with deep attention Adding them together yields the attention fusion features. and RGB attention fusion features Deep attention fusion features and attention fusion features Input the multimodal fusion module (MFM). Figure 4 (b) and Represents the features of the nth stage from different modalities, where and , where r represents the RGB modality and d represents the depth modality. RGB attention fusion features. Deep attention fusion features and attention fusion features The specific expression is as follows: ; ; ; in, This represents the RGB attention fusion feature. This represents the deep attention fusion feature. Indicates attention fusion features, This indicates the operation of the spatial attention mechanism. This indicates the channel attention mechanism operation. This indicates a convolution operation.

[0033] like Figure 4 As shown in (c), the multimodal fusion module MFM is based on RGB attention fusion features. Deep attention fusion features and attention fusion features Generate a confusion matrix HY, and then fuse the confusion matrix HY with the RGB attention features. Features fused with deep attention Perform calculations to generate RGB attention-enhanced features. and deep attention enhancement features Confusion matrix HY, RGB attention enhancement features and deep attention enhancement features The specific expression is as follows: ; ; ; in, Represents the confusion matrix. Indicates shape reshaping. This represents a 1×1 convolution operation. This represents RGB attention enhancement features. Attention-enhancing features.

[0034] Since the input image consists of a set of corresponding RGB and depth images, the two modalities only have a clear correspondence at corresponding locations. Modeling the relationship between all pixels of both modalities would be computationally redundant, and such forced correlation modeling might introduce unnecessary noise. To address this issue, this embodiment of the invention uses a Grouped Attention Unit (GA) to provide global guidance for the interaction process at each location. The GA uses a mask to suppress negative impacts caused by interactions between features from different scales and modalities. Through the GA, the feature information of the two modalities can fully interact under the guidance of a saliency-guided vector. Finally, a linear layer fuses the modal features to form cross-modal features, which are then input into the decoder for decoding.

[0035] The group attention unit (GA) is configured with n units, and the GA is used to fuse the RGB features of the corresponding stage. and depth features The system generates cross-modal features, which are processed by the decoder to generate the decoded features of the current stage. These features are then input into the Group Attention Unit (GA) of the next stage to guide the interactive fusion of RGB and depth features. The input to the GA of the nth stage is the output of the Multi-Level Attention (MAF) module, while the inputs to the GAs of the remaining n-1 stages are the multi-resolution features of the corresponding stage. And enhanced features and decoding features from the previous stage of the decoder output.

[0036] In this embodiment of the invention, n is 4. Therefore, the fourth stage is the group attention unit (GA), i.e. Figure 1 The rightmost group attention unit (GA) receives the output of the multi-level attention mechanism module (MAF) as its input. The other three group attention units (GA) are... Figure 1 The three group attention units (GAs) on the left side of the image are inputs to the multi-resolution features of the corresponding stage. And enhanced features and decoding features from the previous stage of the decoder output.

[0037] S2: Train the salient object detection model, use the trained salient object detection model to detect the input image, and output the decoding features of each stage by the decoder.

[0038] S3 uses a CNN thinning unit to extract edge features from the RGB image in the input image and compares them with the decoded features from the nth stage (i.e., Figure 1 The leftmost decoded feature is fused using spatial attention, and the fused feature is regressed into a salient object image to generate a predicted saliency map. In this embodiment of the invention, the CNN thinning unit extracts edge features of the RGB image in the input image through the first two layers of the VGG network backbone.

[0039] In this embodiment of the invention, a loss function is introduced as a supervision signal during the decoding process of the decoder and the generation of the predicted saliency map by the CNN refinement unit. Binary cross-entropy loss is used as the loss function to measure the SSIM loss and IOU loss of structural similarity. The specific calculation formula is as follows: ; ;

[0040] in, Indicates the overall loss. This represents the combination of SSIM loss and IOU loss. This represents the final predicted saliency map. This indicates a correct label that was manually marked. This represents the saliency map for the prediction at the m-th stage. This represents the correct label after downsampling in the m-th stage. , Indicates the loss of bce. Indicates SSIM loss, This indicates Iou's loss.

[0041] The technical effects of the embodiments of the present invention are further illustrated below through simulation experiments: Tables 1 and 2 present the comparative experimental results of the embodiments of the present invention with 23 other existing RGB-D salient object detection methods on the DUT dataset, LFSD dataset, NLPR dataset, NJU2K dataset, and STERE1000 dataset. Structural similarity metrics were used in the experiments. Maximum comprehensive evaluation index Maximum Enhanced Matching Indicator and mean absolute error These four metrics comprehensively evaluate the scores of each method on the five widely used RGB-D datasets mentioned above. The top three results are displayed in three formats: bold underline, bold text, and " ". ↑ indicates that a higher metric is better, and ↓ indicates that a lower metric is better. "-" indicates that the code or result is unavailable.

[0042] Table 1. Comparative experimental results of the embodiments of the present invention with 23 other existing RGB-D saliency object detection methods on the DUT dataset, LFSD dataset, and NJU2K dataset.

[0043] Table 2. Comparative experimental results of the embodiments of the present invention with 23 other existing RGB-D salient object detection methods on the NLPR and STERE1000 datasets.

[0044] As shown in Tables 1 and 2, out of a total of 20 metrics across the five datasets, the present invention's embodiment exhibits the best performance in 9 metrics, and a total of 19 metrics rank among the top three of all methods. Therefore, the comprehensive performance of the method described in the present invention across the five datasets is superior to existing RGB-D salient object detection methods.

[0045] Table 3 shows the comparative experimental results of the embodiments of the present invention with seven other existing RGB-D salient object detection methods on the COME15K dataset. The experiments used structural similarity metrics. Maximum comprehensive evaluation index Maximum Enhanced Matching Indicator and mean absolute error These four metrics comprehensively evaluate the scores of each method on the five widely used RGB-D datasets mentioned above. The two best results are displayed in both bold underline and bold text formats. ↑ indicates that a higher metric is better, and ↓ indicates that a lower metric is better. "-" indicates that the code or result is unavailable.

[0046] Table 3. Comparative experimental results of the embodiments of the present invention with seven other existing RGB-D salient object detection methods on the COME15K dataset.

[0047] As shown in Table 3, out of a total of 8 indicators, 7 indicators of the present invention are the best, and a total of 8 indicators are ranked in the top two among all methods. Therefore, the comprehensive performance of the method described in the present invention on the COME15K dataset is better than the existing RGB-D salient object detection methods.

[0048] By visually comparing the output results of the embodiments of the present invention with those of existing RGB-D salient object detection methods on the NLPR, LFSD, NJU2K, DUT, STERE1000, and COME15K datasets, the following results can be obtained: Figure 5 The comparison results are shown in the figure. From Figure 5As can be seen, compared with other methods, the output results of the embodiments of the present invention are closest to the manually annotated graph GT.

[0049] Tables 4 and 5 demonstrate the effectiveness of the cross-level attention fusion module ICFM, the deep feature aggregation module DFAM, and the multi-level attention mechanism module MAF proposed in the embodiments of this invention. In the tables, "bold underline" indicates the best performance. FULL represents the salient object detection model described in the embodiments of this invention, w / o indicates removal of the current module, IC represents the cross-level attention fusion module ICFM, DF represents the deep feature aggregation module DFAM, and MA represents the multi-level attention mechanism module MAF. The rows ID1 to ID6 in the tables correspond to the model performance metrics after removing the cross-level attention fusion module ICFM, removing the deep feature aggregation module DFAM, removing the multi-level attention mechanism module MAF, simultaneously removing both the deep feature aggregation module DFAM and the multi-level attention mechanism module MAF, simultaneously removing both the deep feature aggregation module DFAM and the cross-level attention fusion module ICFM, and simultaneously removing both the multi-level attention mechanism module MAF and the cross-level attention fusion module ICFM, respectively. As can be seen from Tables 4 and 5, the salient object detection model described in this embodiment outperforms the model after removing the corresponding modules on all four datasets, indicating that the cross-level cross-attention fusion module ICFM, the deep feature aggregation module DFAM, and the multi-level attention mechanism module MAF proposed in this embodiment are all effective.

[0050] Table 4. Experimental results of the effectiveness of the proposed cross-level attention fusion module ICFM, deep feature aggregation module DFAM, and multi-level attention mechanism module MAF on the SIP and NJU2K datasets.

[0051] Table 5. Experimental results of the effectiveness of the proposed cross-level attention fusion module ICFM, deep feature aggregation module DFAM, and multi-level attention mechanism module MAF on the NLPR and STERE1000 datasets.

[0052] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A salient object detection method based on multimodal feature refinement and fusion, characterized in that, Includes the following steps: S1: Construct a salient object detection model, which includes a Swin-Transformer backbone network, a cross-level attention fusion module ICFM, a deep feature aggregation module DFAM, a multi-level attention mechanism module MAF, a group attention unit GA, and a decoder; The Swin-Transformer backbone network includes an RGB branch and a Depth branch. Both the RGB branch and the Depth branch include a patch embedding module and n stages. The RGB branch is used to extract RGB features from different stages of the input image. The Depth branch is used to extract depth features at different stages of the input image. ; The cross-level cross-attention fusion module (ICFM) has n-1 stages, and the RGB features extracted from the first n-1 stages of the RGB branch are from different stages. The features are refined by inputting them into n-1 cross-level cross-attention fusion modules (ICFM) to generate multi-resolution features. The cross-level cross-attention fusion module ICFM integrates the RGB features from three consecutive stages. , and The coupling is divided into two groups, where the first group consists of the RGB features of the i-th stage. and the RGB features of the (i+1)th stage The second group represents the RGB features of the (i+1)th stage. and the RGB features of the (i+2)th stage Let i = 1, 2, ..., n, and input the two pairs of paired RGB features into the i-th cross-level attention fusion module ICFM. For each pair of paired RGB features, the higher-stage RGB features are copied twice and embedded as key vector K and value vector V, respectively. The lower-stage RGB features are copied once and embedded as query vector Q. Standard-scale dot product attention operations and linear projections are performed on each pair of key vector K, query vector Q, and value vector V to generate attention features ca. The two attention features ca are transformed in terms of channel and spatial resolution and then added to the lower-stage RGB features to generate fused features z. An improved ResNet module is used to enhance the fused features z, generating multi-resolution features. The improved ResNet module includes layer normalization, two consecutive linear MLP layers, and residual connections; The Deep Feature Aggregation (DFAM) module has n-1 stages, and the depth features extracted from the first n-1 stages of the Depth branch are from different stages. The enhanced features are generated by inputting each of the n-1 deep feature aggregation modules (DFAM). The deep feature aggregation module DFAM expands the receptive field of the corresponding stage's depth features through dilated convolutional layers, generating dilated convolutional features X. dconv Simultaneously, the depth features of the corresponding stage are convolved and then segmented into N grouped features X along the channel dimension. These grouped features X are then convolved, and adjacent results are fused during the convolution process to generate N convolutionally fused features Y. These N convolutionally fused features Y are then convolved to generate feature Y'. Finally, the features X are dilated and convolved... dconv The deep features and feature Y' of the corresponding stage are fused to generate enhanced feature Y. d ; The Multi-Level Attention (MAF) module is used to extract RGB features from the nth stage of the RGB branch. and the depth features extracted in the nth stage of the Depth branch Complementary fusion is performed to generate attention-enhanced features h, which enhances the expressive power of high-level semantics of the features; The group attention unit (GA) is configured with n units, and the GA is used to fuse the RGB features of the corresponding stage. and depth features The system generates cross-modal features, which are processed by the decoder to generate the decoded features of the current stage. These features are then input into the Group Attention Unit (GA) of the next stage to guide the interactive fusion of RGB and depth features. The input to the GA of the nth stage is the output of the Multi-Level Attention (MAF) module, while the inputs to the GAs of the remaining n-1 stages are the multi-resolution features of the corresponding stage. and enhanced features And the decoding features of the previous stage output by the decoder; S2: Train the salient object detection model, use the trained salient object detection model to detect the input image, and output the decoding features of each stage by the decoder. S3: Use the CNN thinning unit to extract the edge features of the RGB image in the input image, and fuse them with the decoding features of the nth stage through attention. Regress the fused features into a salient object image to generate a predicted saliency map.

2. The salient object detection method based on multimodal feature refinement and fusion according to claim 1, characterized in that, The specific expression for the feature processing of the i-th cross-level attention fusion module ICFM is: ; ; ; ; ; ; in, Represents the RGB features of the i-th stage The corresponding query vector after copying This represents the weight matrix corresponding to the query vector Q. Represents the RGB features of the (i+1)th stage The corresponding key vector after copying, This represents the weight matrix corresponding to the key vector K. Represents the RGB features of the (i+1)th stage The corresponding value vector after copying This represents the weight matrix corresponding to the value vector V. Represents the query vector Key vector Sum value vector The attention features obtained from the calculation, Represents the query vector Key vector Sum value vector Dimensions Represents the key vector transpose, Indicates a linear operation. This represents the activation function operation. This represents the fusion feature generated by the i-th cross-level attention fusion module ICFM. This indicates an upsampling operation with a multiplier of 2. This indicates an upsampling operation with a multiplier of 4. This represents the multi-resolution feature generated by the i-th cross-level attention fusion module ICFM. This represents two consecutive linear layer operations. Presentation layer normalization operation, This indicates an addition operation.

3. The salient object detection method based on multimodal feature refinement and fusion according to claim 2, characterized in that, The specific expression for the feature processing of the i-th deep feature aggregation module DFAM is: ; ; ; ; ; ; in, Let C represent the depth features of the i-th stage, and let C denote a 1×1 convolution operation. Represents the depth features of the i-th stage The features obtained after convolution are denoted by Split, which represents the segmentation of the features along the channel dimension, and N represents the number of segments. This represents the feature of the j-th group. Represents the j-th convolutional fusion feature. This represents a 3×3 convolution operation. Represents N convolutional fused features Features are generated after convolution. concat represents the concatenation operation of channel dimensions. Represents the depth features of the i-th stage The dilated convolution feature obtained after the dilated convolution operation, Dconv represents the dilated convolution operation. This indicates an addition operation.

4. The salient object detection method based on multimodal feature refinement and fusion according to claim 3, characterized in that, The Multi-Level Attention Mechanism (MAF) module includes two attention fusion modules (AFM) and one multimodal fusion module (MFM). The AFM module comprises a spatial attention mechanism, a channel attention mechanism, and a cross-level attention fusion module (ICFM). The two AFM modules respectively handle the RGB features at the nth stage. and the depth features of the nth stage Enhancement is performed to generate RGB attention fusion features. Features fused with deep attention RGB attention fusion features Features fused with deep attention Adding them together yields the attention fusion features. and RGB attention fusion features Deep attention fusion features and attention fusion features Input the multimodal fusion module (MFM); The multimodal fusion module MFM is based on RGB attention fusion features. Deep attention fusion features and attention fusion features Generate a confusion matrix HY, and then fuse the confusion matrix HY with the RGB attention features. Features fused with deep attention Perform calculations to generate RGB attention-enhanced features. and deep attention enhancement features .

5. The salient object detection method based on multimodal feature refinement and fusion according to claim 4, characterized in that, The specific expression for the feature processing process of the Multi-Level Attention Mechanism (MAF) module is as follows: ; ; ; ; ; ; in, This represents the RGB attention fusion feature. This represents the deep attention fusion feature. Indicates attention fusion features, This indicates the operation of the spatial attention mechanism. This indicates the channel attention mechanism operation. This represents the convolution operation. Represents the confusion matrix. Indicates shape reshaping. This represents a 1×1 convolution operation. This represents RGB attention enhancement features. This indicates a deep attention enhancement feature.

6. The salient object detection method based on multimodal feature refinement and fusion according to claim 5, characterized in that, In the decoding process of the decoder and the generation of the predicted saliency map by the CNN refinement unit, a loss function is introduced as a supervision signal. Binary cross-entropy loss is used as the loss function to measure the SSIM loss and IOU loss of structural similarity. The specific calculation formula is as follows: ; ; in, Indicates the overall loss. This represents the combination of SSIM loss and IOU loss. This represents the final predicted saliency map. This indicates a correct label that was manually marked. This represents the saliency map for the prediction at the m-th stage. This represents the correct label after downsampling in the m-th stage. , Indicates the loss of bce. Indicates SSIM loss, This indicates Iou's loss.

7. A salient object detection method based on multimodal feature refinement and fusion according to any one of claims 1 to 6, characterized in that, The CNN thinning unit uses the first two layers of the VGG network backbone to extract edge features of the RGB image in the input image.