RGB-D salient object detection method
The RGB-D SOD method leverages transformer and convolutional networks to enhance feature extraction and fusion, addressing the challenge of cross-modality integration in conventional methods, resulting in improved accuracy and speed of salient object detection.
Patent Information
- Application Number
- GB2024003824
- Authority / Receiving Office
- GB · GB
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-04-25
- Filing Date
- 2024-03-18
- Publication Date
- 2025-09-03
- Estimated Expiration
- 2044-03-18
AI Technical Summary
Conventional RGB-D salient object detection (SOD) methods struggle with effectively fusing cross-modality features and capturing complementary information between RGB and depth images, leading to suboptimal detection of salient objects, especially in complex scenes.
An RGB-D SOD method utilizing a transformer encoder based on T2T-ViT for RGB images and a lightweight MobileNet V2 for depth images, combined with a cross-modal transformer fusion module (CMTFM) and cross-modal dense cooperative aggregation module (CMDCAM) to enhance feature extraction and fusion, followed by supervised learning with ground truth maps.
The method improves the accuracy and speed of SOD by effectively fusing cross-modality features, capturing global dependencies, and enhancing salient object detection performance.
Smart Images

Figure 00000001_0000 
Figure 00000002_0000 
Figure 00000003_0000
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of image processing, and specifically relates to an RGB-depth (RGB-D) salient object detection (SOD) method. BACKGROUND
[0002] In visual scenes, humans can rapidly shift attention to the most important regions. In SOD in computer vision, computers simulate the vision of human eyes to identify the most salient objects in the scene. As an important preprocessing task in computer vision applications, SOD has been widely used in image understanding, image retrieval, semantic segmentation, image restoration, and object recognition. As Kinect, RealSense, and other depth cameras develop, it becomes easier to capture depth images of various scenes, and depth information and RGB images can complement each other. This helps improve the capability of SOD. Therefore, RGB-D-based SOD has attracted the attention of researchers.
[0003] In the conventional RGB-D SOD method, feature extraction is performed manually and then RGB images and depth images are fused. For example, Lang et al. used a Gaussian mixture model to simulate the distribution of depth-induced saliency. Ciptadi et al. extracted three-dimensional (3D) layout and shape features from depth measurements and measured depth contrast based on depth differences between different regions. Although the conventional RGB-D SOD method is effective, the extracted low-level features restrict the generalization capability of the model and are not suitable for complex scenes.
[0004] One requirement of SOD is to effectively fuse cross-modality information. After the encoding of RGB images and RGB-D images, two learned modal features further need to be fused. An SOD method based on a convolutional neural network (CNN) has achieved many impressive results. However, the existing SOD method based on a CNN is restricted in the receptive field of a convolution and has serious disadvantages in learning global remote dependencies. Second, the early or late fusion strategy used in the prior art is difficult to capture the complementary and interactive information between the RGB and depth images. In this case, higher-level information cannot be learned from the two modalities and integrated fusion rules cannot be mined. As a result, complete salient objects cannot be effectively detected.
[0005] Therefore, there is now a need for a method that can effectively fuse cross-modality features and improve the accuracy of SOD. SUMMARY
[0006] A main objective of the present disclosure is to provide an RGB-D SOD method to resolve the prior-art problem that cross-modality features cannot be effectively fused and the accuracy of SOD is low.
[0007] To achieve the foregoing objective, the present disclosure provides an RGB-D SOD method, specifically including the following steps: SI: inputting an RGB image and a depth image; S2: performing, by a transformer encoder based on a tokens-to-token vision transformer (T2T-ViT), feature extraction on the RGB image, and performing, by an encoder based on a lightweight convolutional network MobileNet V2, feature extraction on the depth image, to obtain salient features of different levels of the RGB image and the depth image; S3: fusing, by a cross-modal transformer fusion module (CMTFM), complementary semantic information between deep-level RGB features and depth features to generate cross-modality joint features; S4: implementing, by a cross-modal dense cooperative aggregation module (CMDCAM) enhanced by dense connections, feature fusion of two different modalities to fuse level by level depth features and RGB features at different scales, and inputting the fused features into an SOD part; and S5: sorting predicted saliency maps in ascending order of resolutions, and performing supervised learning on the network by using a ground truth map to output a final saliency detection result.
[0008] Further, a T2T operation in the transformer encoder based on a T2T-ViT in step S2 includes reorganization and soft segmentation. The reorganization is to reconstruct a token sequence 1 € IK into a 3D tensor c , where is a length of the token sequence is the number of channels of the token sequence and the 3D tensor , represent a height and a width of , respectively, and — .
[0009] The soft segmentation is to softly divide I into sized blocks through an expand operation. The token sequence is obtained after the soft segmentation of 7 e , with a length 0 expressed as:
[0010]
[0011] where represents the number of overlapping pixels between the blocks, represents the number of padding pixels between the blocks, — represents a step in a convolution operation, and when 5 <- h the length of the token is reduced. J
[0012] An original RGB image is s , where X represent a height, a width, and the number of channels of , respectively, the token sequence * obtained after the reorganization undergoes three transformer transformations and two T2T operations to obtain a T' T T T T' multi-level tokens sequence ’ s ’ p 2’ 2 . This process is expressed as: F = k F = k T. ----- F = )
[0013] "w
[0014] Further, in step S2, the encoder based on the lightweight convolutional network MobileNet V2 includes an inverted residual block (IRB) structure.
[0015] Further, in step S3, the CMTFM includes a cross-modality interactive attention module and a transformer layer. The cross-modality interactive attention module is configured to model a remote cross-modality dependency between the RGB image and the depth image and integrate complementary information between RGB data and depth data.
[0016] Further, a formula in which the CMTFM obtains cross-modality interactive information is expressed as: T / ZrwftofX F, K -. IT ) = 0, / Jdl '77 J1 V * • -L- rnnm = n((.J-A.= '' 1(1(1171 - 2 V a. (F (F F-.- F n
[0018] where and are queries in two modalities, and * and are keys of the two modalities, and “ and r' are values of the two modalities.
[0019] Further, in step S4, the CMDCAM includes three feature aggregation modules (FAMs) and one double inverted residual module. The CMDCAM is configured to extend low-resolution encoder features to be consistent with resolution of the input image. The FAM is configured to aggregate features and fuse cross-modality information.
[0020] Further, the FAM includes one convolutional block attention module (CBAM) and two inverted residual blocks (IRBs), and further includes two element multiplication operations and one element addition operation. The process of feature aggregation and fusion of cross-modality information based on the FAM includes the following steps: 7' T
[0021] S4.1: multiplying the RGB feature *£ and the depth feature and then convolving, by one IRB, a result of the multiplication to obtain a transitional RGB-D feature map 2 , where this process is expressed as:
[0022] 7
[0023] S4.2: enhancing, by the CBAM, the depth feature , wherein the enhanced feature is 7" denoted as n , and this process is expressed as: K= Channel) xT^
[0024] T^SpatiaKT’^xT^ T T" T'
[0025] S4.3: multiplying and the depth feature D to enhance semantic features to obtain 1 , where this process is expressed as:
[0026] r! J and the RGB feature R to re-enhance salient features, introducing a T lower-level output feature for element addition, and then obtaining, by the IRB, an RGB-D feature * after cross-modality fusion, where this process is expressed as: Tn = Tn + T R R
[0028] „ , T = IRB(T^+T]X)
[0029] Further, in step S4, reorganized RGB information T , T , and T from J2T-ViT and C r. C- C depth information ’ from the MobileNet V2 are input to a decoder enhanced through dense connections. The dense connections are used to fuse depth features and RGB features at different scales.
[0030] Further, in step S5, the predicted saliency maps are supervised by a correspondingly resized ground truth map. Four losses generated in this stage are denoted as ? $. A calculation formula of a total loss function is as follows:
[0031]
[0032] where ? represents a weight of each loss, four predicted saliency maps are sequentially denoted as ~ kX-f Ti in ascending order of resolutions, represents the supervision from the ground truth map, with a resolution corresponding to G and represents a cross entropy loss function.
[0033] The present disclosure has the following beneficial effects.
[0034] 1. The present disclosure fully considers a difference between the RGB image and the depth image. A T2T-ViT network based on a transformer and a lightweight MobileNet V2 network are used to extract RGB information and depth information, respectively. With this design of an asymmetric dual-stream learning network, compared with other SOD methods, the present disclosure has a reduced number of model parameters, improves an SOD speed, and has excellent SOD performance.
[0035] 2. The decoder designed by the present disclosure includes a CMTFM and a CMDCAM. The CMTFM, as a block of the decoder, can model a remote cross-modality dependency between RGB data and depth data, and implements cross-modality information interaction between the RGB data and the depth data. The present disclosure enhances the decoder by using dense connections. The designed CMDCAM aggregates features at different levels in a dense cooperative fusion manner, and effectively fuses the cross-modality information. The decoder designed by the present disclosure effectively fuses RGB image information and depth information, and improves the accuracy of SOD. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] To describe the technical solutions in the specific implementations of the present disclosure or the prior art more clearly, the accompanying drawings required for describing the specific implementations or the prior art are briefly described below. Apparently, the accompanying drawings in the following description show merely some implementations of the present disclosure, and a person of ordinary skill in the art may still derive other accompanying drawings from these accompanying drawings without creative efforts. In the accompanying drawings:
[0037] FIG. lisa flowchart of an RGB-D SOD method according to the present disclosure;
[0038] FIG. 2 is a schematic structural diagram of an RGB-D SOD method according to the present disclosure;
[0039] FIG. 3 is a schematic structural diagram of a transformer encoder based on a T2T-ViT of FIG. 2; and
[0040] FIG. 4 is a schematic structural diagram of an FAM in a decoder of FIG. 2. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The following clearly and completely describes the technical solutions of the present disclosure with reference to accompanying drawings. Apparently, the described embodiments are some rather than all of the embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0042] An RGB-D SOD method shown in FIG. 1 specifically includes the following steps:
[0043] SI: Input an RGB image and a depth image.
[0044] S2: A transformer encoder based on a T2T-ViT performs feature extraction on the RGB image, and an encoder based on a lightweight convolutional network MobileNet V2 performs feature extraction on the depth image, to obtain salient features of different levels of the RGB image and the depth image.
[0045] A T2T-ViT network is an improvement on a ViT network, and adds a T2T operation on the basis of a ViT, which is equivalent to downsampling in a convolutional neural network, and is used to simultaneously model local structure information and global correlation of an image. The T2T operation can aggregate adjacent tokens into a new token, thereby reducing a length of the token.
[0046] Specifically, a T2T operation in the transformer encoder based on a T2T-ViT in step S2 includes reorganization and soft segmentation. The reorganization is to reconstruct a token sequence 4 t sv into a 3D tensor 1 c , where ' is a length of the token sequence T, is the number of channels of the token sequence T and the 3D tensor , w represent a height and a width of I , respectively, and x .
[0047] The soft segmentation is to softly divide into x sized blocks through an expand operation. The token sequence is obtained after the soft segmentation of1 E -1 , with a length 4 expressed as:
[0048]
[0049] where $ represents the number of overlapping pixels between the blocks, represents the number of padding pixels between the blocks, — represents a step in a convolution operation, and when $ ~ the length of the token is reduced. / G
[0050] An original RGB image is , where represent a height, a width, I 7’ c and the number of channels of , respectively, the token sequence * obtained after the reorganization undergoes three transformer transformations and two T2T operations to obtain a T I V T r multi-level tokens sequence ’ 2 5 2 . This process may be expressed as: T* T. = Uiffo / rH R^hapeiT')}. T7 = TrahsjbrmertTi X F -Utifnldi. ResIwpf(T,')\ 17 =
[0051] ‘
[0052] Specifically, in step S2, the encoder based on the lightweight convolutional network MobileNetV2 includes an IRB structure. Semantic information mainly exists in the RGB image. The depth image conveys information without object details. Compared with the RGB image, the depth image includes undiversified and a smaller amount of information. In addition, a part with the darkest color in the depth image is usually a salient object to be found in an SOD task. Therefore, the present disclosure can well extract the information of the depth image by using the lightweight MobileNet V2 network. MobileNet V2 is an improvement of MobileNet VI and proposes an IRB structure. The IRB structure is exactly the opposite of a residual structure in which dimensions are first reduced and then expanded, which is more conducive to feature learning. As shown in FIG. 2, c c c c four-level depth feature images output on a side of the MobileNet V2 are denoted as *' x ’ ■ *.
[0053] S3: A CMTFM fuses complementary semantic information between deep-level RGB features and depth features to generate cross-modality joint features.
[0054] Specifically, in step S3, the CMTFM includes a cross-modality interactive attention module and a transformer layer. The cross-modality interactive attention module is configured to model a remote cross-modality dependency between the RGB image and the depth image and integrate complementary information between RGB data and depth data, thereby improving the accuracy of saliency prediction. The CMTFM is based on an RGB-D converter in a visual saliency transformer (VST). To save parameters and computing resources, a self-attention part of the RGB-D converter is removed. Tr C
[0055] Specifically, as shown in FIG. 2, in the CMTFM, 2 and 4 are fused to integrate the complementary information between the RGB and depth data. is converted through three linear F projection operations to generate query , key , and value 5 . Similarly, use three other C O K F linear projection operations to convert 4 into query key n , and value . A formula for cross-modality interactive information can be obtained from a formula of "scaling dot product attention" in multi-head attention in the transformer layer and expressed as: J''- / 'x / X?Fr, mnr / i , A\ , Fs = .SQ.ft / Ji.X’.'iO- AT A / jA )K [UUOOj ~ - -' - — - * / V ■> A
[0057] In this way, an information flow from an RGB block mark and a depth block mark passes through the cross-modality interactive attention module four times for cross-modality information interaction, and then is enhanced by a four-level transformer layer to obtain a token sequence .
[0058] RGB and depth sequences from the encoder need to pass through a linear projection layer to convert their embedded dimensions from 384 into 64, to reduce computation and parameters.
[0059] S4: A CMDCAM enhanced by dense connections performs feature fusion of two different modalities to fuse level by level depth features and RGB features at different scales, and inputs the fused features into an SOD part.
[0060] Specifically, in step S4, the CMDCAM includes three FAMs and one double inverted residual module. The CMDCAM is configured to extend low-resolution encoder features to be consistent with resolution of the input image for pixel-level classification. The FAM not only can serve as a component of a decoder network and perform a function of aggregating features, but also can effectively fuse cross-modality information.
[0061] Specifically, the FAM includes one CBAM and two IRBs, and further includes two element multiplication operations and one element addition operation. The depth image conveys only an a priori region and lacks object details. Therefore, the semantic features of RGB are enhanced through two multiplications. The process of feature aggregation and fusion of cross-modality information based on the FAM includes the following steps: T
[0062] S4.1: multiplying the RGB feature and the depth feature u, and then convolving, by T one IRB, a result of the multiplication to obtain a transitional RGB-D feature map 1 , where this process is expressed as:
[0063] T = IRBCI'-.KTri)
[0064] S4.2: enhancing, by the CBAM, the depth feature 'D , where the enhanced feature is denoted as 75, and this process is expressed as: = Channel (TD ^TD
[0065] , PT' !
[0066] S4.3: multiplying 1 and the depth feature ” to enhance semantic features to obtain where this process is expressed as: T T'*'
[0067] and
[0068] S4.4: adding T and the RGB feature ^R to re-enhance salient features, introducing a lower-level output feature for element addition, and then obtaining, by the IRB, an RGB-D feature 7 after cross-modality fusion, where this process is expressed as: T' =Tr+T' K K
[0069] „ . T= IRBIT, + Tf F
[0070] Specifically, in step S4, reorganized RGB information 1 * 1 ’ *s from the T2T-ViT and G C, C- A depth information 5 ’ '3 - ’ from the MobileNet V2 are input to a decoder enhanced through dense connections. The dense connections are used to fuse depth features and RGB features at different scales.
[0071] S5: Sort predicted saliency maps in ascending order of resolutions, and perform supervised learning on the network by using a ground truth map to output a final saliency detection result.
[0072] Specifically, as shown in FIG. 1, in step S5, ¼ single-channel convolution and a Sigmoid activation function are sequentially added to an output of each decoder module to perform saliency mapping. During training, the predicted saliency maps are supervised by a correspondingly resized ground truth map. Four losses generated in this stage are denoted as . A / calculation formula of a total loss function 'W’ is as follows:
[0073]
[0074] where 7 represents a weight of each loss, four predicted saliency maps are sequentially PG =J 1 4) G denoted as 1 A* E in ascending order of resolutions, ! represents the supervision from P BCE l ) the ground truth map, with a resolution corresponding to i, and 1 -7 represents a cross entropy loss function. L
[0075] The total loss function can be calculated by using a formula of a binary cross entropy (BCE) loss function, and a calculation formula is follows:
[0076]
[0077] where * represents the weight of each loss.
[0078] In the SOD method, using a pre-trained model based on image classification as a backbone network facilitates loss convergence during training, thereby effectively improving the accuracy of SOD. The present disclosure uses a pre-trained transformer encoder based on a T2T-ViT and an encoder based on a lightweight convolutional network MobileNet V2 as a backbone network to extract features.
[0079] The present disclosure designs a CMDCAM. The module is based on an inverted residual module and has advantages of a small number of computation parameters and a small amount of computation. The module not only can fuse information of two modalities: RGB information and depth information, but also can aggregate different levels of feature information. This model can significantly improve SOD performance while reducing the computation amount of the detection method, and improve the accuracy of SOD.
[0080] It should be noted that the above description is not intended to limit the present disclosure, and the present disclosure is not limited to the above examples. Changes, modifications, additions or replacements made by those of ordinary skill in the art within the essential range of the present disclosure should fall within the protection scope of the present disclosure.
Claims
1. An RGB-depth (RGB-D) salient object detection (SOD) method, specifically comprising the following steps:SI: inputting an RGB image and a depth image;S2: performing, by a transformer encoder based on a tokens-to-token vision transformer (T2T-ViT), feature extraction on the RGB image, and performing, by an encoder based on a lightweight convolutional network MobileNet V2, feature extraction on the depth image, to obtain salient features of different levels of the RGB image and the depth image;S3: fusing, by a cross-modal transformer fusion module (CMTFM), complementary semantic information between deep-level RGB features and depth features to generate cross-modality joint features;S4: implementing, by a cross-modal dense cooperative aggregation module (CMDCAM) enhanced by dense connections, feature fusion of two different modalities to fuse level by level depth features and RGB features at different scales, and inputting the fused features into an SOD part;S5: sorting predicted saliency maps in ascending order of resolutions, and performing supervised learning on the CMDCAM by using a ground truth map to obtain a supervised SOD network; andS6: detecting, by the supervised SOD network, a salient object in a scene to output a final saliency detection result.
2. The RGB-D SOD method according to claim 1, wherein a T2T operation in the transformer encoder based on a T2T-ViT in step S2 comprises reorganization and soft segmentation, wherein the reorganization is to reconstruct a token sequence R into a three-dimensional (3D) tensor1 c , wherein 1 is a length of the token sequence 1 , L is the number of channels of thet / n w / token sequence 1 and the 3D tensor 1 , represent a height and a width of 1 , respectively,•>the soft segmentation is to softly divide I into x sized blocks through an expandoperation, and the token sequence is obtained after the soft segmentation ofAxwxc, with alength 0 expressed as:w+2p-k k-s, whereinrepresents the number ofoverlapping pixels between the blocks, represents the number of padding pixels between the blocks, S represents a step in a convolution operation, and when s< , the length of the token sequence is reduced; andan original RGB image is , wherein R represent a height, a width, andthe number of channels of input , respectively, the token sequence * € M obtained after the reorganization undergoes three transformer transformations and two T2T operations to obtain a multilevel tokens sequenceand the process is expressed as:rf tn r / nn x= I ransjormer{ 1),t t S' 1 1 / J 1 / nn f x x, =LJnjola(Resnape^ )),= Transformer(Tf),T2 = Unfold(Reshape(Tf)\T2 = Transformer (If)3. The RGB-D SOD method according to claim 1, wherein in step S2, the encoder based on the lightweight convolutional network MobileV2Net comprises an inverted residual block (IRB) structure.
4. The RGB-D SOD method according to claim 1, wherein in step S3, the CMTFM comprises a cross-modality interactive attention module and a transformer layer, and the cross-modality interactive attention module is configured to model a remote cross-modality dependency between the RGB image and the depth image and integrate complementary information between RGB data and depth data.
5. The RGB-D SOD method according to claim 4, wherein a formula in which the CMTFM obtains cross-modality interactive information is expressed as:Attention(QR,KD,VD) = softmax(QRKTD / Attention(QD ,KR,VR) = softmax(QDKTR / )VRwherein and are queries in two modalities, and are keys of the two modalities,V Vand R and D are values of the two modalities.
6. The RGB-D SOD method according to claim 1, wherein in step S4, the CMDCAM comprises three feature aggregation modules (FAMs) and one double inverted residual module, the CMDCAM is configured to extend low-resolution encoder features to be consistent with resolution of the input image, and the FAM is configured to aggregate features and fuse cross-modality information.
7. The RGB-D SOD method according to claim 6, wherein the FAM comprises one convolutional block attention module (CBAM) and two inverted residual blocks (IRBs), and further comprises two element multiplication operations and one element addition operation; and the process of feature aggregation and fusion of cross-modality information based on the FAM comprises the following steps:TS4.1: multiplying the RGB feature R and the depth feature D, and then convolving, by one IRB, a result of the multiplication to obtain a transitional RGB-D feature map T wherein the process is expressed as:TS4.2: enhancing, by the CBAM, the depth feature D, wherein the enhanced feature is denoted y nas 1D , and the process is expressed as:Tp = Channel (TDK = Spatial(T’)xT’T T" T’S4.3: multiplying 1 and the depth feature D to enhance semantic features to obtain ' , wherein the process is expressed as:Tf = Dxr;-andT’ TnS4.4: adding 1 and the RGB feature A to re-enhance salient features, introducing a lower-level output feature for element addition, and then obtaining, by the IRB, an RGB-D feature after cross-modality fusion, wherein the process is expressed as:
8. The RGB-D SOD method according to claim 1, wherein in step S4, reorganized RGB, J„ £< £-,information ' , , and *3 from the T2T-ViT and depth information s’ ““ 4 from theMobileNet V2 are input to a decoder enhanced through dense connections, and the dense connections are used to fuse depth features and RGB features at different scales.
9. The RGB-D SOD method according to claim 1, wherein in step S5, the predicted saliency maps are supervised by a correspondingly resized ground truth map, four losses generated in the stage£ / = 1234 Lare denoted as ’ , and a calculation formula of a total loss function total is asfollows:4LM«I = Tl^BCE(<Pl’Gl) / =1wherein 1 represents a weight of each loss, four predicted saliency maps are sequentiallyP(i = l 2 3 4) gdenoted as 1 v 5 ’ ’ 7 in ascending order of resolutions, represents the supervision fromP BCE( )the ground truth map, with a resolution corresponding to ', and ’ ' ’ represents a cross entropy loss function.
Citation Information
Patent Citations
RGB-D saliency target detection method
CN111582316A
RGB-D saliency target detection method based on interactive attention guidance and trapezoidal pyramid fusion
CN114283315A
Target detection method based on multi-source information fusion, thermal infrared and three-dimensional depth map
CN115713679A
Cross-modal feature fusion and asymptotic decoding saliency target detection method and device
CN115908789A