An object detection method based on depth-aware RGB multi-scale fusion network

By designing the depth-aware RGB feature optimization module, multi-scale attention enhancement fusion module and dual attention guidance module in RGB-D object detection, the problem of insufficient detection accuracy under complex backgrounds and variable lighting conditions in the prior art is solved, and more efficient and robust significance detection is achieved.

CN118781326BActive Publication Date: 2025-06-06ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410906096.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-08
Publication Date
2025-06-06
Estimated Expiration
2044-07-08

AI Technical Summary

Technical Problem

The existing RGB-D object detection methods perform poorly when dealing with complex backgrounds and variable lighting conditions, and the multi-scale feature fusion is not effective enough, resulting in a decrease in detection accuracy.

Method used

A U-Net-based significance detection framework is designed, including the depth-aware RGB feature optimization module (DARFOM), the multi-scale attention-enhanced fusion module (MSAEFM) and the dual attention-guiding module (DAGM). These modules effectively fuse RGB and depth information to enhance feature extraction and fusion capabilities.

Benefits of technology

Improve the accuracy of significance detection under complex backgrounds and different lighting conditions, enhance the robustness of complex backgrounds and noise, and achieve more accurate detection without adding too many parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118781326B_ABST
    Figure CN118781326B_ABST
Patent Text Reader

Abstract

The present invention discloses a target detection method based on a depth-perceived RGB multi-scale fusion network, comprising: step 1, in the RGB branch and the depth branch, feature downsampling is performed on the RGB image and the depth image to obtain RGB features and depth features; after downsampling in the RGB branch, a depth-perceived RGB feature optimization module is inserted to enhance the RGB features; step 2, RGB features and depth features are fused based on a multi-scale attention enhancement fusion module; step 3, a dual attention guidance module uses deeper features to guide the fused features for further filtering; step 4, sigmoid function is used to optimize the salient region detection to obtain the final detection target. The present invention can filter out the current features more accurately without adding too many parameters, perform more efficient extraction and fusion, and improve the accuracy of saliency detection under complex backgrounds or different lighting conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep RGB-D salient target detection, and in particular relates to a target detection method based on a depth-aware RGB multi-scale fusion network. Background Art

[0002] RGB-D is a special image format that combines traditional RGB (red, green, blue) color images and depth images. RGB images provide color information of the scene, while depth images provide distance information from each point in the scene to the camera. This combination gives RGB-D images significant advantages in applications such as three-dimensional perception, object recognition, and pose estimation. However, how to detect salient RGB-D is an important task in computer vision, which is crucial in many applications such as image retrieval, video segmentation, person re-identification, and visual tracking. Traditional salient object detection methods mainly rely on low-level features such as color and texture, but often perform poorly when dealing with complex backgrounds and changing lighting conditions.

[0003] With the development of deep learning and multimodal data processing technology, researchers have begun to explore how to combine the information of RGB images and depth images to improve the accuracy of salient object detection. RGB images provide rich color and texture information, and depth images provide important clues about the spatial position and geometric structure of objects in the scene. The fusion of multimodal information enables the model to understand the scene more comprehensively and locate salient objects more accurately.

[0004] Simply fusing the features of RGB and depth images is not enough to solve all problems. Features of different scales have different importance for salient object detection. Small-scale features can capture the detailed information of the object, while large-scale features focus more on the overall structure and contextual information of the object. Therefore, how to effectively fuse features of different scales so that the model can make full use of information of various scales has become a key issue.

[0005] When fusing RGB and depth information, there are three main fusion strategies: early fusion, multi-scale fusion, and late fusion. Both early fusion and late fusion only concatenate RGB and depth data once, failing to effectively utilize the correlation between the two. More RGB-D object detection methods adopt multi-scale fusion strategies. However, most of these methods intuitively fuse RGB and depth features while ignoring the features of saliency tasks and are not targeted when selecting complementary information. At the same time, the process of encoding the input is accompanied by downsampling, and a lot of information is lost in the process. Summary of the invention

[0006] The technical problem to be solved by the present invention is to provide a target detection method based on a depth-aware RGB multi-scale fusion network in view of the deficiencies of the above-mentioned prior art, propose a new saliency detection framework based on U-Net, design a depth-aware RGB feature optimization module (DARFOM) to clearly eliminate background interference, and enhance RGB features through depth prior knowledge, and use a multi-scale attention enhancement fusion module (MSAEFM) to more effectively fuse RGB and depth information. At the same time, a dual attention guidance module (DAGM) is proposed in the decoder network, and a multi-scale fusion network guided by dual attention is adopted. By utilizing the complementarity and correlation between RGB and depth information, the current features can be filtered out more accurately without adding too many parameters, and more efficient extraction and fusion can be performed, thereby improving the accuracy of saliency detection under complex backgrounds or under different lighting conditions.

[0007] In order to achieve the above technical objectives, the technical solution adopted by the present invention is:

[0008] A target detection method based on a depth-aware RGB multi-scale fusion network, comprising:

[0009] Step 1: In the RGB branch and the depth branch, VGG19 is used as the backbone network to perform feature downsampling of different levels and scales on the RGB image and the depth image to obtain the corresponding RGB features and depth features; and a depth-aware RGB feature optimization module DARFOM is inserted after each level of downsampling in the RGB branch to enhance the RGB features using the original depth image to obtain the final output features;

[0010] Step 2: Fusion of RGB features and depth features based on the multi-scale attention enhancement fusion module MSAEFM;

[0011] Step 3, the dual attention guidance module DAGM uses deeper features to guide the features fused in step 2 for further filtering;

[0012] Step 4: Based on step 3, the sigmoid function is used to optimize the salient region detection to obtain the final detection target.

[0013] To optimize the above technical solutions, the specific measures taken also include:

[0014] In step 1 above, DARFOM uses the original depth map to enhance the RGB features, and the process of obtaining the final output features is as follows:

[0015] (1) Decompose the original depth map into T+1 regions and generate T+1 spatial attention masks. The steps are as follows:

[0016] (2) Assign T+1 spatial attention masks to T+1 sub-branches in the RGB branch. Each sub-branch processes image information related to the corresponding depth region to accurately integrate RGB features and depth information to obtain enhanced RGB features.

[0017] (3) The final output features are obtained through residual connection:

[0018]

[0019] in is the RGB feature map in the i-th level of the RGB branch.

[0020] The above (1) includes:

[0021] (1.1) Convert the original depth map into a depth histogram, and then select T most significant depth distribution patterns from the histogram, each of which corresponds to a depth range, i.e., T depth interval windows;

[0022] (1.2) The original depth map is divided into T different regions according to T depth interval windows, and the remaining unselected parts in the histogram constitute an additional background region;

[0023] (1.3) Each region is normalized and its value range is limited to [0, 1], thereby generating T+1 spatial attention masks.

[0024] The above (2) includes:

[0025] (2.1) Use the maximum pooling operation to combine the mask and Size alignment:

[0026] p t =MaxPool(b t ) (1)

[0027] in is the RGB feature map in the i-th level of the RGB branch; b t is the t-th spatial attention mask;

[0028] (2.2) Using the mask p after alignment t With RGB features Get enhanced RGB features:

[0029]

[0030] in, Represents element-wise multiplication.

[0031] In the above step 2, for the first-level features, the RGB features and the depth features are combined by a series operation, and the obtained combined features are the first-level fused features;

[0032] For other level features, the multi-scale attention enhancement fusion module MSAEFM is used for fusion, as follows:

[0033] First, we combine the RGB features and the depth features to obtain the combined features.

[0034] Then, the combined feature c i After passing through the convolutional layer with a 3×3 kernel, it enters four branches and the channel attention branch, and then fuses to obtain the fused features;

[0035] For the last three of the four branches, a first convolution with a 1×1 kernel is set in each branch to change the number of feature channels. Asymmetric convolution and dilated convolution are also set. The asymmetric convolution approximates the square kernel convolution layer through a two-layer sequence with 1×d and d×1 kernels. The dilation rate of the dilated convolution of each branch is different.

[0036] The fused features obtained in the above MSAEFM block are:

[0037]

[0038] Where CA represents the channel attention branch and C is a convolutional layer with a 3×3 kernel; These are the results obtained through the four branches.

[0039] The dual attention guidance module DAGM described in step 3 above uses deeper features to guide the features fused in step 2 for further filtering, as follows:

[0040] M i =Cat(L i ×f i ,UL i ×f i ) (6)

[0041] d i =C(M i ×CA(M i )) (7)

[0042] Where L i and UL i denote the learned salient regions and the unlearned regions respectively, CA denotes the channel attention module, and d i is the output of the dual attention guidance module, C is a 1×1 kernel convolutional layer, and M i Contains the learned salient regions and The positions that have not been detected in , i represents the level of the feature.

[0043] The above step 4 uses the sigmoid function to optimize the detection of salient regions as follows:

[0044]

[0045] UL i =1-L i (9)

[0046] Where L i and UL i They represent the learned salient areas and the unlearned areas respectively, and i represents the level of the feature.

[0047] The present invention has the following beneficial effects:

[0048] 1) A new DARFOM module is proposed, which first decomposes the input depth map into multiple regions. Then, based on the results of the depth decomposition, these regions are regarded as spatial attention maps. After the RGB feature map passes through the pooling layer, the corresponding deep attention mask is used to weight the RGB features to ensure that the RGB features are depth-sensitively enhanced under the guidance of the DARFOM module, thereby paying more attention to the area of ​​salient objects and suppressing background interference.

[0049] 2) A new MSAEFM module is proposed, which first extracts features from local details to global levels through network layers of different depths, and then fuses the extracted features through a dual attention mechanism to ensure that features of different scales can complement each other, so that the network can capture contextual information of different scales and better adapt to salient objects of different sizes and shapes, which helps to understand the image content more comprehensively and improve the detection performance.

[0050] 3) A new DAGM module is proposed, which first performs effective feature extraction on the input depth image, and then generates an attention map by introducing an attention mechanism, so that the model can focus on important depth information and suppress irrelevant noise information, so that the model can automatically learn and emphasize deep features related to salient objects, thereby enhancing the robustness to complex backgrounds and noise. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 The overall flow chart of the method of the present invention is

[0052] Figure 2 This is a schematic diagram of the principle structure of the present invention;

[0053] Figure 3 It is the flow chart of DARFOM module of the present invention;

[0054] Figure 4 It is the multi-scale attention enhancement fusion module of the present invention;

[0055] Figure 5 This is the dual attention guiding module of the present invention. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0057] Although the steps in the present invention are arranged with numbers, they are not used to limit the order of the steps. Unless the order of the steps is clearly stated or the execution of a certain step requires other steps as a basis, the relative order of the steps can be adjusted. It is understood that the term "and / or" used in this article involves and covers any and all possible combinations of one or more of the associated listed items.

[0058] The RGB salient object detection method based on the multi-scale fusion network of depth perception of the present invention is mainly divided into three parts:

[0059] The first part inserts a depth-aware RGB feature optimization module (DARFOM) after each downsampling of the RGB branch. Depth information is crucial for identifying the three-dimensional structure and relative position of objects in the scene. This module enhances and adjusts the RGB features by utilizing depth information to better highlight salient objects and suppress background information. The optimized RGB features can more accurately reflect the boundaries and internal details of salient objects, thereby improving the performance of salient object detection;

[0060] The second part uses a multi-scale attention enhancement fusion module (MSAEFM), which extracts features of different scales and preliminarily processes them through convolution. It then uses the attention mechanism to generate attention weight maps of different scales, and uses the weight maps to perform weighted fusion of features of different scales.

[0061] The third part, the dual attention guidance module (DAGM), takes into account the importance of different regions of the transmitted information and uses deeper features to guide the output of the previous MSAEFM block, which means that it uses two different attention mechanisms to enhance feature representation. The attention mechanism allows the model to focus on important regions or features in the feature space, thereby improving the performance of the model. This module can effectively improve the accuracy of the network without adding too many parameters.

[0062] The present invention is an RGB-D salient target detection method based on a dual attention multi-scale fusion network. Compared with the existing RGB-D salient target detection method, the present invention introduces a depth-aware RGB feature optimization module. The module achieves depth-sensitive enhancement of RGB features through steps such as deep decomposition, generation of spatial attention maps, and extraction of RGB features. Salient objects appear at different scales in the image. The multi-scale attention enhancement fusion module can process information of different scales at the same time, and by introducing the attention mechanism, the network can automatically learn and emphasize features related to salient objects, thereby enhancing robustness to complex backgrounds and noise. By combining multi-scale fusion and attention mechanisms, and making full use of RGB-D information, significant beneficial effects are brought to the salient object detection task, and the detection accuracy and robustness are improved. The flow chart of the method is shown as follows. Figure 1 As shown, the specific process is as follows:

[0063] Step 1: In the RGB branch and the depth branch, VGG19 is used as the backbone network to perform feature downsampling of different levels and scales on the RGB image and the depth image to obtain the corresponding RGB features and depth features; and a depth-aware RGB feature optimization module DARFOM is inserted after each level of downsampling in the RGB branch to enhance the RGB features using the original depth image to obtain the final output features;

[0064] This step uses VGG19 as the backbone network. rgb and C d They are RGB and depth features extracted by VGG19;

[0065] The depth-aware RGB feature optimization module (DARFOM) is used to extract RGB features in different regions. DARFOM enhances RGB features and suppresses background interference by capturing deep geometric prior information. Specifically, DARFOM uses depth information to generate an attention map to guide the RGB feature extraction process, so that the model pays more attention to areas related to salient objects.

[0066] Step 2: Fusion of RGB features and depth features based on the multi-scale attention enhancement fusion module MSAEFM;

[0067] This step utilizes the new Multi-Scale Attention Enhanced Fusion Module (MSAEFM) to capture different levels of feature information from low-level to high-level. By fusing these features of different scales, the model can better understand the structure and context of the target, thereby improving the detection accuracy.

[0068] Step 3, the dual attention guidance module DAGM uses deeper features to guide the features fused in step 2 for further filtering;

[0069] This step uses the Dual Attention Guided Module (DAGM) to enhance the ability to locate salient objects. The dual attention mechanism includes channel attention and spatial attention. Channel attention helps the model selectively focus on important feature channels, while spatial attention helps the model accurately locate the position of salient objects.

[0070] Step 4: Based on step 3, the sigmoid function is used to optimize the salient region detection to obtain the final detection target.

[0071] This step uses the sigmoid function to calculate the learned salient areas and other unlearned locations after sampling each layer of features in the decoder stage, thereby improving the accuracy of detection.

[0072] Finally, the effectiveness of the designed module is verified through ablation experiments. After the model training is completed, the test set data is input into the model. Thanks to the feature extraction and fusion capabilities, as well as the attention mechanism, the model can effectively find the accurate salient target area.

[0073] Through the above steps, VGG19 is used as the backbone network, combined with the depth perception RGB feature optimization module and the depth information enhancement RGB feature to suppress background interference. The multi-level features are integrated through the multi-scale attention enhancement fusion module to improve the detection accuracy. The dual attention guidance module enhances the target positioning ability. Finally, the sigmoid function is used to optimize the salient area detection. The specific principle architecture diagram of the present invention is shown in the figure. Figure 2 As shown:

[0074] Depth-aware RGB feature optimization module:

[0075] In order to better extract RGB features, the present invention designs a depth-aware RGB feature optimization module (Depth-Aware RGB Feature Optimization Module, DARFOM).

[0076] In the RGB and depth branches, VGG19 is used as the backbone network to get C from large to small. 1 ,C 2 ,C 3 ,C 4 ,C 5 Five-level feature,deep branch is a lightweight deep network that can obtain deep features of different scales.

[0077] A depth-aware RGB feature optimization module is inserted after each downsampling in the RGB branch.

[0078] Each DARFOM uses the original depth map to enhance the RGB features. Specifically, the original depth map is decomposed into T+1 regions. The steps are as follows:

[0079] In order to make full use of the depth information, the original depth map is first converted into a depth histogram, and then T most significant depth distribution patterns are selected from the histogram. Each of these patterns corresponds to a depth range, that is, T depth interval windows.

[0080] Subsequently, the original depth map is segmented into T different regions according to these depth interval windows, while the remaining unselected parts of the histogram constitute an additional background region.

[0081] For the convenience of subsequent processing, each region is normalized and its value range is limited to [0,1], thereby generating a spatial attention mask. These masks are then used in the RGB branch in DARFOM. Specifically, they are assigned to T+1 sub-branches in the RGB branch, and each sub-branch will focus on processing image information related to a specific depth region. In this way, DARFOM can more accurately integrate RGB images and depth information to improve the accuracy of saliency detection. Figure 3 As shown, in form, is the RGB feature map in the i-th level of the RGB branch, where C i , H i and W i Represent the number, height and width of the channel respectively. t is represented as the tth attention mask obtained in the above-mentioned deep decomposition process. The present invention uses the maximum pooling operation to combine the mask with C rgbi Size alignment:

[0082] p t =MaxPool(b t ) (1)

[0083] in

[0084] Next, using the resized mask {p 1 ,p 2, …,p T+1} to extract depth-sensitive features in T+1 parallel sub-branches.

[0085] Specifically, each mask p t With RGB features Each channel of is multiplied, and a 3×3 convolution layer is used as a transition layer in the t-th sub-branch, because the 3×3 convolution kernel can capture more fine-grained local features and retain more spatial information to refine RGB features from various depth intervals. Afterwards, all depth-sensitive features from the T+1 sub-branches are aggregated by element-wise summation:

[0086]

[0087] in, is the enhanced RGB feature, and Represents element-wise multiplication.

[0088] Finally, a residual connection is introduced to obtain the final output features:

[0089]

[0090] In this way, DARFOM not only provides geometric prior knowledge in depth for RGB features, but also removes intractable background distractions (e.g., cluttered objects or similar textures).

[0091] Multi-scale attention enhancement fusion module:

[0092] In order to better fuse RGB and depth features, the present invention designs a Multi-Scale Attention-Enhanced Fusion Module (MSAEFM), where f i is to c r and c d After fusion, for the last four levels of features, c is concatenated using a concatenation operation. r and c d Merge together and then pass through a convolutional layer to reduce the dimension of the feature. The combined feature is expressed by the following formula:

[0093]

[0094] For the first level features, it mainly contains spatial details. Directly add c r and c d More such information can be retained, which can refine the final saliency prediction. This is because there is a certain overlap or difference between the information provided by RGB and depth. Therefore, channel attention is introduced to assign higher weights to channels that are more relevant to salient objects, thus achieving information selection from the perspective of attention. The structure of the MSAEFM block can be found in Figure 4The first convolution with a 1×1 kernel in each branch is mainly used to change the number of feature channels in order to save parameters and computational overhead. The use of asymmetric convolution is also based on this consideration, which approximates the square kernel convolution layer by a two-layer sequence with 1×d and d×1 kernels. In addition, dilated convolutions are introduced in this block, and the dilation rates of the dilated convolutions of each branch are different. Therefore, the receptive field of each branch is different, and the next stage can extract features from multiple scales, and features with a wider scale range are conducive to making more accurate predictions. Although RGB mainly provides 2D information and depth can additionally supplement geometric clues, there is still some of the same information between them. Due to the differences in the importance of information, channel attention is used to assign weights. The MSAEFM block is defined as follows:

[0095]

[0096] Where CA represents channel attention and C is a convolutional layer with a 3×3 kernel, which is used to reduce the dimension of convolution. These are the results obtained through these four branches. After this step of fusion, the features at all levels will be input into the dual attention module for further selection.

[0097] Dual attention guidance module:

[0098] In order to make up for the information lost in the downsampling process as much as possible, U-Net combines low-level and high-level features through skip connections. However, this strategy does not consider the transmission information of the importance of different fields, nor does it fully utilize the characteristics of the saliency detection task. Inspired by reverse attention, this paper proposes a Dual Attention Guidance Module (DAGM), whose structure is as follows Figure 5 As shown in Figure 2. This block uses deeper features to guide the output of the MSAEFM block in the previous level for further filtering. Therefore, the information contained in L and UL is combined through channel attention to correctly guide the upper layer features. The process can be expressed as:

[0099] M i =Cat(L i ×f i ,UL i ×f i ) (6)

[0100] d i =C(M i ×CA(M i )) (7)

[0101] Where L and UL represent the learned salient regions and the unlearned regions respectively, CA represents the channel attention module, d is the output of the dual attention block, and C is a 1×1 kernel convolutional layer. i Completely contains the learned salient regions and The channel attention plays a measurement role, and more valuable information will be given a greater weight, so only M i By introducing the attention mechanism, it can be used to enhance the different features of the position object. Under the guidance of L and UL, the network's ability to continuously correct errors is enhanced.

[0102] Sigmoid function:

[0103] Because a dual attention mechanism is needed to guide multi-scale fusion. In this mechanism, the sigmoid function is used to generate attention weights, thereby emphasizing important features and suppressing irrelevant features. In this way, the model can better focus on areas related to salient targets and improve the accuracy of detection. Specifically: for each level of features in the decoder stage, after upsampling them, the present invention uses a sigmoid function to calculate the learned salient areas and other unlearned locations, which can be formulated as:

[0104]

[0105] UL i =1-L i (9)

[0106] Where L and UL represent the learned salient regions and the unlearned regions, respectively, and i represents the level of the feature in the decoder level. L obtained by performing S-shaped operation i can be intuitively considered as learned salient regions, where UL can supplement the information provided by L. Although most of the regions in UL are backgrounds, they still contain some salient regions that have not been detected.

[0107] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.

[0108] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.

Claims

1. A target detection method based on a depth-aware RGB multi-scale fusion network, characterized in that: include: Step 1: Use VGG19 as the backbone network in the RGB branch and the depth branch to perform feature downsampling at different levels and scales on the RGB image and the depth image to obtain the corresponding RGB features and depth features; A depth-aware RGB feature optimization module DARFOM is inserted after each level of downsampling in the RGB branch to enhance the RGB features using the original depth map to obtain the final output features; The DARFOM uses the original depth map to enhance the RGB features, and the process of obtaining the final output features is as follows: (1) Decompose the original depth map into T+1 regions and generate T+1 spatial attention masks. The steps are as follows: (2) Assign T+1 spatial attention masks to T+1 sub-branches in the RGB branch. Each sub-branch processes image information related to the corresponding depth region to accurately integrate RGB features and depth information to obtain enhanced RGB features. (3) The final output features are obtained through residual connection: in is the RGB feature map in the i-th level of the RGB branch; Step 2: Fusion of RGB features and depth features based on the multi-scale attention enhancement fusion module MSAEFM; Step 3, the dual attention guidance module DAGM uses deeper features to guide the features fused in step 2 for further filtering; Step 4: Based on step 3, the sigmoid function is used to optimize the salient region detection to obtain the final detection target.

2. According to claim 1, a target detection method based on depth-aware RGB multi-scale fusion network is characterized in that: The (1) includes: (1.1) Convert the original depth map into a depth histogram, and then select T most significant depth distribution patterns from the histogram, each of which corresponds to a depth range, i.e., T depth interval windows; (1.2) The original depth map is divided into T different regions according to T depth interval windows, and the remaining unselected parts in the histogram constitute an additional background region; (1.3) Each region is normalized and its value range is limited to [0, 1], thereby generating T+1 spatial attention masks.

3. The object detection method based on depth-aware RGB multi-scale fusion network according to claim 1, characterized in that: The (2) includes: (2.1) Use the maximum pooling operation to combine the mask and Size alignment: p t =MaxPool(b t ) (1) in is the RGB feature map in the i-th level of the RGB branch; b t is the t-th spatial attention mask; (2.2) Using the mask p after alignment t With RGB features Get enhanced RGB features: in, Represents element-wise multiplication.

4. The object detection method based on depth-aware RGB multi-scale fusion network according to claim 1, characterized in that: In step 2, for the first-level features, the RGB features and the depth features are combined by a series operation, and the obtained combined features are the first-level fused features; For other level features, the multi-scale attention enhancement fusion module MSAEFM is used for fusion, as follows: First, we combine the RGB features and the depth features to obtain the combined features. Then, the combined feature c i After passing through the convolutional layer with a 3×3 kernel, it enters four branches and the channel attention branch, and then fuses to obtain the fused features; For the last three of the four branches, a first convolution with a 1×1 kernel is set in each branch to change the number of feature channels. Asymmetric convolution and dilated convolution are also set. The asymmetric convolution approximates the square kernel convolution layer through a two-layer sequence with 1×d and d×1 kernels. The dilation rate of the dilated convolution of each branch is different.

5. The object detection method based on depth-aware RGB multi-scale fusion network according to claim 4 is characterized in that: The fused features obtained in the MSAEFM block are: Where CA represents the channel attention branch and C is a convolutional layer with a 3×3 kernel; These are the results obtained through the four branches respectively. Cat means that the results are fused through convolution.

6. The object detection method based on depth-aware RGB multi-scale fusion network according to claim 1, characterized in that: The dual attention guidance module DAGM in step 3 uses deeper features to guide the features fused in step 2 for further filtering, as follows: M i =Cat(L i ×f i ,UL i ×f i ) (6) d i =C(M i ×CA(M i )) (7) Where L i and UL i denote the learned salient regions and the unlearned regions respectively, CA denotes the channel attention module, and d i is the output of the dual attention guidance module, C is a 1×1 kernel convolutional layer, and M i Contains the learned salient regions and The position that has not been detected in , i represents the level of the feature. In the DAGM module, Represents the result after upsampling of each layer of features in the decoder stage.

7. The object detection method based on depth-aware RGB multi-scale fusion network according to claim 1, characterized in that: The method of optimizing the salient region detection by using the sigmoid function in step 4 is as follows: the i =1-L i (9) Where L i and UL i They represent the learned salient areas and the unlearned areas respectively, and i represents the level of the feature.

Citation Information

Patent Citations

  • RGBD saliency detection method based on feature aggregation

    CN111931787A

  • RGB-D saliency target detection method based on boundary deformable convolution guidance

    CN115830420A