Three-primary-color deep image saliency detection system

By using a combination of feature encoder, attention enhancement module, cross-modal feature fusion module and cascade correction decoder in the significance detection system of three primary color depth images, the problem that the prior art is difficult to accurately capture significant targets in complex backgrounds is solved, and higher recognition segmentation accuracy and better feature representation capabilities are achieved.

CN119625289BActive Publication Date: 2025-05-23SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510162492.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-23
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

The prior art is difficult to accurately capture significant targets in three primary color depth images when dealing with complex backgrounds, resulting in edge blur and reduced segmentation accuracy.

Method used

Using a combination of feature encoder, attention enhancement module, cross-modal feature fusion module and cascade correction decoder, multi-level features are extracted and enhanced through multi-head attention mechanism and cross-modal fusion to generate more accurate significance maps.

Benefits of technology

The recognition and segmentation accuracy of the significance detection of three primary color depth images is improved, the accuracy and expression ability of feature representation are enhanced, and the calculation amount is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625289B_ABST
    Figure CN119625289B_ABST
Patent Text Reader

Abstract

The present invention discloses a saliency detection system for three-primary color depth images. The system of the present invention includes a feature encoder, an attention enhancement module, a cross-modal feature fusion module, and a cascade correction decoder; the feature encoder is used to extract the original color and depth multi-level features of the three-primary color depth image; the attention enhancement module is used to enhance the original multi-level features, and perform spatial alignment processing and channel recalibration processing on the enhanced depth features; the cross-modal feature fusion module is used to perform multi-scale feature fusion processing on the depth features after spatial alignment and channel recalibration and the enhanced color features to obtain corresponding multi-scale fusion features; the cascade correction decoder is used to use the multi-scale fusion features as new multi-level features of the three-primary color depth image, and decode to generate the corresponding saliency map. The present invention has higher recognition and segmentation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a three-primary-color depth image saliency detection system. Background Art

[0002] The three primary colors, namely red, green and blue (RGB), are the three basic colors that cannot be decomposed any further. The three primary color light model (RGB colormodel) is also called the RGB color model. It is mainly used in computer vision systems to represent and display color images (RGBMap). It is a common mode for computers to process image colors. The three primary colors can form corresponding color features. The depth image (Depth Map) is a grayscale image that contains the distance information of each pixel from the camera when shooting. It is another commonly used image mode in computer vision systems and is used to describe the three-dimensional structure of the scene. The three primary color depth image (RGB-DImage) is a combination of color image and depth map. It refers to the color image and the corresponding depth image obtained by the same camera in the same frame. These two types of images can be obtained at the same time when the shooting is completed. At present, monocular cameras, binocular cameras, etc. can directly capture and obtain three primary color depth images. Saliency detection is a technology that simulates the attention mechanism of the human visual system to identify salient areas in an image.

[0003] The goal of saliency detection on three-primary-color depth images is to detect salient objects or regions in a given color image and the corresponding depth image. Salient objects or regions refer to those image features that are more visually attractive or more prominent to humans. By combining RGB and depth information, saliency detection on three-primary-color depth images can more accurately identify salient objects in complex scenes and solve problems caused by factors such as lighting and occlusion. Due to its strong versatility and scalability, saliency detection has a wide range of applications in multiple computer vision fields, including visual recognition and tracking, human-computer interaction, intelligent driving, medical imaging, multi-task learning, etc.

[0004] In the prior art, saliency detection of three-primary-color depth images is performed using deep learning, specifically based on the combination of a backbone network and a decoder. However, the existing methods still have certain limitations when dealing with complex backgrounds. In particular, when the background is complex or the difference between it and the foreground is small and the boundaries are blurred, the network may have difficulty accurately capturing salient targets, resulting in blurred edges and reduced segmentation accuracy. Summary of the invention

[0005] In order to overcome the defects and shortcomings of the prior art, an object of the present invention is to provide a three-primary-color depth image saliency detection system for achieving saliency detection of three-primary-color depth images with higher recognition and segmentation accuracy.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] A three-primary-color deep image saliency detection system, comprising a feature encoder, an attention enhancement module, a cross-modal feature fusion module, and a cascade correction decoder;

[0008] The feature encoder is connected to the attention enhancement module and the cross-modal feature fusion module respectively;

[0009] The cross-modal feature fusion module is connected to the cascaded rectification decoder;

[0010] The feature encoder is used to extract the original color and depth multi-level features from the three-primary-color depth image;

[0011] The attention enhancement module is used to enhance the original multi-level features and perform spatial alignment and channel recalibration on the enhanced deep features;

[0012] The cross-modal feature fusion module is used to perform multi-scale feature fusion processing on the depth features after spatial alignment and channel recalibration and the enhanced color features to obtain the corresponding multi-scale fusion features;

[0013] The cascaded correction decoder is used to decode the multi-scale fusion features as new multi-level features of the three-primary color depth image to generate the corresponding saliency map.

[0014] Preferably, it also includes a loss function module;

[0015] The loss function module is connected to the cascade correction decoder, feature encoder, attention enhancement module, and cross-modal feature fusion module respectively;

[0016] The loss function module is used to perform the sum operation of weighted binary cross entropy loss and weighted intersection-over-union loss according to the true result of the three-primary-color depth image saliency and the saliency map output by the cascade correction decoder, and then back-propagate the gradient of the result of the summation operation to the feature encoder, attention enhancement module, and cross-modal feature fusion module.

[0017] Preferably, the feature encoder comprises two block partitioning modules and two multi-level shift window transform networks;

[0018] One block partitioning module is connected to a multi-level shift window transform network, and another block partitioning module is connected to another multi-level shift window transform network;

[0019] A multi-level shift window transformation network is connected to the attention enhancement module;

[0020] The block division module is used to receive the color image and / or the depth image in the three-primary-color depth image, and then perform feature segmentation on the color image and / or the depth image;

[0021] A multi-level shifted window transform network is used to extract the original multi-level features of color and / or depth, which are then input into the attention enhancement module.

[0022] Further, the multi-level shifted window transform network includes a linear embedding module, three block merging modules, and four shifted window transform network blocks;

[0023] The linear embedding module is connected with the block partitioning module;

[0024] A linear embedding module, a first shift window transform network block, a first block merging module, a second shift window transform network block, a second block merging module, a third shift window transform network block, a last block merging module, and a last shift window transform network block are connected in sequence;

[0025] The output of each shift window transformation network block is connected to the attention enhancement module respectively;

[0026] The linear embedding module is used to perform data dimension reduction mapping on the segmented features;

[0027] The block merging module is used to merge the features of the dimension reduction mapping to achieve hierarchical features;

[0028] The shifted window transform network block is used to extract color features and / or depth features.

[0029] Preferably, the attention enhancement module includes two attention feature enhancement modules and four channel space attention enhancement modules;

[0030] Two attention feature enhancement modules are connected to the feature encoder respectively;

[0031] The two attention feature enhancement modules are respectively connected to one of the channel space attention enhancement modules;

[0032] The other three channel spatial attention enhancement modules are connected to the feature encoder respectively;

[0033] The four channel spatial attention enhancement modules are connected to the cross-modal feature fusion module respectively;

[0034] The attention feature enhancement module is a multi-head self-attention mechanism network, which is used to enhance the long-range dependencies of color features and / or depth features;

[0035] The channel spatial attention enhancement module is used to perform spatial alignment and channel recalibration on deep features.

[0036] Furthermore, the channel spatial attention enhancement module includes a spatial attention module and two channel attention modules;

[0037] The output of the feature encoder and / or the output of the two attention feature enhancement modules are connected to the input of the spatial attention module after element-wise multiplication.

[0038] The output of the spatial attention module is connected to the input of one of the channel attention modules after element-wise multiplication with the output of the feature encoder and / or the output of one of the feature enhancement modules;

[0039] The output of the spatial attention module is connected to the input of another channel attention module after element-wise multiplication with the output of the feature encoder and / or the output of another feature enhancement module;

[0040] The output of one of the channel attention modules is connected to the cross-modal feature fusion module after element-wise multiplication with the output of the feature encoder and / or the output of one of the feature enhancement modules;

[0041] The output of another channel attention module is connected to the cross-modal feature fusion module after element-by-element multiplication with the output of the feature encoder and / or the output of another feature enhancement module;

[0042] The spatial attention module is used for spatial alignment processing;

[0043] The channel attention module is used to perform channel recalibration processing.

[0044] Preferably, the cross-modal feature fusion module includes a depth feature enhancement module, a color feature enhancement module, and a position attention enhancement module;

[0045] The inputs of the depth feature enhancement module and the color feature enhancement module are respectively connected to the attention enhancement module;

[0046] The outputs of the depth feature enhancement module and the color feature enhancement module are concatenated and then connected to the position attention enhancement module.

[0047] The position attention enhancement module is connected to the cascaded correction decoder;

[0048] The depth feature enhancement module and the color feature enhancement module are respectively connected to the attention enhancement module;

[0049] The deep feature enhancement module is used to further enhance the enhanced deep features by combining the semantic information of the color features with the spatial alignment and channel recalibration mechanisms.

[0050] The color feature enhancement module is used to combine the enhanced color features with the depth features, and further enhance them by spatial alignment and channel recalibration.

[0051] Furthermore, the position attention enhancement module includes a spatial attention module and a channel attention module;

[0052] The outputs of the depth feature enhancement module and the color feature enhancement module are concatenated, and then sequentially split, convolve, split, and element-by-element multiplication operations are performed before being connected to the input of the channel attention module.

[0053] The output of the channel attention module is connected to the input of the spatial attention module;

[0054] The output of the spatial attention module, the output of the deep feature enhancement module, and the output of the color feature enhancement module are concatenated and connected to the cascaded correction decoder through element-by-element multiplication.

[0055] Preferably, the cascade correction decoder is provided with a first decoder, a second decoder, and a down-sampling module;

[0056] The input of the first decoder is connected to the cross-modal feature fusion module to output a rough saliency map;

[0057] The output of the cross-modal feature fusion module and the output of the first decoder are added element by element and then connected to the input of the downsampling module;

[0058] The output of the downsampling module and the output of the first decoder are added element by element and then connected to the input of the second decoder;

[0059] The output of the second decoder is concatenated with the output of the feature encoder to output a refined saliency map.

[0060] Preferably, the feature encoder further comprises an edge generation module;

[0061] The edge generation module is connected to the cascade correction decoder;

[0062] The edge generation module is used to perform edge generation processing according to one or more original multi-level features extracted by the multi-level shift window transformation module, and then input the edge information into the cascade correction decoder.

[0063] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0064] The present invention avoids the shortcomings of traditional algorithms that require a large number of manually designed features and can often only mark rough areas. It is necessary to input three-primary color depth image data into the system for multiple rounds to obtain high-quality saliency detection results; the diversity of multi-level features is improved through feature enhancement, and the depth features are spatially aligned and channels are recalibrated to improve the accuracy and expressiveness of feature representation; it covers the processing methods of coordinate attention, spatial attention, and channel attention, and combines deep feature enhancement, RGB feature enhancement, one-dimensional coordinate attention and spatial attention to achieve the acquisition of a wider range of attention information, resulting in less computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 A schematic diagram of the overall structural framework of a three-primary-color depth image saliency detection system according to the present invention;

[0066] Figure 2 for Figure 1 Schematic diagram of the structural framework of the mid-channel spatial attention enhancement module;

[0067] Figure 3 for Figure 1 Schematic diagram of the structural framework of the mid- and cross-modal attention fusion module;

[0068] exist Figures 1 to 3 middle, is an element-by-element addition operation, is an element-by-element multiplication operation, For splicing operation, Represents multi-level features, represents the downsampled low-order features, represents the deep features, Indicates color features, represents multi-scale fusion features, Represents the splicing feature, represents the horizontal coordinate attention feature, represents the vertical coordinate attention feature, Represents the intermediate attention feature. DETAILED DESCRIPTION

[0069] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0070] It should be noted that the directions or positional relationships indicated by terms such as “center”, “up”, “down”, “left”, “right”, “vertical”, “horizontal”, “inside” and “outside” are based on the directions or positional relationships shown in the accompanying drawings and are only for the convenience of describing the present disclosure and simplifying the description. They do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction. Therefore, they should not be understood as limitations on the present disclosure.

[0071] In addition, the terms "first", "second", and "third" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance. Similarly, words such as "one", "an", or "the" do not indicate a quantitative limitation, but rather indicate the presence of at least one. Words such as "include" or "comprise" and the like mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "electrically connected" or "connected" and the like are not limited to physical or mechanical electrical connections, but may include electrical connections, whether direct or indirect.

[0072] Example 1

[0073] like Figures 1 to 3 As shown, this embodiment provides a three-primary-color depth image saliency detection system, including a feature encoder, an attention feature enhancement module, a cross-modal feature fusion module (COT), a cascade correction decoder, and a loss function module. The feature encoder is respectively connected to the attention feature enhancement module, the cross-modal feature fusion module, and the cascade correction decoder, and the cross-modal feature fusion module is connected to the cascade correction decoder. The loss function module is respectively connected to the cascade correction decoder, the feature encoder, the attention enhancement module, and the cross-modal feature fusion module.

[0074] The feature encoder is used to extract the original color and depth multi-level features from the three-primary color depth image; the attention enhancement module is used to enhance the original multi-level features, and perform spatial alignment and channel recalibration on the enhanced depth features; the cross-modal feature fusion module is used to perform multi-scale feature fusion on the depth features after spatial alignment and channel recalibration and the enhanced color features to obtain the corresponding multi-scale fusion features; the cascade correction decoder is used to decode the multi-scale fusion features as new multi-level features of the three-primary color depth image to generate the corresponding saliency map. The loss function module is used to perform the sum operation of the weighted binary cross entropy loss and the weighted intersection-over-union loss according to the true result of the saliency of the three-primary color depth image and the saliency map output by the cascade correction decoder, and then back-propagate the gradient of the result of the sum operation to the feature encoder, the attention enhancement module, and the cross-modal feature fusion module. The feature encoder is also used to perform edge generation processing on the extracted original multi-level features to obtain the edge information of the corresponding saliency area of ​​the three-primary color depth image, and then input the edge information into the cascade correction decoder.

[0075] Combination Figure 1 As shown, the feature encoder includes two block partition modules (Patch Partition), two multi-level shift window transformation networks, and an edge generation module.

[0076] A block division module is connected to a multi-level shift window transformation network, and another block division module is connected to another multi-level shift window transformation network; the multi-level shift window transformation network is connected to the attention enhancement module; the block division module is used to receive the color image and / or depth image in the three-primary color depth image, and then perform feature segmentation on the color image and / or depth image; the multi-level shift window transformation network is used to extract the original color and / or depth multi-level features, and then input the original color and / or depth multi-level features into the attention enhancement module.

[0077] The multi-level shifted window transform network includes a linear embedding module (Linear Embedding), three block merging modules (Patch Merging), and four shifted window transform network blocks (Swin-Transformer Block, and set to feature_only mode); the linear embedding module is connected to the block partitioning module; the linear embedding module, the first shifted window transform network block, the first block merging module, the next shifted window transform network block, the next block merging module, the next shifted window transform network block, the last block merging module, and the last shifted window transform network block are connected in sequence; the output of each shifted window transform network block is respectively connected to the attention enhancement module; the linear embedding module is used to perform spatial dimensionality reduction mapping on the segmented features; the block merging module is used to merge the features after dimensionality reduction mapping to realize hierarchical features; the shifted window transform network block is used to extract color features and / or depth features.

[0078] The edge generation module is connected to the cascade correction decoder; the edge generation module is used to perform edge generation processing according to one or more original multi-level features extracted by the multi-level shift window transformation module, and then input the edge information into the cascade correction decoder.

[0079] The attention enhancement module includes two attention feature enhancement modules (AFEM) and four channel space attention enhancement modules (HF).

[0080] Two attention feature enhancement modules are respectively connected to the feature encoder; two attention feature enhancement modules are respectively connected to one of the channel space attention enhancement modules; the other three channel space attention enhancement modules are respectively connected to the feature encoder; four channel space attention enhancement modules are respectively connected to the cross-modal feature fusion module; the attention feature enhancement module is a multi-head self-attention mechanism (MHSA) network, which is used to enhance the long-range dependencies of color features and / or depth features; the channel space attention enhancement module is used to perform spatial alignment processing and channel recalibration processing on the depth features.

[0081] Combination Figure 2 As shown, the channel spatial attention enhancement module includes a spatial attention module (SpatialAttention is preferably used in this embodiment) and two channel attention modules (ChannelAttention is preferably used in this embodiment).

[0082] The output of the feature encoder and / or the output of the two attention feature enhancement modules are connected to the input of the spatial attention module after element-by-element multiplication operation; the output of the spatial attention module and the output of the feature encoder and / or the output of one of the feature enhancement modules are connected to the input of one of the channel attention modules after element-by-element multiplication operation; the output of the spatial attention module and the output of the feature encoder and / or the output of another feature enhancement module are connected to the input of another channel attention module after element-by-element multiplication operation; the output of one of the channel attention modules and the output of the feature encoder and / or the output of one of the feature enhancement modules are connected to the cross-modal feature fusion module after element-by-element multiplication operation; the output of the other channel attention module and the output of the feature encoder and / or the output of another feature enhancement module are connected to the cross-modal feature fusion module after element-by-element multiplication operation; the spatial attention module is used for spatial alignment processing; the channel attention module is used for channel recalibration processing.

[0083] Combination Figure 3 As shown, the cross-modal feature fusion module includes a deep feature enhancement module (DFEM), a color feature enhancement module (RFEM), and a position attention enhancement module (PA).

[0084] The inputs of the depth feature enhancement module and the color feature enhancement module are respectively connected to the attention enhancement module; the outputs of the depth feature enhancement module and the color feature enhancement module are connected to the position attention enhancement module after splicing operation; the position attention enhancement module is connected to the cascade correction decoder; the depth feature enhancement module and the color feature enhancement module are respectively connected to the attention enhancement module; the depth feature enhancement module is used to combine the enhanced depth features with the semantic information of the color features, and further enhance them by using spatial alignment and channel recalibration mechanism; the color feature enhancement module is used to combine the enhanced color features with the depth features, and further enhance them by using spatial alignment and channel recalibration mechanism.

[0085] The position attention enhancement module includes a spatial attention module (SAM is preferably used in this embodiment) and a channel attention module (CAM is preferably used in this embodiment); the outputs of the depth feature enhancement module and the color feature enhancement module are respectively connected to the input of the channel attention module after the splicing operation, and then sequentially undergo splitting, convolution, splitting, and element-by-element multiplication operations; the output of the channel attention module is connected to the input of the spatial attention module; the output of the spatial attention module, and the outputs of the depth feature enhancement module and the color feature enhancement module are respectively connected to the cascade correction decoder through element-by-element multiplication operations after the splicing operation.

[0086] The cascade correction decoder is provided with a first decoder (Decoder1), a second decoder (Decoder2), and a downsampling module. The input of the first decoder is connected to the cross-modal feature fusion module to output a rough saliency map; the output of the cross-modal feature fusion module is connected to the input of the downsampling module after element-by-element addition operation with the output of the first decoder; the output of the downsampling module is connected to the input of the second decoder after element-by-element addition operation with the output of the first decoder; the output of the second decoder is concatenated with the output of the feature encoder to output a fine saliency map.

[0087] Combination Figure 1 , Figure 2 and Figure 3 As shown, in this embodiment, the feature encoder preferably uses two parallel multi-level shift window transform networks as the backbone, and cooperates with two block division modules and one edge generation module to complete the feature extraction task;

[0088] The first partitioning module is used as an input port of the color image; the first partitioning module, the first linear embedding module, the first shifted window transform network block, the first merging module, the second shifted window transform network block, the second merging module, the third shifted window transform network block, the third merging module, and the fourth shifted window transform network block are sequentially connected in a manner of using the output of the previous module as the input of the next module, so as to extract the original multi-level color features;

[0089] The second partitioning module is used as an input port of the depth image; the second partitioning module, the second linear embedding module, the fifth shifted window transform network block, the fourth merging module, the sixth shifted window transform network block, the fifth merging module, the seventh shifted window transform network block, the sixth merging module, and the eighth shifted window transform network block are sequentially connected in a manner that the output of the previous module is used as the input of the next module to extract the original multi-level depth features;

[0090] The outputs of the fifth shift window transformation network block and the sixth shift window transformation network block are connected to the edge generation module for generating edge information;

[0091] In a further preferred embodiment of the present invention, the edge generation module is calibrated using the real result in the feature encoder.

[0092] In this embodiment, the output of the fourth shift window transformation network block is preferably connected to the input of the first attention feature module of the attention feature enhancement module, and the first attention feature module is used to perform diversity enhancement on the color features extracted by the fourth shift window transformation network block;

[0093] In this embodiment, the output of the eighth shift window transform network block is preferably connected to the input of the second attention feature module of the attention feature enhancement module, and the second attention feature module is used to perform diversity enhancement on the deep features extracted by the eighth shift window transform network block;

[0094] The attention feature enhancement module is provided with a first channel spatial attention enhancement module, a second channel spatial attention enhancement module, a third channel spatial attention enhancement module, and a fourth channel spatial attention enhancement module in parallel;

[0095] The features extracted by the first shifted window transformation network block and the fifth shifted window transformation network block are respectively input into the first channel spatial attention enhancement module; the features extracted by the second shifted window transformation network block and the sixth shifted window transformation network block are respectively input into the second channel spatial attention enhancement module; the features extracted by the third shifted window transformation network block and the seventh shifted window transformation network block are respectively input into the third channel spatial attention enhancement module; the color features output by the first attention feature module and the depth features output by the second attention feature module are respectively input into the fourth channel spatial attention enhancement module for spatial alignment and channel recalibration;

[0096] In the channel space attention enhancement module, the original color features of the corresponding level and deep features After element-by-element multiplication, it is input into the first spatial attention module; the output of the first spatial attention module and the deep feature After element-by-element multiplication, it is used as the input of the first channel attention module; the other output of the first spatial attention module and the color feature After element-by-element multiplication, it is used as the input of the second channel attention module; the output of the first channel attention module and the deep features Perform element-by-element multiplication to form enhanced depth features , corresponding to the input phase level cross-modal feature fusion module; the output of the second channel attention module and the color feature Perform element-by-element multiplication to form enhanced color features , corresponding to the cross-modal feature fusion module of the input phase level (i represents the corresponding level number, r represents color, and d represents depth).

[0097] In this embodiment, the cross-modal feature fusion module is preferably provided with four parallel levels; the input of the first cross-modal feature fusion module is connected to the output of the first channel spatial attention enhancement module, the input of the second cross-modal feature fusion module is connected to the output of the second channel spatial attention enhancement module, the input of the third cross-modal feature fusion module is connected to the output of the third channel spatial attention enhancement module, and the input of the fourth cross-modal feature fusion module is connected to the output of the first channel spatial attention enhancement module;

[0098] In each level of the cross-modal feature fusion module, there is a deep feature enhancement module, a color feature enhancement module, and a position attention enhancement module;

[0099] The input of the deep feature enhancement module is the diverse, spatially aligned, and channel-calibrated color features output by the channel space attention enhancement module at the corresponding level. and deep features ; The input of the color feature enhancement module is the color features after spatial alignment and channel calibration of the channel space attention enhancement module output at the corresponding level and deep features ;

[0100] In the deep feature enhancement module, there are the second spatial attention module and the third channel attention module; the color feature and deep features After the element-by-element multiplication operation, it is input into the second spatial attention module; the output and deep features of the second spatial attention module Output after element-by-element multiplication, and deep features After element-by-element addition, the output is sent to the third channel attention module; the output and deep features of the third channel attention module After element-by-element multiplication, it is used as the final output of the deep feature enhancement module;

[0101] The color feature enhancement module includes the third spatial attention module and the fourth channel attention module; the color feature and deep features After element-by-element multiplication, the third spatial attention module is input; the output of the third spatial attention module and the color feature Output after element-by-element multiplication, and color features After element-by-element addition, the output is sent to the fourth channel attention module; the output of the fourth channel attention module and the color feature The final output of the color feature enhancement module is obtained after element-by-element multiplication.

[0102] The position attention enhancement module includes a fifth channel attention module and a fourth spatial attention module;

[0103] The final output of the depth feature enhancement module and the final output of the color feature enhancement module are spliced ​​to obtain the original spliced ​​feature , then the original splicing features The input position attention enhancement module; the original splicing feature After splitting, convolution, and splitting, we get the horizontal coordinate attention feature and vertical coordinate attention features ; Original splicing features , horizontal coordinate attention features and vertical coordinate attention features After performing element-by-element multiplication, the three get a new concatenation feature. ; New splicing features Input the fifth channel attention module, the output of the fifth channel attention module is connected to the fourth spatial attention module; the output of the fourth spatial attention module is connected to the original splicing feature After performing element-by-element multiplication, we get the multi-scale fusion feature ;

[0104] The final output of the cross-modal feature fusion module at each level is multi-scale fusion features As a new multi-level feature of three-primary color depth image , respectively input into the cascade correction decoder.

[0105] In this embodiment, preferably, the first decoder of the cascade correction decoder includes a first multi-scale aggregation module, a second multi-scale aggregation module, and a third multi-scale aggregation module, and the second decoder includes a fourth multi-scale aggregation module, a fifth multi-scale aggregation module, and a sixth multi-scale aggregation module;

[0106] For the first decoder, the multi-level features output by the cross-modal feature fusion module , multi-level features The third multi-scale aggregation module is input respectively, and the output of the third multi-scale aggregation module and the multi-level features are The second multi-scale aggregation module is input respectively, and the output of the second multi-scale aggregation module and the multi-level features are The first multi-scale aggregation module is respectively inputted, and the first multi-scale aggregation module outputs a rough saliency map;

[0107] For the downsampling module, the output of the first multi-scale module is downsampled to obtain the corresponding downsampled low-order features. ;

[0108] For the second decoder, multi-level features and downsample low-level features After element-by-element addition, the sixth multi-scale aggregation module is input, and the output of the third multi-scale aggregation module and the downsampled low-order features are After element-by-element addition, the output of the second multi-scale aggregation module and the downsampled low-order features are input into the sixth multi-scale aggregation module. After performing element-by-element addition operation, the output of the sixth multi-scale aggregation module is connected to the input of the fifth multi-scale aggregation module; the output of the first multi-scale aggregation module and the down-sampled low-order features are connected to the fifth multi-scale aggregation module. After performing element-by-element addition operation, the output of the fifth multi-scale aggregation module is input into the fourth multi-scale aggregation module, and the output of the fifth multi-scale aggregation module is connected to the input of the fourth multi-scale aggregation module; after the output of the fourth multi-scale aggregation module and the edge information output by the edge generation module are spliced, a refined saliency map is obtained.

[0109] In this embodiment, the loss function module preferably performs a sum operation of weighted binary cross entropy loss and weighted intersection-over-union loss according to the true result of the three-primary-color depth image saliency, the rough saliency map and the fine saliency map output by the cascade correction decoder, and then back-propagates the gradient to each module of the feature encoder, the attention enhancement module, and the cross-modal feature fusion module;

[0110] Loss Function The specific definition is shown as follows:

[0111] ,

[0112] in, represents the first weighted loss, represents the second weighted loss; , The respective calculations are shown below:

[0113] ,

[0114] ,

[0115] in, represents the weighted binary cross entropy loss function, represents the weighted intersection-over-union loss function, represents the predicted contour, Represents the true contour, represents the predicted image, Represents a real image;

[0116] , The respective calculations are shown below:

[0117] ,

[0118] ,

[0119] ,

[0120] in, express The similarity between the pixel and the surrounding pixels. The larger the value, the lower the similarity with the surrounding pixels. represents the predicted probability, Represents pixels The surrounding pixels are Indicates background or foreground, represents all the parameters of the model, represents the hyperparameter, H represents the height of the input feature, W represents the width of the input feature, Represents pixels The true value of the position, Represents pixels Surrounding pixels The true value of the position, Represents pixels The predicted value for the location.

[0121] Compared with the prior art, this embodiment has the following beneficial effects:

[0122] This embodiment uses a multi-head attention mechanism in feature extraction to obtain longer-range dependencies in the three-primary-color depth image, better focus on the salient area, and enhance the accuracy of feature recognition; share feature representation with the edge detection strategy to achieve more accurate segmentation accuracy; and better integrate color information and depth information in the three-primary-color depth image through a cross-modal fusion strategy for multiple features;

[0123] This embodiment avoids the shortcomings of traditional algorithms that require a large number of manually designed features and can often only mark rough areas, and requires inputting three-primary color depth image data into the system for multiple rounds to obtain high-quality saliency detection results; the diversity of multi-level features is improved through feature enhancement, and the depth features are spatially aligned and channels are recalibrated to improve the accuracy and expressiveness of feature representation; it covers the processing methods of coordinate attention, spatial attention, and channel attention, and combines depth feature enhancement, RGB feature enhancement, one-dimensional coordinate attention and spatial attention to achieve the acquisition of a wider range of attention information, resulting in less computational complexity.

[0124] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.

Claims

1. A three-primary-color depth image saliency detection system, characterized in that: It includes feature encoder, attention enhancement module, cross-modal feature fusion module, and cascade correction decoder; The feature encoder is connected to the attention enhancement module and the cross-modal feature fusion module respectively; A cross-modal feature fusion module is connected to the cascade correction decoder; The feature encoder is used to extract the original color and depth multi-level features from the three-primary-color depth image; The attention enhancement module is used to enhance the original multi-level features and perform spatial alignment and channel recalibration on the enhanced deep features; The cross-modal feature fusion module is used to perform multi-scale feature fusion processing on the depth features after spatial alignment and channel recalibration and the enhanced color features to obtain the corresponding multi-scale fusion features; The cascade correction decoder is used to decode the multi-scale fusion features as new multi-level features of the three-primary color depth image to generate the corresponding saliency map; The attention enhancement module includes two attention feature enhancement modules and four channel space attention enhancement modules; The two attention feature enhancement modules are respectively connected to the feature encoder; Two attention feature enhancement modules are respectively connected to one of the channel space attention enhancement modules; The other three channel spatial attention enhancement modules are connected to the feature encoder respectively; The four channel spatial attention enhancement modules are connected to the cross-modal feature fusion module respectively; The attention feature enhancement module is a multi-head self-attention mechanism network, which is used to enhance the long-range dependencies of color features and / or depth features; The channel space attention enhancement module is used to perform spatial alignment and channel recalibration on deep features; The channel spatial attention enhancement module includes a spatial attention module and two channel attention modules; The output of the feature encoder and / or the output of the two attention feature enhancement modules are connected to the input of the spatial attention module after an element-by-element multiplication operation; The output of the spatial attention module is connected to the input of one of the channel attention modules after element-wise multiplication with the output of the feature encoder and / or the output of one of the feature enhancement modules; The output of the spatial attention module is connected to the input of another channel attention module after element-wise multiplication with the output of the feature encoder and / or the output of another feature enhancement module; The output of one of the channel attention modules is connected to the cross-modal feature fusion module after element-wise multiplication with the output of the feature encoder and / or the output of one of the feature enhancement modules; The output of another channel attention module is connected to the cross-modal feature fusion module after element-wise multiplication with the output of the feature encoder and / or the output of another feature enhancement module; The spatial attention module is used for spatial alignment processing; The channel attention module is used to perform channel recalibration processing; The feature encoder also includes an edge generation module; The edge generation module is used to perform edge generation processing according to one or more original multi-level features extracted by the multi-level shift window transformation module, and then input the edge information into the cascade correction decoder.

2. The three-primary-color depth image saliency detection system according to claim 1, characterized in that: It also includes a loss function module; The loss function module is respectively connected to the cascade correction decoder, the feature encoder, the attention enhancement module, and the cross-modal feature fusion module; The loss function module is used to perform the sum operation of weighted binary cross entropy loss and weighted intersection-over-union loss according to the true result of the three-primary-color depth image saliency and the saliency map output by the cascade correction decoder, and then back-propagate the gradient of the result of the summation operation to the feature encoder, attention enhancement module, and cross-modal feature fusion module.

3. The three-primary-color depth image saliency detection system according to claim 1, characterized in that: The feature encoder includes two block partitioning modules and two multi-level shift window transformation networks; One of the block partitioning modules is connected to one of the multi-level shift window transform networks, and another block partitioning module is connected to another multi-level shift window transform network; A multi-level shift window transformation network is connected to the attention enhancement module; The block division module is used to receive the color image and / or the depth image in the three-primary-color depth image, and then perform feature segmentation on the color image and / or the depth image; A multi-level shifted window transform network is used to extract the original multi-level features of color and / or depth, which are then input into the attention enhancement module.

4. The three-primary-color depth image saliency detection system according to claim 3, characterized in that: The multi-level shift window transform network includes a linear embedding module, three block merging modules, and four shift window transform network blocks; The linear embedding module is connected to the block partitioning module; A linear embedding module, the first shift window transform network block, the first block merging module, the next shift window transform network block, the next block merging module, the next shift window transform network block, the last block merging module, and the last shift window transform network block are connected in sequence; The output of each shift window transformation network block is connected to the attention enhancement module respectively; The linear embedding module is used to perform spatial dimension reduction mapping on the segmented features; The block merging module is used to merge the features after dimensionality reduction mapping to achieve hierarchical features; The shifted window transform network block is used to extract color features and / or depth features.

5. The three-primary-color depth image saliency detection system according to claim 1, characterized in that: The cross-modal feature fusion module includes a depth feature enhancement module, a color feature enhancement module, and a position attention enhancement module; The inputs of the depth feature enhancement module and the color feature enhancement module are respectively connected to the attention enhancement module; The outputs of the depth feature enhancement module and the color feature enhancement module are connected to the position attention enhancement module after being spliced; The position attention enhancement module is connected to the cascaded correction decoder; The depth feature enhancement module and the color feature enhancement module are respectively connected to the attention enhancement module; The deep feature enhancement module is used to further enhance the enhanced deep features by combining the semantic information of the color features with the spatial alignment and channel recalibration mechanisms. The color feature enhancement module is used to combine the enhanced color features with the depth features, and further enhance them by spatial alignment and channel recalibration.

6. The three-primary-color depth image saliency detection system according to claim 5, characterized in that: The position attention enhancement module includes a spatial attention module and a channel attention module; The outputs of the depth feature enhancement module and the color feature enhancement module are concatenated, and then sequentially split, convoluted, split, and element-by-element multiplication operations are performed to connect with the input of the channel attention module; The output of the channel attention module is connected to the input of the spatial attention module; The output of the spatial attention module, the output of the deep feature enhancement module, and the output of the color feature enhancement module are concatenated and connected to the cascaded correction decoder through element-by-element multiplication.

7. The three-primary-color depth image saliency detection system according to claim 1, characterized in that: The cascade correction decoder is provided with a first decoder, a second decoder, and a down-sampling module; The input of the first decoder is connected to a cross-modal feature fusion module for outputting a rough saliency map; The output of the cross-modal feature fusion module and the output of the first decoder are added element by element and then connected to the input of the downsampling module; The output of the downsampling module and the output of the first decoder are added element by element and then connected to the input of the second decoder; The output of the second decoder is concatenated with the output of the feature encoder to output a refined saliency map.

Citation Information

Patent Citations

  • Cross-modal feature fusion and asymptotic decoding saliency target detection method and device

    CN115908789A

  • Salient target detection method and device

    CN116310394A