Panoramic image saliency target detection method, system and device based on depth information fusion and medium

Through the panoramic image saliency object detection network that fuses color images and depth image features, the problem of poor detection effects caused by the lack of depth information in the prior art is solved, and higher detection accuracy and robustness are achieved.

CN120259873APending Publication Date: 2025-07-04XIDIAN UNIV +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510280788.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing panoramic image significance object detection technology lacks real labels of depth maps and ignores depth information, resulting in poor detection results, especially in scenarios similar to background and significant targets.

Method used

A panoramic image depth map generation method is constructed, and an end-to-end significance target detection network model is designed. By fusing color image features and depth image features, it captures significant graph flow characteristics and pays attention to the depth map flow information in the image imaging spatial structure.

Benefits of technology

It improves the accuracy and robustness of panoramic image significance object detection, effectively suppresses noise and redundant information, and improves the object detection capability in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259873A_ABST
    Figure CN120259873A_ABST
Patent Text Reader

Abstract

The invention discloses a panoramic image saliency target detection method, system and device based on depth information fusion and a medium, belongs to the field of computer vision, and is suitable for a panoramic image analysis scene of high-precision target detection. The method comprises the following steps: performing depth estimation and saliency target detection on a panoramic image to obtain a depth image and a saliency target image, extracting features from the depth image and the saliency image respectively to obtain a depth image flow feature and a saliency image flow feature, fusing the depth image flow feature and the saliency image flow feature to generate a fusion feature, and decoding the fusion feature to obtain a depth image flow feature and a saliency image flow feature; obtaining a detection result; performing joint training on depth estimation, saliency target detection and feature fusion according to a detection result, and finally inputting a panoramic image into the trained network model, detecting a saliency target, and obtaining a saliency target result map; the system, the equipment and the medium are used for implementing the method. According to the method, noise and redundant information are suppressed, and the accuracy of saliency target detection of the panoramic image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly relates to a panoramic image salient object detection method, system, device and medium based on depth information fusion, which is applicable to panoramic image analysis scenarios requiring high-precision object detection. Background Art

[0002] Panoramic image salient object detection is an important basic task in the field of computer vision technology. Its main goal is to detect the most attention-grabbing objects or regions in panoramic image data that conform to the human eye vision mechanism. By selecting regions of interest through salient object detection, it can guide the reasonable allocation of limited computing resources, thereby reasonably reducing the redundant information obtained by human eye vision or machine vision from images and increasing the efficiency of the transmission and storage processes. In recent years, deep learning methods have become the mainstream research direction in the field of panoramic image salient object detection. Currently, there have been many studies on panoramic image salient object detection in the prior art.

[0003] Piao, Yongri et al. proposed an adaptive and attention-driven depth distiller for RGB-D salient object detection tasks in the literature "Adaptive and Attentive Depth Distiller for Efficient RGB-D Salient Object Detection" (Piao, Yongri, et al. "Adele: Adaptive and attentive depth distiller for efficient RGB-D salient object detection." Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2020.). This technical solution transfers depth information from the depth stream to the RGB stream through two bridges, so that object detection can not only rely on RGB images during model inference. And by minimizing the difference between the predictions generated by the depth stream and the RGB stream, it realizes the adaptive control transfer of pixel-level depth knowledge of the RGB stream. Finally, by unifying the consistency between the dilated predictions of the depth stream and the attention map of the RGB stream, the localization knowledge of the salient object is transferred to the RGB features to generate the final salient object result map. However, since there are no depth map ground truths in the panoramic salient object detection dataset. Secondly, for the feature extraction of the RGB stream, the panoramic projection distortion characteristics cannot be considered, resulting in poor performance in model salient object detection.

[0004] Huang, Mengke et al. proposed a feature adaptation method for omnidirectional salient object detection in the literature "Features Adaptation Network for 360° Omnidirectional Salient Object Detection" (Huang, Mengke, et al. "FANet: Features Adaptation Network for 360° Omnidirectional Salient Object Detection." IEEE Signal Processing Letters 27 (2020): 1819-1823.): By simultaneously inputting the equirectangular projection (ERP) image and the cubemap projection (CMP) image of the panorama, the method uses the feature extraction ability of the convolutional neural network to capture global object information. The equirectangular projection image and six cubemap projection images are used as the inputs of the network simultaneously, aiming to utilize the respective advantages of these two projections and combine the globality of the equirectangular projection and the small distortion of the cubemap projection. And different levels of features are adaptively integrated through different weights and common spatial attention to generate the final saliency map. However, since this method only considers the feature information such as color, texture, illumination, and composition in the color image space that guides the user's attention, while ignoring the spatial structure information in the real physical world, that is, depth information, the detection of this method is inaccurate in scenarios where the background and the salient object are similar. Summary of the Invention

[0005] In order to overcome the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a method, system, device and medium for omnidirectional salient object detection in panoramic images based on depth information fusion, construct a method for generating a depth map of panoramic images, design an end-to-end salient object detection network model for depth flow and saliency map flow, and by fusing color image features and depth image features, it can capture the characteristics of the saliency map flow while paying attention to the depth map flow information in the image imaging spatial structure to obtain a better detection effect.

[0006] In order to achieve the above purpose, the technical solution adopted by the present invention is:

[0007] A method for omnidirectional salient object detection in panoramic images based on depth information fusion, comprising the following steps:

[0008] S1: Obtain the corresponding depth image from the panoramic image through a depth estimation method;

[0009] S2: Obtain the corresponding preliminary salient object image from the panoramic image through a salient object detection method;

[0010] S3: Extract features from the depth image obtained in step S1 and the preliminary saliency image obtained in step S2 respectively to obtain depth map flow features and saliency map flow features;

[0011] S4: Fuse the depth map flow features and saliency map flow features extracted in step S3 to generate fused features;

[0012] S5: Decode the fused features generated in step S4 to obtain the final detection result; meanwhile, jointly train the depth estimation network, saliency object detection network and feature fusion network according to the detection result to optimize the performance of the entire system;

[0013] S6: Input the panoramic image into the network model after joint training in step S5, detect the saliency object in the image, and infer the saliency object result map of the panoramic image.

[0014] The specific steps of step 1 include:

[0015] S1.1: Construct a depth estimation network model, using BiFuse++ as the basic network architecture. The BiFuse++ network includes a dual-path feature extraction module, a multi-scale feature fusion module and a depth regression output module;

[0016] S1.2: Use the DepthAnything model to perform depth estimation on the training samples to generate virtual ground truth labels. The DepthAnything model is a pre-trained large-scale depth estimation model, and the training samples are panoramic images without manually annotated depth values;

[0017] S1.3: Network model training. Use the virtual ground truth labels generated in step S1.2 and the dataset with ground truth labels as supervision to perform end-to-end training on the BiFuse++ network adopted in step S1.1, and optimize the network parameters. The training process uses the L1 loss function and the gradient descent optimization algorithm;

[0018] S1.4: Apply the BiFuse++ network trained in step S1.3 to the panoramic saliency object detection dataset, obtain the depth information of each panoramic image through network inference, and construct a depth dataset for saliency object detection. The depth dataset includes panoramic images, depth images and corresponding saliency annotation information.

[0019] The specific steps of step 2 include:

[0020] S2.1: First project the panoramic image into an equidistant cylindrical projection format. Among them, the resolution of the projected panoramic image is 1024×512 pixels. The projection of the panoramic image into an equidistant cylindrical format is shown in formulas (2-5), (2-6), (2-7):

[0021]

[0022] y = sin(θ) (2 - 6)

[0023]

[0024] wherein, represents the longitude and latitude coordinates on the spherical surface, and (x, y, z) represents a point in three-dimensional space;

[0025] S2.2: Input the panoramic image projected in step 2.1 into the saliency object detection network to obtain a preliminary saliency object image. The input panoramic image successively passes through a feature extraction module that uses a residual convolutional network as the backbone network; the saliency object detection network includes sub-modules for channel attention and spatial attention; a multi-scale feature fusion module; and finally, through a saliency prediction module, a saliency object image is output.

[0026] The method for extracting depth map flow features in step 3 includes:

[0027] S3.1: Use the depth image to extract panoramic depth map features through the VGG16 network to obtain depth map flow features;

[0028] The VGG16 network is used as the basic architecture. This VGG16 network encodes the input image into five levels of feature vectors, respectively denoted as F1, F2, F3, F4, and F5;

[0029] S3.2: Take the feature vectors F1 and F2 encoded in step 3.1 as the shared features of the branch structure. Based on the shared features F1 and F2, perform convolution through the VGG16 network to obtain the feature vectors F3, F4, and F5 of the attention branch. Then input the feature vectors F3, F4, and F5 into the max-pooling unit to extract global context information, and finally output the attention feature vector through the Gaussian kernel convolution unit. The Gaussian kernel convolution unit is specifically shown in Equation (3 - 1):

[0030]

[0031] where G(i, j) represents the Gaussian convolution kernel, I(x + i, y + j) represents the pixel value at the pixel coordinate (x + i, y + j), (x, y) represents the pixel coordinate, and (i, j) represents the convolution kernel weight coordinate;

[0032] S3.3: In the detection branch, take the output feature F att of the attention branch in step S3.2 as the input, and obtain the features and

[0033] The significant graph flow feature extraction method in step 3 includes:

[0034] Step 3.1: Using the preliminary saliency target image, extract the panoramic image saliency features through the VGG16 network to obtain the significant graph flow features;

[0035] The VGG16 network is used as the basic architecture. The VGG16 network encodes the input image into five levels of feature vectors, denoted as F1, F2, F3, F4, and F5 respectively;

[0036] Step 3.2: Take the feature vectors F1 and F2 encoded in step 3.1 as the shared features of the branch structure. Based on the shared features F1 and F2, perform convolution through the VGG16 network to obtain the feature vectors F3, F4, and F5 of the attention branch. Then input the feature vectors F3, F4, and F5 into the max pooling unit to extract the global context information, and finally output the attention feature vector F through the Gaussian kernel convolution unit att , and the Gaussian kernel convolution unit is specifically shown in Equation (3-1):

[0037]

[0038] Among them, G(i,j) represents the Gaussian convolution kernel, I(x+i,y+j) represents the pixel value at the pixel coordinate (x+i,y+j), (x,y) represents the pixel coordinate, and (i,j) represents the convolution kernel weight coordinate;

[0039] Step 3.3: In the detection branch, take the output feature F of the attention branch in step 3.2 att as the input, and obtain the features after being processed by the third, fourth, and fifth convolutional layers of the VGG16 network and

[0040] Step 4 specifically includes:

[0041] S4.1: Respectively extract the attention vectors of the depth flow and the significant graph flow features. Use the attention module to extract the attention vectors of the depth flow and the significant graph flow respectively, specifically including: first perform global average pooling on the depth flow features and the significant graph flow features respectively; then perform feature transformation through the fully connected layer FC; then use the softmax activation function to calculate the channel attention weights; finally output the depth map flow attention vector and the significant graph flow attention vector, and the specific forms are shown in Equations (3-2) and (3-3):

[0042]

[0043] Among them, and respectively represent the input features of the saliency map stream and the depth map stream feature extraction modules, AvgPooling represents the average pooling operation, W i and b i respectively represent the weights and biases of the convolutional operation, and respectively represent the attention vectors of the saliency map stream and the depth map stream features;

[0044] S4.2: Cross - feature fusion, which realizes the effective fusion of the saliency map features and the depth map features through the CFM module. The CFM module includes: a global statistical information extraction unit for obtaining the global features of the saliency map stream and the depth map stream; a channel attention calculation unit for generating a channel attention vector; a feature enhancement unit for channel feature selection and enhancement; and a cross - modal fusion unit for generating a cross - referenced channel attention feature vector Its specific form is shown in Equation (3 - 5):

[0045]

[0046] Among them, and are respectively the attention vectors of the saliency map stream and the depth map stream features obtained in step S4.1, represents the channel attention feature, Max represents the maximum - value operation, represents the non - linear activation operation;

[0047] S4.3: Using the channel attention feature vector in step S4.2 to generate the fusion feature, which specifically includes: first generating the channel - enhanced feature through the channel multiplication operation; then aggregating the attention feature vectors of the saliency map stream and the depth map stream; then performing the feature normalization process; and finally using a 1×1 convolutional layer to merge the enhanced feature and output the fusion feature Its specific form is shown in Equation (3 - 6):

[0048]

[0049] Among them, and are respectively the attention vectors of the saliency map stream and the depth map stream features obtained in step S4.1, represents the channel attention feature, and respectively represent the input features of the saliency map stream and the depth map stream feature extraction modules, Concat represents the concatenation operation.

[0050] The specific content of step 5 includes:

[0051] S5.1: The input fused feature enters the multi-scale feature decoding module, and the fused feature is decoded by the multi-scale feature decoding module. The multi-scale feature decoding module includes: four parallel processing branches, each branch contains a 1×1 convolutional layer for channel number compression; an asymmetric convolutional layer for enhancing the multi-scale expression ability of features; an atrous convolutional layer for expanding the receptive field; a feature integration unit for multi-level feature fusion; and finally, the final panoramic image saliency object detection result map is obtained.

[0052] S5.2: The saliency object detection result map obtained through step S5.1 is used to supervise the network training. The BCELoss is used as the loss function for network training, which specifically includes: calculating the BCELoss for the saliency map branch, depth map branch, and fusion branch respectively; summing up the loss functions of each branch. The BCELoss is used to measure the probability distribution difference between the saliency map predicted by the model and the true saliency map, and its mathematical form is shown in Equation (5-1):

[0053]

[0054] where, represents y i the label of the i-th pixel in the true saliency map, p i represents the saliency probability of the i-th pixel predicted by the model, and N is the total number of pixels.

[0055] A panoramic image saliency object detection system based on depth information fusion includes:

[0056] Feature encoding unit: From the panoramic image, through the depth estimation method and the saliency object detection method, the corresponding depth image and the preliminary saliency object image are obtained. Then, features are extracted from the depth image and the preliminary saliency image respectively to obtain the depth map stream feature and the saliency map stream feature.

[0057] Feature processing unit: Fuse the depth map stream feature and the saliency map stream feature to generate a fused feature.

[0058] Feature decoding unit: Decode the fused feature to obtain the final detection result; meanwhile, according to the detection result, jointly train the depth estimation network, the saliency object detection network, and the feature fusion network to optimize the performance of the entire system; after joint training, input the panoramic image into the network model, detect the saliency object in the image, and infer the panoramic image saliency object result map.

[0059] A panoramic image saliency object detection device based on depth information fusion includes:

[0060] Memory: Used to store the computer program for implementing the panoramic image saliency object detection method based on depth information fusion.

[0061] Processor: When executing the computer program, it implements the method for panoramic image salient object detection based on depth information fusion as described above.

[0062] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the method for panoramic image salient object detection based on depth information fusion as described above.

[0063] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0064] 1. The present invention uses the DepthAnything model to generate virtual ground truth labels for panoramic images and trains the BiFuse++ network on a combined set of labeled images and pseudo-labeled images. By combining the advantages and limitations of the DepthAnything model and the BiFuse++ network, a depth dataset is constructed, improving the accuracy of depth estimation.

[0065] 2. The present invention fuses the depth map stream features and the saliency map stream features, and combines the attention mechanism to accurately integrate the most discriminative channels in the saliency map features and the depth features, and generates enhanced features through channel multiplication operations, effectively highlighting the key salient features, while suppressing noise and redundant information, and realizing multi-feature fusion.

[0066] 3. The present invention proposes a feature extraction method for depth maps and saliency maps. This method performs multi-level salient feature extraction on depth maps and saliency maps through a convolutional neural network. On this basis, an attention mechanism based on a Gaussian convolution kernel is constructed, which can adaptively suppress the noise interference regions in the image while enhancing the feature expression of the salient objects. By combining the extracted attention features with salient object detection, the feature extraction efficiency and detection accuracy of depth maps and saliency maps are improved, providing reliable technical support for object detection in complex scenarios.

[0067] 4. The present invention constructs a multi-branch joint loss function (BCELoss) based on binary cross-entropy for the training supervision of the salient object detection network. This loss function can not only effectively measure the probability distribution difference between the predicted saliency map and the true saliency map of the model, but also jointly optimize the network outputs of the saliency map branch, the depth map branch, and the fusion branch through a multi-task learning mechanism. This design significantly improves the network's ability to extract features from each modality, and further improves the accuracy and robustness of salient object detection in panoramic images through a collaborative optimization strategy.

[0068] In summary, the present invention proposes a panoramic image saliency target detection network based on depth information fusion. The detection network includes panoramic depth map acquisition at the input end, feature extraction and processing of the depth map stream and the saliency map stream, fusion of the depth map stream features and the saliency map stream features, and multi-scale feature fusion at the decoding end, suppressing noise and redundant information, and improving the accuracy and robustness of panoramic image saliency target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 It is a schematic diagram of the detection network architecture described in the present invention.

[0070] Figure 2 It is a network architecture diagram of the panoramic image depth map generation model of the present invention.

[0071] Figure 3 It is a structural diagram of the feature fusion network of the present invention.

[0072] Figure 4 It is a schematic diagram of the visualization result of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0073] The following describes the present invention in detail with reference to the drawings.

[0074] This embodiment was carried out on a computer equipped with an AMD(R) Ryzen 9 5900x CPU with 32GB of memory and an NVIDIA GeForce GTX 3060 GPU. The neural network was built using the deep learning framework pytorch. The operating system used was Ubuntu 20.04, the CUDA version was 11.1, and the PyTorch version was 1.8.0. During network training, we used the Stochastic Gradient Descent (SGD) optimizer, which is a classic optimization algorithm widely used for training deep learning models. The initial learning rate was set to 0.0005, which is an empirical value that performs well in many vision tasks. The batch size was set to 4, and 100 epochs were trained.

[0075] As Figure 1 shown, a panoramic image saliency target detection algorithm based on depth information fusion extracts the depth information and saliency map features of the panoramic image using a deep learning-based method. A feature extraction module with attention feature extraction and saliency feature branches is constructed through VGG16, and a multi-information feature fusion module based on the attention mechanism is constructed to efficiently fuse the panoramic image depth features and saliency map features to further improve the accuracy of the panoramic image saliency target detection algorithm. Finally, the supervised network is trained with BCE loss as the loss function to obtain a network model that can be used for panoramic image saliency target detection, specifically including the following steps:

[0076] S1: Obtain the corresponding depth image from the panoramic image through a depth estimation method;

[0077] S2: Obtain the corresponding preliminary saliency target image from the panoramic image through a saliency target detection method;

[0078] S3: Extract features from the depth image obtained in step S1 and the preliminary saliency image obtained in step S2 respectively to obtain depth map flow features and saliency map flow features;

[0079] S4: Fuse the depth map flow features and saliency map flow features extracted in step S3 to generate fused features;

[0080] S5: Decode the fused features generated in step S4 to obtain the final detection result; meanwhile, jointly train the depth estimation network, saliency target detection network, and feature fusion network according to the detection result to optimize the performance of the entire system;

[0081] S6: Input the panoramic image into the network model after joint training in step S5, detect the saliency target in the image, and infer the saliency target result map of the panoramic image.

[0082] As Figure 2 shown, step 1 specifically includes:

[0083] S1.1: Construct a depth estimation network model, using BiFuse++ as the basic network architecture. The BiFuse++ network includes a dual-path feature extraction module, a multi-scale feature fusion module, and a depth regression output module;

[0084] S1.2: Use the DepthAnything model to perform depth estimation on the training samples to generate virtual ground truth labels. The DepthAnything model is a pre-trained large-scale depth estimation model;

[0085] S1.3: Network model training. Use the virtual ground truth labels generated in step S1.2 and the Stanford2D3D dataset as supervision to perform end-to-end training on the BiFuse++ network used in step S1.1, and optimize the network parameters. The training process uses the L1 loss function and the gradient descent optimization algorithm. The specific form of the L1 loss function is as shown in the formula:

[0086]

[0087] where P(i, j) is the predicted saliency value at pixel position (i, j), G(i, j) represents the true label at the corresponding pixel position, and W and H represent the width and height of the image respectively;

[0088] S1.4: Apply the trained BiFuse++ network in step S1.3 to the panoramic saliency object detection dataset 360-SOD, obtain the depth information of each sample through network inference, and construct a depth dataset for saliency object detection. The depth dataset includes panoramic images, depth images, and corresponding saliency annotation information.

[0089] The specific steps of step 2 are as follows:

[0090] S2.1: First, project the panoramic image into an equidistant cylindrical projection format. Among them, the resolution of the projected panoramic image is 1024×512 pixels. The projection of the panoramic image into an equidistant cylindrical format is shown in equations (2-5), (2-6), and (2-7):

[0091]

[0092] y = sin(θ) (2-6)

[0093]

[0094] Among them, represents the longitude and latitude coordinates on the spherical surface, and (x, y, z) represents the points in three-dimensional space;

[0095] S2.2: Input the panoramic image projected in step 2.1 into the saliency object detection network FANet to obtain a preliminary saliency object image. The input panoramic image sequentially passes through a feature extraction module that uses ResNet50 as the backbone network. The saliency object detection network includes sub-modules for channel attention and spatial attention, a multi-scale feature fusion module, and finally outputs a saliency object image through a saliency prediction module.

[0096] Among them, the specific steps for ResNet50 to extract four layers of features are as follows:

[0097] For example, if the input image size is 3×H×W, then the output feature size is

[0098] The depth map flow feature extraction method in step 3 includes:

[0099] S3.1: Use the depth image to extract panoramic depth map features through the VGG16 network to obtain depth map flow features;

[0100] The VGG16 network is used as the basic architecture. This VGG16 network encodes the input image into five levels of feature vectors through five convolutional layers Conv1_1, Conv2_1, Conv3_1, Conv4_1, and Conv5_1, which are respectively represented as F1, F2, F3, F4, and F5;

[0101] S3.2: Take the feature vectors F1 and F2 obtained by encoding in step 3.1 as the shared features of the branch structure. Based on the shared features F1 and F2, perform convolution through the VGG16 network to obtain the feature vectors F3, F4, and F5 of the attention branch. Then input the feature vectors F3, F4, and F5 into the max-pooling unit to extract global context information, and finally output the attention feature vector F through the Gaussian kernel convolution unit att Gaussian kernel convolution is a commonly used technique in image processing and computer vision. It is based on the Gaussian distribution to achieve a smoothing or blurring effect. By using a Gaussian convolution kernel to more effectively suppress the noise regions in the image, the Gaussian kernel convolution unit is specifically shown in Equation (3-1):

[0102]

[0103] where G(i,j) represents the Gaussian convolution kernel, I(x+i,y+j) represents the pixel value at the pixel coordinate (x+i,y+j), (x,y) represents the pixel coordinate, and (i,j) represents the convolution kernel weight coordinate;

[0104] S3.3: In the detection branch, take the output feature F of the attention branch in S3.2 att as the input, and process it through the third, fourth, and fifth convolutional layers of the VGG16 network to obtain the features and

[0105] The significant map flow feature extraction method in step 3 includes:

[0106] Step 3.1: Use the preliminary significant object image to extract the panoramic image significant features through the VGG16 network to obtain the significant map flow features;

[0107] Taking the VGG16 network as the basic architecture, this VGG16 network encodes the input image into five levels of feature vectors through five convolutional layers Conv1_1, Conv2_1, Conv3_1, Conv4_1, and Conv5_1, which are respectively represented as F1, F2, F3, F4, and F5;

[0108] Step 3.2: Take the feature vectors F1 and F2 obtained by encoding in step 3.1 as the shared features of the branch structure. Based on the shared features F1 and F2, perform convolution through the VGG16 network to obtain the feature vectors F3, F4, and F5 of the attention branch. Then input the feature vectors F3, F4, and F5 into the max-pooling unit to extract global context information, and finally output the attention feature vector F through the Gaussian kernel convolution unit att, Gaussian kernel convolution is a commonly used technique in image processing and computer vision. It is based on the Gaussian distribution to achieve a smoothing or blurring effect. By using a Gaussian convolution kernel to more effectively suppress the noise regions in an image, the Gaussian kernel convolution unit is specifically shown in Equation (3-1):

[0109]

[0110] where G(i,j) represents the Gaussian convolution kernel, I(x+i,y+j) represents the pixel value at pixel coordinates (x+i,y+j), (x,y) represents the pixel coordinates, and (i,j) represents the convolution kernel weight coordinates;

[0111] Step 3.3: In the detection branch, using the output feature F of the attention branch in Step 3.2 att as the input, after being processed by the third, fourth, and fifth convolutional layers of the VGG16 network, the features and

[0112] as Figure 3 shown, the specific steps of Step 4 include:

[0113] S4.1: Respectively extract the attention vectors of the features of the depth stream and the saliency map stream. Use the attention module to extract the attention vectors of the depth stream and the saliency map stream respectively, which specifically includes: first perform global average pooling on the depth stream features and the saliency map stream features respectively; then perform feature transformation through the fully connected layer FC; then use the softmax activation function to calculate the channel attention weights; finally, output the attention vectors of the depth map stream and the saliency map stream, and the specific forms are shown in Equations (3-2) and (3-3):

[0114]

[0115] where and respectively represent the input features of the feature extraction modules of the saliency map stream and the depth map stream, AvgPooling represents the average pooling operation, W i and b i respectively represent the weights and biases of the convolutional operation, and respectively represent the attention vectors of the saliency map stream and the depth map stream features;

[0116] S4.2: Cross - feature fusion, which realizes the effective fusion of saliency map features and depth map features through the CFM module. The CFM module includes: a global statistics extraction unit for obtaining the global features of the saliency map stream and the depth map stream; a channel attention calculation unit for generating a channel attention vector; a feature enhancement unit for channel feature selection and enhancement; and a cross - modal fusion unit for generating a cross - referenced channel attention feature vector Its specific form is shown in Equation (3 - 5):

[0117]

[0118] Wherein, and are the attention vectors of the saliency map stream and the depth map stream features obtained in step S4.1 respectively, represents the channel attention feature, Max represents the maximum - value operation, represents the non - linear activation operation;

[0119] S4.3: Using the channel attention feature vector in step S4.2 to generate the fusion feature. This process not only enhances the feature expression ability but also promotes the effective fusion of RGB and depth information. Specifically, it includes: first generating the channel - enhanced feature through the channel multiplication operation; then aggregating the attention feature vectors of the saliency map stream and the depth map stream; then performing the feature normalization process; and finally using the 1×1 convolutional layer to merge the enhanced feature and output the fusion feature Its specific form is shown in Equation (3 - 6):

[0120]

[0121] Wherein, and are the attention vectors of the saliency map stream and the depth map stream features obtained in step S4.1 respectively, represents the channel attention feature, and represent the input features of the saliency map stream and the depth map stream feature extraction modules respectively, and Concat represents the concatenation operation.

[0122] Step 5 specifically includes:

[0123] S5.1: Input the fusion feature into the multi - scale feature decoding module, and decode the fusion feature through the multi - scale feature decoding module. Specifically, we designed an efficient context module. This module consists of four branches {b m , m = 1,..., 4}, and each branch first reduces the number of channels to 32 through the 1x1 convolutional layer. For non - first branches {b m, where m > 1, we also introduced asymmetric convolutional layers and dilated convolutional layers of different sizes to enhance the multi-scale representation ability of features. Specifically, in the m-th branch, an asymmetric convolutional layer of 1×(2m - 1) and (2m - 1)×1 and a 3×3 dilated convolutional layer with a dilation rate of (2m - 1) are added after the 1×1 convolutional layer. Finally, multi-level features are integrated through the upsampling and feature connection strategies, and the final feature map is generated using a 3×3 convolutional layer and a 1×1 convolutional layer, and at the same time, it is adjusted to the target size to obtain the final panoramic image saliency object detection result map;

[0124] S5.2: Supervise the network training with the saliency object detection result map obtained in step S5.1, and use BCELoss as the loss function for network training, which specifically includes: calculating BCELoss for the saliency map branch, depth map branch, and fusion branch respectively; summing up the loss functions of each branch. BCELoss is used to measure the difference in probability distribution between the saliency map predicted by the model and the true saliency map, and its mathematical form is shown in Equation (5-1):

[0125]

[0126] where, represents y i the label of the i-th pixel in the true saliency map, and p i represents the saliency probability of the i-th pixel predicted by the model, and N is the total number of pixels; by minimizing the loss function, the model can better learn the boundary and detail information of the saliency object.

[0127] Finally, to unify the training of the three-way branches, the design of the loss function takes into account the process features of the RGB branch and the depth map branch. The final loss function is shown in Equation (5-2) and is used to supervise the end-to-end network training.

[0128] L Total = L RGB +L Depth +L Fusion (5-2)

[0129] where, L RGB represents the loss between the saliency object map obtained from the RGB stream and the true label, and L Depth represents the loss between the saliency object map obtained from the depth stream and the true label, and L Fusion represents the loss between the saliency object map obtained after fusing the RGB stream and the depth stream and the true label. Their sum gives the final loss L Total to supervise the network training.

[0130] As Figure 4 shown, step 6 specifically includes:

[0131] Figure 4 Some representative visualization results are listed. The first column is the RGB image, the second column is the label, the third column is the result of the method of the present invention, the fourth column is the result of the depth flow branch detection, and the fifth column is the result of the saliency map flow detection branch. It can be seen that the method proposed by the present invention has a more refined and accurate detection effect for salient object segmentation.

[0132] Verified by experiments, the mean absolute error (MAE) of the present invention in the salient object detection task is 0.0181, the Max-F metric reaches 0.7959, the Max-E metric is 0.9020, and the S-measure metric is 0.8467. The salient object detection method of the present invention can not only independently analyze the panoramic image through the depth information branch and the image spatial domain information branch, but also effectively combine the advantages of the two through the feature fusion module to generate more accurate salient object detection results, providing more reliable object detection support for panoramic image analysis and computer vision systems.

[0133] The present invention respectively analyzes the performance of the depth map branch and the saliency map branch and the network after feature fusion in the panoramic image salient object detection task, compares the performance of different branches, and verifies the feasibility of introducing depth information in the panoramic salient object detection task and the effectiveness of the fusion module designed in this paper. Among them, the MAE metric is improved by 23.3% compared with the depth information branch, the Max-F metric is improved by 5.16% compared with the depth information branch, the Max-E metric is improved by 0.83% compared with the depth information branch, and the S-measure metric is improved by 2.42% compared with the depth information branch.

[0134] A panoramic image salient object detection system based on depth information fusion, comprising:

[0135] Feature encoding unit: Through the depth estimation method and the salient object detection method for the panoramic image, the corresponding depth image and the preliminary salient object image are obtained, and then features are extracted from the depth image and the preliminary salient image respectively to obtain the depth map flow feature and the saliency map flow feature, which are used to implement steps 1 to 3 of a panoramic image salient object detection based on depth information fusion;

[0136] Feature processing unit, which fuses the depth map flow feature and the saliency map flow feature to generate a fusion feature, which is used to implement step 4 of a panoramic image salient object detection based on depth information fusion;

[0137] Feature decoding unit: Decode the fused features to obtain the final detection results; at the same time, jointly train the depth estimation network, the saliency object detection network, and the feature fusion network according to the detection results to optimize the performance of the entire system; after the joint training, input the panoramic image into the network model, detect the saliency object in the image, and infer the saliency object result map of the panoramic image, which is used to implement steps 5-6 of a panoramic image saliency object detection method based on depth information fusion.

[0138] A panoramic image saliency object detection device based on depth information fusion, comprising:

[0139] Memory: Used to store the computer program for implementing the panoramic image saliency object detection method based on depth information fusion described above;

[0140] Processor: Used to implement the panoramic image saliency object detection method based on depth information fusion described above when executing the computer program.

[0141] The present invention also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the panoramic image saliency object detection method based on depth information fusion described above are implemented.

Claims

1. A panoramic image salient object detection method based on depth information fusion, characterized in that, It includes the following steps: S1: Obtain the corresponding depth image from the panoramic image through a depth estimation method; S2: Obtain the corresponding preliminary saliency target image from the panoramic image through a saliency target detection method; S3: Extract features from the depth image obtained in step S1 and the preliminary saliency image obtained in step S2 respectively to obtain depth map stream features and saliency map stream features; S4: Fuse the depth map stream features and saliency map stream features extracted in step S3 to generate fused features; S5: Decode the fused features generated in step S4 to obtain the final detection result; meanwhile, jointly train the depth estimation network, saliency target detection network, and feature fusion network according to the detection result to optimize the performance of the entire system; S6: Input the panoramic image into the network model after joint training in step S5, detect the saliency target in the image, and infer the saliency target result map of the panoramic image.

2. The panoramic image saliency target detection method based on depth information fusion according to claim 1, wherein, The specific steps of step 1 include: S1.1: Construct a depth estimation network model, using BiFuse++ as the basic network architecture. The BiFuse++ network includes a dual-path feature extraction module, a multi-scale feature fusion module, and a depth regression output module; S1.2: Use the DepthAnything model to perform depth estimation on the training samples to generate virtual ground truth labels. The DepthAnything model is a pre-trained large-scale depth estimation model, and the training samples are panoramic images without manually annotated depth values; S1.3: Network model training. Use the virtual ground truth labels generated in step S1.2 and the dataset with ground truth labels as supervision to perform end-to-end training on the BiFuse++ network used in step S1.1, and optimize the network parameters. The training process uses the L1 loss function and the gradient descent optimization algorithm; S1.4: Apply the BiFuse++ network trained in step S1.3 to the panoramic saliency target detection dataset, obtain the depth information of each panoramic image through network inference, and construct a depth dataset for saliency target detection. The depth dataset includes panoramic images, depth images, and corresponding saliency annotation information.

3. A panoramic image salient object detection method based on depth information fusion according to claim 1, characterized in that The specific steps of step 2 include: S2.1: First project the panoramic image into an equidistant cylindrical projection format. Among them, the resolution of the projected panoramic image is 1024×512 pixels. The projection of the panoramic image into an equidistant cylindrical format is shown in formulas (2-5), (2-6), and (2-7): y = sin(θ) (2-6) Among them, represents the longitude and latitude coordinates on the spherical surface, and (x, y, z) represents a point in three-dimensional space; S2.2: Input the panoramic image projected in step 2.1 into the saliency target detection network to obtain a preliminary saliency target image. The input panoramic image sequentially passes through a feature extraction module using a residual convolutional network as the backbone network. The saliency target detection network includes sub-modules with channel attention and spatial attention; a multi-scale feature fusion module; and finally, through a saliency prediction module, output the saliency target image.

4. A panoramic image salient object detection method based on depth information fusion according to claim 1, characterized in that, The depth map stream feature extraction method in step 3 includes: S3.1: Using the depth image, extract the panoramic depth map features through the VGG16 network to obtain the depth map stream features; The VGG16 network is used as the basic architecture. This VGG16 network encodes the input image into five levels of feature vectors, denoted as F1, F2, F3, F4, and F5 respectively; S3.2: Take the feature vectors F1 and F2 encoded in step 3.1 as the shared features of the branch structure. Based on the shared features F1 and F2, perform convolution through the VGG16 network to obtain the feature vectors F3, F4, and F5 of the attention branch. Then input the feature vectors F3, F4, and F5 into the max pooling unit to extract the global context information, and finally output the attention feature vector through the Gaussian kernel convolution unit. The Gaussian kernel convolution unit is specifically shown in Equation (3-1): Among them, G(i,j) represents the Gaussian convolution kernel, I(x+i,y+j) represents the pixel value at the pixel coordinate (x+i,y+j), (x,y) represents the pixel coordinate, and (i,j) represents the convolution kernel weight coordinate; S3.3: In the detection branch, using the output feature F of the attention branch in step S3.2 att as the input, and obtaining features after being processed by the third, fourth, and fifth convolutional layers of the VGG16 network and 5. A panoramic image salient object detection method based on depth information fusion according to claim 1, characterized in that, The significant map stream feature extraction method in step 3 includes: Step 3.1: Using the preliminary significant object image, extract the panoramic image saliency features through the VGG16 network to obtain the significant map stream features; The VGG16 network is used as the basic architecture. This VGG16 network encodes the input image into five levels of feature vectors, denoted as F1, F2, F3, F4, and F5 respectively; Step 3.2: Take the feature vectors F1 and F2 encoded in Step 3.1 as the shared features of the branch structure. Based on the shared features F1 and F2, perform convolution through the VGG16 network to obtain the feature vectors F3, F4, and F5 of the attention branch. Then input the feature vectors F3, F4, and F5 into the max pooling unit to extract global context information, and finally output the attention feature vector F through the Gaussian kernel convolution unit att , and the Gaussian kernel convolution unit is specifically shown in Equation (3-1): Among them, G(i,j) represents the Gaussian convolution kernel, I(x+i,y+j) represents the pixel value at the pixel coordinate (x+i,y+j), (x,y) represents the pixel coordinate, and (i,j) represents the convolution kernel weight coordinate; Step 3.3: In the detection branch, using the output feature F of the attention branch in Step 3.2 att as the input, and obtaining features after being processed by the third, fourth, and fifth convolutional layers of the VGG16 network and 6. A panoramic image salient object detection method based on depth information fusion according to claim 1, characterized in that Step 4 specifically includes: S4.1: Extract the attention vectors for the features of the depth stream and the significant map stream respectively. Use the attention module to extract the attention vectors of the depth stream and the significant map stream respectively, specifically including: first perform global average pooling on the depth stream features and the significant map stream features respectively; then perform feature transformation through the fully connected layer FC; then use the softmax activation function to calculate the channel attention weights; finally output the depth map stream attention vector and the significant map stream attention vector, and the specific forms are shown in Equations (3-2) and (3-3): Among them, and respectively represent the input features of the saliency map stream and depth map stream feature extraction modules, AvgPooling represents the average pooling operation, W i and b i respectively represent the weights and biases of the convolutional operation, and respectively represent the attention vectors of the saliency map stream and depth map stream features; S4.2: Cross - feature fusion, which effectively fuses the salient map features and depth map features through the CFM module. The CFM module includes: a global statistics extraction unit for obtaining the global features of the salient map stream and the depth map stream; a channel attention calculation unit for generating a channel attention vector; a feature enhancement unit for channel feature selection and enhancement; and a cross - modal fusion unit for generating a cross - referenced channel attention feature vector Its specific form is shown in Equation (3 - 5): Among them, and are the attention vectors of the salient graph stream and the depth graph stream features obtained in step S4.1 respectively, represents the channel attention feature, and Max represents the operation of taking the maximum value, represents the non-linear activation operation; S4.3: Utilize the channel attention feature vector in step S4.2 to generate fused features, specifically including: first generating channel-enhanced features through channel multiplication operations; then aggregating the attention feature vectors of the saliency map stream and the depth map stream; then performing feature normalization processing; and finally using a 1×1 convolutional layer to merge the enhanced features and output the fused features Its specific form is shown in Equation (3-6): Among them, and are the attention vectors of the saliency map stream and the depth map stream features obtained in step S4.1 respectively, represents the channel attention feature, and represent the input features of the saliency map stream and the depth map stream feature extraction modules respectively, Concat represents the concatenation operation.

7. A method for panoramic image salient object detection based on depth information fusion according to claim 1, characterized in that Step 5 specifically includes: S5.1: Input the fusion features into the multi-scale feature decoding module, and decode the fusion features through the multi-scale feature decoding module. The multi-scale feature decoding module includes: four parallel processing branches, each branch contains a 1×1 convolutional layer for channel number compression; an asymmetric convolutional layer for enhancing the multi-scale expression ability of the features; a dilated convolutional layer for expanding the receptive field; a feature integration unit for multi-level feature fusion; finally obtain the final panoramic image significant object detection result map; S5.2: Supervise the network training with the significant object detection result map obtained in step S5.1, and use BCELoss as the loss function for network training, which specifically includes: calculating BCELoss for the significant map branch, depth map branch, and fusion branch respectively; summing up the loss functions of each branch. BCELoss is used to measure the probability distribution difference between the significant map predicted by the model and the true significant map, and its mathematical form is shown in Equation (5-1): where, y is represented i the label of the i-th pixel in the true saliency map, p i represents the saliency probability of the i-th pixel predicted by the model, and N is the total number of pixels.

8. A panoramic image saliency target detection system based on depth information fusion according to the method described in claims 1 to 7, characterized in that, Including: Feature encoding unit: Through the depth estimation method and significant object detection method for the panoramic image, obtain the corresponding depth image and preliminary significant object image, and then extract features from the depth image and preliminary significant image respectively to obtain the depth map stream feature and significant map stream feature; Feature processing unit: Fuse the depth map stream feature and significant map stream feature to generate a fused feature; Feature decoding unit: Decode the fused feature to obtain the final detection result; at the same time, jointly train the depth estimation network, significant object detection network, and feature fusion network according to the detection result to optimize the performance of the entire system; after joint training, input the panoramic image into the network model to detect the significant object in the image and infer the significant object result map of the panoramic image.

9. A panoramic image salient object detection device based on depth information fusion, characterized in that, Including: Memory: Used to store the computer program for implementing a panoramic image significant object detection method according to any one of claims 1 to 7; Processor: Used to implement a panoramic image significant object detection method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the steps of a panoramic image significant object detection method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Elevator landing door lock engagement depth detection device and detection method

    CN121025999A