An intelligent transportation monocular depth estimation method based on orthogonal planes

By using orthogonal plane characteristics and mixed Gaussian estimation in the intelligent traffic monocular depth estimation method, the excessive smoothing and edge artifact problems of target boundary depth in monocular depth estimation are solved, and the depth prediction with high accuracy and robustness is achieved.

CN117893590BActive Publication Date: 2025-06-24TONGJI UNIV

Patent Information

Application Number
CN202410063481.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-06-24
Estimated Expiration
2044-01-16

AI Technical Summary

Technical Problem

Existing monocular depth estimation methods have problems with excessive smoothing and edge artifacts when dealing with target boundary depth, and require a large amount of labeled data and expensive sensors, and the dataset collection is complex.

Method used

Using an intelligent traffic monocular depth estimation method based on orthogonal planes, a global and local feature extraction module is constructed in the encoder, combining a non-moving window Transformer and a lightweight Transformer, pixel-level depth prediction is performed, and the final depth value is determined through mixed Gaussian estimation.

Benefits of technology

It effectively improves the depth estimation accuracy of the target edge, solves the aliasing problem when estimating depth in convolutional networks, reduces depth artifacts, and realizes real-time and stable depth positioning under unsupervised signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117893590B_ABST
    Figure CN117893590B_ABST
Patent Text Reader

Abstract

The present invention discloses a monocular depth estimation method for intelligent transportation based on an orthogonal plane, including: S1, obtaining an RGB image and an image dataset with the same resolution as its lower frame; S2, constructing a global and local feature extraction module to obtain encoded features at four resolutions; S3, sending the encoder output features into a non-mobile window Transformer for generating pixel depth to obtain decoded features at three different resolutions; S4, by sending the decoder depth probability representation into a bins classification network, outputting an interval bins group through a lightweight Transformer combination, and performing a Hadamard product on the bins interval group and the depth representation to obtain a classified depth value; S5, determining the final depth value through a mixture Gaussian estimation of the imaging plane depth probability output by the decoder and the optical axis multi-plane depth probability output by the multi-plane bins. According to the present invention, the present invention provides instant and accurate target depth information for intelligent driving vehicles in the traffic field, providing strong support for the accurate positioning of targets and the three-dimensional scene reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent transportation, and particularly relates to a monocular depth estimation method for intelligent transportation based on orthogonal planes. Background Art

[0002] In recent years, significant progress has been made in monocular depth estimation. Researchers have developed new algorithms and models, achieving excellent performance on various benchmarks. However, most of the existing state-of-the-art methods rely on supervised learning, which requires deploying expensive sensors and incurring a large cost for manual annotation. The collection of datasets is a complex and difficult challenge. As an alternative, researchers have adopted photometric consistency and self-supervised learning methods, treating depth estimation as the reconstruction between stereo images, monocular videos, or a combination of both to eliminate the need for ground truth. However, the number of three-dimensional scenes projected onto the same two-dimensional image is infinite, and depth map estimation is an ill-posed problem.

[0003] Traditional self-supervised monocular depth estimation methods use consecutive frames as input to regress a dense depth map pixel by pixel. Generally, these methods involve warping the reference view image to the target view and then using a reprojection loss to constrain the regression model to learn the geometric consistency between adjacent frames. Although these methods show potential in addressing data requirements, classical regression networks are affected by smoothing bias, especially depth value artifacts at the truncation of object boundaries. In addition, inconsistent pixel differences between moving foreground and background objects pose challenges for accurately matching occluded background boundary pixels. Therefore, these edge pixels are warped the same distance as the foreground pixels so that the foreground pixels of the source view and the target view are aligned as much as possible to achieve the minimum photometric loss, but these incorrect matching pixels further exacerbate the prediction distortion.

[0004] The commonly used depth feature extraction structure is an encoder-decoder. In this architecture, the encoder uses backbone networks such as ResNet, VGG, HRNet, and PackNet, etc. The decoder reconstructs the depth map based on these features, using UNet or GCN. The reprojection consistency between the reference view and the target view can be used as a constraint for depth estimation. However, these methods have a limitation in terms of the correlation between global and local features. Although the encoder incorporates multi-scale feature extraction to enhance the overall perception ability of the convolutional network, it fails to effectively integrate local and global features. Therefore, traditional convolutional regression networks have difficulty handling depth discontinuities at the boundaries of foreground objects when predicting object depth, resulting in problems such as over-smoothing and edge artifacts near object edges.

[0005] Another challenge in depth estimation is the depth discontinuity at the truncated object boundaries, which can lead to depth value artifacts. Although some methods use semantic segmentation or edge ground truth as prior information, take edge information as input, add branch constraints or post-processing to solve the problem of blurred object boundaries, these methods still require accurate labels, which do not conform to unsupervised depth estimation methods. At the same time, the depth values between the predicted and real boundaries are all assigned to the nearest neighbor background depth, which is obviously unreasonable. To solve this problem, a discretized multi-plane depth interval is introduced to transform depth regression into a pixel classification problem, and this method has been proven to achieve higher performance. However, it is worth noting that depth discretization will lead to a decrease in visual quality and obvious depth discontinuities. Summary of the Invention

[0006] Aiming at the deficiencies in the prior art, the purpose of the present invention is to provide a monocular depth estimation method for intelligent transportation based on orthogonal planes, which solves the problems of over-smoothing and edge artifacts when estimating the depth of an object and its boundary, so as to effectively extract the depth information of a monocular image. To achieve the above object and other advantages according to the present invention, a monocular depth estimation method for intelligent transportation based on orthogonal planes is provided, including the following steps:

[0007] S1. Obtain an RGB image with a resolution of H×W and an image dataset with the same resolution as its next frame;

[0008] S2. By constructing global and local feature extraction modules in the encoder, obtain encoded features of four resolutions;

[0009] S3. Feed the encoder output features into a non-shifting window Transformer for the generation of pixel depth, and obtain decoded features of three different resolutions;

[0010] S4. By feeding the decoder depth probability representation into a bins classification network, output an interval bins group through a lightweight Transformer combination, and perform a Hadamard product on the bins interval group and the depth representation to obtain a classified depth value;

[0011] S5. Determine the final depth value by a mixture of Gaussians estimation for the imaging plane depth probability output by the decoder and the optical axis multi-plane depth probability output by the multi-plane bins.

[0012] Preferably, in step S2, a convolutional network is used to extract features from the color image to generate features as small as H / 2×W / 2. After that, in each layer of global and local feature extraction, convolution and Transformer are used for parallel interaction, and the resolution decreases at a scale of 2 times. The extracted features are fed into a pyramid parsing module to extract multi-scale features.

[0013] Preferably, in step S3, the spatial resolution output disparity is used as the scene depth representation through cross-layer and cross-scale connections, and pixel-level depth prediction is performed through non-shifted Swin-Transformer blocks. The query and key vectors of each patch in the window are calculated using the feature map output by the encoder. The query and key vectors are the same, and then the dot product between the two is calculated to obtain the similarity score. Channel conversion is performed in the decoder, and at the same time, the moving window depth prediction is removed, reducing the artifacts of the target edge depth truncation.

[0014] Preferably, in step S4, the depth interval center bins are predicted through a simple post-processing transformer module:

[0015]

[0016] where d min and d min represent the minimum and maximum depths respectively, and b i represents the i-th interval distance.

[0017] Preferably, in step S4, the continuous depth is estimated through a plane discretization method. The decoder converts the volume plane into a pixel-level probability map, generates N embeddings as the image feature map, interacts with N bins queries, outputs the depth embeddings, and feeds the embeddings into a linear perceptron and softmax to obtain the normalized depth.

[0018] Preferably, the pixel (u, v) of the input image is parameterized as a Gaussian distribution to accurately locate the sampling selection of the bins. The training depth search space is [μ u,v -βσ u,v , μ u,v +βσ u,v , and the k-th depth hypothesis value is d u,v,k :

[0019] d u,v,k =μ u,v +b k σ u,v (5),

[0020] where μ u,v and σ u,v represent the depth mean and variance of the pixel point (u, v) respectively, and b k represents the probability function.

[0021] Preferably, in step S5, the two Gaussian models of imaging plane regression depth and optical axis multi-plane classification depth make probability predictions for the target depth from two directions respectively, and their means and variances are μ1, μ2, σ1, σ2; the probability density function of their mixture Gaussian distribution is:

[0022]

[0023] where μ depth as the maximum probability depth estimate, x i and μ i are respectively the depth mean and variance of the i-th depth estimate, represents the Gaussian mixture variance.

[0024] Compared with the prior art, the beneficial effects of the present invention are as follows: 1) For the monocular depth estimation scenarios of traffic scenes or unsupervised signals, such as block and field operation scenes, etc., the present invention uses a monocular camera to collect front and rear frame images, without the true value supervision training collected by radar, and provides real-time and stable depth positioning for the intelligent driving system;

[0025] 2) By adopting a pixel-level attention depth estimation system, it can effectively improve the depth estimation accuracy of the target edge, solve the aliasing problem when the convolutional network estimates depth, and effectively alleviate depth artifacts;

[0026] 3) Utilizing the orthogonality characteristics of the imaging plane and the optical axis multi-planes, the cross depth is used for Gaussian mixture distribution, realizing the unification of the regression network and the classification network, and ensuring the accuracy and robustness of depth prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a flowchart of the intelligent transportation monocular depth estimation method based on orthogonal planes according to the present invention;

[0028] Figure 2 is a schematic diagram of the orthogonal planes of the intelligent transportation monocular depth estimation method based on orthogonal planes according to the present invention;

[0029] Figure 3 is a model structure diagram of the intelligent transportation monocular depth estimation method based on orthogonal planes according to the present invention;

[0030] Figure 4 is a Gaussian mixture distribution probability diagram of the intelligent transportation monocular depth estimation method based on orthogonal planes according to the present invention;

[0031] Figure 5 is a depth map and error map on the KITTI dataset of the intelligent transportation monocular depth estimation method based on orthogonal planes according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0033] Referring to Figures 1-5 , an intelligent transportation monocular depth estimation method based on an orthogonal plane, comprising the following steps:

[0034] S1. Obtain an RGB image with a resolution of H×W and an image dataset of the same resolution in the next frame; perform global and local feature extraction on the original data, focusing on improving high-resolution features and accurately estimating small targets and distant targets. Therefore, each layer of global-local feature interaction constructs spatial information interaction within this scale and gradually decreases at a scale of 2 times. The attention features are as follows:

[0035]

[0036]

[0037] Among them, Att j represents the jth global multi-scale patch embedding extracted by each Transformer block, As a feature fusion module, one-dimensional convolution is used to merge the local representation Conv i and the global attention feature Att j . This fusion provides image feature maps of different scales for the decoder module.

[0038] S2. Obtain encoded features of four resolutions by constructing global and local feature extraction modules in the encoder; the depth decoder takes the multi-scale features extracted from the encoder and the pyramid module (PPM) as inputs, and realizes spatial resolution output through cross-layer and cross-scale connections. The disparity is used as the scene depth representation. The present invention proposes a pixel-level depth prediction method based on a non-shifted Swin-Transformer block, which calculates the query and key vectors of each patch in the window using the feature map output by the encoder where the query and key vectors are the same. Then, the dot product between the two calculates the scores between any two nodes, and the relative position embedding P is introduced, and the dot product operation is performed between the output F of the previous layer decoder and the SoftMax output to obtain the output of the current decoder

[0039]

[0040] S3. Feed the encoder output features into a non - moving window Transformer for generating pixel depth, and obtain decoded features of three different resolutions;

[0041] S4. By feeding the decoder depth probability representation into the bins classification network, and outputting an interval bins group through lightweight Transformer combination. Perform Hadamard product on the bins interval group and the depth representation to obtain the classified depth value; The decoder converts the volume plane into a pixel - level probability map, outputs N embeddings as the image mapping map F, interacts with N bins queries, and outputs the embedding Q. Then send Q into the linear perceptron SoftMax to generate N Finally, calculate the linear combination of the probability distribution of pixel points and the bins center c to obtain the final depth value d, as shown in Equation 1. The eigenvalue cost volume V output by the decoder is transformed into a multi - plane probability distribution through Transformer. If the depth intervals are too few, the distribution will be unreasonable; if the depth intervals are too many, too much computing resources will be occupied. Therefore, in the present invention, the pixel (u, v) parameters of the input image are parameterized as a Gaussian distribution. Assume that the depth search space of bins b is [μ u,v -βσ u,v ,μ u,v +βσ u,v , where β is a hyperparameter, and the k - th depth hypothesis value d u,v,k ,

[0042] d u,v,k =μ u,v +b k σ u,v ,

[0043]

[0044] where Φ(·) is the probability function, is the probability mass covered by the interval [μ u,v ±βσ u,v .

[0045] S5. Determine the final depth value through a mixture of Gaussians estimation for the depth probability of the imaging plane output by the decoder and the depth probability of the optical axis multi - plane output by the multi - plane bins. The two Gaussian models of the imaging plane regression depth and the optical axis multi - plane classification depth respectively make probability predictions for the target depth from two directions, and their means and variances are μ1, μ2, σ1, σ2 respectively. The probability density function of the mixture of Gaussians distribution is as shown in Equation 3.

[0046] Example 1

[0047] In the embodiment of the present invention, an original color image is input into a monocular depth estimation network. First, the encoder uses convolution and Transformer to output global-local feature interaction features of four different scales in parallel. Then, the multi-scale features extracted by the encoder and the pyramid pooling module (PPM) are fed into the decoder as inputs, and spatial resolution output is achieved through cross-layer and cross-scale connections. Disparity is used as the scene depth representation. The depth probability representation of the decoder is sent into the bins classification network, and lightweight Transformer is used to combine and output groups of interval bins, which are then multiplied by the Hadamard product of the depth representation to obtain the classified depth value. According to the depth probability of the imaging plane and the Gaussian distribution of the multi-plane classified depth output by the decoder, the final depth value is obtained through a mixture of Gaussian depth probability estimation.

[0048] Referring to Figure 2 As shown, it is a schematic diagram of generating a depth map by the method of the embodiment of the present invention on the KITTI dataset. The depth distribution in the horizontal direction of the imaging plane is represented as a Gaussian distribution diagram, where the jump of the boundary value is the highest point of the Gaussian distribution, which is expressed as:

[0049]

[0050] where μ and σ represent the mean and variance of the depth jump. The depth distribution mapping in the optical axis direction can discretize the image depth and more accurately predict the depth information.

[0051] Referring to Figure 3 As shown, each global and local feature fusion and encoding module contains a convolution and three Transformer modules, and the attention is as shown in the formula:

[0052]

[0053]

[0054] Att j represents the global feature output by the j-th Transformer, Conv i represents the local feature extracted by the convolutional network, and X i+1 represents the output fusion feature.

[0055] The depth decoder takes the multi-scale features extracted by the encoder and the pyramid pooling module (PPM) as inputs, and achieves spatial resolution output through cross-layer and cross-scale connections. Disparity is used as the scene depth representation. The pixel-level depth prediction of the non-shifted Swin-Transformer block calculates the query and key vectors of each patch in the window using the feature map output by the encoder. The query and key vectors are the same. Then, the dot product between the two calculates the scores between any two nodes. The relative position embedding P is introduced and undergoes a dot product operation with the output F of the previous decoder layer and the SoftMax output to obtain the output of the current decoder.

[0056] Refer to Figure 4 As shown, the decoder converts the volume plane into a pixel-level probability map, outputs N embeddings as the image mapping map F, interacts with N bins queries, and outputs the embedding Q. Then, Q is fed into the linear perceptron SoftMax to generate N Finally, the linear combination of the probability distribution of the pixel points and the bin center c is calculated to obtain the final depth value d.

[0057] Refer to Figure 5 As shown, it is a schematic diagram of the method of the embodiment of the present invention for generating images and error comparisons on the KITTI dataset. RMSE represents the root mean square error, and the method of the embodiment of the present invention is superior to other methods.

[0058] The number of devices and the processing scale described here are used to simplify the description of the present invention. The applications, modifications, and variations of the present invention are obvious to those skilled in the art.

[0059] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and the illustrated and described examples here.

Claims

1. A monocular depth estimation method for intelligent transportation based on orthogonal planes, characterized in that: The following steps are involved: S1, obtain an RGB image with a resolution of H×W and an image dataset with the same resolution as its next frame; S2, by building global and local feature extraction modules in the encoder, obtain encoding features of four resolutions; S3, the encoder output features are sent to the non-shift window Transformer for pixel depth generation to obtain three decoding features of different resolutions; S4, by sending the decoder depth probability representation into the bins classification network, combining the output bins interval group through the lightweight Transformer, and performing Hadamard product on the bins interval group and the depth representation to obtain the classification depth value; S5. Determine the final depth value by using mixed Gaussian estimation to estimate the imaging plane depth probability output by the decoder and the optical axis multi-plane depth probability output by the multi-plane bins.

2. The method for intelligent traffic monocular depth estimation based on orthogonal planes according to claim 1, characterized in that: In step S2, a convolutional network is used to extract color images to generate H / 2×W / 2 small features. After that, each layer of global and local feature extraction uses convolution and Transformer in parallel, and the scale resolution is reduced by 2 times. The extracted features are sent to the pyramid parsing module to extract multi-scale features.

3. The method for intelligent traffic monocular depth estimation based on orthogonal planes according to claim 1, characterized in that: In step S3, the spatial resolution output disparity is realized through cross-layer and cross-scale connections as the scene depth representation, and the pixel-level depth prediction is performed through the non-shifted Swin-Transformer block. The query and key vectors of each patch in the window are calculated using the feature map output by the encoder. The query and key vectors are the same, and then the dot product between the two is calculated to obtain the similarity score. The relative position embedding P is introduced, and the dot product operation is performed between the output F of the previous layer decoder and the SoftMax output to obtain the output of the current decoder, and channel conversion is performed in the decoder. At the same time, the moving window depth prediction is removed, reducing the artifacts of target edge depth truncation.

4. The method for intelligent traffic monocular depth estimation based on orthogonal planes according to claim 1, characterized in that: In step S4, the depth interval center bins are predicted by a simple post-processing transformer module: where d min and dmax represent the minimum and maximum depths respectively, b i represents the i-th interval distance.

5. The method for intelligent traffic monocular depth estimation based on orthogonal planes according to claim 4, characterized in that: In step S4, the continuous depth is estimated by the plane discretization method. The decoder converts the volume plane into a pixel-level probability map, generates N embeddings as image feature maps, interacts with N binsqueries, outputs depth embeddings, and feeds the embeddings into a linear perceptron and softmax to obtain the normalized depth.

6. The method for intelligent traffic monocular depth estimation based on orthogonal planes according to claim 1, characterized in that: The pixels (u, v) of the input image are parameterized as Gaussian distributions, the sampling selection of bins is precisely located, and the training depth search space is [μ u,v -βσ u,v ,μ u,v +βσ u,v ], the kth depth is assumed to be d u,v,k : d u,v,k =μ u,v +b k s u,v (2), where μ u,v and σ u,v Respectively represent the depth mean and variance of pixels (u, v), b k represents the probability function.

7. The method for intelligent traffic monocular depth estimation based on orthogonal planes according to claim 1, characterized in that: In step S5, the two Gaussian models of imaging plane regression depth and optical axis multi-plane classification depth make probability predictions of the target depth from two directions respectively, and their means and variances are μ1, μ2, σ1, σ2 respectively; the probability density function of the mixed Gaussian distribution is: where μ depth As the maximum probability depth estimate, x i and μ i The depth mean and variance of the i-th depth estimate, respectively, represents the variance of a mixture of Gaussians.

Citation Information

Patent Citations

  • Monocular depth estimation method, apparatus, terminal, and storage medium

    CN109087349A

  • Indoor scene monocular image depth estimation method based on deep learning

    CN114638870A

Cited By

  • Monocular self-supervision depth estimation method fusing multi-resolution features and global context

    CN120976282A