Binocular scene depth estimation method and system based on multistage pyramid parallax optimization
Through the binocular scene depth estimation method with multi-level pyramid parallax optimization, the feature extraction and disparity calculation are used to use a two-dimensional convolutional network to solve the problem of high computing resource consumption in the prior art, and efficient binocular depth estimation in low energy consumption and low hardware cost environments are achieved.
Patent Information
- Application Number
- CN202510411473.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-08
AI Technical Summary
The existing binocular scene depth estimation method based on deep learning has high computing resource consumption and is difficult to effectively apply in practical applications with limited computing resources.
The multi-level pyramid disparity optimization method is adopted, and a two-dimensional convolution network is used to perform feature extraction and disparity calculation. The multi-level pyramid structure is fused with multi-scale feature information, and the three-dimensional convolution module is used to perform disparity calculation and optimization.
It significantly reduces computing resource consumption and running time, improves the application performance of the algorithm in low-energy and low hardware cost environments, and has strong generalization capabilities.
Smart Images

Figure CN120279076A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of binocular scene depth estimation, and specifically, to a binocular scene depth estimation method and system based on multi-level pyramid disparity optimization. Background Art
[0002] With the rapid development of related technologies in the field of artificial intelligence, scene depth estimation, as one of the key technologies for multiple computer vision tasks such as scene understanding, 3D scene reconstruction, and autonomous driving scene perception, has been widely studied and applied.
[0003] Scene depth estimation can be divided into three mainstream methods according to the different sensors used: time-of-flight algorithm, structured light, and image-based depth estimation method. Among them, the time-of-flight algorithm has relatively high requirements for the cost of hardware, can handle relatively low resolutions, and currently has certain limitations in applications outside industrial production; the structured light algorithm is greatly restricted by the usage environment, is mainly applied indoors, and is difficult to adapt to outdoor environments with strong light and diverse scene changes; compared with the first two methods, the image-based scene depth estimation method has low cost and strong environmental adaptability, and has unique advantages in applications in various actual environments. The image-based scene depth estimation method is mainly divided into monocular scene depth estimation method and binocular scene depth estimation method. At present, the binocular scene depth estimation method has higher accuracy and higher generalization ability than the monocular method.
[0004] The traditional methods of binocular depth estimation are mainly divided into four steps: matching cost calculation, matching cost aggregation, disparity calculation, and disparity optimization. Matching cost calculation is a process of calculating the similarity cost between all pixel points within a certain offset in the right image and a left image pixel point with the left image pixel point as the reference; matching cost aggregation is a process of considering the matching cost situation of adjacent pixel points in the neighborhood of a pixel point in the left image, and aggregating to obtain a better matching cost matrix based on constraints such as depth continuity in the real scene; disparity calculation is a process of calculating the disparity of each pixel point in the left image based on the result of matching cost aggregation; disparity optimization is a process of denoising, smoothing, etc. according to the global effect of the obtained disparity map to make it closer to the actual scene depth distribution. According to the different calculation regions in matching cost calculation and aggregation, traditional depth estimation algorithms can be divided into global stereo matching and semi-global stereo matching algorithms.
[0005] With the rapid development of deep learning algorithms, the binocular scene depth estimation method based on deep learning has replaced the traditional method in terms of performance leadership. The binocular scene depth estimation method based on deep learning has formed a transition from the four steps of the traditional method to the method of feature extraction by 2D convolutional network, disparity feature aggregation, 3D convolutional network disparity calculation and optimization. However, due to the large amount of computation and numerous parameters of the 3D convolutional network itself, most of the current binocular scene depth estimation methods based on deep learning consume a large amount of computing resources and take a long time, making it difficult to be applied in practical applications with limited computing resources such as embedded systems.
[0006] Xu H, Zhang J. AANet: Adaptive Aggregation Network for Efficient Stereo Matching[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2020:1959 - 1968. discloses the entry point based on the fact that the most time-consuming module in the current binocular depth estimation algorithm is the 3D convolutional network module. It is considered that the 3D convolutional disparity calculation module learns the feature information connection of different pixel points in the feature map in the disparity dimension during the network training process, and this method can be completed based on deformable convolution. Therefore, a 2D convolutional module for feature extraction is built based on deformable convolutional kernels, enabling the network to focus more attention on the pixel points that belong to the same object as the current pixel point, so that the network can fully learn the disparity smoothness information based on the object distribution characteristics in real-world data and output a continuous and stable disparity map without relying on a 3D convolutional network. In addition, this method designs a disparity calculation module based on multi-scale 2D convolution to fuse the context information from disparity maps at different scales. However, the difference between this method and the method of the present invention is that this method includes 3 disparity calculation modules composed of feature maps with different resolutions, and the outputs at different levels are restored to the same resolution through upsampling or downsampling and cascaded as the input of the second-layer disparity calculation, and finally a multi-scale estimated disparity map is obtained to establish a loss function with the ground truth. The multi-scale feature information in its network structure is in a parallel relationship, while the multi-scale feature information in the present invention is in a serial relationship in the network structure, and the feature information at different scales is fused through a bridging structure, gradually restoring the rough disparity map at low resolution to obtain the disparity map at high resolution, and establishing a multi-scale ground truth loss function with the ground truth.
[0007] Patent document CN116485864A (application number: 202310397767.1) discloses a method and device for binocular depth estimation based on reparameterization. By providing a feature extractor based on a reparameterization module, a cost volume constructed in stages, a cost aggregation network with two-dimensional convolution, and a TensorRT optimized model, a high-precision disparity estimation result can be obtained on edge devices.
[0008] Therefore, most of the existing binocular scene depth estimation methods based on deep learning generate a four-dimensional matching cost volume and use a three-dimensional convolutional network to regress to obtain a dense disparity map. The network structure is complex, the calculation is time-consuming, the memory resources are occupied highly, and it is difficult to meet the requirements of low energy consumption and low hardware cost in practical applications. To solve this problem, in the present invention, an initial disparity map is obtained by fusing the features of the left and right images, and a dense disparity map is gradually regressed through a two-dimensional convolutional network. The network structure is simple, the calculation is less time-consuming, the memory resources are occupied lowly, and it better meets the usage requirements of practical applications. Summary of the Invention
[0009] Aiming at the defects in the prior art, the purpose of the present invention is to provide a binocular scene depth estimation method and system based on multi-level pyramid disparity optimization.
[0010] A binocular scene depth estimation method based on multi-level pyramid disparity optimization provided by the present invention includes:
[0011] Step S1: Input the left image and the right image into a feature encoding network respectively to obtain a left initial multi-level pyramid feature map and a right initial multi-level pyramid feature map;
[0012] Step S2: Input the left initial multi-level pyramid feature map and the right initial multi-level pyramid feature map into a feature decoding network respectively to obtain a left multi-level pyramid feature map and a right multi-level pyramid feature map;
[0013] Step S3: Obtain an initial disparity calculation result through a disparity calculation module based on the left multi-level pyramid feature map and the right multi-level pyramid feature map;
[0014] Step S4: The initial disparity calculation result passes through an N-level disparity optimization unit to obtain a disparity map;
[0015] Step S5: Perform binocular scene depth estimation based on the disparity map.
[0016] Preferably, the step S1 includes: inputting the left image matrix or the right image matrix into the feature encoding network respectively, and the feature encoding network gradually extracts features in multiple convolutional layers and gradually reduces the image resolution in multiple average pooling layers to obtain an initial multi-level pyramid feature map.
[0017] Preferably, inputting the left image matrix or the right image matrix into the feature encoding network, and gradually extracting features by the feature encoding network in multiple two-dimensional convolutional networks and gradually reducing the image resolution by multiple average pooling layers to obtain an initial multi-level pyramid feature map, including:
[0018] Denote the left image or the right image as a matrix data format of (3, H, W); where H and W are the number of pixels of the image height and width respectively, and 3 is the number of channels of the color RGB image;
[0019] After passing the left image matrix through a convolutional layer, a regularization layer, and a rectified linear unit, repeating the convolutional layer - regularization layer - rectified linear unit m times, obtain the output of the current two-dimensional convolutional network, and then reduce the resolution of the feature map through an average pooling layer as the input of the next-level two-dimensional convolutional network;
[0020] There are N groups of the two-dimensional convolutional network and the average pooling layer, and N is the number of levels of the multi-level pyramid disparity optimization;
[0021] The set composed of the feature maps output by each average pooling layer is used as the output of the feature encoding network.
[0022] Preferably, the step S2 includes: inputting the initial multi-level pyramid feature map into the feature decoding network, and fusing the multiple resolution features in the initial multi-level pyramid feature map by the feature decoding network in multiple convolutional layers to obtain a multi-level pyramid feature map.
[0023] Preferably, the feature encoding network is alternately composed of a two-dimensional convolutional network and an upsampling layer using bilinear interpolation; after passing the initial multi-level pyramid feature map through a depth convolutional layer, a regularization layer, and a rectified linear unit, repeating the structure of convolutional - regularization layer - rectified linear unit m times, obtain the output of the current two-dimensional convolutional network unit, and gradually restore the resolution of the initial multi-level pyramid feature map through the upsampling layer as the input of the next-level two-dimensional convolutional network; starting from the next two-dimensional convolutional network unit, merge the output of the previous upsampling layer and the feature map output by the feature encoding network with the same resolution in the feature channel dimension through a bridging structure to achieve feature fusion as the input of the current two-dimensional convolutional network unit; finally, obtain the set composed of the feature maps output by each two-dimensional convolutional network unit as the output of the feature decoding network, that is, the multi-level pyramid feature map;
[0024] The feature decoding network has N two-dimensional convolutional networks and N - 1 upsampling levels, and there is no upsampling layer connected after the last two-dimensional convolutional network.
[0025] Preferably, the step S3 includes: selecting the feature maps with the lowest resolution in the left multi-level pyramid feature map and the right multi-level pyramid feature map to perform disparity calculation to obtain an initial disparity calculation result;
[0026] Record the maximum disparity search range of the left and right images as D. Translate the right feature map 0 to D - 1 units respectively along the epipolar line direction after binocular image rectification to obtain a total of D right feature map matrix data. Then subtract each of them from the left feature map with the same resolution to obtain a total of D disparity feature matrix data. Subsequently, perform normalization and weighted sum calculation on the disparity feature matrix data in the D disparity dimension to achieve dimensionality reduction in the disparity dimension and obtain the initial disparity calculation result.
[0027] Preferably, each level of the N-level disparity optimization unit includes: an upsampling layer, two two-dimensional convolutional neural networks with a preset convolutional kernel for adjusting the number of channels, and two two-dimensional convolutional neural networks with a preset convolutional kernel for not adjusting the number of channels;
[0028] First, upsample the initial disparity calculation result through the upsampling layer and merge it with the feature map with the same output resolution as the left and right feature decoding networks on the feature channels. Subsequently, obtain the disparity feature map through two two-dimensional convolutional neural networks with a preset convolutional kernel for adjusting the number of channels, and then obtain the disparity-optimized feature map through two two-dimensional convolutional neural networks with a preset convolutional kernel for not adjusting the number of channels, which is used as the input for the next-level disparity optimization unit;
[0029] And so on, obtain the input for the final N-level disparity optimization. After upsampling through the upsampling layer, then pass through two two-dimensional convolutional neural networks with a preset convolutional kernel for adjusting the number of channels and two two-dimensional convolutional neural networks with a preset convolutional kernel for not adjusting the number of channels to obtain the disparity-optimized feature. Finally, calculate through the normalization exponential function and weighted average to obtain the disparity map optimization result as the disparity map output.
[0030] Preferably, perform multi-level downsampling on the disparity map ground truth and calculate the average pixel error with the disparity map optimization result with the same resolution as the loss function:
[0031]
[0032] where N is the number of levels of the current multi-level pyramid optimization unit, P is the pixel set of the disparity map at the current resolution, is the disparity estimation value of any pixel point, d gt is the disparity ground truth of the same pixel point. The final total loss function is the weighted sum of the loss functions of all levels of the multi-level pyramid disparity optimization units:
[0033] L = ∑ N k N L N
[0034] where k is the weight hyperparameter set for different levels of pyramid disparity optimization.
[0035] Preferably, the step S5 includes: for the disparity value d of any pixel point in the disparity map, calculating the depth through the formula z = Bf / d;
[0036] where B is the baseline length of the binocular system for capturing the left and right images, f is the focal length of the left and right cameras of the binocular system, and z is the depth of the pixel point.
[0037] A binocular scene depth estimation system based on multi-level pyramid disparity optimization provided by the present invention includes:
[0038] Module M1: Inputting the left image and the right image into the feature encoding network respectively to obtain a left initial multi-level pyramid feature map and a right initial multi-level pyramid feature map;
[0039] Module M2: Inputting the left initial multi-level pyramid feature map and the right initial multi-level pyramid feature map into the feature decoding network respectively to obtain a left multi-level pyramid feature map and a right multi-level pyramid feature map;
[0040] Module M3: Obtaining an initial disparity calculation result through a disparity calculation module based on the left multi-level pyramid feature map and the right multi-level pyramid feature map;
[0041] Module M4: Obtaining a disparity map from the initial disparity calculation result through an N-level disparity optimization unit;
[0042] Module M5: Performing binocular scene depth estimation based on the disparity map.
[0043] Compared with the prior art, the present invention has the following beneficial effects:
[0044] 1. The present invention reconstructs a dense disparity map according to the original disparity hypothesis at low resolution and multi-scale feature fusion at high resolution, effectively utilizes the multi-scale feature information extracted by the feature network, so as to maintain good performance when facing image inputs of different resolutions and actual image inputs with high depth estimation difficulty, and has strong generalization performance in actual scenes containing diverse environments;
[0045] 2. The present invention is a pyramid multi-level disparity feature fusion network based on a two-dimensional convolutional network, which replaces the three-dimensional convolutional module in most deep learning-based binocular depth estimation methods for disparity calculation and optimization, significantly speeds up the running speed of the algorithm and reduces the consumption of computing resources; it has obvious advantages in practical applications with low energy consumption and low hardware costs, such as embedded algorithm deployment. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present invention will become more apparent:
[0047] Figure 1Schematic diagram of a binocular scene depth estimation system based on multi-level pyramid disparity optimization.
[0048] Figure 2 Schematic diagram of a feature encoding network.
[0049] Figure 3 Schematic diagram of a feature decoding network. Specific implementation manners
[0050] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.
[0051] Embodiment 1
[0052] According to a binocular scene depth estimation method and system based on multi-level pyramid disparity optimization provided by the present invention, as Figures 1 to 3 shown, it includes: using a two-dimensional convolutional neural network to replace the mainstream three-dimensional convolutional neural network for disparity calculation and optimization, and using a multi-level pyramid structure for multi-level disparity optimization to make full use of the multi-scale feature information extracted by the two-dimensional convolutional neural network.
[0053] Among them, the two-dimensional convolutional network has the advantage of fully extracting the multi-scale feature information of the image in image feature extraction and visual task calculation, and has a smaller calculation amount compared with higher-dimensional networks such as three-dimensional convolutional networks. Therefore, the sub-networks of the algorithm of the present invention are all based on the two-dimensional convolutional neural network as the main body of the algorithm.
[0054] The image input of the present invention is a binocular image synchronously captured by a general binocular sensor, or an existing binocular image file that has been binocular corrected, denoted as the left image and the right image.
[0055] The binocular scene depth estimation method based on multi-level pyramid disparity optimization includes:
[0056] Step 1: The left and right images are input into the feature encoding network to obtain an initial multi-level pyramid feature map; the initial multi-level pyramid feature map is obtained by the feature encoding network gradually extracting features in multiple convolutional layers and gradually reducing the map resolution in multiple average pooling layers;
[0057] Step 2: The initial multi-level pyramid feature map is input into the feature decoding network to obtain a multi-level pyramid feature map; the multi-level pyramid feature map is obtained by the multiple convolutional layers of the feature decoding network fusing the features at multiple resolutions in the initial multi-level pyramid feature map;
[0058] Step 3: The multi-level pyramid feature map is input into the disparity calculation module to obtain an initial disparity calculation result, which is a relatively rough disparity calculation result obtained by calculating with fewer convolutional layers.
[0059] Step 4: The initial disparity calculation result is optimized through N-level disparity optimization to obtain a disparity map.
[0060] Step 5: The disparity map is a grayscale map with a data dimension of (x, y, 1). Here, x and y are the pixel length and width values of the original left and right images, and the dimension of 1 is the disparity value. For the disparity value d of any pixel point, the depth z can be calculated through the formula z = Bf / d, where B is the baseline length of the binocular system for taking the left and right images, f is the focal length of the left and right cameras of the binocular system, both are fixed parameters of the binocular system, and z is the depth of this pixel point. Therefore, the depth map can be obtained by calculating each pixel point of the disparity map through the above formula, that is, the binocular scene depth estimation task is completed.
[0061] Specifically, Step 1 includes:
[0062] Step 1: Input the left and right images into the feature encoding network respectively. The overall structure of the feature encoding network is as Figure 2 shown. Denote the input of the left image or the right image as a matrix data format of (3, H, W), where H and W are the number of pixels of the image height and width respectively, and 3 is the number of channels of the color RGB image.
[0063] The feature encoding network structure is composed of alternating two-dimensional convolutional networks and average pooling layers. The convolutional kernel of each two-dimensional convolutional network is a 3*3 depth convolution. The output feature map passes through a regularization layer and a rectified linear unit. After repeating the structure of convolution - regularization layer - rectified linear unit 3 times, the output of the current two-dimensional convolutional network is obtained, and the resolution of the feature map is reduced through the average pooling layer as the input of the next-level two-dimensional convolutional network.
[0064] There are N groups of two-dimensional convolutional networks and average pooling layers. N is the number of levels of the multi-level pyramid disparity optimization in this embodiment, which can be adjusted according to actual application requirements such as network depth, hardware computing resources, and operation time requirements. Generally, the larger N is, the higher the accuracy of the finally output disparity map, but the longer the overall operation time and the greater the consumption of computing resources of the method.
[0065] The matrix data format of the image input (3, H, W) becomes a feature map matrix of dimension (32, H, W) after passing through the first two-dimensional convolutional network. After average pooling, it becomes a feature map matrix of dimension (32, H / 2, W / 2). Thereafter, the output feature map of each two-dimensional convolutional network is twice the input in the first-dimensional feature channel, and the output feature map of each average pooling layer is 1 / 2 of the input in the second and third-dimensional feature map sizes. Finally, a set composed of the feature map outputs of each average pooling layer is obtained as the feature encoding output:
[0066]
[0067] Specifically, step 2 includes:
[0068] Taking the initial multi-level pyramid feature map obtained in step 1 as the input of the feature decoding network, and the overall structure of the feature decoding network is as Figure 3 shown. The feature encoding network structure is alternately composed of a two-dimensional convolutional network and an upsampling layer using bilinear interpolation. The convolutional kernel of each two-dimensional convolutional network is a 3*3 depth convolution. The output feature map passes through a regularization layer and a rectified linear unit. After repeating the structure of convolution-regularization layer-rectified linear unit 3 times, the output of the current two-dimensional convolutional network unit is obtained. The resolution of the feature map is restored layer by layer through the upsampling layer and used as the input of the next-level two-dimensional convolutional network.
[0069] The feature decoding network has N two-dimensional convolutional network units and N - 1 upsampling levels, and there is no upsampling layer connected only after the last two-dimensional convolutional network unit.
[0070] Taking the feature map with the lowest resolution output by the feature encoding network (32×2^(N - 1), H / 2^N, W / 2^N) as the initial input, a feature map matrix of (32×2^(N - 2), H / 2^N, W / 2^N) is obtained after passing through the first two-dimensional convolutional network unit, and a feature map matrix of (32×2^(N - 2), H / 2^(N - 1), W / 2^(N - 1)) is obtained through the upsampling layer. Starting from the second two-dimensional convolutional network unit, through a bridging structure, the output of the previous upsampling layer and the feature map output by the feature encoding module with the same resolution are merged in the feature channel dimension to achieve feature fusion, which is used as the input of the current two-dimensional convolutional network unit, that is:
[0071]
[0072] The operations of the remaining feature decoding networks are the same. Finally, a set composed of the feature map outputs of each two-dimensional convolutional network unit is obtained as the feature decoding output, denoted as the multi-level pyramid feature:
[0073]
[0074] In this embodiment, according to actual usage requirements, it is replaced with a feature encoding and decoding structure with a lower or higher number of layers, and a feature extraction module that can extract multi-resolution feature maps, so as to achieve the effect of adjusting the number of network parameters and the amount of computation.
[0075] Specifically, step 3 includes: performing disparity calculation on the feature map with the lowest resolution among the multi-level pyramid features extracted from the left and right images through steps 1 and 2 to obtain an initial disparity map. The specific calculation method is as follows:
[0076] Denote the maximum disparity search range of the left and right images as D, which is jointly determined by the baseline of the sensor for collecting binocular images and the image resolution; translate the right feature map along the epipolar line direction after rectifying the binocular images, that is, in the W dimension of the image features, by 0 to D - 1 units respectively, to obtain a total of D matrix data of right feature maps with dimensions of (32×2^(N - 1), H / 2^N, W / 2^N). Then subtract each of them from the left feature map with the same resolution to obtain a total of D matrix data of disparity feature maps with dimensions of (32×2^(N - 1), H / 2^N, W / 2^N), denoted as (D, 32×2^(N - 1), H / 2^N, W / 2^N). Subsequently, perform normalization and weighted sum calculation on the four-dimensional disparity feature matrix data in the D disparity dimension to realize the reduction of the disparity dimension, and obtain the initial disparity calculation result with dimensions of (32×2^(N - 1), H / 2^N, W / 2^N).
[0077] Specifically, step 4 includes:
[0078] Each level of disparity optimization unit consists of an upsampling layer + two two-dimensional convolutional neural networks with 3*3 convolutional kernels for adjusting the number of channels + two two-dimensional convolutional neural networks with 3*3 convolutional kernels for not adjusting the number of channels.
[0079] First, perform upsampling on the initial disparity calculation result with dimensions of (32×2^(N - 1), H / 2^N, W / 2^N) to obtain (32×2^(N - 1), H / 2^(N - 1), W / 2^(N - 1)), and merge it with the feature map with the same resolution as the output of the left and right feature decoding on the feature channels to obtain the input of the first-level disparity optimization unit, that is:
[0080]
[0081] Subsequently, the number of channels is adjusted through two - dimensional convolutional neural networks with 3*3 convolutional kernels. The number of channels of the input feature dimensions is calculated and adjusted to 1 / 2 and 1 / 4 of the input respectively. For example, if the input of the current - level disparity optimization unit is (32×2^N, H / 2^(N - 1), W / 2^(N - 1)), after adjusting the number of channels through two - dimensional convolutional neural networks with 3*3 convolutional kernels, a disparity feature map of (32×2^(N - 2), H / 2^(N - 1), W / 2^(N - 1)) is obtained. Then, through two - dimensional convolutional neural networks with 3*3 convolutional kernels that do not adjust the number of channels, a disparity - optimized feature map of (32×2^(N - 2), H / 2^(N - 1), W / 2^(N - 1)) is obtained, which is used as the input of the next - level disparity optimization unit.
[0082] And so on, the input of the last N - level disparity optimization is (32, H / 2, W / 2). After upsampling, an input of (32, H, W) dimension is obtained. Specifically, since there is no left - right feature decoding output of the same size, no additional feature fusion is performed. Subsequently, through two - dimensional convolutional neural networks with 3*3 convolutional kernels to adjust the number of channels and two - dimensional convolutional neural networks with 3*3 convolutional kernels that do not adjust the number of channels, a disparity - optimized feature of (8, H, W) is obtained. Finally, on the 8 - feature dimension, through the normalized exponential function (Softmax) calculation and weighted average, a disparity - optimized result of (1, H, W) is obtained, which is used as the disparity map output. According to the internal and external calibration parameters of the binocular vision system, the depth map can be calculated according to the formula (depth = focal length×baseline length / disparity value).
[0083] During training on a public dataset with disparity map ground truth such as KITTI, the output of each level of the disparity optimization unit is adjusted through a two - dimensional convolutional neural network to calculate the number of channels and through the normalized exponential function (Softmax), and then weighted average to obtain the disparity map optimization result at the current resolution. The disparity map ground truth is downsampled at multiple levels, and the average pixel error is calculated with the disparity map optimization result of the same resolution as the loss function:
[0084]
[0085] where N is the number of levels of the current multi - level pyramid optimization unit, P is the pixel set of the disparity map at the current resolution, is the disparity estimate value of any pixel point, d gt is the disparity ground truth of the same pixel point. The final total loss function is the weighted sum of the loss functions of all levels of the multi - level pyramid disparity optimization units:
[0086] L = ∑ N k N L N
[0087] Where k is a weight hyperparameter for optimizing the parallax of different levels of the pyramid.
[0088] In this embodiment, the algorithm network completes pre-training and overall training on the Scene Flow dataset and the KITTI dataset, and is tested on the KITTI dataset. During the training process, the number of levels N for optimizing the pyramid parallax is set to 4, and the weight hyperparameters of the corresponding loss function are 0.1, 0.3, 0.6, and 1.0 from level 1 to level 4 respectively. The parallax search range is 192 pixels. The incorrect rate of the parallax optimization result of the training network (pixels with an error between the estimated value and the true value exceeding 3 pixels are regarded as incorrect) is less than 5%, which can quickly and accurately meet the binocular depth estimation requirements in the actual scene.
[0089] In this embodiment, for the multi-level pyramid parallax optimization, according to actual usage needs, based on the existing multi-level calculation structure, bridging structures and residual structures and other side branches can be added between the parallax optimizations at all levels to achieve the multi-scale feature fusion effect of enhancing the algorithm performance without significantly increasing more computational consumption.
[0090] The present invention also provides a binocular scene depth estimation system based on multi-level pyramid parallax optimization. The binocular scene depth estimation system based on multi-level pyramid parallax optimization can be implemented by executing the process steps of the binocular scene depth estimation method based on multi-level pyramid parallax optimization. That is, those skilled in the art can understand the binocular scene depth estimation method based on multi-level pyramid parallax optimization as the preferred implementation manner of the binocular scene depth estimation system based on multi-level pyramid parallax optimization.
[0091] Those skilled in the art know that in addition to implementing the systems, devices, and their respective modules provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the systems, devices, and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, the systems, devices, and their respective modules provided by the present invention can be regarded as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structure within the hardware component; the modules for implementing various functions can also be regarded as either software programs for implementing the method or the structure within the hardware component.
[0092] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A binocular scene depth estimation method based on multi-level pyramid parallax optimization, characterized in that Including: Step S1: Input the left image and the right image into the feature encoding network respectively to obtain the left initial multi-level pyramid feature map and the right initial multi-level pyramid feature map; Step S2: Input the left initial multi-level pyramid feature map and the right initial multi-level pyramid feature map into the feature decoding network respectively to obtain the left multi-level pyramid feature map and the right multi-level pyramid feature map; Step S3: Obtain the initial disparity calculation result based on the left multi-level pyramid feature map and the right multi-level pyramid feature map through the disparity calculation module; Step S4: The initial disparity calculation result passes through the N-level disparity optimization unit to obtain the disparity map; Step S5: Perform binocular scene depth estimation based on the disparity map.
2. The binocular scene depth estimation method based on multi-level pyramid parallax optimization according to claim 1, characterized in that, The said Step S1 includes: Input the left image matrix or the right image matrix into the feature encoding network respectively, and the feature encoding network gradually extracts features in multiple convolutional layers and gradually reduces the map resolution in multiple average pooling layers to obtain the initial multi-level pyramid feature map.
3. The binocular scene depth estimation method based on multi-level pyramid parallax optimization according to claim 2, wherein The input of the left image matrix or the right image matrix into the feature encoding network, and the feature encoding network gradually extracts features in multiple two-dimensional convolutional networks and gradually reduces the map resolution in multiple average pooling layers to obtain the initial multi-level pyramid feature map, includes: Denote the left image or the right image as a matrix data format of (3, H, W); where, H and W are the number of pixels of the image height and width respectively, and 3 is the number of channels of the color RGB image; The left image matrix passes through the convolutional layer, the regularization layer and the rectified linear unit, and after repeating the convolutional layer - regularization layer - rectified linear unit structure m times, the output of the current two-dimensional convolutional network is obtained, and then the resolution of the feature map is reduced through the average pooling layer as the input of the next-level two-dimensional convolutional network; There are N groups of the said two-dimensional convolutional network and average pooling layer, and N is the number of levels of multi-level pyramid disparity optimization; The set composed of the feature maps output by each average pooling layer is used as the output of the feature encoding network.
4. The binocular scene depth estimation method based on multi-level pyramid parallax optimization according to claim 1, characterized in that, The said Step S2 includes: Input the initial multi-level pyramid feature map into the feature decoding network, and the feature decoding network fuses the multiple resolution features in the initial multi-level pyramid feature map in multiple convolutional layers to obtain the multi-level pyramid feature map.
5. The binocular scene depth estimation method based on multi-level pyramid parallax optimization according to claim 4, characterized in that, The feature encoding network is alternately composed of a two-dimensional convolutional network and an upsampling layer using bilinear interpolation; the initial multi-level pyramid feature map passes through the depth convolutional layer, the regularization layer and the rectified linear unit, and after repeating the convolutional - regularization layer - rectified linear unit structure m times, the output of the current two-dimensional convolutional network unit is obtained, and the resolution of the initial multi-level pyramid feature map is restored layer by layer through the upsampling layer as the input of the next-level two-dimensional convolutional network; starting from the next two-dimensional convolutional network unit, the output of the previous upsampling layer and the feature map output by the feature encoding network with the same resolution are merged in the feature channel dimension through the bridging structure to achieve feature fusion as the input of the current two-dimensional convolutional network unit; finally, the set composed of the feature maps output by each layer of two-dimensional convolutional network units is used as the output of the feature decoding network, that is, the multi-level pyramid feature map; The said feature decoding network has N two-dimensional convolutional networks and N - 1 upsampling levels, and there is no upsampling layer connected after the last two-dimensional convolutional network.
6. The binocular scene depth estimation method based on multi-level pyramid parallax optimization according to claim 1, characterized in that Step S3 includes: selecting the feature map with the lowest resolution among the left multi-level pyramid feature map and the right multi-level pyramid feature map to perform disparity calculation to obtain an initial disparity calculation result; Denote the maximum disparity search range of the left and right images as D. Translate the right feature map by 0 to D-1 units respectively along the epipolar line direction after binocular image rectification to obtain a total of D right feature map matrix data. Then subtract them from the left feature map with the same resolution respectively to obtain a total of D disparity feature matrix data. Subsequently, perform normalization and weighted sum calculation on the disparity feature matrix data in the D disparity dimension to realize the reduction of the disparity dimension and obtain the initial disparity calculation result.
7. The binocular scene depth estimation method based on multi-level pyramid parallax optimization according to claim 1, characterized in that Each level of the N-level disparity optimization unit includes: an upsampling layer, two two-dimensional convolutional neural networks with preset convolutional kernels for adjusting the number of channels, and two two-dimensional convolutional neural networks with preset convolutional kernels for not adjusting the number of channels; First, upsample the initial disparity calculation result through the upsampling layer and merge it with the feature map with the same output resolution of the left and right feature decoding networks on the feature channels. Then obtain the disparity feature map through two two-dimensional convolutional neural networks with preset convolutional kernels for adjusting the number of channels, and then obtain the disparity optimized feature map through two two-dimensional convolutional neural networks with preset convolutional kernels for not adjusting the number of channels, which is used as the input of the next-level disparity optimization unit; And so on, obtain the input of the final N-level disparity optimization, perform upsampling through the upsampling layer, then obtain the disparity optimized feature through two two-dimensional convolutional neural networks with preset convolutional kernels for adjusting the number of channels and two two-dimensional convolutional neural networks with preset convolutional kernels for not adjusting the number of channels. Finally, calculate the disparity map optimization result through normalization exponential function calculation and weighted average as the disparity map output.
8. The binocular scene depth estimation method based on multi-level pyramid parallax optimization according to claim 7, characterized in that, Perform multi-level downsampling on the disparity map ground truth and calculate the average pixel error with the disparity map optimization result of the same resolution as the loss function: where N is the number of levels of the current multi-level pyramid optimization unit, and P is the set of pixels of the disparity map at the current resolution, is the estimated disparity value of any pixel point, d gt is the true disparity value of the same pixel point, and the final total loss function is the weighted sum of the loss functions of the multi-level pyramid disparity optimization units at all levels: L = Σ N k N L N where k is the weight hyperparameter set for the disparity optimization of different levels of pyramids.
9. The binocular scene depth estimation method based on multi-level pyramid parallax optimization according to claim 1, wherein Step S5 includes: for the disparity value d of any pixel point in the disparity map, calculate the depth through the formula z = Bf / d; where B is the baseline length of the binocular system for taking the left and right images, f is the focal length of the left and right cameras of the binocular system, and z is the depth of this pixel point.
10. A binocular scene depth estimation system based on multi-level pyramid parallax optimization, characterized in that, Includes: Module M1: Input the left image and the right image into the feature encoding network respectively to obtain the left initial multi-level pyramid feature map and the right initial multi-level pyramid feature map; Module M2: Input the left initial multi-level pyramid feature map and the right initial multi-level pyramid feature map into the feature decoding network respectively to obtain the left multi-level pyramid feature map and the right multi-level pyramid feature map; Module M3: Based on the left multi-level pyramid feature map and the right multi-level pyramid feature map, obtain the initial disparity calculation result through the disparity calculation module; Module M4: The initial disparity calculation result passes through the N-level disparity optimization unit to obtain the disparity map; Module M5: Perform binocular scene depth estimation based on the disparity map.
Citation Information
Patent Citations
Three-stage binocular depth estimation method and device based on re-parameterization
CN116485864A