Universal deep learning method for binocular parallax estimation
By introducing custom disparity range and Galerkin attention mechanism into the binocular parallax estimation neural network, the problem of insufficient adaptability of binocular parallax estimation neural network in the prior art is solved, and efficient disparity estimation and computational efficiency improvement in complex scenarios is achieved.
Patent Information
- Application Number
- CN202510052711.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-06
AI Technical Summary
Existing binocular parallax estimation neural networks are not adaptable when migrating between different domains, making it difficult to maintain efficient matching performance in complex and changeable scenarios.
A general deep learning method is proposed, by learning the space-to-disparity space mapping of binocular image pairs, customizing the parallax range parameters, supporting multi-scale training, and introducing the Galerkin attention mechanism to optimize computing efficiency and resource occupancy.
It realizes efficient parallax estimation in various scenarios, significantly improving the adaptability and generalization capabilities of the model, and reducing the computational complexity and memory footprint.
Smart Images

Figure CN119941818A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer vision, binocular stereo matching and disparity estimation, and relates to a universal deep learning method for binocular disparity estimation. Background Art
[0002] In the field of computer vision, binocular disparity estimation is a key technology for recovering depth information from planar images. Its core principle is based on the geometric triangulation formula, and the depth value is calculated through the disparity formula d=D / f. This technology simulates the binocular visual perception mechanism of the human eye. It only needs to install two cameras on the same horizontal line and perform stereo correction to achieve efficient depth estimation. This makes binocular disparity estimation widely used in fields such as autonomous driving, robot navigation, and three-dimensional reconstruction. However, the actual application scenarios are complex and changeable, and the image resolution differences and disparity ranges vary significantly, which greatly reduces the performance of the model when migrating between different scenes.
[0003] Traditional binocular disparity estimation methods, such as semi-global matching algorithms, combine the advantages of global and local matching and achieve good matching performance in specific scenes. However, these methods are highly dependent on manually designed similarity measurement formulas and image prior knowledge, and show limitations in complex and changing scenes. With the rapid development of deep learning, deep neural networks are gradually applied to binocular disparity estimation tasks. Zbontar and LeCun first proposed using convolutional neural networks to calculate the matching cost between image blocks. Subsequent researchers further proposed end-to-end training network architectures, such as group correlation networks, to improve the accuracy of similarity measurement. These deep learning-based methods significantly improve matching performance, but are also accompanied by a sharp increase in computational costs. In addition, the current mainstream deep learning models mostly use fixed scales and fixed step sizes for reasoning. When processing high-resolution images, they are limited by GPU memory and are difficult to adapt to complex scenes with different disparity ranges and resolutions. Designing a disparity estimation network with stronger generalization ability and robustness can not only solve the applicability problem of the model in large-scale, high-resolution scenes, but also provide more efficient solutions for practical applications. It is an important direction of current research. Summary of the invention
[0004] The present invention aims to make up for the shortcomings of the prior art, and proposes a universal deep learning method for the adaptability of binocular disparity estimation neural networks when migrating between different domains. The network structure can customize the disparity range parameters according to the characteristics of the data set during the training or testing phase by learning the mapping from the space of the binocular image pair to the disparity space, and supports multi-scale training. At the same time, the network can dynamically construct cost geometry with variable scale and disparity range perception, and introduce the Galerkin attention mechanism with excellent computational efficiency and video memory occupancy, making it a universal model suitable for a variety of scenarios.
[0005] The general deep learning method for binocular disparity estimation designed by the present invention includes the following core stages: data processing stage, multi-level feature extraction stage, cost geometry construction stage, and feature aggregation and disparity prediction stage. The specific operation steps are as follows:
[0006] Step 1) Preprocess the input left and right view color images using a variable scale strategy: crop B pairs of images with a resolution of {s (i) H×s (i) W}, the HR images of different scales are uniformly adjusted to 4H×4W resolution and then input into the network for training, and n points are randomly sampled from the true disparity value of each HR image to ensure that the LR images in each batch of data have a consistent number of supervision points; where H×W is the height and width of the low-resolution image LR obtained by the initial downsampling of the model inference, s (i) From uniform distribution Sample B random scales, where scale_min and scale_max are the custom scale ranges in the network parameters, i∈[1,B].
[0007] Step 2) Input the adjusted resolution 4H×4W image as the input image to the U-Net module to extract shallow features at multiple resolutions: first, perform a preliminary downsampling of the input image through the convolution layer to obtain the LR image, and then continue to downsample to 1 / 32 times the original resolution through convolution and pooling operations; after the downsampling is completed, the transposed convolution is used to gradually upsample, and after each upsampling, the result is fused with the feature map of the corresponding resolution of the downsampling to obtain the shallow features F of the left and right views at each resolution. l and F r .
[0008] Step 3) Scale the image according to s (i) And the dataset disparity range [start_disp, end_disp], construct the correlation cost geometry of the dynamic step size: According to the left boundary of the disparity range start_disp, calculate the right view feature F r The distance from the midpoint A′ to the first point B involved in the correlation calculation is start_distant = start_disp / scale, and then based on start_distant and the dynamic step size step in F r Find the features of B, C, …, z, a total of P non-integer coordinate points, and use cosine similarity to calculate the left view features F l Midpoint A and F rThe correlation between the P points in the image is used to construct a cost geometry, where step = (end_disp-start_disp) / (P×scale), and P is the number of points that need to be matched; the cost geometry refers to a 4D matrix that stores the disparity values and coordinate information between the left and right view pixels.
[0009] Step 4) For the constructed cost geometry, the attention mechanism is applied to optimize it in the disparity dimension, and the matching cost information is integrated through the aggregation module of the hourglass model. Then, the bimodal Laplace mixture distribution is used to predict the preliminary disparity map of the aggregated features.
[0010] Step 5) The gated recurrent unit iterative training mechanism is used to achieve recursive optimization of the disparity estimation results: in each iteration, the model indexes the disparity dimension of the cost geometry through linear interpolation based on the current disparity estimate, generates a new disparity information feature matrix, and inputs it into the gated recurrent unit to calculate the updated disparity value, thereby gradually improving the accuracy and reliability of the disparity estimation.
[0011] Step 6) According to the normalized coordinates of the n supervision points in the data preprocessing stage, the linear interpolation method is used on the disparity map generated by the model reasoning to calculate the disparity information of these supervision points, and the disparity results output by each iterative optimization are compared with the corresponding disparity true values. The difference values of each supervision point are weighted averaged according to specific weights, and finally the loss value loss is obtained, which is combined with the learning rate driven model for reverse reasoning optimization.
[0012] Compared with the prior art, the method of the present invention has at least the following beneficial effects:
[0013] (1) Multi-scale training and arbitrary resolution generation: During the data processing stage, the input image is randomly scaled to achieve multi-scale training, which effectively reduces the GPU resource consumption in high-resolution image processing and significantly enhances the adaptability and generalization ability of the model.
[0014] (2) Dynamic cost geometry construction: By customizing the disparity range and random scaling parameters, the cost geometry is dynamically constructed to optimize resource allocation, meet various disparity matching requirements, and improve the adaptability and stability of the model in complex scenarios.
[0015] (3) Edge-based sampling strategy: A sampling strategy based on the edge of an object is adopted to increase the proportion of pixels in the key area, efficiently learn edge information, and improve training efficiency and the model's accuracy in capturing object boundaries.
[0016] (4) Galerkin attention mechanism: Introducing the Galerkin attention mechanism reduces the computational complexity to O(nd 2), reducing video memory usage and improving computing efficiency and matching performance in high-resolution image scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is the overall flow chart of the method of the present invention.
[0018] Figure 2 It is a cost geometry construction diagram of dynamic parallax perception of the method of the present invention.
[0019] Figure 3 It is a structural diagram of the Galerkin attention mechanism of the method of the present invention.
[0020] Figure 4 It is a coarse-grained cost body information iterative optimization structure diagram of the method of the present invention.
[0021] Figure 5 It is a model binocular disparity estimation diagram of the method of the present invention. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention is further described in detail below with reference to the accompanying drawings. It should be noted that the specific embodiments are only used to explain the technical content of the present invention, but not to limit the present invention.
[0023] The present invention discloses a general deep learning method for binocular disparity estimation, comprising the following steps:
[0024] Step 1) Use the Scene Flow dataset as the training set, the left and right view color image pairs and the corresponding disparity true value d of the left view gt As training data pairs. Set the batch size B to 8, the resolution of the low-resolution image LR to 64×128, the scale range of random cropping is scale_min=4 to scale_max=7, use the object edge priority sampling strategy, the number of model supervision points is 60,000, and use the image pairs normalized to a resolution of 4H×4W and the normalized coordinates of the supervision points as the input of the model network. Figure 1 The overall structure and workflow of the neural network of the present invention are shown. Specifically, during a model forward propagation process, the network performs model inference on the input image pair and finally completes the disparity prediction of the left view pixel point. The main links are as follows:
[0025] S1-1. In the data processing stage, the image is first cropped into image regions of different sizes using a random scale method, and then they are resized into image pairs with a fixed size of 4H×4W resolution to achieve multi-scale training. The pixel differences of the input image in the vertical and horizontal directions are then calculated based on the ground truth, and the threshold is used to determine the boundary pixels. N points are randomly sampled from the sampled boundary pixels and background pixels in a 1:1 ratio as supervision signals.
[0026] S1-2. Input the image pair with a resolution of 4H×4W into the U-Net network to extract shallow features at multiple resolutions. First, the input image is preliminarily downsampled through two convolutional layers with a stride of 2 to obtain a low-resolution LR image with a resolution of H×W. Subsequently, convolution and pooling operations are used to further downsample to 1 / 32 of the original resolution. After downsampling, the feature map is gradually upsampled through transposed convolution, and it is fused with the feature map of the corresponding resolution each time it is upsampled to obtain the shallow features F of the left and right views at each resolution. l and F r .
[0027] S1-3. Based on the shallow features of the left and right views, a correlation matching cost geometry that supports variable disparity step length is constructed. This cost geometry can customize the disparity range and dynamically adjust the step length of the correlation calculation according to the characteristics of different data sets during training or fine-tuning. The calculation formula for the step length is as follows: Among them, d l and d u are the left and right boundaries of the custom disparity range, respectively, and P is the number of disparity levels limited by computing resources, that is, the disparity dimension in the cost geometry, which can also be understood as the number of traversal steps. i is the scaling scale of the i-th image, Δs i It is the step length for constructing the group-related cost geometry containing global information and the local information cascade cost aggregate. Unlike the traditional method that starts traversal from the parallax offset of 0px, the traversal starting point of this method is adjusted according to the left boundary of the custom parallax range. The specific formula is as follows: Figure 2 The specific processing flow is shown, where point A in the right image is a point with a parallax offset of 0px relative to point T in the left image. First, calculate the distance from point B in the right image to point A (denoted as start), and determine the step size Δs in the matching process. Then, starting from point B, traverse to the left in sequence according to the step size Δs to find P points from point B to point Z. Finally, by interpolating the right view feature map, the features of these P non-integer coordinate points are obtained, and the correlation between the reference point T and these P points in the right image is calculated, and finally the corresponding cost geometry is constructed.
[0028] S1-4. For the constructed cost geometry, apply the Galerkin attention mechanism once to optimize the network's expression in the disparity dimension in multi-scale scenes. Figure 3 As shown, represents the eigenvector (z(x1),…,z(x n )) T , whose corresponding query, key and value functions are Q = zW q ,K=zW k ,V=zW v ∈R n×d Their columns represent the vector representation of the learning basis function, and the output feature z out The expression is as follows: in Ln(·) is the layer normalization operation. Compared with the traditional attention mechanism, the traditional method uses QK T Calculate the correlation between n points, the time complexity is O(n 2 d), while the Galerkin attention mechanism focuses on the correlation of the feature channel dimension, and the time complexity is O(nd 2 ), and n>>d, which is more suitable for the disparity prediction task that focuses on the channel dimension information at each coordinate point, and effectively improves the reasoning speed of the network.
[0029] S1-5. In the feature aggregation stage, the matching cost information is integrated through the aggregation module of the hourglass model to help the model capture richer contextual information. At the same time, the bimodal Laplace mixture distribution is used to perform preliminary disparity map prediction based on the aggregated features to improve the accuracy and stability of the prediction.
[0030] S1-6. In the disparity optimization stage, the disparity estimation result is recursively optimized using the iterative training mechanism of the gated recurrent unit. Figure 4 As shown in the figure, during each iteration, the model indexes the disparity level of the relevant geometry according to the current disparity estimate, and constructs a feature structure of length 2r+1 in the disparity dimension through linear interpolation. Subsequently, the extracted feature values are connected into a unified feature map, and the updated disparity value is calculated through the gated recurrent unit, thereby gradually improving the accuracy of the disparity estimation.
[0031] Step 2) Extract the corresponding disparity true value d according to the sampling position coordinates obtained in step S1-1 gt , collect its values at these positions. The F1 loss function is used as a supervisory signal to guide the update of model parameters by calculating the loss. During the back-propagation process, the Adam optimizer is used to update the gradient, and the learning rate is dynamically optimized in combination with the learning rate adjustment strategy OneCycleLR function to ensure that the training process is more efficient and stable.
[0032] Step 3) Alternate the forward propagation and back propagation of step 1 and step 2 to train the model. The entire training process lasts for 200k rounds, and finally the model parameter optimization is completed.
[0033] Step 4) In the test phase, gridded position coordinates are generated according to the target resolution, and predictions are performed using the model trained in step 3 according to the forward process of step 1. Finally, the disparity prediction results of the test image pair are generated. Figure 5 Shows the disparity map predicted by the model on the Scene Flow test set.
Claims
1. A general deep learning method for binocular disparity estimation aims to solve the problems of insufficient migration and generalization of existing models due to disparity prediction of fixed-size images and fixed disparity ranges, and to improve the adaptability of binocular disparity estimation networks when migrating between different domains, thereby achieving more general model training and application; the method is characterized in that by learning the mapping from image space to disparity space, a functional relationship between image pairs and disparity is established, and dynamic adjustment of image resolution and disparity range is achieved, specifically including the following steps: Step 1) Preprocess the input left and right view color images using a variable scale strategy: crop B pairs of images with a resolution of {s (i) H×s (i) W}, the HR images of different scales are uniformly adjusted to 4H×4W resolution and then input into the network for training, and n points are randomly sampled from the true disparity value of each HR image to ensure that the LR images in each batch of data have a consistent number of supervision points; where H×W is the height and width of the low-resolution image LR obtained by the initial downsampling of the model inference, s (i) From uniform distribution Sample B random scales, where scale_min and scale_max are the custom scale ranges in the network parameters, i∈[1, B]. Step 2) Input the adjusted resolution 4H×4W image as the input image to the U-Net module to extract shallow features at multiple resolutions: first, perform a preliminary downsampling of the input image through the convolution layer to obtain the LR image, and then continue to downsample to 1 / 32 times the original resolution through convolution and pooling operations; after the downsampling is completed, the transposed convolution is used to gradually upsample, and after each upsampling, the result is fused with the feature map of the corresponding resolution of the downsampling to obtain the shallow features F of the left and right views at each resolution. l and F r . Step 3) Scale the image according to s (i) And the dataset disparity range [start_disp, end_disp], construct the correlation cost geometry of the dynamic step size: According to the left boundary of the disparity range start_disp, calculate the right view feature F r The distance from the midpoint A′ to the first point B involved in the correlation calculation is start_distant = start_disp / scale, and then based on start_distant and the dynamic step size step in F r Find the features of B, C, ..., Z, a total of P non-integer coordinate points, and use cosine similarity to calculate the left view features F l Midpoint A and F r The correlation between the P points in the image is used to construct a cost geometry, where step = (end_disp-start_disp) / (P×scale), and P is the number of points that need to be matched; the cost geometry refers to a 4D matrix that stores the disparity values and coordinate information between the left and right view pixels. Step 4) For the constructed cost geometry, the attention mechanism is applied to optimize the disparity dimension, and the matching cost information is integrated through the aggregation module of the hourglass model. Then, the bimodal Laplace mixture distribution is used to predict the preliminary disparity map based on the aggregated features. Step 5) The gated recurrent unit iterative training mechanism is used to achieve recursive optimization of the disparity estimation results: in each iteration, the model indexes the disparity dimension of the cost geometry through linear interpolation based on the current disparity estimate, generates a new disparity information feature matrix, and inputs it into the gated recurrent unit to calculate the updated disparity value, thereby gradually improving the accuracy and reliability of the disparity estimation. Step 6) According to the normalized coordinates of the n supervision points in the data preprocessing stage, the linear interpolation method is used on the disparity map generated by the model reasoning to calculate the disparity information of these supervision points, and the disparity results output by each iterative optimization are compared with the corresponding disparity true values. The difference values of each supervision point are weighted averaged according to specific weights, and finally the loss value loss is obtained, which is combined with the learning rate driven model for reverse reasoning optimization.
2. A general deep learning method for binocular disparity estimation as claimed in claim 1, characterized in that: In step 1), a random cropping strategy is adopted for each pair of images in the training set, where the starting position of the cropping (x1, y1) is randomly generated: The images involved in each training have almost no overlap, which enables the neural network to be exposed to training samples of various scales. The model can adapt to image features of different scales according to different scenarios, so that high-resolution images can be accurately predicted even under limited GPU resources by reducing the image resolution. This solves the problem of poor adaptability of traditional methods in cross-scenario applications and improves the adaptability of the network in multiple scenarios and multiple scales.
3. A general deep learning method for binocular disparity estimation as claimed in claim 1, characterized in that: In step 1), to ensure that LR images of different scales have a consistent number of supervision points, the coordinate points of the edge area of the object are calculated according to the true disparity value of each HR image, and half of the coordinate points are randomly and uniformly sampled from the edge area and the non-edge area respectively. The sampled coordinate points are then subjected to s (i) W and S (i) H is normalized and mapped to the standard coordinate space of [-1, 1] to achieve consistency in the number of sampling points at different resolutions.
4. A general deep learning method for binocular disparity estimation as claimed in claim 1, characterized in that: In step 3), the correlation matching cost with variable step size is calculated according to the scaling scale and disparity range of each pair of images, and the disparity range in different scenes is adapted by dynamically adjusting the step size; in the right feature map, P candidate points are selected based on a point in the left feature map to calculate the correlation, and two types of cost geometries containing local correlation information and global context information are constructed, respectively, so that the model can perform disparity prediction in any scale range and disparity range, thereby significantly improving the stability and generalization ability of model training.
5. A general deep learning method for binocular disparity estimation as claimed in claim 1, characterized in that: In step 4), the faster Galerkin attention mechanism is used to help the network perform feature aggregation. represents the input features (z(x1),…,z(x n )) T , Q=zW q , K = zW k , V = zW v ∈R n×d , corresponding to the query, key and value functions respectively, where the column vector contains the expression of the learning basis function and the output feature z out It is expressed as: in Ln(·) is the layer normalization; compared to the traditional attention mechanism, the traditional method relies on QK T Calculate the correlation between n points, the time complexity is O(n 2 d), while the Ga1erkin attention mechanism focuses on the correlation of feature channel dimensions, with a time complexity of O(nd 2 ), and n>>d, which is more suitable for the disparity prediction task that focuses on the channel dimension information at each coordinate point, and effectively improves the reasoning speed of the network.
6. A general deep learning method for binocular disparity estimation as claimed in claim 1, characterized in that: Step 5) A coarse-grained cost volume information iterative optimization training based on a gated recurrent unit is adopted. First, the residual network is used to extract the feature maps of the left and right images, and they are cascaded and spliced into a feature map according to the channel dimension. Then, the integrated feature maps are processed through the Galerkin attention mechanism to obtain richer contextual feature information, and it is used as the initial candidate hidden state input of the gated recurrent unit to iteratively optimize the disparity calculation accuracy and improve the accuracy and robustness of the final disparity estimation result.