A multi-scale real-time binocular stereo matching method based on grouping distance

By using the method of diced convolution feature extraction and grouping distance cost construction in the binocular stereo matching algorithm, the real-time and accuracy problems of disparity estimation on resource-constrained devices in the prior art are solved, and efficient and accurate disparity estimation is achieved.

CN115294365BActive Publication Date: 2025-05-16GUANGZHOU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210700083.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-20
Publication Date
2025-05-16
Estimated Expiration
2042-06-20

AI Technical Summary

Technical Problem

It is difficult for existing binocular stereo matching algorithms to achieve real-time high-precision parallax estimation on devices with power consumption or memory limitations, and traditional feature extraction methods such as U-Net will lose feature information during downsampling, affecting the accuracy of parallax estimation.

Method used

A multi-scale feature extraction network based on tilt convolution is used to construct cost amounts in combination with the packet distance method, and parallax estimation is performed through multi-stage processing methods from coarse to fine. This method retains more feature information during feature extraction and cost construction, and improves inference speed and accuracy through grouping distance calculation.

Benefits of technology

Real-time and high-precision disparity estimation is realized on resource-constrained devices, improving the network's inference speed and accuracy of disparity estimation, and providing important application support in the fields of autonomous driving, robots, etc.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115294365B_ABST
    Figure CN115294365B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-scale real-time binocular stereo matching method based on grouping distance, comprising the following steps: firstly, extracting features from an input stereo image using a feature extractor based on block convolution to obtain high-resolution and low-resolution feature maps; then obtaining a rough disparity map of the first stage with the lowest resolution; in the second stage, correcting the rough disparity map of the first stage, the rough disparity map of the first stage is first upsampled, and the disparity map of the second stage is generated by adding the residual map of the second stage and the rough disparity map of the first stage; in the third stage, the disparity map of the third stage is generated by adding the residual map of the third stage and the disparity map obtained in the second stage; finally, the disparity maps of the three stages are upsampled by bilinear interpolation respectively, and the disparity maps of the three stages are output. The present invention can be deployed in real time on resource-constrained devices through a lightweight design structure, and can perform disparity estimation with high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of graphic image processing technology, and specifically relates to the field of binocular stereo matching technology for performing disparity estimation in a multi-stage processing method from coarse to fine, and in particular to a multi-scale real-time binocular stereo matching method based on grouping distance. Background Art

[0002] Binocular stereo matching algorithms are widely used in autonomous driving, robotics, smart homes, smart cities, augmented reality, and 3D reconstruction. Therefore, high-precision, lightweight, and real-time binocular depth estimation algorithms are of great significance for devices with limited memory or power consumption. The accuracy of traditional binocular stereo matching algorithms is usually affected by some pathological areas of stereo images, such as textureless, reflection, and occlusion. In recent years, as deep convolutional networks have shown powerful feature extraction capabilities and can extract high-dimensional feature information, the accuracy of binocular stereo matching algorithms based on deep convolutional networks has been significantly improved. However, since the current high-precision binocular stereo matching algorithms often have problems such as high computational cost, high power consumption, and delay, the current methods are difficult to be deployed in real time on devices with limited power consumption or memory.

[0003] Stereo matching algorithms are mainly divided into four steps: unary feature extraction, cost volume construction, matching cost aggregation, and disparity regression. Among them, unary feature extraction and cost volume construction play a key role in the accuracy and inference speed of the network. For feature extraction, many stereo matching algorithms use the general U-Net architecture to extract the features of the input image. U-Net has a symmetrical encoding and decoding structure and can output all resolutions of the image at the same time. However, the feature information of the input image is usually lost during the downsampling process of the U-Net encoding structure. For the construction of the cost volume, Zbontar et al. were the first to propose the use of deep learning methods to predict the matching cost (reference J. Zbontar, and Y. LeCun, "Computing the stereo matching cost with a convolutional neural network," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2015, pp. 1592-1599). Luo et al. use correlated features and product layers to speed up the calculation of matching costs (reference W. Luo, A. G. Schwing, and R. Urtasun, "Efficient deep learning for stereo matching," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2016, pp. 5695-5703). Mayer et al. proposed using related cost quantities to calculate matching costs (reference N. Mayer, E. Ilg, P. Hausser, P. Fischer, D. Cremers, A. Dosovitskiy, and T. Brox, "A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2016, pp. 4040-4048).Knobelreiter et al. used the features learned by neural networks to match the cost of conditional random fields (CRFs), and the calculated matching cost became the single cost of CRFs (reference P. Knobelreiter, C. Reinbacher, A. Shekhovtsov, and T. Pock, "End-to-end training of hybrid CNN-CRF models for stereo," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2017, pp. 2339-2348). In addition, some researchers proposed connection-based feature quantities and three-dimensional aggregation networks to construct four-dimensional cost quantities. Guo et al. proposed GWCNet, which constructs the cost volume by exploiting group correlation, which retains more information and provides a good similarity measurement (reference X.Guo, K.Yang, W.Yang, X.Wang, and H.Li, "Group-wise correlation stereo network," in Proc. IEEE / CVF Conf. Comput. Vis. Pattern Recognit., 2019, pp. 3273-3282). Although the above network can estimate disparity with high accuracy, it cannot run in real time on devices with limited power or memory.

[0004] There are many existing methods for correcting the rough disparity estimation results with a small disparity offset at such a high resolution. For example, the paper Y. Wang, Z. Lai, G. Huang, BH Wang, LV D Maaten, M. Campbell, and KQ Weinberger, "Anytime stereo image depth estimation on mobile devices," in Proc. IEEE Int. Conf. Robot. Autom., 2019, pp. 5893-5900. proposed a method of using step-by-step refinement of disparity estimation. The specific implementation steps of this method are as follows: First, a unary feature extractor is used to extract multi-scale features of the corrected input image to obtain a feature map relative to the original image. Figure 1 / 4, 1 / 8 and 1 / 16 resolution feature maps; in the first stage, the network uses the 1 / 16 resolution feature map to construct the cost volume, and obtains a rough disparity map after 3D convolution and disparity regression operations; in the second stage, the rough disparity map obtained in the first stage is upsampled to 1 / 8 resolution, and the 1 / 8 resolution features of the feature extractor are used to estimate the disparity to generate a more accurate disparity map; in the third stage, the disparity map generated in the second stage is upsampled to 1 / 4 resolution, and the 1 / 4 resolution in the feature extractor is used to estimate the disparity, which is a similar process to the second stage; in the fourth stage, the spatial propagation network (SPN) is used to refine the disparity map obtained in the third stage to generate an accurate disparity map. The algorithm performs disparity estimation according to the step-by-step optimization method, and first constructs the cost volume with a low-resolution feature map, which can greatly reduce the network's computing resources and speed up the model's reasoning. In the paper X.Guo, K.Yang, W.Yang, X.Wang, and H.Li, "Group-wise correlationstereo network," in Proc.IEEE / CVF Conf.Comput.Vis.Pattern Recognit., 2019, pp.3273-3282., it is proposed to estimate disparity based on group correlation. The general implementation steps of the algorithm are: first, use the feature extractor to extract multi-level left-right features from the rectified input stereo image; divide the left-right unigram features into multiple groups along the channel dimension, and calculate the correlation map between each group; the above operations can obtain multiple matching cost schemes, and then merge these matching cost schemes into a group correlation amount; combine the obtained group correlation amount with a connection amount to form the final group correlation cost amount, which further improves the disparity prediction result; the final group correlation cost amount can generate an accurate disparity map after being optimized by the 3D aggregation network. The algorithm proposes a group correlation method for disparity estimation, and combines the group correlation amount and the connection amount to construct the cost amount, which provides a good similarity measure.

[0005] In order to reduce the computational resources for constructing the cost volume, some lightweight algorithms use a coarse-to-fine approach to predict disparity. These methods usually use low-resolution features to construct the cost volume first, which generates a relatively coarse disparity map, and then the disparity map is upsampled to match the resolution of the next level, and the coarse disparity estimation result is corrected with a small disparity offset at a higher resolution. This method can significantly improve the reasoning speed of the network, but its low-resolution cost volume cannot generate a very accurate disparity map. Although some algorithms can use a multi-stage coarse-to-fine processing method to further improve the results of disparity estimation, as the number of stages increases, the reasoning speed of the algorithm will be significantly reduced, which is still difficult to obtain accurate disparity estimation results in real time on devices with limited power or memory. Existing algorithms are still difficult to generate high-precision disparity maps in real time on devices with limited power or memory, so the problem that the present invention needs to solve is to obtain high-precision disparity estimation results in real time on resource-constrained devices.

[0006] In summary, the shortcomings of the prior art are as follows:

[0007] (1) The feature information of the input image is usually lost during the downsampling process of the U-Net encoding structure;

[0008] (2) Most existing networks require long inference time for disparity estimation;

[0009] (3) Existing algorithms cannot strike a good balance between inference speed and accuracy and cannot generate accurate disparity maps in real time on low-power devices. Summary of the invention

[0010] In view of the above-mentioned defects of the prior art, the object of the present invention is to provide a multi-scale real-time binocular stereo matching method based on grouping distance, which can be deployed in real time on resource-constrained devices through a lightweight design structure and perform disparity estimation with high accuracy.

[0011] The present invention provides the following technical solutions:

[0012] A multi-scale real-time binocular stereo matching method based on grouping distance, characterized in that it comprises the following steps:

[0013] First, a feature extractor based on block convolution is used to extract features from the input stereo image to obtain high-resolution and low-resolution feature maps;

[0014] Then, in the first stage, a low-resolution feature map is used to construct a group distance cost, and a rough disparity map of the first stage with the lowest resolution is obtained through a 3D convolutional network and disparity regression; in the second stage, the rough disparity map of the first stage is corrected, the rough disparity map of the first stage is first upsampled, and then dynamically offset with the left image feature to construct the group distance cost, and the second stage residual map is obtained through a 3D convolutional network and disparity regression, and the disparity map of the second stage is generated by adding the second stage residual map and the coarse disparity map of the first stage; in the third stage, the disparity map of the second stage is corrected, the disparity map is first upsampled, and then dynamically offset with the left image feature to construct a new group distance cost, and the third stage residual map is obtained through a 3D convolutional network and disparity regression, and the third stage disparity map is generated by adding the third stage residual map and the disparity map obtained in the second stage;

[0015] Finally, the disparity maps of the three stages are upsampled by bilinear interpolation, and the disparity maps of the three stages are output, namely, the first stage disparity map, the second stage disparity map and the third stage disparity map.

[0016] The feature extractor is a unary feature extractor, which uses block convolution for downsampling. The input stereo image passes through the feature extractor and can generate three feature maps with resolutions of 1 / 16, 1 / 8 and 1 / 4 relative to the original input image.

[0017] Furthermore, the block convolution step includes: first using convolution kernels and convolutions with the same step size for block convolution, then using Leak Relu activation function to improve the expression ability of the model, and finally using Batch Normalization to speed up the convergence of the model.

[0018] Furthermore, the unary feature extractor extracts three features with channel numbers of 2c, 4c and 8c, where c is a hyperparameter of the feature extraction network. The three features correspond to feature maps with scales of 1 / 4, 1 / 8 and 1 / 16 of the original image resolution, respectively. The feature extractor quickly obtains clear output by continuously increasing the number of c.

[0019] Furthermore, the group distance cost includes a group distance and a connection cost.

[0020] Furthermore, the group distance cost is expressed as:

[0021]

[0022] Where Cgwd represents the group distance cost value of the similarity between the pixels in the left image and the right image, d is the disparity, x and y are the feature vectors of the feature map, g is the number of the feature group, fl is the feature of the left image and fr is the feature of the right image, and the distance of all feature groups g with disparity d is calculated.

[0023] Furthermore, the 3D convolutional network is composed of six 3D convolutional layers with a convolution kernel size of 3*3*3.

[0024] Furthermore, the cost volume is converted to a probability volume by taking the negative of each disparity value, and the selected disparity value di is the softmax weighted combination of all disparity values:

[0025]

[0026] Where M is the maximum disparity of the cost volume, Ci(d') is the cost value of the disparity level d' in the current pixel i, d is all disparity levels, and Ci(d) is the cost value of the disparity level d in the current pixel i.

[0027] Furthermore, the maximum disparity M of the cost volume is selected to be 12 at a low resolution of 1 / 16.

[0028] Furthermore, the three stages are jointly trained with different loss weights. The corresponding weight coefficients of the first stage, the second stage and the third stage are set to λ1=0.25, λ2=0.5 and λ3=1 respectively, so the loss function is expressed as:

[0029]

[0030] in and They represent the loss functions of the first stage, the second stage, and the third stage respectively.

[0031] The beneficial effects of the present invention are:

[0032] (1) In the feature extraction step, this method proposes a multi-scale feature extraction network based on block convolution, which uses block convolution for downsampling, which is conducive to retaining more feature information. Due to its lightweight design, the feature extractor can extract more feature information from stereo images at a faster inference speed.

[0033] (2) In the step of cost volume construction, this method proposes a group distance method to construct the cost volume. The group distance divides the left and right unary features into multiple groups along the channel dimension and calculates the L1 distance map between each group. This method can obtain multiple matching cost schemes, which are then merged into a group distance volume. Compared with group correlation, group distance can obtain more accurate and robust results at a faster inference speed.

[0034] (3) The network uses a multi-stage coarse-to-fine processing method to perform disparity estimation, which significantly improves the network's reasoning speed. Although the network first uses a 1 / 16 resolution relative to the input stereo image to construct the cost volume, the novel feature extractor and grouping distance method enable the network to recover very fine details. The present invention can achieve 2.52% and 2.18% prediction results (D1-all%) on the KITTI 2012 and KITTI 2015 datasets. On power or memory-constrained devices (such as NVIDIA Jetson Nano), the present invention can perform disparity estimation in real time, which is of great significance in the fields of autonomous driving, robotics, smart homes, smart cities, and augmented reality. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a schematic diagram of the overall composition structure and process flow of the multi-scale real-time binocular stereo matching method based on grouping distance of the present invention.

[0036] Figure 2 It is a feature extractor based on block convolution of the present invention, which extracts feature diagrams with 1 / 16, 1 / 8 and 1 / 4 resolutions.

[0037] Figure 3 It is a detailed schematic diagram of the present invention, which is divided into 4 groups based on the cost amount of grouping distance;

[0038] Figure 4 This is a schematic diagram of the three-dimensional convolutional network of the present invention, with 6 3D convolutional layers in each stage. DETAILED DESCRIPTION

[0039] The embodiments of the present invention are described in detail below. The following embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation methods and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.

[0040] Example

[0041] The multi-scale real-time binocular stereo matching method based on grouping distance provided by the present invention has an overall structure and process as follows: Figure 1 The method of this embodiment is a structure similar to a pyramid design, which is divided into three stages to estimate the disparity map from coarse to fine.

[0042] See also Figure 1-Figure 4 The multi-scale real-time binocular stereo matching method based on grouping distance provided in this embodiment includes the following steps:

[0043] First, a feature extractor based on block convolution is used to extract features from the input stereo image to obtain high-resolution and low-resolution feature maps;

[0044] Then, in the first stage, a low-resolution feature map is used to construct a group distance cost, and a rough disparity map of the first stage with the lowest resolution is obtained through a 3D convolutional network and disparity regression; in the second stage, the rough disparity map of the first stage is corrected, the rough disparity map of the first stage is first upsampled, and then dynamically offset with the left image feature to construct the group distance cost, and the second stage residual map is obtained through a 3D convolutional network and disparity regression, and the disparity map of the second stage is generated by adding the second stage residual map and the coarse disparity map of the first stage; in the third stage, the disparity map of the second stage is corrected, the disparity map is first upsampled, and then dynamically offset with the left image feature to construct a new group distance cost, and the third stage residual map is obtained through a 3D convolutional network and disparity regression, and the third stage disparity map is generated by adding the third stage residual map and the disparity map obtained in the second stage;

[0045] Finally, the disparity maps of the three stages are upsampled by bilinear interpolation, and the disparity maps of the three stages are output, namely, the first stage disparity map, the second stage disparity map and the third stage disparity map.

[0046] Specifically, in the first stage of this embodiment, the network first uses a feature map with a resolution of 1 / 16 to construct a cost volume, thereby obtaining a coarse disparity map with a lower resolution. Since the input resolution is 1 / 16 of the original image, the first stage only requires a small computational cost and inference time.

[0047] In stage 2, the network needs to correct the coarse disparity map obtained from stage 1, which is first upscaled to match the resolution of stage 2. The network then corrects the disparity for each pixel by computing a residual map.

[0048] In the third stage, the network enlarges the 1 / 8 disparity map obtained in the second stage to 1 / 4 resolution for disparity estimation, which is similar to the method in the second stage.

[0049] The specific steps include the following:

[0050] 1) Build a multi-scale unary feature extractor based on block convolution

[0051] This paper proposes a lightweight and high-precision feature extraction network that uses a block convolutional architecture to extract effective features from input stereo images, such as Figure 2As shown in the figure. The feature extractor uses block convolution for downsampling, which can obtain a large receptive field and retain a lot of effective feature information. The input stereo image passes through the feature extractor, which can generate three feature maps with 1 / 16, 1 / 8 and 1 / 4 resolution relative to the original input image. Each block convolution in the feature extractor represents a series of steps: convolution with the same kernel and stride size, activation function (Leaky ReLU) and batch normalization (BatchNorm).

[0052] Table 1 Feature extraction network based on block convolution, as shown below:

[0053]

[0054] Specifically, the feature extraction network first uses a 3×3 convolution with a stride of 1 and a 3×3 convolution with a stride of 2 to reduce the resolution of the feature map by half. Three consecutive block convolutions and ordinary convolutions are used to continuously reduce the size of the feature map resolution and gradually increase the number of feature map channels. This method extracts three features with 2c, 4c, and 8c channels (c: a hyperparameter of the feature extraction network), corresponding to feature maps with a scale of 1 / 4, 1 / 8, and 1 / 16 of the original image resolution, respectively. By continuously increasing the number of c, the feature extractor is able to learn effective features at very low resolutions, thereby quickly obtaining clear output.

[0055] 2) Constructing the group distance cost

[0056] The cost based on grouped distance provided by the present invention consists of two parts: grouped distance and connection. The distance cost is calculated at each disparity level d, and then the disparity estimation is performed by combining the connection cost. When the feature extraction network extracts the left image feature fl and the right image feature fr, the ungrouped distance cost can be defined as the distance between the two features:

[0057] C dist (d, x, y) = || f l (x, y)-f r (xd,y)||1

[0058] Among them, x and y are the feature vectors of the feature map, representing the horizontal and vertical coordinates of the stereo image respectively, d is the disparity, ||·-·||1 is the L1 distance between the two features, and Cdist represents the similarity of the pixels in the left and right images. Then, by establishing a group distance, the matching cost of each pixel at all possible disparity levels is calculated. The basic idea of ​​group distance is: divide the left image feature f1 and the right image feature fr into several groups, and then calculate the distance map group by group, which can obtain multiple cost matching schemes, and finally merge the matching schemes into a group distance, such as Figure 3The calculation method of group distance is as follows:

[0059]

[0060] Where x and y represent feature vectors, and g is the number of the feature group. The distance between all disparities d and feature group g is calculated. In addition, the connection-based feature quantity is used to predict the disparity, which can achieve better context aggregation. A 4D feature quantity is formed by connecting the left and right feature maps. The calculation method of the connection cost is:

[0061] C concat (d, x, y, ) = Concat{f l (x,y),f r (xd,y)}

[0062] Where x and y are feature vectors, and d represents all disparity levels. The present invention combines the grouped distance amount with the connection amount to form a final grouped distance cost amount, which can improve the accuracy of disparity estimation. The network effectively improves the performance in inference time and accuracy by using a very low-resolution cost amount to predict disparity.

[0063] 3) Building a 3D convolutional network

[0064] The cost amount may be affected by uncertain factors in the input stereo image, such as fuzzy matching, weak texture and occlusion. To solve this problem, the present invention uses a 3D convolutional network to refine the group distance amount and the connection amount, such as Figure 4 As shown in the figure, the 3D convolutional network consists of six 3D convolutional layers with a convolution kernel size of 3*3*3, where the first convolutional layer is used to increase the dimension, the middle four convolutional layers are used to optimize the cost, and the last convolutional layer is used to reduce the dimension.

[0065] 4) Parallax regression

[0066] After the cost volume is refined by the 3D convolutional network, disparity regression is performed, which can generate a smooth disparity map. When filtering the cost volume, an argmin operation is usually performed to estimate the minimum disparity in the cost. If the cost volume is accurate, the disparity value d i The argmin function for all disparity values ​​is calculated:

[0067] d i =argmin d C i (d)

[0068] Where Ci(d) is the cost value of the disparity level d in the current pixel i. However, argmin is a non-differentiable function and cannot be trained using the back-propagation method. Therefore, this patent uses a soft argmin operation to estimate the minimum disparity value d i The cost volume is converted to a probability volume by taking the negative value of each disparity value. The selected disparity value d i is the softmax weighted combination of all disparity values:

[0069]

[0070] Where M is the maximum disparity of the cost volume, and Ci(d') is the cost value for disparity level d' at the current pixel i. The soft-argmin operation is fully differentiable, easy to optimize and converges in a short time. The soft-argmin weighted average can be used to recover the minimum disparity of the cost match.

[0071] 5) Residual prediction

[0072] The number of parameters is usually proportional to the resolution of the input image, so the disparity estimation in the first stage starts from a low-resolution cost (1 / 16). If the reasoning time is not strictly limited, the residual map generated in the second and third stages can be used to further predict a more accurate disparity map. When predicting disparity with high resolution, the maximum disparity M between two pixels is usually large. The present invention selects the maximum disparity M=12 (192 / 16) at low resolution (1 / 16). By predicting the residual and correcting the existing disparity, the disparity range in the cost is limited to between -2 and 2. In order to predict the residual of the second and third stages, the network enlarges the disparity map obtained in the first stage and performs a distortion operation on the input features. The value of each pixel in the disparity map represents the offset of the right image relative to the left image. Therefore, the disparity offset obtained from the distortion operation is very small. The extracted left image features and the right image features after the distortion operation are used to establish a new cost, thereby greatly improving the reasoning speed of the network. The low-resolution feature map may cause several pixels to match incorrectly, and the network needs to use the residual map to correct the disparity. The residual map obtained from disparity regression is added to the prediction of the first or second stage.

[0073] 6) Calculate the loss function

[0074] Since the L1 loss function is insensitive to outliers, smaller gradient changes can prevent gradient explosion. Therefore, the present invention uses a smooth L1 function as a loss function to minimize the depth value difference between the ground truth disparity map and the predicted disparity map. The calculation method of the smooth L1 loss function is as follows:

[0075]

[0076] Where x represents the input variable of the loss function. The gradient value of the smooth L1 loss function is small enough when the predicted value and the ground truth value are not much different. The network provided by the present invention can output three stages of disparity maps, and the corresponding loss function Lk of each stage is defined as:

[0077]

[0078] where d is an estimated value, represents the ground truth disparity at the same pixel and N is the number of candidate disparities. Depend on and They represent the loss functions of the first stage, the second stage, and the third stage, respectively. In order to match the ground truth resolution, the predicted disparity map of each stage needs to be upsampled by bilinear interpolation. In the task of the present invention, all three stages are jointly trained with different loss weights. The weight coefficients corresponding to the first, second, and third stages are set to λ1=0.25, λ2=0.5, and λ3=1, respectively.

[0079]

[0080] Where Lsum represents the final loss function of the model. and The loss function is then calculated by summing up the different coefficients.

[0081] The multi-scale feature extraction method based on block convolution proposed in the above embodiment of the present invention can extract effective feature information with resolutions of 1 / 16, 1 / 8 and 1 / 4 for disparity estimation. The feature extractor uses block convolution for downsampling to obtain a relatively large receptive field. In addition, the present invention proposes a method based on group distance to construct a cost amount. The group distance divides the left and right features into multiple groups along the channel dimension, and calculates the L1 distance map between each group, which can obtain multiple matching cost schemes, which are then merged into a group distance amount. Compared with group correlation, group distance can obtain more accurate and robust results at a faster reasoning speed. Because the present invention has a lightweight design structure, it can be deployed in real time on a resource-constrained device (NVIDIA Jetson Nano) and perform disparity estimation with higher accuracy.

[0082] In summary, the present invention focuses on a multi-scale real-time binocular stereo matching method based on grouping distance, which performs disparity estimation by extracting more effective feature information. The feature extraction network uses block convolution for downsampling, which can obtain a relatively large receptive field. In order to achieve fast reasoning, some technologies usually only use simple convolutional network layers to extract features, which cannot generate high-precision disparity maps. The feature extraction network provided by the present invention has a lightweight design architecture. Although the number of channels of the feature extraction network is gradually deepened, the network can still perform disparity estimation at a fast reasoning speed.

[0083] In addition, the full correlation method only generates a single-channel correlation map for each disparity level, which usually loses a lot of information. Group correlation divides the left and right features along the channel dimension and calculates the correlation map between each group, which can obtain multiple matching cost schemes, which are then merged into a group correlation amount. In contrast, the present invention proposes a cost volume construction method based on group distance. The left and right features of this method are divided along the channel dimension, and the distance map between each group is calculated to obtain several matching cost schemes, which are then merged into a group distance amount. Compared with group correlation, group distance can generate a higher-precision disparity map with faster inference time. Existing technologies are generally unable to perform fast and high-precision disparity estimation on power- or memory-constrained devices. In contrast, the present invention can be deployed in real time on resource-constrained devices (NVIDIA Jetson Nano) and perform disparity estimation with higher accuracy.

[0084] The preferred specific embodiments of the present invention are described in detail above. It should be understood that ordinary technicians in the field can make many modifications and changes based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by technicians in the technical field based on the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A multi-scale real-time binocular stereo matching method based on grouping distance, characterized in that: The method comprises: First, a feature extractor based on block convolution is used to extract features from the input stereo image to obtain high-resolution and low-resolution feature maps; Then, in the first stage, a low-resolution feature map is used to construct a group distance cost, and a rough disparity map of the first stage with the lowest resolution is obtained through a 3D convolutional network and disparity regression; in the second stage, the rough disparity map of the first stage is corrected, the rough disparity map of the first stage is first upsampled, and then dynamically offset with the left image feature to construct the group distance cost, and the second stage residual map is obtained through a 3D convolutional network and disparity regression, and the disparity map of the second stage is generated by adding the second stage residual map and the coarse disparity map of the first stage; in the third stage, the disparity map of the second stage is corrected, the disparity map is first upsampled, and then dynamically offset with the left image feature to construct a new group distance cost, and the third stage residual map is obtained through a 3D convolutional network and disparity regression, and the third stage disparity map is generated by adding the third stage residual map and the disparity map obtained in the second stage; Finally, the disparity maps of the three stages are upsampled by bilinear interpolation, and the disparity maps of the three stages are output, namely, the first stage disparity map, the second stage disparity map, and the third stage disparity map; The packet distance cost includes a packet distance and a connection cost; The group distance cost is expressed as: Among them C gwd The group distance cost value representing the similarity of the pixels in the left and right images, d is the disparity, x and y are the feature vectors of the feature map, g is the number of the feature group, and f is the distance cost value representing the similarity of the pixels in the left and right images. l The left figure features and f r For the features in the right image, calculate the distance of the feature group g of all disparity d; The cost volume is converted to a probability volume by taking the negative of each disparity value, and the selected disparity value di is the softmax-weighted combination of all disparity values: Where M is the maximum disparity of the cost volume, C i (d') is the cost value of the disparity level d' in the current pixel i, d is the disparity level, C i (d) is the cost value of the disparity level d in the current pixel i.

2. The multi-scale real-time binocular stereo matching method based on grouping distance according to claim 1, characterized in that: The feature extractor is a unary feature extractor, which uses block convolution for downsampling. The input stereo image passes through the feature extractor and can generate three feature maps with resolutions of 1 / 16, 1 / 8 and 1 / 4 relative to the original input image.

3. The multi-scale real-time binocular stereo matching method based on grouping distance according to claim 2, characterized in that: The block convolution step includes: firstly, using convolution kernels and convolutions with the same step size for block convolution, then using Leak Relu activation function to improve the expression ability of the model, and finally using Batch Normalization to speed up the convergence of the model.

4. The multi-scale real-time binocular stereo matching method based on grouping distance according to claim 2, characterized in that: The unary feature extractor extracts three features with channel numbers of 2c, 4c and 8c, where c is a hyperparameter of the feature extraction network. The three features correspond to feature maps with scales of 1 / 4, 1 / 8 and 1 / 16 of the original image resolution, respectively. The feature extractor quickly obtains clear output by continuously increasing the number of c.

5. The multi-scale real-time binocular stereo matching method based on grouping distance according to claim 1, characterized in that: The 3D convolutional network consists of six 3D convolutional layers with a convolution kernel size of 3*3*3.

6. The multi-scale real-time binocular stereo matching method based on grouping distance according to claim 1, characterized in that: The maximum disparity M=12 of the cost volume is selected at the low resolution 1 / 16.

7. The multi-scale real-time binocular stereo matching method based on grouping distance according to claim 1, characterized in that: The three stages are all jointly trained with different loss weights. The corresponding weight coefficients of the first stage, the second stage and the third stage are set to λ1=0.25, λ2=0.5 and λ3=1 respectively, so the loss function is expressed as: in and They represent the loss functions of the first stage, the second stage, and the third stage respectively.

Citation Information

Patent Citations

  • Real-time binocular stereo matching method based on adaptive candidate parallax prediction network

    CN112435282A

  • Binocular stereo matching method

    CN114170311A