A method for obtaining multi-scale disparity maps based on a stereo matching deep neural network

By modeling the pixel matching relationship as the optimal transmission problem in the stereo matching deep neural network and introducing global constraints, the noise and mismatch problems existing in the initial disparity map estimation of the stereo matching deep neural network in the prior art are solved, and the accuracy and efficiency of the network are improved.

CN114723801BActive Publication Date: 2025-06-13XIHUA UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210415618.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-20
Publication Date
2025-06-13
Estimated Expiration
2042-04-20

AI Technical Summary

Technical Problem

The existing stereo matching depth network has a large number of noise points and mismatch situations during the initial estimation of the disparity map, resulting in increased network complexity and reduced output time efficiency.

Method used

By modeling the pixel matching relationship on the image pole line as an optimal transmission problem, the stereo matching uniqueness constraint and global optimal transmission constraint of pixels are introduced to obtain an initial disparity map that is better than the traditional distance metric or correlation metric.

Benefits of technology

This reduces the subsequent regularization and refinement process of the initial disparity map, reduces network complexity, and improves the prediction accuracy and speed of the stereo matching network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114723801B_ABST
    Figure CN114723801B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for obtaining a multi-scale disparity map based on a stereo matching deep neural network. Binocular stereo vision cameras are used to capture left and right images, which are sent to the stereo matching deep neural network after stereo rectification and epipolar alignment processing. In the stereo matching deep neural network, first, the features of the left and right images are extracted, and then the matching cost calculation module calculates the matching cost between pixels row by row using distance metrics for the left and right image features at each resolution level, and finds the corresponding optimal pixel transfer matrix. Then, based on the optimal pixel transfer matrix, a multi-scale initial disparity map is predicted and refined, so as to output a multi-scale disparity map with the same resolution as the initial disparity map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning and computer vision. More specifically, it relates to a method for obtaining a multi-scale disparity map based on a stereo matching deep neural network. Background Art

[0002] As one of the main representation forms of three-dimensional information, the depth map provides basic information for the three-dimensional application direction of computer vision. The computer vision acquisition technology of the depth map can be divided according to the usage form of depth prediction with different numbers of views. Among them, binocular stereo depth estimation has a wide range of practical applications, including robot navigation, three-dimensional reconstruction, industrial control, remote sensing images, and autonomous driving, etc., due to its simple and low-cost device requirements and high-efficiency passive prediction of the depth value corresponding to each pixel of the input image, and has important research significance and application value.

[0003] Binocular stereo depth estimation simulates the way of human eyes perceiving depth based on disparity. By matching the pixel correspondence between stereo images, high-precision disparity values can be obtained, and then the dense pixel depth, that is, the depth map, can be calculated according to the triangle similarity relationship, which has received continuous attention from researchers. Among them, the stereo matching technology in the core link has roughly gone through three generations of technological evolution. The first-generation stereo matching technology generally uses distance metrics for pixel-level matching of stereo images. Although it has good hardware real-time performance after optimization, it pays too much attention to the local feature differences of images and is difficult to handle complex actual situations. The second-generation stereo matching technology models the stereo matching problem as an optimization objective function, and introduces a global matching cost term and a disparity map smoothing regularization term for optimization. There are many types of such methods, but usually the optimization process is complex and the algorithm efficiency is not high. The third-generation stereo matching technology combines the powerful feature representation ability and non-linear data fitting ability of deep learning technology, and regards stereo matching as a data-driven network parameter optimization learning task, and has become the preferred solution for estimating depth maps in stereo vision in recent years.

[0004] However, in the existing stereo matching deep networks, the method of calculating the distance metric for each pair of pixels is still used for the matching relationship between pixels, lacking global constraints on the pixel matching relationship. There are still a large number of noise points and mismatching situations in the initial estimation process of the stereo matching deep network for the disparity map. It is necessary to set complex cost regularization neural network layers and disparity refinement neural network layers in the stereo matching deep neural network, which not only increases the complexity of the stereo matching network but also reduces the time efficiency of outputting a dense depth map. Summary of the Invention

[0005] The object of the present invention is to overcome the deficiencies of the prior art and provide a method for obtaining a multi-scale disparity map based on a stereo matching deep neural network. By modeling the pixel matching relationship on the epipolar line of an image as an optimal transport problem and introducing the uniqueness constraint of stereo matching and the global optimal transport constraint of pixels, an initial disparity map superior to traditional distance metrics or correlation metrics is obtained, thereby reducing the subsequent regularization and refinement processes of the initial disparity map, reducing the network complexity, and further improving the prediction accuracy and speed of the stereo matching network.

[0006] To achieve the above object of the invention, a method for obtaining a multi-scale disparity map based on a stereo matching deep neural network according to the present invention is characterized by comprising the following steps:

[0007] (1) Image acquisition and preprocessing;

[0008] Use a binocular stereo vision camera to capture left and right images with a resolution of W×H; then, after stereo rectification and epipolar alignment processing, obtain the left image I L and the right image I R .

[0009] (2) Use the stereo matching deep neural network to extract a multi-scale disparity map;

[0010] (2.1) Feature extraction;

[0011] Input the left image I L and the right image I R into the stereo matching deep neural network. In the feature extraction module, first perform an operation of gradually halving the resolution by a two-dimensional convolutional downsampling block with a stride of 2 to obtain the contracted features at each level;

[0012] (e L ,e R ) m = Down m {(e L ,e R ) m-1 ; θ m},m = 1, 2,..., M

[0013] where (e L ,e R ) m represents the feature after the m-th level of contraction, e L represents the contracted left image feature, e R represents the contracted right image feature; θ m represents the convolutional kernel parameter in the m-th level two-dimensional convolutional downsampling block; Down represents the downsampling operation; M represents the total number of downsampling levels;

[0014] Next, a two-dimensional transposed convolutional upsampling block with a step size of 2 is used to perform an operation of doubling the resolution of each level of the contracted features step by step to obtain each level of expanded features;

[0015] (u L ,u R ) m-1 =Up m {(u L ,u R ) m ,Skip m [((e L ,e R ) m-1 ;η m )];μ m},m=M,M - 1,…,1

[0016] Among them, (u L ,u R ) m-1 represents the feature after expansion at the m-th level, u L represents the left image feature of the expansion, u R represents the right image feature of the expansion; Skip represents the two-dimensional convolutional operation of the skip connection; η m represents the convolutional kernel parameter in the skip connection, μ m represents the convolutional kernel parameter in the two-dimensional transposed convolutional upsampling block at the m-th level; Up represents the upsampling operation;

[0017] (2.2) Calculate the initial matching cost;

[0018] In the matching cost calculation module, the pixel - to - pixel matching cost is calculated row by row for the left and right image features at each resolution level using a distance metric. Among them, the initial matching cost C m (y,x L ,x R ) of the pixels on the y - th row at the m - th resolution level is calculated by the formula:

[0019]

[0020] Among them, represents the feature value of the pixel point (y,x L ) in the left image feature at the m - th level, represents the feature value of the pixel point (y,x R ) in the right image feature at the m - th level; the value range of y is x L and the value ranges of x R are both

[0021] (2.3) Obtain the optimal pixel transfer matrix π corresponding to the initial matching cost among the pixels in the y-th row at the m-th resolution level. m ;

[0022]

[0023]

[0024] Among them, π m (y, x L , x R ) is an element in π m , representing the optimal transfer probability value that the pixel point (y, x L ) in the y-th row of the left image feature at the m-th resolution level is transferred to the pixel point (y, x R ) in the y-th row of the right image feature. And π m (y, x L , x R ) needs to satisfy that the sum of each row is 1 and the sum of each column is also 1;

[0025] (2.4) Output the multi-scale disparity map;

[0026] (2.4.1) Traverse each row coordinate x m in π L , obtain the column coordinate x L corresponding to the maximum value in the optimal pixel transfer matrix π m with the row coordinate being x R , then calculate the coordinate difference between x L and the corresponding x R , and use it as the disparity value;

[0027]

[0028] Among them, d m (y, x L ) represents the disparity value of the y-th row and the x L -th column in the initially predicted disparity map at the m-th resolution level;

[0029] (2.4.2) Concatenate all the disparity values d m (y, x L ) in the width direction to form the disparity vector d m (y) of the y-th row at the m-th resolution level;

[0030]

[0031] (2.4.3) Concatenate the disparity vectors of all rows at the m-th resolution level in the height direction to form the initially predicted disparity map d m ;

[0032]

[0033] (2.4.4) Send the predicted initial disparity map d m into the disparity refinement structure of the stereo matching depth neural network, and perform disparity refinement processing according to the following formula, so as to output a multi-scale disparity map with the same resolution as the initial disparity map

[0034]

[0035] where Resblocks represents the residual convolution module, Up represents bilinear upsampling, represents the refined disparity map output by the (m + 1)-th level disparity refinement structure, represents the left image feature at the m-th level.

[0036] The object of the present invention is achieved as follows:

[0037] The method for obtaining a multi-scale disparity map based on a stereo matching depth neural network in the present invention uses a binocular stereo vision camera to capture left and right images, which are sent into the stereo matching depth neural network after stereo rectification and epipolar alignment processing; in the stereo matching depth neural network, first extract the left and right image features, then use the matching cost calculation module to calculate the pixel-by-pixel matching cost row by row for the left and right image features at each resolution level using a distance metric, and find the corresponding optimal pixel transfer matrix; then predict a multi-scale initial disparity map according to the optimal pixel transfer matrix, and perform refinement processing, so as to output a multi-scale disparity map with the same resolution as the initial disparity map.

[0038] At the same time, the method for obtaining a multi-scale disparity map based on a stereo matching depth neural network in the present invention also has the following beneficial effects:

[0039] (1) Using the hourglass-shaped convolutional neural network structure can obtain the context space features of pixels from the multi-scale image space and minimize the amount of calculation to the greatest extent; secondly, setting skip connections is to introduce the feature details lost due to resolution downsampling and increase the information flow between deep and shallow features, avoiding the phenomenon of gradient disappearance during training;

[0040] (2) The present invention designs a disparity refinement network structure based on hierarchical feature guidance in the stereo matching depth neural network, which alleviates the problem of false matching noise introduced by using local matching cost metrics in the stereo matching network. By introducing global pixel transfer constraint relationships, the prediction accuracy of the initial disparity map is improved; in addition, after obtaining a more accurate initial disparity map, the complexity of the network structure for refining the initial disparity map is effectively reduced, and the time efficiency of the disparity map output by the stereo matching network is improved;

[0041] (3) By embedding the optimal transport theory into the stereo matching network structure, the present invention can specifically learn the image context features with optimal transport matching during the network training process, thereby reducing the network training difficulty and improving the training iteration convergence speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 is a flowchart of the method for obtaining a multi-scale disparity map based on the stereo matching deep neural network of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The following describes the specific embodiments of the present invention with reference to the drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may obscure the main content of the present invention, these descriptions will be omitted here.

[0044] Embodiment

[0045] Figure 1 is a flowchart of the method for obtaining a multi-scale disparity map based on the stereo matching deep neural network of the present invention.

[0046] In this embodiment, as Figure 1 shown, a method for obtaining a multi-scale disparity map based on a stereo matching deep neural network of the present invention includes the following steps:

[0047] S1. Image acquisition and preprocessing;

[0048] Use a binocular stereo vision camera to capture left and right images with a resolution of W×H; then, after stereo rectification and epipolar alignment processing, obtain the left image I L and the right image I R ;

[0049] S2. Use the stereo matching deep neural network to extract a multi-scale disparity map;

[0050] S2.1 Feature extraction;

[0051] In this embodiment, the feature extraction network samples a hourglass-shaped convolutional neural network structure that first shrinks and then expands in the image resolution and has skip connections. The specific operation is as follows:

[0052] Input the left image I L and the right image I R into a stereo matching deep neural network with shared parameters and dual-path feature extraction. In the feature extraction module, first perform the operation of gradually halving the resolution by a two-dimensional convolutional downsampling block with a stride of 2 to obtain each level of contracted features;

[0053] (eL , e R ) m = Down m {(e L , e R ) m-1 ; θ m}, m = 1, 2,..., M

[0054] Among them, (e L , e R ) m represents the feature after contraction at the m-th level, e L represents the left image feature of contraction, e R represents the right image feature of contraction; θ m represents the convolution kernel parameter in the two-dimensional convolutional downsampling block at the m-th level; Down represents the downsampling operation; M represents the total number of downsampling levels;

[0055] Next, use a two-dimensional transposed convolutional upsampling block with a stride of 2 to double the resolution of each level of the contracted feature to obtain each level of the expanded feature;

[0056] (u L , u R ) m-1 = Up m {(u L , u R ) m , Skip m [((e L , e R ) m-1 ; η m )]; μ m}; m = M, M - 1,..., 1

[0057] Among them, (u L , u R ) m-1 represents the feature after expansion at the m-th level, u L represents the left image feature of expansion, u R represents the right image feature of expansion; Skip represents the two-dimensional convolutional operation of the skip connection; η m represents the convolution kernel parameter in the skip connection, μ m represents the convolution kernel parameter in the two-dimensional transposed convolutional upsampling block at the m-th level; Up represents the upsampling operation;

[0058] In this embodiment, two convolutional layers are set in the downsampling block, and four convolutional layers are set in the upsampling block. A non-linear activation unit Leaky ReLU is connected after each convolutional layer. In this feature extraction network structure, a batch normalization layer (BatchNorm Layer) is not set to avoid the problem of decreased accuracy after network domain migration due to the inconsistency between the local statistical features of the images in the training domain and the local statistical features of the images in the actual test domain. For the skip connection method of each level of resolution features, the feature layers in the shrinking part are cascaded to the feature layers of the same resolution in the expanding part using 2D convolution with a kernel size of 1×1, and then a 2D convolutional layer with a kernel size of 3×3 is connected for feature fusion. Set M = 4, then the left and right image features of four scales are finally output, and the corresponding resolutions are

[0059] In summary, the hourglass convolutional neural network structure used in this embodiment can obtain the context space features of pixels from the multi-scale image space and minimize the amount of computation to the greatest extent. The skip connection is set to introduce the feature details lost due to resolution downsampling and increase the information flow between deep and shallow features, avoiding the phenomenon of gradient disappearance during training.

[0060] S2.2. Calculate the initial matching cost;

[0061] For the vast majority of stereo matching methods, the matching cost between pixels on the same row scan line y of the left and right image features at the m-th scale is calculated using the L p distance (p is usually taken as 1 or 2) as:

[0062]

[0063] In this embodiment, the matching cost calculation module is used to calculate the matching cost between pixels row by row using the distance metric for the left and right image features at each resolution level. Among them, the initial matching cost C m (y, x L , x R ) of the pixels on the y-th row at the m-th resolution is calculated by the formula:

[0064]

[0065] Among them, represents the feature value of the pixel point (y, x L ) in the left image feature at the m-th level, represents the feature value of the pixel point (y, x R ) in the right image feature at the m-th level; the value range of y is x L and the value ranges of x R are both

[0066] S2.3. Obtain the optimal pixel transfer matrix π corresponding to the initial matching cost among the pixels in the y-th row at the m-th resolution level m ;

[0067]

[0068]

[0069] where π m (y, x L , x R ) is an element in π m , representing the optimal transfer probability value that the pixel point (y, x L ) in the y-th row of the left image feature at the m-th resolution level is transferred to the pixel point (y, x R ) in the y-th row of the right image feature, and π m (y, x L , x R ) needs to satisfy that the sum of each row is 1 and the sum of each column is also 1;

[0070] S2.4. Output the multi-scale disparity map;

[0071] S2.4.1. Combining with step S2.2, for C m (y, x L , d), usually directly select the minimum cost as the initial disparity output, that is:

[0072]

[0073] However, in this embodiment, C m (y, x L , x R ) is used to construct the optimal pixel transfer optimization objective formula, solve the optimal pixel transfer matrix π m according to step S2.3, then traverse each row coordinate x m in π L , obtain the column coordinate x L corresponding to the maximum value in the optimal pixel transfer matrix π m with the row coordinate x R , then calculate the coordinate difference between x L and the corresponding x R , and use it as the disparity value;

[0074]

[0075] where d m (y, x L ) represents the disparity value of the y-th row and the x L -th column in the initial disparity map predicted at the m-th resolution level;

[0076] S2.4.2. Cascade all disparity values ​​d in the width direction m (y,x L ), forming the disparity vector d of the yth row of the mth resolution m (y);

[0077]

[0078] S2.4.3. Concatenate the disparity vectors of all rows at the m-th resolution in the height direction to form the initial disparity map d predicted at the m-th resolution m ;

[0079]

[0080] S2.4.4. In this embodiment, the stereo matching deep neural network includes a disparity refinement network structure based on hierarchical feature guidance, which transforms the predicted initial disparity map d m The disparity refinement structure sent to the stereo matching deep neural network, for each level of resolution scale, the estimated initial disparity map and the left image feature of the scale are cascaded, followed by a set of residual convolution modules for predicting the disparity residual map, and then the residual is attached to the initial disparity map to obtain the refined predicted disparity map at the resolution scale. For a higher level of resolution scale, it is necessary to cascade not only the initial disparity map of this level, but also the refined disparity map sampled from the lower level. By refining layer by layer, the refined disparity map of the original resolution scale is finally output, that is:

[0081]

[0082] Among them, Resblocks represents the residual convolution module, which consists of 6 2D convolution residual blocks with a convolution kernel size of 3×3. Each residual block consists of 2 2D convolution layers. All convolution layers are followed by LeakReLU nonlinear activation layers. Batch normalization layer BatchNorm is not set. The convolution expansion rate of each residual block is set to [1,1,2,4,8,1] respectively; Up means using bilinear upsampling, and the input disparity map is upsampled by one image resolution through bilinear interpolation. Through step-by-step upward iteration, the stereo matching network finally outputs a disparity map with full resolution.

[0083] Although the above describes the illustrative specific embodiments of the present invention to facilitate those skilled in the art to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations using the concept of the present invention are protected.

Claims

1. A method for obtaining a multi-scale disparity map based on a stereo matching deep neural network, characterized in that, it includes the following steps: (1), Image acquisition and preprocessing; The left and right images are captured by a binocular stereo camera with a resolution of W×H. After stereo correction and epipolar alignment, the left image I is obtained. L , right image I R ; (2), Using a stereo matching deep neural network to extract a multi-scale disparity map; (2.1), Feature extraction; Input the left image I L and the right image I R into the stereo matching depth neural network. In the feature extraction module, first perform the operation of gradually halving the resolution by the two-dimensional convolutional downsampling blocks with a stride of 2 respectively to obtain the contracted features at each level; ( e L ,e R ) m = Down m {(e L ,e R ) m-1 ; θ m}, m = 1, 2, …, M Among them, (e L , e R ) m represents the feature after the m-th level of contraction, e L represents the left image feature of the contraction, e R represents the right image feature of the contraction; θ m represents the convolution kernel parameter in the two-dimensional convolutional downsampling block at the m-th level; Down m represents the downsampling operation; M represents the total number of downsampling levels; Next, use a two-dimensional transposed convolutional upsampling block with a stride of 2 to perform an operation of doubling the resolution of each level of contracted features to obtain each level of expanded features; (u L ,u R ) m-1 = Up m {(u L ,u R ) m ,Skip m [((e L ,e R ) m-1 ; η m )]; μ m}, m = M, M - 1, …, 1 Among them, (u L , u R ) m-1 represents the feature after the (m - 1)-th level of expansion, u L represents the left image feature of the expansion, u R represents the right image feature of the expansion; Skip m represents the two-dimensional convolutional operation of the skip connection; η m represents the convolutional kernel parameter in the skip connection, μ m represents the convolutional kernel parameter in the two-dimensional transposed convolutional upsampling block at the m-th level; Up m represents the upsampling operation; (2.2), Calculate the initial matching cost; In the matching cost calculation module, the matching cost between pixels is calculated row by row for the left and right image features at each resolution level using a distance metric. Among them, the initial matching cost C m (y, x L , x R ) of the pixel at the y-th row of the m-th level resolution is calculated according to the following formula: Among them, represents the eigenvalue of the pixel point (y, x L ) in the left image feature of the m-th level, represents the eigenvalue of the pixel point (y, x R ) in the right image feature of the m-th level; the value range of y is x L and the value ranges of both x R are (2.3) Obtain the optimal pixel transfer matrix π corresponding to the initial matching cost among the pixels in the y-th row at the m-th resolution level m ; Among them, π m (y, x L , x R ) is an element in π m , representing the optimal transport probability value that the pixel point (y, x L ) in the y-th row of the left image feature at the m-th resolution level is transported to the pixel point (y, x R ) in the y-th row of the right image feature, and π m (y, x L , x R ) needs to satisfy that the sum of each row is 1 and the sum of each column is also 1; (2.4), Output a multi-scale disparity map; (2.4.1), in π m Traverse each row coordinate x L , get the row coordinate as x L The optimal pixel transfer matrix π m The column coordinate x corresponding to the maximum value R , then calculate x L The corresponding x R The coordinate difference is used as the disparity value; where d m (y, x L ) represents the disparity value at the y-th row and x-th L column of the initial disparity map predicted at the m-th level of resolution; (2.4.2) Cascade all the disparity values d in the width direction m (y, x L ) to form the disparity vector d of the y-th row at the m-th resolution m (y); (2.4.3) Concatenate the disparity vectors of all rows at the m-th level of resolution in the height direction to form the initial disparity map d predicted at the m-th level of resolution m ; (2.4.4) Send the predicted initial disparity map d m to the disparity refinement structure of the stereo matching depth neural network, and perform disparity refinement processing according to the following formula to output a multi-scale disparity map with the same resolution as the initial disparity map Among them, Resblocks represents the residual convolution module, and Up represents bilinear upsampling. represents the refined disparity map output by the (m + 1)-th level disparity refinement structure. represents the left image feature at the m-th level.

Citation Information

Patent Citations

  • Binocular stereo matching method based on joint up-sampling convolutional neural network

    CN111402129A

  • Binocular parallax matching method and system based on shared features and attention upsampling

    CN111915660A