Stereo Matching Method and System Based on Iterative Geometric Coding Volume
By iterating the geometric code structure, combining multi-scale feature extraction and GRU iterative update, the pathological regional fuzzy problem of parallax estimation in stereo matching is solved, more efficient parallax estimation is achieved, and powerful cross-data set adaptability and rapid inference capabilities are provided.
Patent Information
- Application Number
- CN202310002146.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-03
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-01-03
AI Technical Summary
The existing stereo matching methods have local blurring problems in pathological regions in parallax estimation, especially in occlusion areas, textureless areas and reflection areas, and the high computational cost limits its application in high-resolution images and large-scale scenarios.
The stereo matching method based on iterative geometric codes is adopted to construct geometric codes through multi-scale feature extraction networks and three-dimensional regularized networks, and iteratively update the disparity with GRU, and use global and local information to perform disparity estimation to generate a full resolution disparity map.
It improves the accuracy of parallax estimation, has strong ability to generalize across data sets and efficient inference speed, and has become the first in KITTI 2015 evaluation method, and has reached an EPE of 0.47 on the Scene Flow data set, showing better generalization across data sets and faster inference speed.
Smart Images

Figure CN116051739B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision three-dimensional reconstruction, and more particularly relates to a stereo matching method and system based on an iterative geometric encoding volume. Background Art
[0002] Inferring the geometric structure of a three-dimensional scene from images is a fundamental task in computer vision and graphics, and its application scope includes three-dimensional reconstruction, robotics, and autonomous driving. Stereo matching, which uses two calibrated images to reconstruct a dense three-dimensional representation, is a key technology for reconstructing the geometric structure of a three-dimensional scene.
[0003] Currently, the methods in this field are mainly divided into aggregation-based methods and iterative methods, which have complementary advantages and limitations. The former can encode sufficient non-local geometric structures and context information in the cost volume, which is crucial for disparity prediction, especially in challenging regions, but the required computational cost and memory cost are often high. The latter can avoid the high computational cost and memory cost of three-dimensional cost aggregation, but due to being only based on full correlation, its adaptability in pathological regions is low.
[0004] To improve the expressive power of the cost volume, most existing learning-based stereo matching methods use powerful convolutional neural network (CNN) features to construct a cost volume. However, in occluded regions, large textureless regions, reflective regions, and repetitive structures, there are still ambiguity problems in the cost volume. Three-dimensional convolutional networks show great potential in regularizing or filtering the cost volume, and they can propagate repeatable sparse matches to ambiguous and noisy regions. GCNet first uses a three-dimensional encoder-decoder structure to regularize a four-dimensional connection volume. PSMNet proposes a stacked hourglass three-dimensional CNN and combines intermediate supervision to regularize the connection volume. GwcNet and ACVNet respectively propose group correlation volume and attention connection volume to improve the expressive power of the cost volume and ultimately improve the performance in ambiguous regions. However, due to the high computational cost, these aggregation-based methods are difficult to be applied to high-resolution images or large-scale scenes.
[0005] Recently, iterative methods have shown attractive performance in high-resolution images and standard scenarios. Different from existing methods, iterative methods bypass the computationally expensive cost aggregation operation and gradually update the disparity map by repeatedly obtaining information from the high-resolution thought cost volume. For example, RAFT-Stereo uses an update module based on a multi-level GRU (Gated Recurrent Unit) to cyclically update the disparity field from full-correlation feature retrieval. Summary of the Invention
[0006] The problem to be solved by the present invention is the local blurring problem in the ill-conditioned region of disparity estimation. The proposed iterative geometric encoding volume structure aggregates the cost volume of feature construction, effectively captures the global geometric structure features, combines non-local and local information, infers and propagates geometric information, and uses surrounding pixels to assist in estimating the disparity of the blurred region. The accuracy of disparity estimation ranks first among all evaluation methods on 2015 KITTI at present and has strong cross-dataset generalization ability and high inference efficiency.
[0007] To achieve the above object, according to one aspect of the present invention, a stereo matching method based on an iterative geometric encoding volume is provided, and the method includes the following steps:
[0008] (1) Take a pair of left and right images, call one image the source image and the other image the reference image for calculating the disparity estimation between the reference image and the source image; use a multi-scale feature extraction network to extract the image features of the source image and the reference image, and at the same time use different networks to extract the multi-scale features of the reference image as context features;
[0009] (2) Select feature map pairs of the target size to calculate the correlation in groups to construct a correlation volume, input it into a three-dimensional regularization network, introduce the reference image features to guide the cost volume aggregation, infer and propagate the scene geometric information, and generate a geometric encoding volume; and calculate the correlation of all feature map pairs to form a local feature correlation volume, and combine the geometric encoding volume with the local feature correlation volume to form a combined geometric encoding volume;
[0010] (3) Regress the disparity estimation from the foregoing geometric encoding volume as the initial value of iterative update. Starting from the initial disparity estimation, use the disparity estimation result of the previous level, the context features, and the foregoing combined geometric encoding volume as inputs, and use a three-level GRU to iteratively update the disparity to obtain a more accurate disparity prediction;
[0011] (4) Upsample the predicted disparity spatially and use weighted combination to output a full-resolution disparity map.
[0012] In one embodiment of the present invention, in the step (1), the use of the multi-scale feature extraction network to extract the image features of the source image and the reference image specifically includes:
[0013] Use the pre-trained mobilenetv2_100 feature extraction network backbone to extract features. The feature extraction network backbone is an encoder-decoder architecture. The encoder uses the mobilenetv2_100 model pre-trained on ImageNet to extract the input image into feature maps with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 through convolutional layers and bottleneck layers. The decoder uses transposed convolutional layers to upsample the feature maps and then performs feature fusion with a 3x3 convolutional kernel. The feature extraction network finally obtains a series of fused 1 / 4, 1 / 8, 1 / 16, and 1 / 32 feature maps.
[0014] In one embodiment of the present invention, in the step (1), the context features of the image are extracted through a context feature extraction network branch. The context feature extraction network consists of a series of residual blocks and downsampling layers. Taking the reference image as the input, it finally outputs multi-scale context features with a resolution of 1 / 4, 1 / 8, and 1 / 16 of the original image and a feature dimension of 128. This multi-scale context feature is used to initialize the hidden layer of the GRU module and is inserted into the GRU module in each iteration.
[0015] In one embodiment of the present invention, in the step (1), the geometric encoding volume is combined with the local feature correlation volume to form a combined geometric encoding volume, specifically:
[0016] The 1 / 4 feature map obtained through the feature extraction network is used to calculate a group correlation volume of H x W x D x G using the group correlation calculation formula, where H represents height, W represents width, D represents the disparity dimension, and G represents the channel dimension; then a lightweight three-dimensional regularization network is used to regularize this volume to obtain a geometric encoding volume.
[0017] In one embodiment of the present invention, the three-dimensional regularization network uses the cost aggregation network of the model, that is, the features of different resolutions before are transformed into weights through a convolution and a Sigmoid function and embedded as attention into the encoder-decoder architecture of the cost aggregation network to guide the network to regularize the cost volume.
[0018] In one embodiment of the present invention, in order to increase the receptive field, one-dimensional average pooling is used to pool the full correlation cost volume and the geometric encoding volume, and finally a two-layer geometric encoding volume pyramid and a full correlation cost volume pyramid are obtained. Then the two are connected to obtain a combined geometric encoding volume.
[0019] In one embodiment of the present invention, in the step (3), the GRU iteratively updates the disparity part by taking regression to obtain the initial value and jointly updates the input disparity, specifically including:
[0020] To facilitate the backpropagation of the network and accelerate the inference speed, a GRU cell is implemented using convolution. The disparity map regressed from the previous geometric encoding volume is used as the initial disparity field, and three layers of GRU are used to iteratively update the disparity. The hidden layers of the three layers of GRU are initialized with multi-scale context features; in each iteration, the current disparity field is used to index the local geometric volume from the combined geometric encoding volume, and linear interpolation is used for sampling. Assuming the search radius is r, finally, a local geometric volume with a range of 2r + 1 is indexed. Then, these geometric volumes and the current disparity field pass through two encoding layers and are connected to the current disparity field to obtain the input x of the GRU. k , and then GRU is used to update the hidden state; based on the current hidden state, the present invention decodes a disparity field residual using two convolutional layers, and then uses the residual to update the current disparity field.
[0021] In one embodiment of the present invention, in the step (4), context feature-guided spatial upsampling is adopted, which specifically includes:
[0022] A complete-resolution disparity map is output through the weighted combination of the predicted disparity fields at 1 / 4 resolution. The hidden state is convolved to generate features, and then they are upsampled to 1 / 2 resolution. The upsampled features are connected to the 1 / 2 feature map of the left image to generate the weight W, and the full-resolution disparity map is output through the weighted combination of its coarse-resolution neighbors.
[0023] In one embodiment of the present invention, high-resolution context features are used to obtain weights in the step (4).
[0024] According to another aspect of the present invention, a stereo matching system based on iterative geometric encoding volume is further provided, including a feature extraction module, a joint geometric encoding volume construction network module, a GRU update network module, and a spatial upsampling network module, wherein:
[0025] The feature extraction module is used to input the left image and the right image into the feature extraction network, extract feature maps through multiple layers of convolution, and output multi-scale feature maps , whose sizes are 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image resolution respectively. The original right image is input into the context feature extraction network, and after a series of residual blocks and downsampling processes, multi-scale context feature maps are output, whose resolution sizes are 1 / 4, 1 / 8, and 1 / 16 of the original image, and the number of channels is 128;
[0026] The feature map with a size of 1 / 4 output by the feature extraction module is , , represents the position of the feature map, d represents the disparity value, Denote the number of feature channels, and group the overall features. Denote the number of groups, perform dot multiplication on the left and right features at the corresponding positions, and construct the group correlation body of [H, W, D] through grouped fusion:
[0027] , (1)
[0028] Input the obtained group correlation body into the 3D regularization network for cost aggregation. In the 3D convolution, use the left feature map to endow the cost body with weight information to supplement the only relevant information of the image features. For the cost body of size , use the corresponding left feature map to endow weight values: ,(2), perform cost body aggregation through feature map-guided 3D convolution, infer and propagate scene geometry information, and generate a geometric encoding body;
[0029] The joint geometric encoding body construction network module is used to calculate the full feature pair correlation based on the left and right feature maps, obtain the local feature correlation body, perform pooling on the disparity dimension with the above geometric encoding body respectively, form a two-level geometric encoding body pyramid and a full feature pair correlation body pyramid, and then fuse them into a joint geometric encoding body;
[0030] The GRU update network module is used to regress the initial disparity estimation map from the cost body output by the previous 3D regularization network , and use three-level GRU to iteratively update it. For the k-th iteration, denote the current disparity map, which, together with the context features output by the context feature extraction network and the geometric features , are used as the input of the GRU. Among them, is obtained by indexing the combined geometric encoding body through linear interpolation of the current disparity , and r is a manually specified search range. The calculation method of
[0031] , (3)
[0032] The three-level GRUs hidden state is determined by the context features, and the hidden state update method is as follows:
[0033]
[0034]
[0035] (4)
[0036]
[0037]
[0038] According to the hidden state Decode the residual through two convolutional layers , and then update the current disparity estimate to obtain an optimized disparity estimate ;
[0039] The spatial upsampling network module is used to output a full-resolution disparity map through the weighted combination of the iteratively updated disparity estimate map with a resolution of 1 / 4 by predicting the disparity of the weighted combination, generate features by convolving the hidden state and upsample it to a resolution of 1 / 2, and then combine it with the feature map extracted from the left image to obtain the required weights, and finally output the full-resolution disparity estimate through the weighted combination of the coarse-resolution neighborhood;
[0040] Constraining the predicted disparity result with the true label, the overall objective function consists of two parts. The first part calculates the smooth L1 loss on the initial disparity regressed from the GEV, and the second part calculates the L1 loss on all disparity predictions . Use an exponentially increasing weight sum, and the specific calculation method is as follows:
[0041] (5)
[0042] (6)
[0043] Among them, represents the true disparity, represents the 1-norm, which calculates the sum of the absolute values of each element in the vector *, takes an empirical value.
[0044] Compared with the prior art through the above technical solutions conceived by the present invention, the following beneficial effects are obtained:
[0045] (1) The present invention provides more comprehensive and concise information for GRU update, generates more effective optimization in each iteration, and thus significantly reduces the number of GRU iterations;
[0046] (2) The present invention regresses an initial disparity map from the geometric encoding volume through soft argmin, provides a more accurate starting point for the GRU-based update operation, and thus enables the iterative process to converge quickly;
[0047] (3) Based on IGEV, the present invention constructs IGEV - stereo for stereo matching tasks, achieving an EPE (Endpoint Error) of 0.47 on the Scene Flow dataset. Compared with the published related methods, it ranks first on the KITTI2015 leaderboard. In terms of inference speed, IGEV - Stereo is the fastest method among the top ten methods on the KITTI2015 list. In addition, compared with most existing stereo networks, IGEV - stereo exhibits better cross - dataset generalization ability. Description of the Drawings
[0048] Figure 1 It is a schematic diagram of the stereo matching IGEV - Stereo structure constructed based on the IGEV structure in the embodiment of the present invention;
[0049] Figure 2 It is the qualitative result of the KITTI test set in the embodiment of the present invention. The first two columns show the results of KITTI 2012, and the last two columns show the results of KITTI 2015. Detailed Embodiment
[0050] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and embodiments. The specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0051] The problem to be solved by the present invention is to implement a stereo matching system that can handle local ambiguity problems. Based on the per - pixel optical flow prediction network, this system aggregates the cost volume, enabling it to aggregate global and non - global information to jointly assist network prediction, so that more accurate optical flow estimates can be obtained for ambiguous parts. The present invention proposes an iterative geometric encoding volume structure to aggregate the cost volume constructed by features, effectively capturing global geometric structure features, combining non - local and local information, inferring and propagating geometric information, and using surrounding pixels to assist in estimating the disparity of the ambiguous region. The accuracy of disparity estimation ranks first among all evaluation methods on 2015 KITTI at present and has strong cross - dataset generalization ability and high inference efficiency.
[0052] The present invention provides a stereo matching deep neural network method based on an iterative geometric encoding volume (IGEV, Iterative Geometry Encoding Volume), including:
[0053] (1)Take the left and right image pairs, and call one of the images the source image and the other the reference image for calculating the disparity estimation between the reference image and the source image; use a multi-scale feature extraction network to extract the image features of the source image and the reference image, and at the same time use different networks to extract the multi-scale features of the reference image as context features;
[0054] Specifically, the use of the multi-scale feature extraction network to extract the image features of the source image and the reference image specifically includes: using a pre-trained mobilenetv2_100 feature extraction network backbone to extract features, and the feature extraction network backbone is an encoder-decoder architecture. The encoder uses the mobilenetv2_100 model pre-trained on ImageNet, and extracts the input image into feature maps with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 through convolutional layers and bottleneck layers. In the decoder, use a transposed convolutional layer to upsample the feature map. For example, upsample the 1 / 32 feature map by transposed convolution, then connect it with the 1 / 16 feature map, and then perform feature fusion with a 3x3 convolutional kernel, and the output is the fused 1 / 16 feature map. Similarly, the feature extraction network finally obtains a series of fused 1 / 4, 1 / 8, 1 / 16, and 1 / 32 feature maps.
[0055] In addition, since the present invention uses a Gated Recurrent Unit (GRU) module, a context feature extraction network branch is required to extract the context features of the image. The context feature extraction network consists of a series of residual blocks and downsampling layers, takes the reference image (the left image for stereo matching tasks and the reference image for multi-view stereo matching) as the input, and finally outputs multi-scale context features with resolutions of 1 / 4, 1 / 8, and 1 / 16 and a feature dimension of 128. This multi-scale context feature is used to initialize the hidden layer of the GRU module and is inserted into the GRU module in each iteration.
[0056] (2)Select feature map pairs of the target size to calculate the correlation to construct a correlation volume, input it into a three-dimensional regularization network, introduce the reference image feature to guide the cost volume aggregation, perform the inference and propagation of the scene geometric information, and generate a geometric encoding volume; and calculate the correlation of all feature map pairs to form a local feature correlation volume, and combine the geometric encoding volume with the local feature correlation volume to form a combined geometric encoding volume;
[0057] Among them, the combination of the geometric encoding volume and the local feature correlation volume to form a combined geometric encoding volume specifically includes:
[0058] The 1 / 4 feature map obtained by the feature extraction network is used to calculate a group correlation volume of H x W x D x G using the group correlation calculation formula, where H represents height, W represents width, D represents the disparity dimension, and G represents the channel dimension. Then, a lightweight 3D regularization network is used to regularize this volume to obtain a geometric encoding volume. Note that in theory, any cost aggregation network can be used for this 3D regularization network. In the present invention, the cost aggregation network of the model is to convert the features of different resolutions before into weights through a convolution and a sigmoid function, and embed them as attention into the encoder-decoder architecture of the cost aggregation network to guide the network to regularize the cost volume. If a general cost aggregation network is used, then the previous feature extraction part does not need to output multi-scale feature maps, and only the 1 / 4 feature map is required.
[0059] From the regularized geometric encoding volume, an initial disparity map can be obtained through disparity regression, which serves as the intermediate supervision and the initial disparity field for the subsequent GRU. However, this geometric encoding volume cannot well obtain long-distance disparity information because the cost aggregation network propagates disparity information through convolution. Therefore, the present invention also constructs an H x W x W full correlation cost volume to provide long-distance disparity information by imitating RAFT-Stereo.
[0060] To increase the receptive field, the present invention uses one-dimensional average pooling to pool the full correlation cost volume and the geometric encoding volume, and finally obtains two layers of geometric encoding volume pyramids and full correlation cost volume pyramids. Then, the two are connected to obtain a combined geometric encoding volume.
[0061] (3) The disparity estimate is regressed from the aforementioned geometric encoding volume as the initial value for iterative update. Starting from the initial disparity estimate, the disparity estimate result of the previous level, the context features, and the aforementioned combined geometric encoding volume are used as inputs, and a three-level GRU is used to iteratively update the disparity to obtain a more accurate disparity prediction;
[0062] The part of the GRU iteratively updating the disparity adopts regression to obtain the initial value and joint input disparity update, specifically including:
[0063] To facilitate the backpropagation of the network and accelerate the inference speed, the present invention uses convolution to implement the GRU unit. The disparity map regressed from the previous geometric encoding volume serves as the initial disparity field, and three layers of GRU are used to iteratively update the disparity (the same as RAFT-Stereo). The hidden layers of the three layers of GRU are initialized with multi-scale context features.
[0064] In each iteration, the present invention uses the current disparity field to index local geometric volumes from the combined geometric encoded volumes, samples using linear interpolation, assumes a search radius of r, and finally indexes a local geometric volume with a range of 2r + 1. Then these geometric volumes and the current disparity field pass through two encoding layers and are concatenated with the current disparity field to obtain the input x of the GRU. k , and then uses the GRU to update the hidden state. Based on the current hidden state, the present invention decodes a disparity field residual using two convolutional layers, and then updates the current disparity field with the residual.
[0065] (4) Upsample the predicted disparity spatially, and use weighted combination to output a full-resolution disparity map, where weights are obtained using high-resolution context features.
[0066] Context feature-guided spatial upsampling is adopted, specifically including:
[0067] We output a full-resolution disparity map through the weighted combination of the 1 / 4-resolution predicted disparity field. Different from RAFT-Stereo, we use higher-resolution context features to obtain weights. We convolve the hidden state to generate features, and then upsample them to 1 / 2 resolution. The upsampled features are concatenated with the 1 / 2 feature map of the left image to generate the weight W. We output the full-resolution disparity map through the weighted combination of its coarse-resolution neighbors.
[0068] Further, as Figure 2 shown, the present invention provides a stereo matching system based on iterative geometric encoded volumes, which mainly has four components: a feature extraction module, a joint geometric encoded volume construction network module, a GRU update network module, and a spatial upsampling network module. The specific implementation is as follows:
[0069] (1) Feature extraction: Input the left and right images into the feature extraction network, extract feature maps through multi-layer convolution, and output multi-scale feature maps , whose sizes are 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image resolution. Input the original right image into the context feature extraction network, and after a series of residual blocks and downsampling processes, output multi-scale context feature maps with a resolution of 1 / 4, 1 / 8, and 1 / 16 of the original image and 128 channels.
[0070] The 1 / 4-sized feature map output by the feature extraction module is , , represents the position of the feature map, d represents the disparity value, represents the number of feature channels, group the overall features, Indicates the number of groups, multiplies the left and right features at the corresponding positions, and constructs a group-related volume of [H, W, D] through grouped fusion:
[0071] , (1)
[0072] Input the obtained group-related volume into a 3D regularization network for cost aggregation. In 3D convolution, use the left feature map to endow the cost volume with weight information to supplement the only relevant information of the image features. For the cost volume of size , , use the corresponding left feature map to endow weight values:
[0073] , (2)
[0074] Conduct cost volume aggregation through feature map-guided 3D convolution, effectively infer and propagate scene geometry information, and generate a geometric encoding volume.
[0075] (2) Joint geometric encoding volume construction network: Calculate the full feature pair correlation of the left and right feature maps to obtain a local feature correlation volume. Pool the disparity dimension with the above geometric encoding volume respectively to form a two-level geometric encoding volume pyramid and a full feature pair correlation volume pyramid, and then fuse them into a joint geometric encoding volume (CGEV, Combined Geometry EncodingVolume).
[0076] (3) GRU update network: Regress the initial disparity estimation map from the cost volume output by the aforementioned 3D regularization network, and use three-level GRU to iteratively update it to finally obtain a more accurate disparity estimation. For the k-th iteration, represents the current disparity map (i.e., the prediction result of the previous iteration), and is used as the input of GRU together with the context feature output by the context feature extraction network and the geometric feature . Among them, is obtained by indexing the combined geometric encoding volume through linear interpolation of the current disparity , r is a manually specified search range, and the calculation method of is:
[0077] , (3)
[0078] The hidden state of the three-level GRUs is determined by the context feature, and the hidden state update method is as follows:
[0079]
[0080]
[0081] (4)
[0082]
[0083]
[0084] According to the hidden state Decode the residual through two convolutional layers , and then update the current disparity estimate to obtain an optimized disparity estimate .
[0085] (4) Spatial upsampling network: Upsample the iteratively updated disparity estimate map with a final resolution of 1 / 4 through predicted disparities to output a full-resolution disparity map through a weighted combination. Convolve the hidden state to generate features and upsample them to a resolution of 1 / 2, and then combine them with the feature map extracted from the left image to obtain the required weights, and finally output the full-resolution disparity estimate through a weighted combination of the coarse-resolution neighborhood
[0086] Constrain the predicted disparity result with the ground truth label. The overall objective function consists of two parts. The first part calculates the smooth L1 loss on the initial disparity regressed from GEV, and the second part calculates the L1 loss on all disparity predictions . Use an exponentially increasing weighted sum, and the specific calculation method is as follows
[0087] (5)
[0088] (6)
[0089] Among them, represents the ground truth disparity, represents the 1-norm, which calculates the sum of the absolute values of the elements in the vector *, takes an empirical value (for example, 0.9).
[0090] Table 1 shows the quantitative evaluation of the method of the present invention and other methods in KITTI 2012 and KITTI 2015 in the embodiments of the present invention. It can be seen from Table 1 that the method of the present invention is superior to other published methods in almost all indicators of KITTI 2012 and 2015. On KITTI 2012, the Out-Noc of IGEVStereo at the 2-pixel error threshold is 10.00% and 10.93% higher than that of lestereo and RAFT-Stereo respectively. On KITTI 2015, IGEV-Stereo exceeds CRE Stereo and RAFT-Stereo by 5.92% and 12.64% respectively in the D1-all index. Compared with other iterative methods such as CRE Stereo and RAFT-Stereo, IGEV-Stereo not only performs better in terms of performance, but also is more than twice as fast.
[0091] Table 1
[0092]
[0093] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A stereo matching method based on an iterative geometric coding volume, characterized in that, The method includes the following steps: (1) Take a pair of left and right images. One image is called the source image, and the other image is called the reference image, which is used to calculate the disparity estimation between the reference image and the source image. Use a multi-scale feature extraction network to extract the image features of the source image and the reference image. At the same time, use different networks to extract the multi-scale features of the reference image as context features; (2) Select a pair of feature maps of the target size, group them to calculate the correlation to construct a correlation volume, input it into a three-dimensional regularization network, introduce the reference image features to guide the cost volume aggregation, infer and propagate the scene geometric information, and generate a geometric encoding volume; and calculate the correlation of all pairs of feature maps to form a local feature correlation volume, and combine the geometric encoding volume with the local feature correlation volume to form a combined geometric encoding volume; (3) Regress the disparity estimation from the aforementioned geometric encoding volume as the initial value of iterative update. Starting from the initial disparity estimation, use the disparity estimation result of the previous level, the context features, and the aforementioned combined geometric encoding volume as inputs, and use a three-level GRU to iteratively update the disparity to obtain a more accurate disparity prediction; In the step (3), the part of the GRU iteratively updating the disparity adopts regression to obtain the initial value and joint input disparity update, which specifically includes: To facilitate the backpropagation of the network and accelerate the inference speed, use convolution to implement the GRU unit. The disparity map regressed from the previous geometric encoding volume is used as the initial disparity field. Use three layers of GRU to iteratively update the disparity. The hidden layers of the three layers of GRU are initialized with multi-scale context features; in each iteration, use the current disparity field to index the local geometric volume from the combined geometric encoding volume, and perform sampling using linear interpolation. Assume the search radius is r, and finally index to obtain a local geometric volume with a range of 2r + 1. Then these geometric volumes and the current disparity field pass through two encoding layers and are connected to the current disparity field to obtain the input xk of the GRU, and then use the GRU to update the hidden state; based on the current hidden state, use two convolutional layers to decode a disparity field residual, and then use the residual to update the current disparity field; (4) Upsample the predicted disparity spatially and use weighted combination to output a full-resolution disparity map.
2. The stereo matching method based on an iterative geometric coding body according to claim 1, wherein In the step (1), the use of the multi-scale feature extraction network to extract the image features of the source image and the reference image specifically includes: Use the pre-trained mobilenetv2_100 feature extraction network backbone to extract features. The feature extraction network backbone is an encoder-decoder architecture. The encoder uses the mobilenetv2_100 model pre-trained on ImageNet to extract the input image into feature maps with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 through convolutional layers and bottleneck layers. The decoder uses deconvolutional layers to upsample the feature maps and then uses a 3x3 convolutional kernel for feature fusion. The feature extraction network finally obtains a series of fused 1 / 4, 1 / 8, 1 / 16, and 1 / 32 feature maps.
3. The stereoscopic matching method based on an iterative geometric coding body according to claim 2, wherein In step (1), context features of the image are extracted through a context feature extraction network branch. The context feature extraction network consists of a series of residual blocks and downsampling layers. Taking the reference image as the input, it finally outputs multi-scale context features with resolutions of 1 / 4, 1 / 8, and 1 / 16 and a feature dimension of 128. This multi-scale context feature is used to initialize the hidden layer of the GRU module and is inserted into the GRU module in each iteration.
4. The stereoscopic matching method based on an iterative geometric coding body according to claim 1 or 2, characterized in that, In step (1), the geometric encoding volume is combined with the local feature correlation volume to form a combined geometric encoding volume, specifically: For the 1 / 4 feature map obtained through the feature extraction network, a group correlation volume of H x W x D x G is calculated using the group correlation calculation formula, where H represents height, W represents width, D represents the disparity dimension, and G represents the channel dimension; then a lightweight 3D regularization network is used to regularize this volume to obtain a geometric encoding volume.
5. The stereo matching method based on an iterative geometric coding body according to claim 4, wherein The 3D regularization network uses the cost aggregation network of the Coex model, that is, the features of different resolutions before are transformed into weights through a convolution and a Sigmoid function and used as attention to be embedded in the encoder-decoder architecture of the cost aggregation network to guide the network to regularize the cost volume.
6. The stereo matching method based on an iterative geometric coding body according to claim 4, wherein To increase the receptive field, one-dimensional average pooling is used to pool the full correlation cost volume and the geometric encoding volume, and finally a two-layer geometric encoding volume pyramid and a full correlation cost volume pyramid are obtained. Then the two are connected to obtain a combined geometric encoding volume.
7. The stereo matching method based on an iterative geometric coding body according to claim 1 or 2, characterized in that In step (4), context feature-guided spatial upsampling is adopted, specifically including: A complete-resolution disparity map is output through the weighted combination of the predicted disparity fields at 1 / 4 resolution. The hidden state is convolved to generate features, and then they are upsampled to 1 / 2 resolution. The upsampled features are connected to the 1 / 2 feature map of the left image to generate weight W, and the full-resolution disparity map is output through the weighted combination of its coarse-resolution neighbors.
8. The stereo matching method based on an iterative geometric coding body according to claim 1 or 2, characterized in that, In step (4), high-resolution context features are used to obtain weights.
9. A stereo matching system based on an iterative geometric coding body, characterized in that, It includes a feature extraction module, a combined geometric encoding volume construction network module, a GRU update network module, and a spatial upsampling network module, where: The feature extraction module is used to input the left image and the right image into the feature extraction network, extract the feature map through multiple layers of convolution, and output multi-scale feature maps , with sizes of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image resolution respectively. The original right image is input into the context feature extraction network, and after a series of residual blocks and downsampling processes, multi-scale context feature maps are output, with a resolution size of 1 / 4, 1 / 8, 1 / 16 of the original image, and the number of channels is 128; The feature map with a size of 1 / 4 output by the feature extraction module is f l,4 , f r,4 , where x and y represent the positions of the feature map, d represents the disparity value, and N c represents the number of feature channels. The overall features are grouped, and N g represents the number of groups. The left and right features at the corresponding positions are multiplied point by point, and grouped fusion is used to construct a group correlation body of [H, W, D]: , (1) The obtained group-related body is input into a three-dimensional regularization network for cost aggregation. In the three-dimensional convolution, the left feature map is used to endow the cost volume with weight information to supplement the only correlation information of the image features. For the cost volume of size , , the corresponding left feature map f l,i is used to endow weight values: , (2), through the feature map-guided three-dimensional convolution for cost volume aggregation, infer and propagate the scene geometry information, and generate a geometric coding volume; The combined geometric encoding volume construction network module is used to calculate the full feature pair correlation based on the left and right feature maps to obtain the local feature correlation volume, and the above geometric encoding volume and the local feature correlation volume are respectively pooled in the disparity dimension to form a two-level geometric encoding volume pyramid and a full feature pair correlation volume pyramid, and then they are fused into a combined geometric encoding volume; The GRU update network module is used to regress the initial disparity estimation map d0 from the cost volume output by the aforementioned three-dimensional regularization network through softargmin, and update it iteratively using a three-level GRU. For the k-th iteration, d k represents the current disparity map, and together with the context feature c k , c r , c h and the geometric feature G f are used as the inputs of the GRU, where G f is obtained by indexing the combined geometric coding volume through linear interpolation using the current disparity d k , r is a manually specified search range, and the calculation method of G f is as follows: , (3) The hidden state h0 of the three-level GRUs is determined by the context features, and the hidden state update method is as follows: (4) According to the hidden state h k Decode the residual Δd through two convolutional layers k , and then update the current disparity estimate to obtain the optimized disparity estimate d k+1 ; The spatial upsampling network module is used to output a full-resolution disparity map through the weighted combination of the disparity estimation map with a 1 / 4 resolution finally obtained by iterative update, and predict the disparity d k of the weighted combination, generate features by convolving the hidden state h k and upsample it to a 1 / 2 resolution, then combine it with the feature map f extracted from the left image l,2 to obtain the required weights, and finally output the full-resolution disparity estimation through the weighted combination of the coarse-resolution neighborhood; Constraining the predicted disparity results with the ground truth labels, the overall objective function consists of two parts. The first part calculates the smooth L1 loss on the initial disparity d0 regressed from the GEV, and the second part calculates the L1 loss on all disparity predictions using an exponentially increasing weighted sum, and the specific calculation method is as follows: (5) (6) Among them, represents the true parallax, represents the 1-norm, which calculates the sum of the absolute values of each element in the vector *, takes an empirical value.
Citation Information
Patent Citations
Binocular stereo matching method based on convolutional neural network
CN110533712A
Disparity estimation optimization method based on up-sampling and accurate rematching
WO2021138992A1