Stereo Matching Method Based on Dual Spatial Pooling Pyramid

By constructing a dual-space pooled pyramid network for feature extraction and disparity regression, the problem of poor matching effect of the stereo matching method on semantic edges and texture complex areas is solved, and the accuracy and calculation efficiency of the disparity map are improved.

CN115375746BActive Publication Date: 2025-07-08XIDIAN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210336322.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-07-08
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

The existing stereo matching method has poor matching effect with complex texture areas at semantic edges, and there are outliers, resulting in low parallax map accuracy and high computational complexity.

Method used

A three-dimensional matching method based on dual space pooled pyramids is adopted. By constructing a dual pyramid network connected in parallel with the average pooled pyramid and the largest pooled pyramid, combining the initial convolutional neural network, feature fusion network and three-dimensional convolutional neural network, feature extraction and parallax regression are performed to reduce outliers and improve matching accuracy.

Benefits of technology

Improve the accuracy of stereo matching results at semantic edges and texture complex areas, reduce outliers, and achieve more efficient disparity map output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115375746B_ABST
    Figure CN115375746B_ABST
Patent Text Reader

Abstract

The present invention discloses a stereo matching method based on a dual spatial pyramid, which mainly solves the problems of missing multi-scale semantic information and inability to fully extract the texture information and overall information of features in the existing binocular stereo matching method. The implementation scheme is as follows: obtain a binocular stereo matching training sample set T and a validation sample set V; respectively establish a feature extraction network E, a cost aggregation network R, and a disparity regression network G, and cascade them to form a binocular stereo matching network S; use the training sample set T to train the binocular stereo matching network S by the stochastic gradient descent method; input the validation sample set V into the trained binocular stereo matching network S t to obtain the stereo matching result I i pred . The present invention improves the matching accuracy of the stereo matching method at semantic edges and in complex texture regions, effectively reduces the outliers in the matching results, and effectively improves the accuracy of the stereo matching results, and can be used for disparity estimation of images and videos obtained by binocular cameras.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This method belongs to the technical field of computer vision, and further relates to a stereo matching method, which can be used for disparity estimation of images and videos obtained by a binocular camera. Background Art

[0002] Binocular vision is an important direction in the field of computer vision, and the stereo matching task therein is one of the most important subtasks in binocular vision. Stereo matching refers to rectifying the epipolar lines of the two-dimensional binocular images obtained by a binocular camera, and obtaining the disparity value and depth information in a three-dimensional scene by performing feature point matching on the rectified two-dimensional binocular images. Its basic principle is to find the corresponding points of the binocular images in the left and right eye images, and obtain the absolute value of the difference in the relative abscissas of the corresponding points in the left and right eye images, then the disparity value can be obtained. Using the disparity value, the focal length of the camera, and the baseline distance of the binocular camera, the three-dimensional depth information can be further obtained. The stereo matching method is based on the spatial geometric structure and has advantages such as fast solving speed and low hardware cost.

[0003] Traditional stereo matching methods are divided into two categories: global methods and semi-global methods. The global method constructs a global energy function based on the smoothness assumption and uses an optimization method to solve the minimum value point of the energy function. This method has a large amount of computation, a long time consumption, and poor real-time performance. The semi-global method uses local information, calculates the total matching cost within a matching window of a specific size, and uses the winner-takes-all (WTA) strategy to find the minimum value to finally obtain the disparity value. This method is greatly affected by the window size. If the window is too small, it will cause the inability to include all the texture information of the object, resulting in ambiguity. If the window is too large, it will produce a dilation effect in the depth discontinuous area and increase the amount of calculation.

[0004] With the development of computer technology, deep learning in the field of artificial intelligence has developed rapidly. Compared with traditional stereo matching methods, the stereo matching method based on deep learning has stronger feature expression ability and feature learning ability, and has the advantage of being parallelizable. The stereo matching method using deep learning can complete the establishment of the stereo matching network through the construction of a convolutional neural network, continuous training, and gradient update based on backpropagation, and realize the function of completing stereo matching at one time.

[0005] Beijing Institute of Technology proposed an "Approximate Nearest Neighbor Propagation Stereo Matching Method Based on Image Pyramid Distance Metric" in the patent document with the application number 202110550638.2 and the publication number 113177565A. The implementation steps are as follows: 1) Extract features from the stereo image pair composed of the left and right images, and pair them using distance metric; 2) Continuously downsample the image by a factor of two, and perform matching on the downsampled image; 3) Select the correctly matched pixel pairs and propagate the matching results to adjacent pixels. This method uses the constraints constructed by feature matching to find matching points. However, since it only constructs an image pyramid and obtains features for the input left and right eye image pairs, it cannot fully extract deeper semantic, texture, and other information of the image, resulting in poor accuracy of the obtained matching results.

[0006] Shandong University proposed a "Stereo Matching Feature Extraction and Post-Processing Method for Disparity Map" in the patent document with the application number 202110281313.9 and the publication number 112991420A. The implementation steps are as follows: 1) Use a convolutional kernel to extract stereo matching features and use the convolutional kernel pyramid feature extraction method; 2) Use the 4-path cost aggregation method to aggregate the features output by the convolutional network; 3) Use the internal left-right consistency check to perform left-right consistency detection on the disparity map and fill in the invalid disparity points to obtain the final disparity map. The deficiencies of this method are as follows: The convolutional neural network is only used in the feature extraction part, and an end-to-end network is not constructed for learning and training, which is not conducive to the algorithm's further learning of disparity and reducing the training difficulty of the model; at the same time, the cost aggregation algorithm it uses has a high complexity and is not easy to perform parallel processing and acceleration, resulting in a low-precision and time-consuming final output disparity map.

[0007] Beijing Institute of Technology disclosed an "Binocular Vision Position Measurement System and Method Based on Deep Learning" in the patent application document with the application number 202110550638.2 and the publication number 113177565A for binocular vision position measurement. The system includes a system of image capture, deep learning object recognition, image segmentation module, fitting module, and binocular point cloud module. The implementation steps of the method are as follows: 1) Use a convolutional neural network to build a disparity calculation sub-module to obtain the disparity map on the binocular left camera; 2) Use the point cloud calculation sub-module to obtain the three-dimensional point cloud information of the scene through the internal and external parameters of the binocular camera for the obtained disparity map. Since the semi-global matching method is used in disparity aggregation in this method, the obtained disparity map is prone to many outliers, resulting in more noise in the final output.

[0008] Jiaren Chang et al. proposed an end-to-end binocular stereo matching network in the paper "Pyramid Stereo Matching Network" (Conference on Computer Vision and Pattern Recognition, 2018). This network first uses a convolutional neural network with a spatial pooling pyramid to extract features from binocular images. Then, it constructs a four-dimensional convolutional cost volume and uses a stacked three-dimensional convolutional neural network to learn and aggregate the cost volume, finally obtaining a disparity map. Since the feature extraction network only uses a spatial average pooling pyramid to extract average features at multiple scales, it cannot retain the edge information and texture information in the original image and feature map, easily resulting in the failure of the final left and right image matching, and poor matching effects at semantic edges and in regions with complex textures. Summary of the Invention

[0009] The object of the present invention is to overcome the deficiencies of the above-mentioned prior art, and propose a stereo matching method based on a dual spatial pooling pyramid to fully extract the feature information of the image, improve the matching effect at semantic edges and in regions with complex textures, reduce the outliers in the matching results at the same time, improve the accuracy of the final stereo matching, and obtain a more accurate disparity map.

[0010] The technical key of the present invention is to use a dual pooling pyramid structure in the feature extraction module of the end-to-end binocular stereo matching network to improve the accuracy of the finally output disparity map. Its implementation scheme includes the following:

[0011] (1) Obtain the binocular stereo matching training sample set T and the validation sample set V

[0012] Obtain N + M pairs of images I from the existing binocular stereo matching dataset i ={I i L , I i R , I i GT}, where I i L and I i R are the left-eye image and the right-eye image of the i-th image pair respectively, and I i GT is the true disparity image of the i-th image pair;

[0013] Take N image pairs in the dataset as the training sample set T = {I 1 , I 2 , …, I j , …, I N}, 1 ≤ j ≤ N; M image pairs as the validation set V of the training result = {I N+1 , I N+2 , …, I N+k , …, I N+M},1 ≤ k ≤ M, N ≥ 180, M ≥ 20;

[0014] (2) Construct a binocular stereo matching network S based on a dual - space pooling pyramid

[0015] (2a) Establish a dual - pyramid pooling network E2 composed of an average - pooling pyramid E2 AVG and a max - pooling pyramid E2 MAX in parallel;

[0016] (2b) Establish a feature extraction network E composed of a cascaded initial convolutional neural network E1, dual - pooling pyramid network E2, and feature fusion network E3;

[0017] (2c) Establish a cost aggregation network R composed of three cascaded 3D convolutional neural networks;

[0018] (2d) Cascade the feature extraction network E, cost aggregation network R, and disparity regression network G to form a binocular stereo matching network S;

[0019] (3) Train the binocular stereo matching network S:

[0020] (3a) Set the learning rate η and the maximum number of training times t max , and the current iteration period t' = 0;

[0021] (3b) Use the forward - propagation method for the left - eye image I i L in the training set T and the right - eye image I i R to obtain the predicted disparity map I i pred output by the network;

[0022] (3c) Use the SmoothL1 loss function to calculate the error L i pred between the predicted disparity map I i GT and the true disparity map I i , and use the gradient - descent algorithm with back - propagation to update the convolutional kernel weight parameter ω t in the current network S t and the convolutional kernel bias parameter b t ;

[0023] (3d)Increment the current iteration cycle by 1, i.e., t' = t' + 1, and determine whether the maximum number of iterations is reached:

[0024] If t' = t max , then the maximum number of iterations is reached, and at this time the training ends, and the trained binocular stereo matching network S is obtained t ,

[0025] If t' < t max , then return to (3b);

[0026] (4) Input the validation sample set V into the trained stereo matching network S of the dual spatial pooling pyramid, and successively pass through the feature extraction network E, the cost aggregation network R, and the disparity regression network G therein to obtain the predicted result I of the disparity map i pred , and this predicted result is the result of stereo matching.

[0027] The present invention has the following advantages compared with the existing technologies:

[0028] 1. Since the present invention uses the dual pyramid pooling network E2, it provides multi-scale features with rich texture information and global semantic information for the binocular stereo matching network, and improves the stereo matching accuracy at the semantic edges, in the regions with complex textures and in the regions with missing textures of the stereo matching result.

[0029] 2. Since the present invention uses the feature extraction network E, it realizes the extraction and further fusion of the features of the left-eye and right-eye images, improves the accuracy of the final stereo matching result, and effectively reduces the false outliers in the matching result.

[0030] 3. Since the present invention uses the cost aggregation network R and the disparity regression network G, it realizes the matching of binocular images and completes the output of the predicted disparity map. Description of the Drawings

[0031] Figure 1 is the implementation flowchart of the present invention;

[0032] Figure 2 is the schematic diagram of the pooling pyramid structure in the present invention;

[0033] Figure 3 is the output disparity map after performing stereo matching on the left-eye and right-eye images by simulation using the present invention. Detailed Embodiment

[0034] The following further describes the embodiments and effects of the present invention in detail with reference to the drawings.

[0035] Refer to Figure 1 , the implementation steps of this example are as follows:

[0036] Step 1. Obtain the binocular stereo matching training sample set T and the validation sample set V.

[0037] Obtain the KITTI dataset, which was established by Geiger et al. from the Karlsruhe Institute of Technology. It uses a car equipped with two high-definition color cameras, a Velodyne lidar, and a GPS global positioning system to capture data in three scenarios: urban, rural, and highway. The dataset provides information such as binocular images, disparity maps, and optical flow maps, which can be used for tasks such as stereo matching, optical flow prediction, 3D object detection, and 3D tracking. It is divided into two sub-datasets: KITTI2012 and KITTI2015. The KITTI2012 sub-dataset provides 194 training image groups containing binocular images and true disparity values, and 195 validation image groups containing only binocular images with the true disparity values not publicly available; KITTI2015 expands on this by adding 200 training image groups and 200 validation image groups.

[0038] Obtain N + M image pairs I with a resolution of W × H from the KITTI binocular stereo matching dataset i ={I i L ,I i R ,I i GT}, where I i L , I i R are the left-eye image and the right-eye image of the i-th image pair respectively, and I i GT is the true disparity image of the i-th image pair. In this embodiment, W = 512 and H = 256.

[0039] Take N = 394 image pairs in the dataset as the training sample set T = {I 1 , I 2 , …, I j , …, I N}, 1 ≤ j ≤ N;

[0040] Take M = 200 image pairs as the validation set V of the training results = {I N+1 , I N+2 , …, I N+k , …, I N+M}, 1 ≤ k ≤ M, N ≥ 180, M ≥ 20.

[0041] Take a total of 394 training image groups from KITTI2012 and KITTI2015 as the training image set T,

[0042] Use 200 validation images in KITTI2012 as the validation set V1.

[0043] Use 200 validation images in KITTI2015 as the validation set V2.

[0044] Step 2. Construct a stereo matching network S based on a dual spatial pooling pyramid.

[0045] 2.1) Refer to Figure 2 to establish a dual pyramid network E2 composed of an average pooling pyramid E2 AVG and a max pooling pyramid E2 MAX in parallel, where:

[0046] The average pooling pyramid E2 AVG contains four average pooling kernels with sizes of 4×4, 8×8, 16×16, and 64×64 respectively, and the number of each type of pooling kernel is 32;

[0047] The max pooling pyramid E2 MAX contains four max pooling kernels with sizes of 4×4, 8×8, 16×16, and 64×64 respectively, and the number of each type of pooling kernel is 32;

[0048] 2.2) Establish a feature extraction network E composed of a cascaded initial convolutional neural network E1, a dual pooling pyramid network E2, and a feature fusion network E3, where:

[0049] The initial convolutional neural network E1 contains three sequentially connected residual units. Each residual unit includes two groups of convolutional layers. The first group of convolutional layers contains two groups of 32 convolutional kernels with a size of 3×3 and a stride of 2 connected in sequence; the second group of convolutional layers contains three groups of 32 convolutional kernels with a size of 3×3 and a stride of 1 connected in sequence;

[0050] The feature fusion network E3 contains four groups of parallel convolutional layers. Each group of convolutional layers includes two layers of convolutional kernels connected in sequence. The size of each layer of convolutional kernel is 3×3, the stride is 1, and the number is 32;

[0051] 2.3) Establish a cost aggregation network R composed of three stacked 3D convolutional neural networks in cascade. Each 3D convolutional neural network is composed of three groups of convolutional layers connected in sequence. Each group of convolutional layers contains 32 3D convolutional kernels with a size of 3×3 and a stride of 1;

[0052] 2.4) Establish a disparity regression network G composed of two groups of 3D convolutional layers connected in sequence. The first group contains 32 3D convolutional kernels with a size of 3×3 and a stride of 1, and the second group contains 1 3D convolutional kernel with a size of 3×3 and a stride of 1;

[0053] 2.5) Cascade the feature extraction network E, the cost aggregation network R, and the disparity regression network G to form the binocular stereo matching network S.

[0054] Step 3. Train the binocular stereo matching network S based on the dual pooling pyramid.

[0055] 3.1) Set the learning rate η and the maximum number of training times t max , the current iteration cycle t' = 0. In this embodiment, η = 0.001 and t max = 500;

[0056] 3.2) For the left-eye image I in the training set T i L and the right-eye image I i R Use the forward propagation method to obtain the predicted disparity map I output by the network i pred :

[0057] 3.2a) Use the initial convolutional neural network E1 to perform feature extraction on the left image I i L and the right image I i R respectively, and obtain the left image feature map F i L and the right image feature map F i R ;

[0058] 3.2b) Use the dual pooling pyramid network E2 to perform pooling operations on the left and right eye feature maps F i L , F i R to obtain the left-eye average pooling feature map group F i L_AVG , the left-eye max pooling feature map group F i L_MAX , and the right-eye average pooling feature map group F i R_AVG , the right-eye max pooling feature map group F i R_MAX , where:

[0059] In the average spatial pooling pyramid E2 AVG Use four average pooling kernels of different sizes to perform average pooling operations on the left and right image feature maps to extract features containing semantic information, and obtain the left-eye average pooling feature map group F i L_AVG , the right-eye average pooling feature map group F i R_AVG :

[0060] F i L_AVG = {F i L_avg1 , F i L_avg2 , F i L_avg3 , F i L_avg4}

[0061] F i R_AVG = {F i R_avg1 , F i R_avg2 , F i R_avg3 , F i R_avg4}

[0062] Among them, F i L_avgi and F i R_avgi are the outputs of the i-th average pooling kernel, where 1 ≤ i ≤ 4;

[0063] In the maximum spatial pooling pyramid E2 MAX , maximum pooling operations are performed on the left image feature map and the right image feature map using four maximum pooling kernels of different sizes to obtain features with texture and edge information, resulting in the left-eye image maximum pooling feature map group F i L_MAX and the right-eye image maximum pooling feature map group F i R_MAX :

[0064] F i L_MAX = {F i L_max1 , F i L_max2 , F i L_max3 , F i L_max4}

[0065] F i R_MAX = {F i R_max1 , F i R_max2 , F i R_max3 , F i R_max4}

[0066] Among them, F i L_maxi and Fi R_maxi is the output of the i-th maximum pooling kernel, where 1 ≤ i ≤ 4;

[0067] 3.2c) Stack the feature maps obtained from the two pooling pyramids and pass them through the feature fusion network E3 to output the left image feature map and the right image multi-scale feature map

[0068] 3.2d) Set the disparity search range D, where D = 192 in this embodiment. First, construct the cost volume for the current disparity value d in a concatenated manner for each disparity value d to obtain a three-dimensional matching cost volume group C i ={C i 1, C i 2, …, C i d , …, C i D}, and then concatenate its d three-dimensional matching cost volumes C i d in sequence to construct a four-dimensional matching cost volume

[0069] 3.2e) Use the cost aggregation network R to perform convolution operations on the four-dimensional matching cost volume C i d to obtain the aggregated cost volume

[0070] 3.2f) Use the disparity regression network G to perform convolution and channel compression operations on the aggregated cost volume in sequence, and use bilinear interpolation to restore its size to obtain the disparity three-dimensional cost volume C i pred ;

[0071] 3.2g) Perform a soft argmin operation on the three-dimensional cost volume C i pred to obtain the predicted disparity map I i pred , and its formula is as follows:

[0072]

[0073] 3.3) Use the SmoothL1 loss function to calculate the error L i pred between the predicted disparity map I i GT and the ground truth disparity map I i , and its calculation formula is as follows:

[0074]

[0075] 3.4) Use the gradient descent algorithm with backpropagation to update the convolutional kernel weight parameter ω t in the current network S t and the convolutional kernel bias parameter b t as follows:

[0076]

[0077]

[0078] where ω t+1 and b t+1 represent the updated convolutional kernel weight parameter and the updated convolutional kernel bias parameter respectively.

[0079] 3.4) Increment the current iteration count by 1, i.e., t' = t' + 1, and determine whether the maximum number of iterations has been reached:

[0080] If t' = t max , then the maximum number of iterations has been reached, and at this point, the training ends, obtaining the trained binocular stereo matching network S t ,

[0081] If t' < t max , then return to 3.2);

[0082] Step 4. Use the trained binocular stereo matching network to obtain the results of the binocular stereo matching validation set.

[0083] Input the validation sample set V into the trained stereo matching network S of the dual spatial pooling pyramid t where it first passes through the feature extraction network E therein, and respectively extracts the left image feature map containing semantic information features and texture edge information features i L for the left-eye image I i R and the right-eye image I and the right image multi-scale feature map Then, use the concatenation method to construct a cost volume with d disparity values for these two features, and then concatenate them to construct a four-dimensional matching cost volume This four-dimensional cost volume undergoes convolutional operations through the cost aggregation network R to obtain the aggregated cost volume; this aggregated cost volume undergoes convolutional and channel compression operations through the disparity regression network G, and its size is restored using bilinear interpolation to obtain the three-dimensional disparity cost volume C i pred , and then use the soft argmin operation on this three-dimensional disparity cost volume to obtain the result of stereo matching, i.e., the predicted disparity map I ipred 。

[0084] The effects of the present invention are further described below in conjunction with simulation experiments:

[0085] 1. Simulation experiment conditions

[0086] The software platform for the simulation experiment is: Ubuntu 16.04 LTS operating system, Python 3.7 programming language, and Pytorch deep learning framework;

[0087] The hardware platform for the simulation experiment is: a host with an Intel i9-7700X central processing unit, 32 GB of memory, an NVIDIA RTX 2080Ti graphics processing unit, and 11 GB of video memory.

[0088] The dataset for the simulation experiment is: a total of 394 training image groups in KITTI2012 and KITTI2015 are used as the training image set T, 200 validation images in KITTI2012 are used as the validation set V1, and 200 validation images in KITTI2015 are used as the validation set V2.

[0089] In the KITTI2012 validation set V1, the area S of the region A where the error value between the predicted disparity I i pred and the true disparity I i GT is greater than e pixels e-all divided by the total area S of the image A Ae-all is defined as the e-all evaluation index: all total area S Aall The ratio of the area S of the region A where the disparity error value in the non-closed disparity region A

[0090]

[0091] is greater than e pixel values to the area S of the non-closed disparity region noc is defined as the e-noc evaluation index: e-noc area S Ae-noc The ratio of the area S of the region A where the disparity error value is greater than e pixel values in the non-closed disparity region A Anoc to the area S of the non-closed disparity region is defined as the e-noc evaluation index:

[0092]

[0093] In this example, e = 2, 3, and four types of evaluation indexes can be obtained: two-pixel global error 2-all, three-pixel global error 3-all, two-pixel non-closed error 2-noc, and three-pixel non-closed error 3-noc;

[0094] In the KITTI2015 validation set V2, the region A where the error between the predicted disparity map and the true disparity map in the region A is greater than 3 pixels or greater than 5% of the true disparity value at that pointt The ratio to the area of Region A is defined as the D1 error evaluation index within Region A:

[0095]

[0096] S At is the area of Region A t and S A is the area of Region A; in the validation set V2, for Region A, the foreground error evaluation index D1-fg is calculated by selecting the foreground region of the image, the background error evaluation index D1-bg is calculated by selecting the background region, and the global error evaluation index D1-all is calculated by selecting the overall region.

[0097] 2. Simulation Content and Result Analysis

[0098] Simulation 1. A total of 394 training image groups from KITTI2012 and KITTI2015 are used as the training image set T to train the constructed binocular stereo matching network S, and the trained stereo matching network S is obtained. t Then, 200 validation images from KITTI2012 are used as the validation set V1 and input into the trained stereo matching network S t to obtain the stereo matching results of the method of the present invention, and the evaluation indexes of two-pixel global error 2-all, three-pixel global error 3-all, two-pixel non-closed error 2-noc, and three-pixel non-closed error 3-noc are calculated;

[0099] Take the paper "Pyramid Stereo Matching Network" as the baseline method, and use the method in the paper to construct a baseline network. The training image set T is used to train the baseline network, and the trained baseline network is obtained. The validation set V1 is input into the trained baseline network to obtain the stereo matching results of the baseline method, and the evaluation indexes of two-pixel global error 2-all, three-pixel global error 3-all, two-pixel non-closed error 2-noc, and three-pixel non-closed error 3-noc are calculated;

[0100] Compare the evaluation indexes obtained from the KITTI2012 test set through the baseline method and the method of the present invention, and the results are shown in Table 1:

[0101] Table 1 Comparison of Validation Indexes between the Method of the Present Invention and the Baseline Method on the KITTI2012 Dataset

[0102]

[0103] As can be seen from Table 1, on the KITTI2012 validation set V1, the evaluation of the method of the present invention relative to the baseline method has been improved. Among them, the 2-noc index has increased by 1.18%, the 2-all index has increased by 0.23%, the 3-noc has increased by 0.07%, and the 3-all index has increased by 0.16%.

[0104] Simulation 2. Use the trained stereo matching network S in Simulation 1 t , and input 200 validation images in KITTI2015 as the validation set V2 into the trained stereo matching network S t , and obtain the stereo matching result of the method of the present invention, as Figure 3 , where Figure 3 (a) is the left-eye image to be matched, Figure 3 (b) is the right-eye image to be matched, Figure 3 (c) is the output disparity map after stereo matching of the left and right-eye images using the simulation of the present invention. Figure 3 (c) shows that the present invention can obtain better matching results at the language edge and in the texture complex area.

[0105] Calculate the foreground error evaluation index D1-fg, background error evaluation index D1-bg, and global error evaluation index D1-all for the stereo matching result of the method of the present invention;

[0106] Use the trained baseline network in Simulation 1, input the validation set V2 into the trained baseline network, obtain the stereo matching result of the baseline method, and calculate its foreground error evaluation index D1-fg, background error evaluation index D1-bg, and global error evaluation index D1-all;

[0107] Compare the evaluation indexes obtained by the KITTI2015 test set through the baseline method and the method of the present invention respectively. The results are shown in Table 2:

[0108] Table 2 Comparison of validation indexes between the method of the present invention and the baseline method on the KITTI2015 dataset

[0109]

[0110] As can be seen from Table 2, when validating on the KITTI2015 dataset V2, the global error index D1-all has increased by 0.19% compared with the baseline method, the background error index D1-bg has increased by 0.06% compared with the baseline method, and the foreground error index D1-fg has increased by 0.67% compared with the baseline method.

[0111] The above simulation results show that, compared with the benchmark method, the error evaluation indexes of the present invention on the KITTI2012 validation dataset V1 and the KITTI2015 validation dataset V2 are improved, which can effectively improve the accuracy of the stereo matching results and obtain a more accurate disparity map.

Claims

1. A stereo matching method based on a dual pooling pyramid structure, characterized in that Including: (1) Obtain the binocular stereo matching training sample set T and the validation sample set V Obtain N + M groups of image pairs I from existing binocular stereo matching datasets i ={I i L , I i R , I i GT}}, where I i L and I i R are the left-eye image and the right-eye image of the i-th image pair respectively, and I i GT is the true disparity image of the i-th image pair; Use N image pairs in the dataset as the training sample set T = {I 1 , I 2 , …, I j , …, I N}, where 1 ≤ j ≤ N; Use M image pairs as the validation set V for the training results = {I N+1 , I N+2 , …, I N+k , …, I N+M}, where 1 ≤ k ≤ M, N ≥ 180, M ≥ 20; (2) Construct a binocular stereo matching network S based on a dual spatial pooling pyramid (2a) Establish a dual pyramid network E2 composed of the parallel connection of the average pooling pyramid E2 AVG and the maximum pooling pyramid E2 MAX ; (2b) Establish a feature extraction network E composed of a cascaded initial convolutional neural network E1, a dual pooling pyramid network E2, and a feature fusion network E3; (2c) Establish a cost aggregation network R composed of three cascaded 3D convolutional neural networks; (2d) Cascade the feature extraction network E, the cost aggregation network R, and the disparity regression network G to form the binocular stereo matching network S; (3) Train the binocular stereo matching network S: (3a) Set the learning rate η and the maximum number of training times t max , and the current iteration cycle t' = 0; (3b) For the left-eye image I in the training set T i L and the right-eye image I i R Use the forward propagation method to obtain the predicted disparity map I output by the network i pred ; (3c) Calculate the predicted disparity map I using the SmoothL1 loss function for the output image i pred The error L i GT between the predicted disparity map I i and the ground truth disparity map I t is calculated, and the convolutional kernel weight parameters ω t and convolutional kernel bias parameters b t in the current network S are updated using the gradient descent algorithm with backpropagation; (3d) Increment the current iteration cycle by 1, i.e., t' = t' + 1, and determine whether the maximum number of iterations has been reached: If t' = t max , the maximum number of iterations is reached, and at this time, the training ends, obtaining a trained binocular stereo matching network S t If t' < t max then return (3b); (4) Input the validation sample set V into the stereo matching network S of the trained dual-space pooling pyramid, and successively pass through the feature extraction network E, cost aggregation network R, and disparity regression network G therein to obtain the predicted result I of the disparity map. i pred This predicted result is the result of stereo matching.

2. The method according to claim 1, characterized in that The left-eye image I in the training set T in (3b) i L and the right-eye image I i R Use the forward propagation method to obtain the predicted disparity map I output by the network i pred , which is implemented as follows: (3b1) Use the initial convolutional neural network E1 to separately perform feature extraction on the left image I i L and the right image I i R to obtain the left image feature map F i L and the right image feature map F i R respectively; (3b2) Use the dual pooling pyramid network E2 to perform pooling operations on the left and right eye feature maps, obtaining the left eye average pooling feature map group F i L_AVG , the left eye maximum pooling feature map group F i L_MAX , and the right eye average pooling feature map group F i R_AVG , the right eye maximum pooling feature map group F i R_MAX ; (3b3) Stack the feature maps obtained from the two pooling pyramids and pass them through the feature fusion network E3 to output the left image feature map and the right image multi-scale feature map (3b4) Set the parallax search range D. First, for each parallax value d, construct the cost volume of the current parallax value in a splicing manner to obtain a three-dimensional matching cost volume group C i = {C i 1, C i 2, …, C i d , …, C i D}, and then sequentially splice its d three-dimensional matching cost volumes C i d to construct a four-dimensional matching cost volume 1 ≤ d ≤ D; (3b5) Use the cost aggregation network R to perform convolution operations on the four-dimensional matching cost volume C i d to obtain the aggregated cost volume (3b6) Use the parallax regression network G to perform convolution and channel compression operations on the aggregated cost volume in sequence, and use bilinear interpolation to restore its size to obtain the disparity three-dimensional cost volume C i pred ; (3b7) For the three-dimensional cost volume C i pred Use a soft argmin operation to obtain the disparity map I predicted by the network i pred .

3. The method according to claim 2, wherein In the above (3b2), the left and right eye feature maps are pooled using the dual pooling pyramid network E2 to achieve the following: In the average spatial pooling pyramid E2 AVG Average pooling operations are performed on the left image feature map and the right image feature map using four average pooling kernels of different sizes to extract features containing semantic information, obtaining the left-eye average pooling feature map group F i L_AVG and the right-eye average pooling feature map group F i R_AVG : F i L_AVG = {F i L_avg1 , F i L_avg2 , F i L_avg3 , F i L_avg4} F i R_AVG = {F i R_avg1 , F i R_avg2 , F i R_avg3 , F i R_avg4} Among them, F i L_avgi and F i R_avgi are the outputs of the i-th average pooling kernel, where 1 ≤ i ≤ 4; In the maximum spatial pooling pyramid E2 MAX Maximum pooling operations are performed on the left image feature map and the right image feature map using four maximum pooling kernels of different sizes to obtain features with texture and edge information, resulting in the left-eye image maximum pooling feature map group F i L_MAX and the right-eye image maximum pooling feature map group F i R_MAX : F i L_MAX = {F i L_max1 , F i L_max2 , F i L_max3 , F i L_max4} F i R_MAX ={F i R_max1 ,F i R_max2 ,F i R_max3 ,F i R_max4} Among them, F i L_maxi , F i R_maxi is the output of the i-th maximum pooling kernel, where 1 ≤ i ≤ 4.

4. The method according to claim 1, wherein The average pooling pyramid E2 in (2a) AVG and the max pooling pyramid E2 MAX , the structure of which is as follows: Average Pooling Pyramid E2 AVG , including four average pooling kernels with sizes of 4×4, 8×8, 16×16, and 64×64 respectively, and the number of each type of pooling kernel is 32; Max Pooling Pyramid E2 MAX , including four max pooling kernels with sizes of 4×4, 8×8, 16×16, and 64×64 respectively, and the number of each type of pooling kernel is 32.

5. The method according to claim 1, wherein The initial convolutional neural network E1 and the feature fusion network E3 in the above (2b) have the following structures: The initial convolutional neural network E1 includes three sequentially connected residual units. Each residual unit includes two groups of convolutional layers. The first group of convolutional layers includes two sequentially connected 32 convolutional kernels of size 3×3 and stride 2; the second group of convolutional layers includes three sequentially connected 32 convolutional kernels of size 3×3 and stride 1; The feature fusion network E3 includes four groups of parallel convolutional layers. Each group of convolutional layers includes two sequentially connected convolutional kernels. Each convolutional kernel has a size of 3×3, a stride of 1, and a quantity of 32.

6. The method according to claim 1, wherein Each 3D convolutional neural network in the above (2c) is composed of three sequentially connected convolutional layers. Each group of convolutional layers includes 32 3D convolutional kernels of size 3×3 and stride 1.

7. The method according to claim 1, characterized in that, The disparity regression network G in the above (2d) is composed of two sequentially connected 3D convolutional layers. The first group includes 32 3D convolutional kernels of size 3×3 and stride 1, and the second group includes 1 3D convolutional kernel of size 3×3 and stride 1.

8. The method according to claim 1, characterized in that The SmoothL1 loss function is used in (3c) to calculate the predicted disparity map I i pred and the ground truth disparity map I i GT to obtain the loss value L i , and its calculation formula is as follows:

9. The method according to claim 1, wherein In the above (3c), the gradient descent algorithm using backpropagation is used to update the convolutional kernel weight parameter ω t in the current network S t and the convolutional kernel bias parameter b t The update is performed according to the following formula: Among them, L i is the loss value between the predicted disparity map I i pred and the true disparity map I i GT , η represents the network update learning rate, 0.0001 ≤ η ≤ 0.001, ω t+1 and b t+1 respectively represent the updated convolutional kernel weight parameter and the updated convolutional kernel bias parameter.

Citation Information

Patent Citations

  • Stereo matching feature extraction and post-processing method of disparity map

    CN112991420A

  • Binocular vision position measurement system and method based on deep learning

    CN113177565A

  • Parallax image generation method and system based on binocular stereo vision matching

    CN110009691A

  • Forging defect detection method based on deep learning

    CN113393439A