Binocular image stereo matching method based on parallax distribution truth value modeling
By integrating the predicted values and disparity truths of the pre-trained network in the stereo matching network and modeling the disparity distribution truths, the problem of insufficient generalization ability of the stereo matching network in real scenarios is solved, and higher cross-domain generalization performance and robustness are achieved.
Patent Information
- Application Number
- CN202510362831.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-27
AI Technical Summary
After the existing stereo matching network is trained on the synthetic domain, it is difficult to generalize in various real scenes, especially in areas such as repeated textures, weak textures, edges and exposures, which are prone to mismatch.
By fusing the predicted value of the pre-trained stereo matching network with the original disparity truth, the disparity distribution truth is modeled, which is used to train the stereo matching network, thereby improving cross-domain generalization performance.
It significantly improves the cross-domain generalization performance of the stereo matching network from synthetic domain to real domain, improves the robustness in challenging areas, and ensures the security and reliability of unmanned systems in practical applications.
Smart Images

Figure CN120219461A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a binocular image stereo matching method in the field of computer vision, and particularly to a binocular image stereo matching method based on disparity distribution true value modeling. Background Art
[0002] Stereo matching, as an important problem in computer vision, aims to reconstruct the depth information of a scene using images from different perspectives and plays an important role in fields such as autonomous driving, augmented reality, and robotics. In these fields where safety is of crucial importance, the reliability of stereo matching networks in dealing with complex and changing scenes has attracted particular attention. However, due to the limitations of hardware acquisition devices, the scale of existing stereo matching real datasets is too small to meet the needs of training stereo matching networks with strong generalization ability. Therefore, the prior art hopes to directly train stereo matching networks with good cross-domain generalization performance on large-scale synthetic datasets.
[0003] To improve the cross-domain generalization performance of stereo matching networks, existing research work often adopts the following several technical strategies: 1) replacing the learning-based feature extractor in the stereo matching network with a manually designed feature descriptor or a feature extractor pre-trained on a large-scale real dataset to enhance the robustness of feature extraction; 2) guiding the network to learn domain-invariant features, such as aligning the features extracted from the original image and the features extracted from the perturbed image, so that the stereo matching network can maintain a consistent feature representation between different domains; 3) designing a more robust network structure, such as introducing an attention mechanism, adaptive convolution, etc., to enhance the adaptability of the stereo matching network to changes in different perspectives, lighting, texture, etc. in cross-domain data.
[0004] The problems of the prior art are as follows: 1) The stereo matching network trained in the synthetic domain only performs well on specific real datasets and cannot generalize to diverse scenes; 2) For challenging regions in stereo matching, such as regions with repetitive texture, weak texture, edges, and exposure, etc., existing stereo matching networks are still prone to false matching. Summary of the Invention
[0005] To solve the problems existing in the background technology, the purpose of the present invention is to provide a binocular image stereo matching method based on disparity distribution true value modeling, which is applicable to most stereo matching networks. The present invention models the disparity distribution true value for training the stereo matching network by fusing the prediction values of the pre-trained stereo matching network and the original disparity true value, and can effectively improve the cross-domain generalization performance of the stereo matching network from the synthetic domain to the real domain. In challenging regions in stereo matching, such as regions with repetitive textures, weak textures, edges, and exposure, etc., the present invention shows high robustness, and can ensure the safety and reliability of the unmanned system based on binocular vision in practical applications. The present invention has strong versatility and can be directly applied to most existing stereo matching networks.
[0006] The steps of the technical solution adopted by the present invention are as follows:
[0007] I. A binocular image stereo matching method based on disparity distribution true value modeling
[0008] 1) Use the synthetic dataset to pre-train the first stereo matching network to obtain several pre-trained stereo matching networks;
[0009] 2) Input each binocular image in the synthetic dataset into several pre-trained stereo matching networks respectively to obtain several disparity distribution prediction values corresponding to the current binocular image;
[0010] 3) Fuse the disparity distribution prediction values corresponding to the current binocular image and the disparity true value map in the Laplace parameter space and model them as the disparity distribution true value corresponding to the current binocular image;
[0011] 4) Repeat steps 2)-3), traverse and process the remaining binocular images in the synthetic dataset to generate the disparity distribution true values corresponding to the remaining binocular images, so as to obtain the true value modeling dataset;
[0012] 5) Use the true value modeling dataset to train the second stereo matching network to obtain the trained second stereo matching network;
[0013] 6) Input the binocular image to be predicted into the trained second stereo matching network and output the corresponding disparity map.
[0014] The structure of the second stereo matching network is the same as or different from that of the first stereo matching network.
[0015] In step 1), convert each disparity true value map in the synthetic dataset into a disparity distribution true value in the form of a unimodal Laplace. During the pre-training process of the first stereo matching network, calculate the cross-entropy between the disparity distribution prediction value output by the first stereo matching network and the corresponding disparity distribution true value in the form of a unimodal Laplace, and use the cross-entropy as the loss function.
[0016] In 2), the disparity distribution predicted by the pre-trained stereo matching network is a multi-modal distribution.
[0017] The 3) is specifically as follows:
[0018] 3.1) Separate multiple unimodal distributions from the predicted values of each disparity distribution in each binocular image, and fit each unimodal distribution to a point in the Laplace parameter space; traverse and process all the predicted values of the disparity distribution of the current binocular image to obtain all the points in the Laplace parameter space;
[0019] 3.2) Convert the ground truth disparity map corresponding to the current binocular image into a point in the Laplace parameter space
[0020] which are the ground truth weight parameter, ground truth position parameter, and ground truth scale parameter respectively; corresponding point of the ground truth disparity map
[0021] is a non-noise point;
[0022] The Laplace parameter space has three dimensions: weight w, position μ, and scale b. Each point (w, μ, b) in the Laplace parameter space corresponds to a unimodal Laplace function p(d; w, μ, b), and the specific formula is as follows:
[0023]
[0024] where the disparity d is the independent variable of the function, D is the maximum search range of the disparity predefined by the stereo matching, || represents the absolute value, and d i represents the disparity candidate value.
[0025] In 3.1), each unimodal distribution is fitted to a point in the Laplace parameter space and the calculation formulas of each parameter are as follows:
[0026]
[0027] where D is the maximum search range of the disparity predefined by the stereo matching network, || represents the absolute value, are the weight, position, and scale of the point coordinates respectively, and d i represents the disparity candidate value.
[0028] In the above (3.4), the true value of the parallax distribution satisfies the following formula:
[0029]
[0030] where p gt (d) represents the true value of the parallax distribution, w k , μ k , b k are respectively the weight, position, and scale of the center coordinates of the k-th clustering cluster, D is the maximum search range of the parallax predefined by the stereo matching network, || represents the absolute value, are respectively the weight, position, and scale of the point coordinates, and d i represents the parallax candidate value.
[0031] II. A computer device
[0032] The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the binocular image stereo matching method based on the true value modeling of the parallax distribution are implemented.
[0033] III. A computer-readable storage medium
[0034] A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, the steps of the binocular image stereo matching method based on the true value modeling of the parallax distribution are implemented.
[0035] IV. A computer program product
[0036] The computer program product includes a computer program / instructions, and when the computer program / instructions are executed by the processor, the steps of the binocular image stereo matching method based on the true value modeling of the parallax distribution are implemented.
[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0038] 1) The present invention can effectively improve the cross-domain generalization performance of the stereo matching network from the synthetic domain to the real domain;
[0039] 2) The present invention shows high robustness in challenging regions in stereo matching, such as regions with repetitive textures, weak textures, edges, and exposure;
[0040] 3) The present invention can be directly applied to training most stereo matching networks and has high generality. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a flowchart of the method proposed by the present invention;
[0042] Figure 2 It is the visualization result of the cross - domain generalization performance of a stereo matching network trained on a synthetic dataset on a real dataset. Specific implementation manners
[0043] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0044] As Figure 1 shown in the flowchart of
[0045] Taking the SceneFlow synthetic dataset as the training set, the four real datasets of KITTI 2015, KITTI 2012, Middlebury, and ETH3D as the test sets, and the three stereo matching network architectures of PSMNet, GwcNet, and PCWNet as the backbone networks as examples, the modeling idea and specific implementation steps of the ground - truth disparity distribution are described.
[0046] The present invention proposes a binocular image stereo matching method based on modeling the ground - truth disparity distribution. The method includes the following steps:
[0047] 1) Pre - train a first stereo matching network using the SceneFlow synthetic dataset to obtain a number of pre - trained stereo matching networks; the SceneFlow synthetic dataset consists of multiple binocular images and corresponding ground - truth disparity maps.
[0048] Among them, convert each ground - truth disparity map in the SceneFlow synthetic dataset into a ground - truth disparity distribution in the form of a unimodal Laplace with a scale parameter of 1. During the pre - training process of the first stereo matching network, calculate the cross - entropy between the predicted disparity distribution values {p1(d), p2(d), … p m , …, p M} output by the first stereo matching network {ε1, ε2, … ε m (d), …, p M (d)} and the corresponding unimodal Laplace - form ground - truth disparity distribution p um (d), and use the cross - entropy as the loss function. The specific formula of the cross - entropy loss is as follows:
[0049]
[0050] Among them, represents the cross - entropy loss, D represents the maximum search range of the predefined disparity. Train the stereo matching network using the cross - entropy loss to obtain a pre - trained stereo matching network. In this embodiment, D is set to 192, M is set to 3, and the three stereo matching network architectures are PSMNet, GwcNet, and PCWNet respectively.
[0051] 2) Input each binocular image in the SceneFlow synthetic dataset into several pre-trained stereo matching networks respectively to obtain several predicted values of the disparity distribution corresponding to the current binocular image; the disparity distribution predicted by the pre-trained stereo matching network is a multi-modal distribution, that is, a multi-modal distribution pseudo-ground truth in a form similar to the multi-modal Laplace form.
[0052] 3) Fuse and model the predicted value of the disparity distribution corresponding to the current binocular image with the disparity ground truth map in the Laplace parameter space as the disparity distribution ground truth of the stereo matching corresponding to the current binocular image;
[0053] 3) Specifically:
[0054] 3.1) Separate multiple unimodal distributions from each predicted value of the disparity distribution of each binocular image, and fit each unimodal distribution to a point in the Laplace parameter space; traverse and process all predicted values of the disparity distribution of the current binocular image to obtain all points in the Laplace parameter space;
[0055] The Laplace parameter space has three dimensions: weight w, position μ, and scale b. Each point (w, μ, b) in the Laplace parameter space corresponds to a unimodal Laplace function p(d; w, μ, b), and the specific formula is as follows:
[0056]
[0057] Among them, the disparity d is the independent variable of the function, D is the maximum search range of the disparity predefined for stereo matching, || represents the absolute value, and d i represents the disparity candidate value.
[0058] Fit each unimodal distribution to a point in the Laplace parameter space The calculation formulas for each parameter are as follows:
[0059]
[0060] Among them, are the weight, position, and scale of the point coordinates respectively.
[0061] 3.2) Convert the disparity ground truth map corresponding to the current binocular image into a point in the Laplace parameter space are the ground truth weight parameter, ground truth position parameter, and ground truth scale parameter respectively, where the ground truth weight parameter and the ground truth scale parameter are predefined hyperparameters, which are set to 1.0 and 0.8 respectively;
[0062] 3.3) Use the DBScan clustering algorithm to cluster all points in the current Laplacian parameter space along the position dimension and filter out noise points to obtain K clusters {Ω1, Ω2, …, Ω K}; where, the distance threshold ∈ and the density threshold minPts in the DBScan algorithm are set to 3 and 2 respectively. The points corresponding to the ground truth disparity map are non-noise points.
[0063] 3.4) After modeling the K clusters, obtain the ground truth disparity distribution corresponding to the current stereo image, which is a multi-modal Laplacian distribution, and the ground truth disparity distribution satisfies the following formula:
[0064]
[0065] where, p gt (d) represents the ground truth disparity distribution, w k , μ k , b k are the weights, positions, and scales of the center coordinates of the k-th cluster respectively, are the predicted weight parameter, predicted position parameter, and predicted scale parameter, that is, the average value of the coordinates of all points in the cluster. Normalize the ground truth disparity distribution p gt (d) to ensure that the sum of probabilities is 1.
[0066] 4) Repeat steps 2)-3), traverse and process the remaining stereo images in the SceneFlow synthetic dataset, generate the ground truth disparity distributions corresponding to the remaining stereo images, and thus obtain the ground truth modeling dataset, that is, composed of multiple stereo images and the corresponding ground truth disparity distributions;
[0067] 5) Use the ground truth modeling dataset to train the second stereo matching network from scratch using the cross-entropy loss to obtain the trained second stereo matching network; in this embodiment, a total of 3 stereo matching networks with different architectures are trained, namely PSMNet, GwcNet, and PCWNet. The training process is as follows: Use one Nvidia 4090 GPU for training, use the Adam optimizer with a momentum of 0.9 and a batch size of 4. Use the cosine annealing learning rate decay strategy, and set the maximum learning rate to 0.001. A total of 80 epochs are trained on the SceneFlow synthetic dataset.
[0068] 6) Input the stereo image to be predicted into the trained second stereo matching network and output the corresponding disparity map.
[0069] On four real datasets, namely KITTI 2015, KITTI 2012, Middlebury, and ETH3D, the performance gain brought by the present invention is tested. The evaluation metric is the percentage of pixels with prediction errors, and the smaller this value is, the more accurate the result is. The error pixels in KITTI2015 / 2012, Middlebury, and ETH3D are respectively defined as pixels where the absolute error of disparity is greater than 3, 2, and 1 pixel.
[0070] Table 1 is a performance comparison table between the method of the present invention and the original backbone network
[0071]
[0072]
[0073] As shown in Table 1 above, PSMNet, GwcNet, and PCWNet trained by applying the method proposed in the present invention have all achieved cross-domain generalization performance far exceeding that of the original backbone network, demonstrating that the present invention can improve the generalization performance of the existing backbone network and has high generality.
[0074] Table 2 is a performance comparison table between the method of the present invention and industry-advanced methods
[0075] Method KITTI 2015 KITTI 2012 Middlebury ETH3D GANet 11.7 10.1 20.3 14.1 DSMNet 6.5 6.2 13.8 6.2 CFNet 5.8 4.7 13.5 5.8 FC-GANet 5.3 4.6 10.2 5.8 Graft-GANet 4.9 4.2 9.8 6.2 ITSA-CFNet 4.7 4.2 10.4 5.1 PSMNet+UMCE 4.7 4.6 9.8 4.2 IGEVStereo 6.0 5.2 7.3 3.6 NMRF 5.1 4.2 7.5 3.8 PSMNet+ADL 4.8 4.2 8.9 3.4 GANet+ADL 4.8 3.9 8.7 2.3 PSMNet+the present invention 4.5 3.7 8.0 3.2 GwcNet+the present invention 4.2 3.7 7.2 2.8 PCWNet+the present invention 4.0 3.6 7.2 2.7
[0076] As shown in Table 2 above, the generalization performance of PCWNet trained by applying the method proposed in the present invention reaches the leading level in the industry on three datasets, namely KITTI 2015, KITTI 2012, and Middlebury, and ranks second on the ETH3D dataset.
[0077] Figure 2 The generalization performance comparison of PSMNet trained by applying UMCE, ADL, and the method proposed in the present invention is visualized, where Figure 2 (a) of it is the input left image, Figure 2 (b) of it is the ground truth disparity, Figure 2 (c) of it is the disparity map output by PSMNet + UMCE, Figure 2 (d) of it is the disparity map output by PSMNet + ADL, Figure 2 (e) of it is the disparity map output by PSMNet using the method proposed in the present invention. The percentage of pixels with prediction errors in each disparity map is shown in the upper left corner. It can be seen from the figure that for weak texture regions such as the sky and carriage, repetitive texture regions such as grass and chair seams, thin and narrow edge regions such as road poles, and exposure regions, the method proposed in the present invention can all predict better disparity results. The present invention shows high robustness when dealing with challenging images, and can ensure the safety and reliability of binocular vision-based unmanned systems in practical applications.
[0078] The above embodiments are used to explain the present invention rather than limit the present invention. Any modifications and changes made to the present invention within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.
Claims
1. A binocular image stereo matching method based on disparity distribution true value modeling, characterized in that: The steps include: 1) Pre-training a first stereo matching network using a synthetic data set to obtain a number of pre-trained stereo matching networks; 2) Input each binocular image in the synthetic data set into several pre-trained stereo matching networks to obtain several disparity distribution prediction values corresponding to the current binocular image; 3) The disparity distribution prediction value corresponding to the current binocular image and the disparity true value map are fused in the Laplace parameter space and modeled as the disparity distribution true value corresponding to the current binocular image; 4) Repeat 2)-3), traverse and process the remaining binocular images in the synthetic data set, generate the true value of the disparity distribution corresponding to the remaining binocular images, and thus obtain the true value modeling data set; 5) Using the true value modeling data set to train the second stereo matching network to obtain a trained second stereo matching network; 6) Input the binocular image to be predicted into the trained second stereo matching network and output the corresponding disparity map.
2. The binocular image stereo matching method based on disparity distribution true value modeling according to claim 1, characterized in that: In the above 1), each disparity truth map in the synthetic data set is converted into a disparity distribution truth value in the form of a unimodal Laplace. During the pre-training process of the first stereo matching network, the cross entropy between the disparity distribution prediction value output by the first stereo matching network and the corresponding unimodal Laplace disparity distribution truth value is calculated and the cross entropy is used as the loss function.
3. The binocular image stereo matching method based on disparity distribution true value modeling according to claim 1, characterized in that: In the above 2), the disparity distribution predicted by the pre-trained stereo matching network is a multi-modal distribution.
4. The binocular image stereo matching method based on disparity distribution true value modeling according to claim 1, characterized in that: The specific aspects of 3) are: 3.1) Separate multiple unimodal distributions from each disparity distribution prediction value of each binocular image, and fit each unimodal distribution to a point in the Laplace parameter space; traverse and process all disparity distribution prediction values of the current binocular image to obtain all points in the Laplace parameter space; 3.2) The disparity truth map corresponding to the current binocular image It also transforms to a point in the Laplace parameter space They are the truth weight parameter, the truth location parameter and the truth scale parameter respectively; 3.3) Cluster all points in the current Laplace parameter space and filter out noise points to obtain K clusters; among them, the disparity true value map The corresponding point is a non-noise point; 3.4) After modeling K clusters, the true value of the disparity distribution corresponding to the current binocular image is obtained.
5. The binocular image stereo matching method based on parallax distribution true value modeling according to claim 4, characterized in that: The Laplace parameter space has three dimensions: weight w, position μ, and scale b. Each point (w, μ, b) in the Laplace parameter space corresponds to a single-peak Laplace function p(d; w, μ, b). The specific formula is as follows: Among them, disparity d is the independent variable of the function, D is the maximum search range of disparity predefined by stereo matching, || represents the absolute value, and d i represents the disparity candidate value.
6. The binocular image stereo matching method based on parallax distribution true value modeling according to claim 4, characterized in that: In 3.1), each unimodal distribution Fitting to a point in Laplace parameter space The calculation formulas for each parameter are as follows: Where D is the maximum search range of disparity predefined by the stereo matching network, || represents the absolute value, are the weight, position and scale of the point coordinates, d i represents the disparity candidate value.
7. The binocular image stereo matching method based on parallax distribution true value modeling according to claim 4, characterized in that: In 3.4), the true value of the disparity distribution satisfies the following formula: Among them, p gt (d) represents the true value of the disparity distribution, w k ,μ k ,b k are the weight, position and scale of the center coordinates of the kth cluster, respectively. D is the maximum search range of the disparity predefined by the stereo matching network. || represents the absolute value. are the weight, position and scale of the point coordinates, d i represents the disparity candidate value.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of a binocular image stereo matching method based on disparity distribution true value modeling as described in any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a binocular image stereo matching method based on disparity distribution true value modeling as described in any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of a binocular image stereo matching method based on disparity distribution true value modeling as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Binocular three-dimensional obstacle avoidance method based on adaptive threshold Census transformation
CN121213630A