A binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction
Patent Information
- Application Number
- CN202610989172.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本发明的目的是提供一种基于跨体积一致性不确定性估计并引导特征融合与对齐校正的双目立体匹配方法,解决上述背景技术中提出的迭代优化过程中因视差估计错误导致特征扭曲失真、误差逐级累积的问题
通过显式建模对齐不确定性并据此动态调节特征融合与扭曲校正,有效抑制了病态区域的误差累积,提升视差估计精度。
Smart Images

Figure CN122821178A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision, and in particular to a binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction. Background Technology
[0002] Binocular stereo vision systems draw upon the fundamental principle of human binocular vision in perceiving distance by processing a pair of stereo images of the same scene acquired from different viewpoints to perceive three-dimensional spatial information. Stereo matching, a core technology of binocular stereo vision, aims to establish dense correspondences between pixels in two stereo images and generate a dense disparity map representing these correspondences. However, various complex factors in real-world scenes, such as image noise, lighting variations, disparity discontinuities, repetitive or lacking texture features, and object occlusion, pose significant challenges to high-precision stereo matching.
[0003] In recent years, although deep learning-based stereo matching methods have made significant progress in terms of accuracy and robustness, distortion of ambiguous regions often leads to feature distortion, which accumulates during iteration and ultimately affects the accuracy of disparity estimation. Summary of the Invention
[0004] The purpose of this invention is to provide a stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction, which solves the problem of feature distortion and error accumulation caused by disparity estimation errors during the iterative optimization process mentioned in the background art.
[0005] To achieve the above objectives, this invention provides a binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction, comprising the following steps: S1. Multi-scale feature extraction: Input left and right binocular images, use a lightweight convolutional network with weight sharing to extract multi-scale feature pyramids, construct a global matching cost volume, and regress to obtain the initial disparity map; S2. Construct a multi-scale cross-volume consistency uncertainty map, utilize the probability distribution difference between the multi-scale global matching cost volume and the local alignment matching cost volume to obtain single-scale uncertainty features, and fuse them to obtain a unified uncertainty map; S3. Perform uncertainty soft fusion and alignment correction on the unified uncertainty map, and input it together with multi-scale context features into the convolutional GRU recurrent update unit to iteratively update the hidden state and predict the disparity residual, thereby realizing the recurrent refinement of the disparity result. S4. Iterate through S2 to S3, recalculate the uncertainty map after each iteration, and upsample the final disparity map to the original resolution after the iteration is complete, and output the depth map.
[0006] Preferably, the specific steps of S1 are as follows: S11. Input left and right binocular images; S12. Construct the multi-scale feature extraction network MobileNetV2 to extract multi-scale feature pyramids from the left and right images; S13. Perform dot product similarity measurement on the multi-scale features of the left and right images along the disparity direction, and construct a multi-scale matching cost pyramid through average pooling.
[0007] Preferably, the specific steps of S2 are as follows: S21. Calculate the global matching cost volume and the local alignment matching cost volume at each scale. The global matching cost volume is independent of the current iteration disparity, and the local alignment matching cost volume is calculated in the local disparity neighborhood with the current iteration disparity as the center. S22. Perform Softmax operation along the disparity dimension on the global matching cost volume and the local alignment matching cost volume respectively to obtain the probability distributions corresponding to the two types of cost volumes; S23. Calculate the difference tensor of the two probability distributions and calculate the variance along the disparity dimension to obtain the single-scale uncertainty map corresponding to each scale. S24. Apply minimum-maximum normalization to each single-scale uncertainty map to normalize the pixel values to the [0,1] interval; S25. The uncertainty maps normalized at each scale are uniformly upsampled to 1 / 4 resolution. Through cross-scale fusion, a unified uncertainty map is generated. The unified uncertainty map is used to constrain the feature soft fusion and alignment correction process of the iterative refinement process.
[0008] Preferably, the specific formula for the unified uncertainty diagram in S25 is as follows: ; in, U For a unified uncertainty diagram, Using 1×1 convolutions to adaptively aggregate uncertainties across scales. This is the uncertainty graph after normalization. This is upsampling to 1 / 4 resolution.
[0009] Preferably, the specific steps of S3 are as follows: S31. Conduct feature soft fusion under uncertainty conditions, downsample and match the unified uncertainty map to each feature scale, use the uncertainty map as a conditional signal, predict pixel-level fusion weights through a lightweight convolutional network, and adaptively complete multi-scale feature fusion. S32. Conduct uncertainty-driven feature alignment correction. At 1 / 4 resolution, use a lightweight convolutional network to predict the correction offset and combine it with the modulation correction intensity to construct the total loss. S33. The features after feature soft fusion and alignment correction are input together with the multi-scale context features into the convolutional GRU recurrent update unit to predict the disparity residuals and iteratively update the disparity estimation results.
[0010] Preferably, the specific steps of S31 are as follows: S311. Obtain a unified uncertainty map at 1 / 4 resolution and downsample it to various scales; S322. Using uncertainty maps at various scales as conditional signals, a lightweight convolutional network is used to output a single-channel pixel-level fusion weight mapping. S323. Based on pixel-level fusion weights, multi-scale features are weighted and fused to obtain fused features.
[0011] Preferably, the specific formula for the fusion feature in S323 is as follows: ; in, This is the fused feature map at resolution s. For a single-channel, pixel-level mapping with uncertain ∈ [0,1] at resolution s, The feature map of the right view at resolution s is warped relative to the feature map of the left view. The left image is the feature map at resolution s, where s∈{1 / 4,1 / 8,1 / 16}, representing three resolutions.
[0012] Preferably, the specific steps of S32 are as follows: S321. Select the 1 / 4 resolution corresponding to the unified uncertainty map, and use the left view feature map, warping feature map and uncertainty map under the 1 / 4 resolution as network input; S322. The correction offset is predicted by a lightweight convolutional network with two parallel 3×3 convolution + ReLU branches. The outputs of each branch are concatenated and connected to the 3×3 output layer. The magnitude of the correction offset is constrained by tanh(·) activation to ensure the stability of the optimization process. S323. Based on a unified uncertainty map with 1 / 4 resolution, the correction offset is modulated to adaptively adjust the correction intensity. S324. At a 1 / 4 resolution scale, construct an uncertainty-weighted correction loss and add an offset amplitude regularization term to constrain the magnitude of the correction offset. S325. The total alignment correction loss is obtained by superimposing the uncertainty weighted correction loss and the offset magnitude regularization term.
[0013] Preferably, the specific formula for the total loss in S325 is as follows: ; in, For the total loss, As the primary disparity regression loss, and These are weighting coefficients. Uncertainty-weighted correction loss, This is the offset amplitude regularization term.
[0014] Therefore, the present invention employs the above-mentioned binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction, which has the following beneficial effects: By explicitly modeling and aligning uncertainties and dynamically adjusting feature fusion and distortion correction accordingly, error accumulation in ill-conditioned regions is effectively suppressed, and disparity estimation accuracy is improved.
[0015] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating an embodiment of a binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction according to the present invention. Figure 2 This is the left eye image of an embodiment of a binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction according to the present invention; Figure 3 This is the right eye image of an embodiment of a binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction according to the present invention; Figure 4 This invention provides an embodiment of a binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction, which is an example of conditional soft fusion (UCSF). Figure 5 This invention provides an embodiment of a binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction, which includes uncertainty-driven alignment correction (UDAR). Figure 6 This is a visualization of the final estimated depth result of an embodiment of the binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction according to the present invention. Detailed Implementation
[0017] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0018] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0019] Example Please see Figures 1-6 This invention provides a binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction. It generates an uncertainty map by comparing global and local cost volumes, guiding UCSF to softly fuse distorted features with the original features, and performs pixel-level correction on the distorted features at 1 / 4 resolution. The fused and corrected features are then fed into a ConvGRU to predict disparity residuals and update the disparity. The uncertainty map U is recalculated after each iteration, forming a closed-loop control. After multiple iterations, the final disparity map is upsampled and output as the final result. The entire process effectively suppresses error accumulation in ill-conditioned regions by explicitly modeling alignment uncertainty and dynamically adjusting feature fusion and distortion correction accordingly. Figure 1 As shown, it includes the following steps: S1. Input left and right binocular images.
[0020] S2. Construct a multi-scale feature extraction network to extract multi-scale feature pyramids from the left and right images. The calculation formula is as follows: ; in, Indicates the features of the left eye image. Indicates the features of the right eye image. This represents a convolution with a stride of 1 and a kernel size of 3. This represents a convolution with a stride of 2 and a kernel size of 3. This indicates a normalization operation. Represents a non-linear activation function. This represents the residual convolution module. In This represents the obtained feature size, when The time is the feature size of the left image. The time represents the feature size of the right image. The calculation formula for the multi-scale feature pyramid extracted from the left and right images represents the process of the multi-scale residual convolution module extracting image matching features.
[0021] Constructing the basic matching cost volume: The left and right features are measured by dot product along the disparity direction, and a multi-scale matching cost pyramid is constructed using average pooling. The calculation formula is as follows: ; in, This indicates the number of levels in the matching cost pyramid. Representing the images respectively Feature map after feature extraction This represents the matrix multiplication operation. The calculated volume of the fully correlated matching cost. This indicates the average pooling operation. This indicates that the matching cost pyramid is obtained by downsampling the matching cost volume. layer, For the overall cost pyramid, For parallax, The parallax retrieval radius is... Using the iteration depth as the origin, depth cues around the origin are extracted. By merging depth clues at different scales, multi-scale matching cost information can be obtained. S3. Construct a cross-volume consistency uncertainty graph (CVCU), such as Figure 4 As shown, the specific steps are as follows: S31. At each scale, calculate the global matching cost volume and the local alignment matching cost volume, where the global matching cost volume is independent of the current iteration disparity, and the local alignment matching cost volume is calculated in the local disparity neighborhood centered on the current iteration disparity: ; in, The new feature map is generated by warping the feature map of the right view at resolution s and then adapting it to the feature map of the left view. For warping operation, This is the feature map of the right view at resolution s. The current disparity is s. In this embodiment, the superscripts of all parameters s∈{1 / 4,1 / 8,1 / 16} represent the three resolution cases.
[0022] S32. Apply a Softmax operation along the disparity dimension to the two cost volumes to obtain the probability distribution: ; ; in, This represents the global cost after softmax optimization. The global cost is calculated for the feature maps of the left and right views. This represents the warpage cost after softmax optimization. The warping cost is generated for the left-view feature map and the warped feature map.
[0023] S33. Calculate the difference tensor between the two probability distributions and calculate the variance along the disparity dimension to obtain the uncertainty plot at this scale: ; in, For the difference tensor, This is the uncertainty plot at resolution s. This is for variance calculation.
[0024] S34. By using minimum-maximum normalization for each image... Normalization to bounded range ∈[0,1]: ; in, This is the uncertainty graph after normalization. It is a very small positive number to prevent the denominator from being close to 0.
[0025] S35. In order to obtain a unified control signal that captures ambiguity from fine-grained texture to global structure, each Upsampled to 1 / 4 resolution and fused to produce a unified uncertainty graph. U : ; in, U For a unified uncertainty diagram, It uses 1×1 convolutions to adaptively aggregate uncertainty cues across scales. Upsampling to 1 / 4 resolution. Unified uncertainty plot. U Used to guide feature fusion (UCSF) and alignment correction (UDAR) in the refinement loop.
[0026] S4. Soft Fusion under Uncertainty Conditions (UCSF), such as Figure 5 As shown.
[0027] Given a unified uncertainty graph at 1 / 4 resolution U In this embodiment, it is downsampled to each scale, such as This is then used as a conditional signal to predict pixel-level fusion weights through a lightweight convolutional network: ; in, For resolution s, the fusion weights are pixel-level. It is the sigmoid activation function. It is a single-channel, pixel-level mapping. The fusion feature is calculated as follows: .
[0028] in, This is the fused feature map at resolution s. The left image is the feature map at resolution s, where s∈{1 / 4,1 / 8,1 / 16}, representing three resolutions.
[0029] S5. Uncertainty-driven alignment correction (UDAR).
[0030] Since the unified uncertainty graph U is fused at 1 / 4 resolution, this embodiment sets... At this scale, UDAR... , and As input, a lightweight convolutional network is used to predict and correct the offset: ; in, and With the same spatial resolution and channel dimension, the tanh(·) activation is constrained to [-1,1] for stable optimization. It is lightweight, shared across iterations, and consists of two parallel 3×3 Conv+ReLU branches whose outputs are concatenated and fed into a 3×3 output layer.
[0031] In obtaining Then, this embodiment uses Modulation correction strength: ; At the 1 / 4 scale, this embodiment defines an uncertainty-weighted correction loss: ; And add an offset magnitude regularization term: ; The specific formula for total loss is as follows: ; in, As the primary disparity regression loss, and These are the weighting coefficients.
[0032] S6. Input the fused and corrected features together with the multi-scale context features into the convolutional GRU recurrent update unit to predict the disparity residuals and iteratively update the disparity estimate.
[0033] S7. Repeat steps S3 to S6 for a total of N iterations, upsampling the final disparity map to the original resolution, and outputting the depth map, as shown below. Figure 6 As shown.
[0034] This invention first inputs left and right corrected images and extracts multi-scale feature pyramids through a lightweight convolutional network with shared weights. Then, a basic global matching cost volume is constructed, and a locally aligned cost volume is built using the disparity estimation of the current iteration. Based on the difference in probability distribution along the disparity dimension between the two cost volumes, cross-scale variance is calculated, and a unified uncertainty map is generated after fusion. This uncertainty map serves as a guiding signal: on the one hand, through an uncertainty-conditional soft fusion (UCSF) module, it adaptively fuses the features of the original left image and the distorted right image, trusting the distorted features in low-uncertainty regions and relying on the left image features in high-uncertainty regions; on the other hand, through an uncertainty-driven alignment correction (UDAR) module, it performs pixel-level correction on the distorted features in high-uncertainty regions. The fused and corrected features, along with the multi-scale context features, are fed into a convolutional GRU unit to iteratively update the hidden state and predict the disparity residuals, thereby achieving cyclic refinement. After each iteration, the uncertainty map is re-estimated, forming a closed-loop control. After multiple iterations, the disparity map is finally upsampled to the original resolution output.
[0035] Therefore, this invention adopts the above-mentioned binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction. By deriving a unified uncertainty map U from multi-scale cross-volume inconsistency, it guides UCSF to perform uncertainty condition feature fusion, UDAR performs alignment correction before cyclic refinement, and iteratively updates the uncertainty map to form a closed-loop control signal to reduce the error accumulation caused by warping, thereby improving the accuracy and robustness of the model.
[0036] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction, characterized in that, Includes the following steps: S1. Multi-scale feature extraction: Input left and right binocular images, use a lightweight convolutional network with weight sharing to extract multi-scale feature pyramids, construct a global matching cost volume, and regress to obtain the initial disparity map; S2. Construct a multi-scale cross-volume consistency uncertainty map, utilize the probability distribution difference between the multi-scale global matching cost volume and the local alignment matching cost volume to obtain single-scale uncertainty features, and fuse them to obtain a unified uncertainty map; S3. Perform uncertainty soft fusion and alignment correction on the unified uncertainty map, and input it together with multi-scale context features into the convolutional GRU recurrent update unit to iteratively update the hidden state and predict the disparity residual, thereby realizing the recurrent refinement of the disparity result. S4. Iterate through S2 to S3, recalculate the uncertainty map after each iteration, and upsample the final disparity map to the original resolution after the iteration is complete, and output the depth map.
2. The binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction according to claim 1, characterized in that, The specific steps of S1 are as follows: S11. Input left and right binocular images; S12. Construct the multi-scale feature extraction network MobileNetV2 to extract multi-scale feature pyramids from the left and right images; S13. Perform dot product similarity measurement on the multi-scale features of the left and right images along the disparity direction, and construct a multi-scale matching cost pyramid through average pooling.
3. The binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction according to claim 1, characterized in that, The specific steps of S2 are as follows: S21. Calculate the global matching cost volume and the local alignment matching cost volume at each scale. The global matching cost volume is independent of the current iteration disparity, and the local alignment matching cost volume is calculated in the local disparity neighborhood with the current iteration disparity as the center. S22. Perform Softmax operation along the disparity dimension on the global matching cost volume and the local alignment matching cost volume respectively to obtain the probability distributions corresponding to the two types of cost volumes; S23. Calculate the difference tensor of the two probability distributions and calculate the variance along the disparity dimension to obtain the single-scale uncertainty map corresponding to each scale. S24. Apply minimum-maximum normalization to each single-scale uncertainty map to normalize the pixel values to the [0,1] interval; S25. The uncertainty maps normalized at each scale are uniformly upsampled to 1 / 4 resolution. Through cross-scale fusion, a unified uncertainty map is generated. The unified uncertainty map is used to constrain the feature soft fusion and alignment correction process of the iterative refinement process.
4. The binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction according to claim 3, characterized in that, The specific formula for the unified uncertainty diagram in S25 is as follows: ; in, U For a unified uncertainty diagram, Using 1×1 convolutions to adaptively aggregate uncertainties across scales. This is the uncertainty graph after normalization. This is upsampling to 1 / 4 resolution.
5. A binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction as described in claim 1, characterized in that, The specific steps of S3 are as follows: S31. Conduct feature soft fusion under uncertainty conditions, downsample and match the unified uncertainty map to each feature scale, use the uncertainty map as a conditional signal, predict pixel-level fusion weights through a lightweight convolutional network, and adaptively complete multi-scale feature fusion. S32. Conduct uncertainty-driven feature alignment correction. At 1 / 4 resolution, use a lightweight convolutional network to predict the correction offset and combine it with the modulation correction intensity to construct the total loss. S33. The features after feature soft fusion and alignment correction are input together with the multi-scale context features into the convolutional GRU recurrent update unit to predict the disparity residuals and iteratively update the disparity estimation results.
6. A binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction as described in claim 5, characterized in that, The specific steps of S31 are as follows: S311. Obtain a unified uncertainty map at 1 / 4 resolution and downsample it to various scales; S312. Using uncertainty maps at various scales as conditional signals, a lightweight convolutional network is used to output a single-channel pixel-level fusion weight mapping. S313. Based on pixel-level fusion weights, multi-scale features are weighted and fused to obtain fused features.
7. A binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction as described in claim 6, characterized in that, The specific formula for the fusion feature in S313 is as follows: ; in, This is the fused feature map at resolution s. For a single-channel, pixel-level mapping with uncertain ∈ [0,1] at resolution s, The feature map of the right view at resolution s is warped relative to the feature map of the left view. The left image is the feature map at resolution s, where s∈{1 / 4,1 / 8,1 / 16}, representing three resolutions.
8. A binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction as described in claim 5, characterized in that, The specific steps of S32 are as follows: S321. Select the 1 / 4 resolution corresponding to the unified uncertainty map, and use the left view feature map, warping feature map and uncertainty map under the 1 / 4 resolution as network input; S322. The correction offset is predicted by a lightweight convolutional network with two parallel 3×3 convolution + ReLU branches. The outputs of each branch are concatenated and connected to the 3×3 output layer. The magnitude of the correction offset is constrained by tanh(·) activation to ensure the stability of the optimization process. S323. Based on a unified uncertainty map with 1 / 4 resolution, the correction offset is modulated to adaptively adjust the correction intensity. S324. At a 1 / 4 resolution scale, construct an uncertainty-weighted correction loss and add an offset amplitude regularization term to constrain the magnitude of the correction offset. S325. The total alignment correction loss is obtained by superimposing the uncertainty weighted correction loss and the offset magnitude regularization term.
9. A binocular stereo matching method based on cross-volume consistency uncertainty estimation and guided feature fusion and alignment correction as described in claim 8, characterized in that, The specific formula for the total loss in S325 is as follows: ; in, For the total loss, As the primary disparity regression loss, and These are weighting coefficients. Uncertainty-weighted correction loss, This is the offset amplitude regularization term.