Lightweight stereo matching method based on weight sharing and channel attention

By combining weighted feature extraction and SE channel attention mechanism with hash coding multi-scale correlation volume construction, the problem of high parameter quantity and computational complexity of stereo matching methods is solved, achieving lightweight and high-precision disparity estimation, which is suitable for scenarios with high real-time requirements such as autonomous driving and robot navigation.

CN121685268APending Publication Date: 2026-03-17GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511874222.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing stereo matching methods have high parameter count and computational complexity, making it difficult to implement real-time deployment on embedded devices. Furthermore, they lack robustness in weak texture regions, repetitive texture regions, and occlusion boundaries, and the computational bottleneck in the construction and indexing process of multi-scale related volumes has not been effectively resolved.

Method used

We employ a feature extraction backbone with weight sharing, insert an SE channel attention mechanism after the residual block for feature calibration, and construct a multi-scale correlation volume through hash coding based on spatial coordinates. We use hash coding for dimensionality reduction indexing and trilinear interpolation sampling to reduce the number of model parameters and improve disparity estimation accuracy.

Benefits of technology

It significantly reduces the number of model parameters by 18%, improves disparity estimation accuracy and inference speed by 28%, and demonstrates stronger robustness and detail preservation in complex scenarios, meeting the needs of real-time applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121685268A_ABST
    Figure CN121685268A_ABST
Patent Text Reader

Abstract

The invention relates to a lightweight stereo matching method based on weight sharing and channel attention. The left image and the right image pass through a weight sharing feature extractor to obtain a feature map, and an SE channel attention mechanism is inserted behind a residual block of the feature extractor for feature calibration; when a multi-scale correlation body is constructed, sparse indexing is carried out on parallax dimensions by adopting Hash coding based on space coordinates, and voxel sampling and trilinear interpolation with constant time complexity are realized; and obtaining a final disparity map through a variable-resolution iterative updating strategy. According to the method, the quantity of model parameters is effectively reduced, EPE and D1 indexes on a Middlebury data set are superior to those of an existing RAFT-Stereo method, and the method is suitable for real-time scenes such as automatic driving and robot navigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision technology and relates to a stereo matching method, specifically a lightweight stereo matching method based on weight sharing and channel attention. This method significantly reduces the number of model parameters and computational complexity by extracting the backbone through weight sharing, inserting an SE channel attention mechanism after each residual block for feature calibration, and constructing a multi-scale correlation volume using hash encoding based on spatial coordinate hash mapping. This approach maintains or improves disparity estimation accuracy while maintaining or enhancing model parameter count. It is particularly suitable for applications sensitive to real-time performance and energy consumption, such as autonomous driving, robot navigation, and 3D reconstruction. Background Technology

[0002] Stereo matching is a core technology in the field of computer vision. It achieves 3D reconstruction by estimating dense disparity maps from left and right stereo image pairs, and has important application value in fields such as autonomous driving, robot navigation, and virtual reality.

[0003] In existing technologies, RAFT-Stereo has achieved high disparity estimation accuracy on datasets such as Middlebury and ETH-3D through an iterative refinement mechanism. However, it has a parameter count as high as 11.2M, resulting in high computational complexity. In particular, it is susceptible to channel redundancy and noise interference in weak texture regions, repetitive texture regions, and occlusion boundaries, leading to insufficient matching robustness and difficulty in meeting the real-time deployment requirements of edge devices.

[0004] Current lightweight methods, such as HitNet and LEAStereo, reduce computational cost through cost volume regularization or depthwise separable convolutions, but usually come with a significant loss of accuracy. Other solutions employ channel pruning or knowledge distillation techniques, which reduce the number of parameters, but generally suffer from weak generalization ability and fail to effectively address the computational bottleneck in the construction and indexing of multi-scale related volumes.

[0005] Current lightweight methods, such as HitNet and LEAStereo, reduce computational cost through cost-volume regularization, but at the cost of accuracy. Furthermore, while some existing solutions attempt to introduce channel attention mechanisms in stereo matching to improve feature saliency, these methods primarily focus on feature extraction optimization and fail to effectively address the significant memory consumption and computational bottlenecks encountered in constructing multi-scale related volumes at high resolutions. In other words, simply introducing attention mechanisms often further increases the model's computational burden, making truly efficient deployment on embedded devices difficult.

[0006] In view of this, this invention proposes a lightweight stereo matching method based on hash-encoded multi-scale correlated bodies. It extracts the backbone through fully weighted feature extraction, inserts an SE channel attention mechanism after each residual block for feature calibration, and uses hash encoding based on spatial coordinate hash mapping to achieve efficient construction of multi-scale correlated bodies and constant time complexity indexing. This reduces the number of model parameters by about 18% (11.2M → 9.2M). On the Middlebury dataset, the endpoint error (EPE) is reduced to 1.185 and the D1-all error is reduced to 7.767%, both of which are significantly better than RAFT-Stereo. At the same time, the inference speed is improved by about 28%. It shows stronger robustness and detail preservation ability in complex scenes such as weak texture and occlusion, and can meet the real-time application requirements without additional hardware acceleration. Summary of the Invention

[0007] This invention discloses a lightweight stereo matching method based on hash-encoded multi-scale correlation, which can significantly reduce the number of model parameters and computational complexity while improving disparity estimation accuracy and inference speed.

[0008] To achieve the above objectives, this invention provides a lightweight stereo matching method based on hash-encoded multi-scale correlations, comprising the following steps:

[0009] Input a pair of epipolar-corrected left and right stereo images, and extract feature maps using a feature extractor with fully shared weights. Insert an SE channel attention module after each residual block to dynamically recalibrate the feature channels to suppress redundant channels and enhance effective channels. Based on the recalibrated left and right feature maps, construct a multi-scale correlation volume of level k=3. In the disparity dimension, hash encoding based on spatial coordinate hash mapping is used for dimensionality reduction indexing, and constant-time-complexity voxel localization and trilinear interpolation sampling are achieved through a pre-built hash table. Input the multi-scale correlation features into the iterative update module, and use a Slow-Fast gated recurrent unit strategy for disparity refinement. Finally, output a full-resolution disparity map.

[0010] The shared feature extractor contains 6 residual blocks. The first two residual blocks maintain the spatial resolution, while the last four residual blocks are downsampled to 1 / 8 of the original image resolution by setting the convolution stride to 2. The number of channels is gradually increased to 256, and the left and right images completely share all convolution weights.

[0011] The SE channel attention module is placed after each residual block. The structure includes global average pooling and two fully connected layers with a dimensionality reduction ratio of r=16 (with ReLU activation in the middle and Sigmoid activation at the end). It outputs channel weights of 0 to 1 and multiplies them back to the original feature map channel by channel.

[0012] The multi-scale correlation volume includes three scales (1 / 8, 1 / 16, and 1 / 32 resolutions, respectively). During the construction process, the features of the right image are pre-mapped based on spatial coordinates and a hash table is established. In the inference stage, after spatial hash encoding of the current disparity hypothesis, the eight vertex features of the corresponding voxel are directly located through the hash table, and then the sub-pixel-level correlation value is obtained through trilinear interpolation. The multi-scale correlation features are then concatenated and sent to the iterative update module.

[0013] The iterative update module adopts a Slow-Fast strategy, iterating 12 times for low resolution and 4 times for high resolution, and finally restoring to full resolution through convex upsampling.

[0014] By combining the aforementioned weight sharing, SE channel attention mechanism, and hash-encoded multi-scale correlation body construction, this invention reduces the number of parameters by approximately 18% compared to RAFT-Stereo on the Middlebury dataset. The endpoint error (EPE) and D1-all metrics are significantly better than existing methods, and the inference speed is improved by approximately 28%. It is particularly suitable for scenarios with high requirements for real-time performance and energy efficiency, such as autonomous driving and robot vision.

[0015] The channel attention module uses SE block calibration: for feature maps The squeeze operation calculates the channel descriptor. : Then, the excitation operation computes the weights s through two fully connected layers:

[0016]

[0017] in For ReLU, for , , The reduction rate is r=16.

[0018] Correlation volume construction utilizes hash-encoded multi-scale concatenation to efficiently represent multi-resolution related information. (Constructing 3D correlation volumes) Where D is the maximum parallax range:

[0019] (Dot product calculation for visual similarity)

[0020] For multi-scale, a pyramid is constructed: k=3 levels are created in the parallax dimension through 1D average pooling (kernel size 2, step size 2), with the resolution of each level being geometrically halved (growth factor r=2), and the lowest resolution being D / 4.

[0021] Hash encoding is used for efficient interpolation: given coordinates Find the eight vertices of the voxel for each scale, and index the feature vector using a hash function.

[0022] Point features are obtained through trilinear interpolation, and multi-scale features are concatenated as input to the GRU.

[0023] The sampling interval increases exponentially: step size Δd = d / 256, interval [1 / 1024, S / 1024] (S is the maximum disparity), maximum number of samples 1024. The sampling process terminates when the transmittance falls below the threshold 1e-4.

[0024] The iterative update module updates the relevant volume and the current disparity estimate through a GRU network, uses a contextualization module to fuse contextual features, iterates 12 times, and finally outputs a high-precision disparity map.

[0025] The iteration starts with the initial disparity d_0=0 and predicts the sequence. (N=12). Each iteration: Use the current d-indexed relevant pyramid to obtain relevant features, concatenate context features (extracted from the left image, similar to a feature encoder but with batch normalization), and input into a multi-level convolutional GRU.

[0026] The GRU maintains the multi-resolution hidden state (1 / 8, 1 / 16, 1 / 32) via upsampling / downsampling cross-connection. The highest resolution GRU outputs a disparity update Δd. .

[0027] For lightweighting, Slow-Fast updates are used: lower resolution updates are more frequent (1 / 32: 12 times, 1 / 16: 8 times, 1 / 8: 4 times), reducing computation.

[0028] The final disparity is upsampled to full resolution via convex upsampling: the full-resolution pixel values ​​are convex combinations of 3x3 coarse grid neighborhoods, with weights predicted by GRU.

[0029] The supervised loss function uses L1 loss to supervise disparity estimation and introduces a smoothing term to ensure continuity. The total loss is a weighted sum of multiple iterations.

[0030] The loss function is a weighted sum of L1 losses for the sequences:

[0031]

[0032] Introduce a second-order smoothing term: This ensures parallax continuity.

[0033] Total loss:

[0034] Specifically, this invention achieves both lightweight and high precision through the following three technical means in synergy:

[0035] 1. The left and right images use a feature extraction backbone network with fully shared weights to eliminate parameter redundancy in traditional dual-channel encoders;

[0036] 2. Insert an SE channel attention module (structure is a two-layer fully connected network with a dimensionality reduction ratio r=16 after global average pooling, and finally obtain the channel weights through Sigmoid) after each residual block to achieve adaptive recalibration of feature channels;

[0037] 3. When constructing multi-scale correlated volumes, hash encoding based on spatial coordinate hash mapping is used to reduce the dimensionality index of the disparity dimension, and a hash table is pre-built to achieve O(1) complexity voxel lookup and trilinear interpolation sampling (e.g. Figure 4 (As shown).

[0038] Experimental Verification: In one embodiment of the present invention, quantitative and qualitative experiments were conducted using the Middlebury dataset to verify the effectiveness and superiority of the method (e.g., Figure 5 (As shown in Table 1). Quantitative indicators include EPE and D1-all (lower values ​​are better). The parameters of this invention are reduced by approximately 18% (9.2M), with EPE decreasing to 1.185 and D1-all decreasing to 7.767%. Qualitative results show that the disparity map is more continuous and the edges are sharper in areas with weak texture / occlusion. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. The present invention can be further understood in conjunction with the content of the accompanying drawings. The components shown in the drawings are not necessarily drawn to scale, but the focus is on illustrating the principles of the embodiments.

[0040] Figure 1 This is an overall framework diagram of a stereo matching system according to an embodiment of the present invention.

[0041] Figure 2 This is a diagram of the feature extraction and channel attention network structure according to an embodiment of the present invention.

[0042] Figure 3 This is a structural diagram of the SE channel attention module according to an embodiment of the present invention.

[0043] Figure 4 Flowchart of hash-encoded multi-scale correlation construct according to an embodiment of the present invention

[0044] Figure 5 This document compares the disparity map performance of the embodiments of the present invention with that of existing technologies on the Middlebury dataset.

[0045] Figure 6This document compares the disparity map performance of the embodiments of the present invention with that of existing technologies on the ETH-3D dataset. Detailed Implementation

[0046] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0047] Step 1: Input standard left and right stereoscopic image pairs

[0048] Obtain the left image I_L and right image I_R after epipolar correction, and perform normalization preprocessing on the images.

[0049] Step 2: Shared Feature Extraction and SE Attention Calibration (e.g.) Figure 2 , Figure 3 (As shown)

[0050] A shared-weight feature extractor is used to process I_L and I_R, with a structure of 6 residual blocks. Each residual block is followed by an SE channel attention module (reduction r=16).

[0051] Each residual block contains two 3×3 convolutional layers (padded with 1), instance normalization, and a ReLU activation function. The number of channels gradually increases from 3 in the input to 256: the first two blocks have 64 channels (preserving resolution), the third and fourth blocks have 128 channels (downsampled with stride=2), and the fifth and sixth blocks have 256 channels (downsampled with stride=2), ultimately outputting a 1 / 8 resolution feature map (H / 8 × W / 8 × 256). Shared weights avoid parameter duplication between the left and right independent encoders, reducing redundancy by approximately 50%.

[0052] The SE module structure is as follows: Figure 3 As shown: First, the channel descriptor is calculated using global average pooling. The formula is:

[0053] in For the c-th channel feature, The feature map space size is then used. Weights are then generated via a two-layer fully connected network. .in δ is The function, σ, is the Sigmoid function, and r=16. Final scaling feature. This operation dynamically emphasizes the effective channel and suppresses noise interference.

[0054] Step 3: Construction of multi-scale hash correlation (e.g.) Figure 4 (As shown)

[0055] Based on the calibrated feature map and Construct a correlation pyramid with k=3 levels. The core formula for the correlation value is as follows: The pyramid is created in the parallax dimension using 1D average pooling (kernel size 2, step size 2), with the resolution geometrically halved at each level (growth factor r=2).

[0056] Hash encoding based on spatial coordinate hash mapping accelerates indexing: hash function

[0057] A pre-built hash table (as shown in the right figure) stores all possible shift features. The sampling interval increases exponentially. The sampling interval is [1 / 1024, S / 1024] (S=192), with a maximum sampling of 1024 points. Sub-pixel-level correlation values ​​are obtained from the vertex features of hash table 8 using trilinear interpolation, achieving efficient multi-scale feature extraction.

[0058] Step 4: Iterative GRU Updates and Slow-Fast Strategy

[0059] Multi-scale relevance features are input into the iterative update module (N=12 times), starting from an initial disparity d_0=0. Each iteration uses a multi-level convolutional GRU to fuse relevance features and contextual features (extracted from the left image, 256 channels), outputting a disparity update. Maintain multi-resolution hidden states (1 / 8, 1 / 16, 1 / 32) through upsampling / downsampling cross-connection.

[0060] A slow-fast strategy is adopted: 12 iterations for low resolution (1 / 32), 8 iterations for medium resolution (1 / 16), and 4 iterations for high resolution (1 / 8) to reduce the computational load for high resolution. The final disparity is recovered to full resolution through convex upsampling: the pixel value is obtained by combining the convex neighborhood of a 3×3 coarse grid, and the combination weight is predicted by the last layer of GRU.

[0061] Step 5: Loss Function and Training

[0062] All iterative outputs are supervised using a weighted L1 loss and a second-order smoothing term: Where γ=0.9 (higher weights in later iterations), λ=0.01 (smoothing intensity), and α is the image gradient sensitivity factor. The optimizer (learning rate 1e-4, weight decay 1e-4), batch size 8, was pre-trained on the SceneFlow dataset and then fine-tuned in Middlebury for a total of 200 epochs.

[0063] Step 6: Experimental verification (e.g.) Figure 5 (As shown in Table 1)

[0064] In one embodiment of the present invention, the model was tested on the Middlebury dataset. The number of model parameters was 9.2M, which is about 18% less than RAFT-Stereo (11.2M→9.2M). The endpoint error (EPE) was 1.185 and the D1 error was 7.767%, both of which are better than RAFT-Stereo and other existing mainstream methods.

[0065] The superior performance mentioned above is due to the synergistic effect of the weight-sharing feature extractor, the SE channel attention mechanism, and the hash-encoded multi-scale correlation body construction, which achieves a good balance between lightweight and high accuracy.

[0066]

[0067] Table 1.

Claims

1. A lightweight stereo matching method based on weight sharing and channel attention, characterized in that, The method comprises the following steps: Step 1: input a set of left and right stereoscopic image pairs after polar correction; Step 2: perform feature extraction on the left and right images through a weight-shared feature extractor to obtain shared feature maps; Step 3: use a channel attention mechanism to calibrate the channel weights of the shared feature maps at preset positions associated with at least part of residual blocks of the feature extractor; Step 4: based on the calibrated shared feature maps, construct a multi-scale correlation body, wherein the multi-scale correlation body uses a spatial hash coding mode to sparsely index the disparity dimension and obtains relevant features through interpolation sampling; Step 5: input the multi-scale correlation body into an iterative update module, refine the disparity through a gated recurrent unit, and output a final disparity map.

2. The lightweight stereo matching method based on weight sharing and channel attention according to claim 1, characterized in that, The channel attention mechanism is a squeeze-excitation (SE) channel attention mechanism, and the positions of the channel weight calibration include preset positions inside each residual block of the feature extractor and an output head of the shared feature extractor.

3. The lightweight stereo matching method based on weight sharing and channel attention according to claim 1, characterized in that, The multi-scale correlation body construction process comprises: setting a scale k=3; encoding a parallax dimension through a hash function based on a hash mapping of a spatial coordinate; the hash function is based on XOR operation and modulo operation of a coordinate (x, y, d), and the hash function is specifically Wherein d is a parallax value, p1 and p2 are preset prime number coefficients, T is the size of a hash table, and represents XOR operation; based on the hash index, exponential sampling and trilinear interpolation are used to obtain a multi-scale correlation feature.

4. The lightweight stereo matching method based on weight sharing and channel attention according to claim 1, characterized in that, The iterative update module adopts a Slow-Fast strategy, iterates 12 times at a low resolution (1 / 32) and iterates 4 times at a high resolution (1 / 8); and the final disparity is obtained through convex upsampling (3x3 grid convex combination, weight predicted by GRU) to obtain a full-resolution output.

5. The lightweight stereo matching method based on weight sharing and channel attention according to claim 1, characterized in that, The training adopts a supervised loss function combining a weighted L1 loss and a smoothing term; the total loss is a weighted sum of the losses calculated for all intermediate disparity prediction results output by the iterative update module, wherein the loss weight corresponding to each iteration increases exponentially with the increase of the iteration number, so as to strengthen the supervision on the fine disparity in the later period.

6. The lightweight stereo matching method based on weight sharing and channel attention according to claim 1, characterized in that, The method comprises the following steps: Step 1: an image input module for acquiring and inputting left and right stereoscopic image pairs; Step 2: a shared feature extraction module for performing feature extraction on the left and right images through a weight-shared residual block and calibrating the channel weights of the extracted features using an SE channel attention mechanism; Step 3: a multi-scale hash correlation body construction module for sparsely indexing the disparity dimension using a spatial hash coding mode and constructing a multi-scale correlation body; Step 4: an iterative GRU update module for iteratively processing the multi-scale correlation body and refining the disparity through a gated recurrent unit; Step 5: a disparity output module for outputting a final disparity map.

7. The modules are sequentially connected to implement the method of any one of claims 1-5.