Binocular stereo matching method based on improved CFNet

CN117635989BActive Publication Date: 2026-09-29JIANGSU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311660788.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-06
Publication Date
2026-09-29
Estimated Expiration
2043-12-06

AI Technical Summary

Benefits of technology

[0043]1、本发明使用改进后的EfficientNetV2来提取特征,能够提取信息更丰富的特征图,为后续步骤提供基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117635989B_ABST
    Figure CN117635989B_ABST
Patent Text Reader

Abstract

The application discloses a binocular stereo matching method based on an improved CFNet, utilizes an optimized EfficientNetV2-M model as a feature extraction network to extract a multi-scale feature map corresponding to an original left-right image pair, utilizes a grouping correlation and a series connection to construct a multi-scale cost volume and fusion, adds a guided cost volume excitation to a 3D convolution module to guide cost aggregation, uses a cost self-reorganization strategy to redistribute the cost quantity in the aggregated cost volume, and then generates an initial disparity map by using disparity regression; a cascaded cost volume is constructed and combined with the initial disparity map to refine the disparity in a coarse-to-fine manner, and finally, a refined disparity map is obtained. Compared with a traditional CFNet method, the application fully improves the feature expression capability of an image, enriches the information contained in the cost volume, and effectively reduces the generation of disparity smoothing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of stereo matching and deep learning technology, specifically to a binocular stereo matching method based on an improved CFNet. Background Technology

[0002] Stereo matching methods aim to construct disparity maps from a pair of calibrated stereo images and have wide applications in autonomous driving, object detection, robotics, and 3D scene reconstruction. Accurate disparity estimation through stereo matching is crucial for improving the safety, efficiency, and decision-making capabilities of autonomous systems. Therefore, much research on stereo matching focuses on exploring various techniques to further improve its accuracy.

[0003] Traditional stereo matching methods are broadly classified into two categories: global methods and semi-global methods. Global methods construct a global energy function based on the smoothness assumption and use optimization methods to find the minimum point of the energy function. This method is computationally intensive, time-consuming, and has poor real-time performance. Semi-global methods utilize local information, calculating the total matching cost within a matching window of a specific size, and using a winner-take-all (WTA) strategy to find the minimum value to obtain the disparity value. This method is significantly affected by the window size; a window that is too small will fail to capture all the texture information of the object, leading to ambiguity; a window that is too large will produce an expansion effect in areas of discontinuous depth and increase computational complexity.

[0004] Over the past decade, Convolutional Neural Networks (CNNs) have been widely applied in computer vision tasks due to their superior ability to automatically learn hierarchical features from raw images. The application of CNNs in end-to-end stereo matching frameworks has made significant progress, outperforming traditional stereo matching methods. In 2015, Jure Zbontar and Yann LeCun first introduced CNNs into stereo matching, designing a deep Siamese network to compute the matching cost and then using 9×9 tiles to learn tile similarity, achieving higher accuracy than traditional methods. In 2017, Amit Shaked and Lior Wolf proposed a high-speed network to compute the matching cost and a global disparity network to predict disparity confidence scores. In the same year, Alex Kendall et al. proposed the end-to-end GC-Net, using a 3D convolutional neural network combining multi-scale features to adjust the matching cost, and finally obtaining a high-precision disparity map through disparity regression. In 2018, Jiaren Chang et al. proposed the Pyramid Stereo Matching Network (PSMNet), which aggregates context at different scales and locations using spatial pyramid pooling (SPP) modules before constructing the cost map, and combines this with a stacked hourglass-shaped 3D convolutional neural network to better utilize contextual information, thereby obtaining an accurate disparity map. In 2021, Zhelun Shen et al. proposed CFNet, which enhances the network's robustness and further improves its accuracy by constructing multi-scale cost convolutions.

[0005] However, many existing stereo matching networks face challenges in extracting accurate features, often failing to extract feature maps rich in structural information; moreover, current stereo matching networks often struggle to effectively extract structural information from the cost convolution during the cost aggregation stage; furthermore, the disparity refinement modules of most current networks often result in excessive disparity smoothing due to the construction of convolutional neural networks.

[0006] For example, patent 2022103753661 discloses a binocular stereo matching method based on multi-scale feature extraction and adaptive aggregation. However, fusing the cost volume after cost aggregation easily leads to the loss of effective information. Furthermore, the use of a convolutional neural network in the refinement module causes excessive smoothing of the disparity map, affecting the accuracy of stereo matching. Patent 2022108273228 discloses a binocular visual stereo matching network system and its construction method. The ResNet network used is not efficient for feature extraction and has a high computational cost, resulting in a low-precision disparity map and long processing time. Patent 202210794705X discloses a binocular stereo matching method based on bilateral grid learning and edge loss. It first calculates a low-resolution cost volume and then restores it to the original resolution through multiple upsampling steps. This process easily loses details, affecting the stereo matching effect. Patent 2023110157770 discloses a stereo matching method and computer storage medium based on multi-scale cost volume. This technical solution only integrates semantic information into the two-dimensional feature map. When constructing the cost volume, it does not make good use of the extracted information, resulting in insufficient structural information in the cost volume, which limits the matching accuracy. Summary of the Invention

[0007] Objective of this invention: The objective of this invention is to address the shortcomings of existing technologies and provide a binocular stereo matching method based on an improved CFNet. The feature extraction network used in this invention can extract feature maps with richer information, and effectively utilizes the feature map information to enrich the cost convolution, making it rich in geometric and structural information, thereby optimizing cost aggregation. Furthermore, this invention optimizes the disparity thinning module, using cost self-reorganization to replace the convolutional network in existing methods, reducing disparity smoothing phenomena.

[0008] Technical solution: The present invention provides a stereo matching method based on an improved CFNet, comprising the following steps:

[0009] Step 1: Construct a feature extraction network. Input the original left and right images with a resolution of H*W*3 into the feature extraction network to extract multi-scale feature maps. The feature extraction network is based on the EfficientNetV2-M network model, and its encoder and decoder are improved and adjusted. First, the improved encoder is used to downsample the original image to 1 / 32 of the original image resolution, and then the improved decoder is used to gradually upsample to 1 / 2 of the original image resolution, thereby obtaining multi-scale feature maps.

[0010] Step 2: Construct and fuse multi-scale cost volumes;

[0011] The multi-scale feature maps obtained in step 1 are used to construct cost volumes using group correlation and concatenation methods. The two cost volumes are then spliced ​​and fused to obtain three combined cost volumes of different scales: 1 / 4, 1 / 8, and 1 / 16. These three cost volumes are then fused into a cost volume with a uniform resolution. The size of the fused cost volume is 1 / 4D*1 / 4H*1 / 4W*32, where 32 refers to the number of channels in the cost volume.

[0012] Step 3: Input the cost volume into the guided cost incentive aggregation module. The guided cost incentive aggregation module includes a 3D hourglass aggregation network. It uses features to generate guiding weights to filter the cost volume. After each scale change during the aggregation process, the 3D hourglass aggregation network is added to guide the cost aggregation.

[0013] Step 4: Use the cost self-reconstruction (CSR) strategy to calculate the offset of each pixel in the original image, adjust the cost distribution in the aggregated cost volume, and then use disparity regression to obtain the initial disparity map.

[0014] Step 5: Construct a cascaded cost volume and combine it with the initial disparity map to refine the disparity in a coarse-to-fine manner, obtaining the final refined disparity map; the specific method is as follows:

[0015] First, the uncertainty estimate is calculated using the initial disparity map. Then, a cost volume is constructed by combining the feature map with a resolution of 1 / 4 of the original image obtained in step 1. The cost volume is then aggregated, and disparity regression is used to obtain an intermediate disparity map. Next, a cost volume with a resolution of 1 / 2 of the original image is constructed to build a cascaded cost volume. The disparity is then refined in a coarse-to-fine manner to obtain the final disparity map.

[0016] Furthermore, in step 1, during the extraction of multi-scale feature maps, the original left image and the original right image share parameters; at the same time, intermediate stage features are retained, including features at 1 / 16, 1 / 8, 1 / 4 and 1 / 2 of the original image resolution.

[0017] Further, in step 2, when constructing the cost volume, grouped correlation and concatenated cost volumes of corresponding scales are constructed using feature maps at three different resolutions: 1 / 16, 1 / 8, and 1 / 4. These two types of cost volumes are then concatenated and fused to obtain a combined cost volume of three scales. Then, 3D convolution is used to fuse the three cost volumes at one encoder-decoder to a size of 1 / 4 of the original image resolution. The size of the fused cost volume is 1 / 4D*1 / 4H*1 / 4W*32. Figure 1 and Figure 2 As shown, when constructing the cost volume, this invention uses the disparity map obtained from the uncertainty estimation of the previous level and the left and right feature maps obtained from feature extraction to construct a larger-scale cost volume again for three-dimensional convolution.

[0018] Furthermore, the specific method for obtaining the initial disparity map using the guided cost-incentive aggregation module in step 3 is as follows:

[0019] For the fused cost volume obtained in step 2, guided cost volume activation is performed by extracting feature maps at the corresponding scale.

[0020] Then, the feature map is fed into a 3D hourglass aggregation network to generate c weights for each pixel, which facilitates filtering of the initial cost volume;

[0021] At scale s, the guided cost volume incentive is represented as:

[0022] A guide =σ(F 2D (I (s) ))

[0023]

[0024] Where σ(·) represents the sigmoid function, F 2D This indicates the use of two-dimensional point-directed convolution, and the symbol × represents broadcast multiplication. A represents the cost volume representing the inputs and outputs before and after a guided cost-incentivized operation. guide I represents the weights generated from the feature map. (s) The feature map represents the s-scale.

[0025] In this step, filtering the initial cost volume refers to extracting a 1 / 4H*1 / 4W*32 feature map from the original image for a cost volume with a scale of s (e.g., s takes values ​​of 1 / 4, 1 / 8, 1 / 16). For example, when s = 1 / 4, the cost volume has a resolution of 1 / 4D*1 / 4H*1 / 4W*32. This feature map is then processed by a sigmoid function to generate weights representing the importance of each channel of the cost volume. A broadcast multiplication is then performed between the cost volume and the weights to obtain a new cost volume. The information contained in this new cost volume is influenced by the feature map. Here, guided cost volume activation enriches the information contained in the cost volume, enabling the network to capture geometric and structural representations from the cost volume through 3D convolution. By incorporating the geometric and structural information derived from the feature map into the cost volume, sharper object edges are obtained in the final disparity map, avoiding the generation of blurred structures.

[0026] Furthermore, the specific execution process of the cost self-reorganization CSR strategy in step 4 is as follows:

[0027] First, contextual structure information is extracted from the input raw left image using a lightweight U-Net network, and dense features are generated. For each pixel in the raw left image, multiple offsets can be predicted, and each offset contains two channels to indicate the pixel coordinates of neighboring points that use these dense features.

[0028] Then, the obtained pixel coordinates are used to sample the adjacent cost distributions, and they are combined to update the central cost distribution. This process is represented as:

[0029]

[0030] Where N is a hyperparameter used to control the number of neighbors participating in the reassembly process; 0 < m i <1 indicates the modulation factor of the i-th neighboring point, which is determined by the output of the sigmoid function, C o and C r They respectively represent having d max The initial and refined cost distribution of the channel, C o and C r They are converted to disparity through disparity regression respectively; the predicted offset (Δx) i Δy i () is a small value, and the sampled value is obtained through bilinear interpolation.

[0031] Furthermore, the above method for disparity regression is as follows:

[0032] First, let the cost c d Take the negative value and use the softmax function σ(·) on -c d Normalization is performed to output the probability corresponding to each disparity d; then, the predicted disparity... The calculation involves multiplying each disparity d by its corresponding probability, and the process is expressed by the following formula:

[0033]

[0034] Using cross-entropy and smooth L1 The loss function is used to train the network, and is defined as follows:

[0035]

[0036]

[0037] Where I is the number of pixels in the image. Let P(d) represent the predicted distribution from softmax, where P(d) is the true distribution and D is the true disparity value. The predicted disparity is calculated using a weighted average, and dmax is the maximum disparity value.

[0038] Finally, the complete loss function is:

[0039]

[0040] Where λ and ε are both balancing weights, and M refers to the number of output disparity maps. In this invention, the two intermediate disparity map outputs and the final disparity map are supervised, so M = 3.

[0041] The aforementioned cost quantity c d The cost convolution is performed multiple times using 3D convolutions to make the number of channels C equal to 1, reflecting the cost value corresponding to each pixel (x, y). This value is used to obtain the maximum disparity value d for each pixel from 0 to 192. max The disparity d is calculated by taking the probability distribution of the values ​​and then summing them by weight to obtain the disparity value of the current pixel. The disparity d ranges from 0 to 192 (the maximum disparity value). Since the probability of each disparity value at each pixel (x, y) is different, these disparity values ​​need to be multiplied by their probabilities and then summed to obtain the disparity value of the current pixel.

[0042] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0043] 1. This invention uses an improved EfficientNetV2 to extract features, which can extract feature maps with richer information, providing a foundation for subsequent steps.

[0044] 2. This invention designs a guided cost volume excitation module and applies it to the cost aggregation stage. It uses feature map information to enrich the cost volume, making it rich in geometric and structural information, thereby optimizing cost aggregation.

[0045] 3. The present invention designs a cost self-reorganization module to redistribute the cost distribution, effectively addressing the problem of excessive smoothing of disparity maps. Attached Figure Description

[0046] Figure 1 This is a schematic diagram of the overall binocular stereo matching network structure of the present invention;

[0047] Figure 2 This is a structural diagram of the multi-scale cost volume fusion module in an embodiment of the present invention;

[0048] Figure 3 This is a diagram of the 3D convolutional network structure in an embodiment of the present invention;

[0049] Figure 4 This is a comparison diagram of the parallax results between the technical solution of the present invention and the prior art. Detailed Implementation

[0050] The technical solution of the present invention will be described in detail below, but the scope of protection of the present invention is not limited to the embodiments described.

[0051] This invention presents a stereo matching method based on an improved CFNet. Using the CFNet stereo matching network as a prototype, it leverages EfficientNetV2 to extract multi-scale feature information, improving the quality of the feature maps. It constructs multi-scale cost volumes and fuses them to a unified resolution. A guided cost volume activation module is designed to enhance the structural information representation capability of the cost volumes. Furthermore, a cost self-reorganization strategy module is designed to redistribute the cost distribution, effectively addressing the problem of excessive smoothing in disparity maps. In summary, this invention effectively improves the matching performance of disparity maps at object details and reduces the mismatch rate of disparity maps.

[0052] like Figure 1 As shown, the stereo matching method based on the improved CFNet in this embodiment includes the following steps:

[0053] Step 1: Construct a feature extraction network. Input the original left and right images into the feature extraction network to extract multi-scale feature maps. The two original images are H*W*3 in size. The feature extraction network is based on the EfficientNetV2-M network model. Its encoder and decoder are improved and adjusted. First, the improved encoder is used to downsample the original image to 1 / 32 of the original image resolution. Then, the improved decoder is used to gradually upsample to 1 / 2 of the original image resolution, thereby obtaining multi-scale feature maps. These multi-scale left and right feature maps play a key role in the subsequent generation of cost volumes of corresponding scales.

[0054] Step 2: Construct and fuse multi-scale cost volumes; Construct cost volumes from the multi-scale feature maps obtained in Step 1 using group correlation and concatenation methods respectively. Concatenate and fuse the two types of cost volumes to obtain combined cost volumes of three scales: 1 / 4, 1 / 8, and 1 / 16. Then fuse the cost volumes of multiple scales into a cost volume of 1 / 4 of the original image resolution. The cost volume size is 1 / 4D*1 / 4H*1 / 4W*32.

[0055] Step 3: Input the cost volume into the guided cost incentive aggregation module. The guided cost incentive aggregation module includes a 3D hourglass aggregation network. It uses features to generate guiding weights to filter the cost volume. After each scale change during the aggregation process, the 3D hourglass aggregation network is added to guide the cost aggregation.

[0056] Step 4: Use the cost self-reassembling (CSR) strategy to calculate the offset of each pixel in the original left image, adjust the cost distribution in the aggregated cost volume, and then use disparity regression to obtain the initial disparity map.

[0057] Step 5: Construct a cascaded cost volume and combine it with the initial disparity map to refine the disparity in a coarse-to-fine manner, obtaining the final refined disparity map. Specifically, first, calculate the uncertainty estimate using the initial disparity map; then, construct a cost volume using a feature map obtained from feature extraction at 1 / 4 the original image resolution, and aggregate this cost volume. Use disparity regression to obtain an intermediate disparity map; next, construct a cost volume at 1 / 2 the original image resolution, thus constructing a cascaded cost volume, and refine the disparity in a coarse-to-fine manner to obtain the final disparity map. Figure 1 As can be seen, the feature maps used in the two constructions of the cascaded cost volume are 1 / 4 and 1 / 2 the size of the original image, respectively.

[0058] The feature extraction network in step 1 above is based on the EfficientNetV2-M model. Before downsampling to 1 / 4 of the original image resolution, the number of Fused-MBConv blocks in the network model is increased, enabling the feature extraction network to acquire more high-resolution feature information. The number of convolutional kernels processing 1 / 4 of the original image resolution is also increased, allowing the network to focus more on high-resolution features. To reduce computational burden and prioritize key regions, this embodiment chooses to reduce the number of channels in the lower-resolution feature map.

[0059] In step 2 above, during the construction and fusion of multi-scale cost volumes, for feature maps with resolutions of 1 / 16, 1 / 8, and 1 / 4 of the original image, grouped correlation cost volumes and concatenated cost volumes at the corresponding scales are constructed using grouped correlation and concatenated methods, respectively. Then, the two cost volumes are concatenated at the corresponding scales to obtain the fused cost volume. For example... Figure 2 As shown, firstly, a 3D convolution with a stride of 2 is used to downsample the cost volume from 1 / 4 to 1 / 8 of the original resolution; then, this downsampled cost volume is concatenated with the 1 / 8 original resolution cost volume along the feature dimension; next, a 3D convolution module is used to adjust the channel dimension of the concatenated cost volume to a fixed value; after this adjustment, a similar process is used to downsample the 1 / 8 original resolution cost volume and then concatenate it with the 1 / 16 original resolution cost volume; after the channel adjustment process, progressive upsampling adjusts the resolution of the cost volume to 1 / 4 of the original image resolution; finally, the fused cost volume is input into a 3D hourglass sub-network for cost aggregation (e.g., ...). Figure 1 (As shown).

[0060] like Figure 3 As shown, the specific method for obtaining the initial disparity map using the guided cost-incentive aggregation module in step 3 is as follows:

[0061] For step 2, the cost volume (the scale of the cost volume is s, which can be 1 / 4, 1 / 8, or 1 / 16 in 3DCNN) is stimulated by extracting feature maps at the corresponding scale.

[0062] Then, the feature map is fed into a 3D hourglass aggregation network to generate c weights for each pixel, which facilitates filtering of the initial cost volume;

[0063] At scale s, the guided cost volume incentive is represented as:

[0064] A guide =σ(F 2D (I (s) ))

[0065]

[0066] Where σ(·) represents the sigmoid function, F 2D This indicates the use of two-dimensional point-directed convolution, and the symbol × represents broadcast multiplication. A represents the cost volume representing the inputs and outputs before and after a guided cost-incentivized operation. guide I represents the weights generated from the feature map. (s) The feature map represents the scale s. Here, guided cost convolution excitation is used to enrich the information contained in the cost convolution, enabling the network to capture geometric and structural representations from the cost convolution through 3D convolution. By incorporating the geometric and structural information derived from the feature map into the cost convolution, sharper object edges are obtained in the final disparity map, avoiding the generation of blurry structures.

[0067] The specific execution process of the cost self-reorganization CSR strategy in step 4 of this embodiment is as follows:

[0068] First, contextual structure information is extracted from the input raw left image using a lightweight U-Net architecture, and dense features are generated. For each pixel in the raw left image, multiple offsets can be predicted, and each offset contains two channels to indicate the pixel coordinates of neighboring points that use these dense features.

[0069] Then, the obtained pixel coordinates are used to sample the adjacent cost distributions, and they are combined to update the central cost distribution. This process is represented as:

[0070]

[0071] Where N is a hyperparameter used to control the number of neighbors participating in the reassembly process; 0 <m i <1 indicates that the modulation factor of the i-th neighbor point is determined by the output of the sigmoid function, C o and C r They respectively represent having d max The initial and refined cost distribution of the channel, C o and C r They are converted to disparity through disparity regression respectively; the predicted offset (Δx)i ,△y i () is a small value, and the sampled value is obtained through bilinear interpolation.

[0072] Example

[0073] This embodiment applies the technical solution of the present invention to a specific dataset for verification. The dataset used is the SceneFlow and KITTI datasets.

[0074] The SceneFlow dataset contains 35,454 pairs of stereo images for training and 4,370 pairs for testing. Each image pair is provided with a dense and detailed real-world disparity map and camera parameter information. All images are 960×540 resolution. A subset consists of three scenes: FlyingThings3D, a scene with randomly typed objects, including numerous floating objects with rich detail; Driving, a street scene captured simulating car driving; and Monkaa, a scene depicting monkeys in a deep forest environment, involving closer targets.

[0075] The KITTI dataset consists of two parts: KITTI 2012 and KITTI 2015. Both datasets provide real-world images captured in driving scenarios. KITTI 2012 includes 194 training and 195 test image pairs, while KITTI 2015 includes 200 training and 200 test image pairs. In both datasets, LiDAR sensors provide sparse, true disparity values ​​for the training images.

[0076] The binocular stereo matching network of this invention runs on the PyTorch deep learning framework, using two NVIDIA A6000 GPUs for training, with a batch size of 8. For all datasets, the resolution of the training stereo image pairs is set to 512×256, and the maximum disparity value d is... max Set to 192. Use the Adam optimizer, with optimization parameters set to: β1 = 0.9, β2 = 0.999.

[0077] The comparative analysis results of the technical solution of this invention with other models are shown in Table 1.

[0078] Table 1

[0079]

[0080] As shown in Table 1, the D1 index is used to analyze the matching accuracy of various neural networks. The smaller the D1 error, the higher the accuracy.

[0081] This invention performs best on the KITTI 2015 dataset, achieving 1.70% on the D1-all metric, which is 9.57% more accurate than CFNet.

[0082] like Figure 4 As shown, for the original left image in the first column of the figure, stereo matching was performed using the technical solution of this invention, the PSMNet network, and the GwcNet network, respectively. The comparison of the performance metrics clearly shows that the technical solution of this invention outperforms PSMNet and GwcNet in areas without texture or with reflective surfaces. Furthermore, because it can effectively aggregate correct matching information to challenging regions to obtain more accurate estimates, the final disparity map of this invention contains more detail and presents a clearer object structure.

Claims

1. A stereo matching method based on an improved CFNet, characterized in that, Includes the following steps: Step 1: Construct a feature extraction network. Input the original left and right images with a resolution of H*W*3 into the feature extraction network to extract multi-scale feature maps. The feature extraction network is based on the EfficientNetV2-M network model. The number of Fused-MBConv blocks is increased in its encoder, and the number of convolutional kernels that can process 1 / 4 of the original image resolution is increased in its decoder. First, the improved encoder is used to downsample the original image to 1 / 32 of the original image resolution, and then the improved decoder is used to gradually upsample to 1 / 2 of the original image resolution, thereby obtaining multi-scale feature maps; Step 2: Construct and fuse multi-scale cost convolutions; The multi-scale feature maps obtained in step 1 are used to construct cost volumes through group correlation and concatenation methods. The two cost volumes are then spliced ​​and fused to obtain three combined cost volumes of different scales: 1 / 4, 1 / 8, and 1 / 16. These are then fused into a cost volume with a uniform resolution. The size of the fused cost volume is 1 / 4D*1 / 4H*1 / 4W*32. Step 3: Input the cost volume into the guided cost-incentive aggregation module. This module includes a 3D hourglass aggregation network that uses features to generate guiding weights to filter the cost volume. The 3D hourglass aggregation network is added after each scale change during the aggregation process to guide cost aggregation. The specific method is as follows: For the cost volume obtained in step 2, guided cost volume activation is performed by extracting feature maps at the corresponding scale. Then, the feature map is fed into a 3D hourglass aggregation network to generate for each pixel. Each weight facilitates filtering of the initial cost volume; At scale s, the guided cost volume incentive is represented as: ; ; in This represents the sigmoid function. This indicates the use of two-dimensional point-directed convolution, and the symbol × represents broadcast multiplication. , The cost volume represents the inputs and outputs before and after a guided cost-incentivized operation. The weights represent those generated from the feature map. A feature map representing the s-scale; Step 4: Calculate the offset of each pixel in the original left image using the cost self-reconstruction (CSR) strategy, adjust the cost distribution in the aggregated cost volume, and then use disparity regression to obtain the initial disparity map; the specific execution process is as follows: Dense features are generated by extracting contextual structure information from the original left image using a lightweight U-Net network. Multiple offsets can be predicted for each pixel in the original left image, and each offset contains two channels, which are used to indicate the pixel coordinates of the adjacent points of the dense features. Then, the adjacent cost distributions are sampled using pixel coordinates, and the cost distributions are combined to update the central cost distribution. This process is represented as: ; in, It is a hyperparameter used to control the number of neighbors participating in the reassembly process; Indicates the first The modulation factor of each neighboring point is determined by the output of the sigmoid function. and They respectively represent having The initial and refined cost distribution of the channel. and The values ​​are converted to disparity via disparity regression; the predicted offset values ​​are... These are small values, obtained through bilinear interpolation to obtain sampled values; Step 5: Construct a cascaded cost volume and combine it with the initial disparity map to refine the disparity in a coarse-to-fine manner, obtaining the final refined disparity map; the specific method is as follows: First, the uncertainty estimate is calculated using the initial disparity map. Then, the cost volume is constructed by combining the feature map with 1 / 4 of the original image resolution obtained in step 1. The cost volumes are then aggregated, and the intermediate disparity map is obtained by disparity regression. Next, a cost volume with 1 / 2 of the original image resolution is constructed to build a cascaded cost volume. The disparity is then refined in a coarse-to-fine manner to obtain the final disparity map.

2. The stereo matching method based on improved CFNet according to claim 1, characterized in that, In step 1, during the extraction of multi-scale feature maps, the original left image and the original right image share parameters; at the same time, the intermediate stage features are retained, including features at 1 / 16, 1 / 8, 1 / 4 and 1 / 2 of the original image resolution.

3. The stereo matching method based on improved CFNet according to claim 1, characterized in that, In step 2, when constructing the cost volume, the corresponding scale grouped correlation and concatenated cost volumes are constructed using three feature maps with different resolutions, and the two cost volumes are spliced ​​and fused to obtain a combined cost volume of three scales. Then, 3D convolution is used to fuse the three cost volumes of resolution to a unified resolution in the form of an encoder-decoder to obtain a fused cost volume. The resolutions of the three feature maps are 1 / 16, 1 / 8, and 1 / 4 of the original image resolution, respectively.

4. The stereo matching method based on improved CFNet according to claim 1, characterized in that, The disparity regression method is as follows: First, the cost amount To take the negative, use the softmax function. right- Perform normalization to output each disparity. The corresponding probability; Then, the predicted disparity Calculate for each disparity Multiplying by the sum of their corresponding probabilities, the calculation process can be expressed by the formula: ; Using cross-entropy and The loss function is used to train the network, and is defined as follows: ; ; in, It is the number of pixels in the image. This represents the prediction distribution from softmax. It is the true distribution. This is the true parallax value. The predicted disparity is calculated using a weighted average. This represents the maximum disparity value. Finally, the complete loss function is: ; in, and They are all balanced weights. This refers to the number of disparity maps output.