A self-supervised stereo matching method based on reconstruction loss

CN115830082BActive Publication Date: 2026-09-25TONGJI UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211418961.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-14
Publication Date
2026-09-25
Estimated Expiration
2042-11-14

AI Technical Summary

Technical Problem

对于大型可标记训练数据集,手动数据标记具有费时低效等缺点

Benefits of technology

[0034](1)本发明提出了更具通用性的立体匹配方法,引入重构损失概念,如果视差预测准确,则重构后的像素RGB值与左图相近;视差预测错误,则重构后的图像与原图差别较大,重构损失较大,模型则因此进行迭代修改,从而优化网络性能,减少了立体匹配网络对视差真值的依赖。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115830082B_ABST
    Figure CN115830082B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of self-supervised stereo matching methods based on reconstruction loss, comprising the following steps: obtaining stereo image pair information, the stereo image pair information includes a left image and a right image, and image depth information;Establish end-to-end stereo matching network PSMNet;Stereo image pair information is input into end-to-end stereo matching network PSMNet, and the predicted disparity value is obtained;Based on the predicted disparity value, reconstruct the right image, and calculate the reconstruction loss between the reconstructed right image and left image;Based on the reconstruction loss, the end-to-end stereo matching network PSMNet is iteratively trained by back propagation;Stereo image is matched using the end-to-end stereo matching network PSMNet trained completely.Compared with prior art, the present application reduces the dependence of stereo network on disparity true value, and improves the accuracy and stability of self-supervised process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot vision, and in particular to a self-supervised stereo matching method based on reconstruction loss. Background Technology

[0002] With the steady development of smart cities and the advancement of transportation networks globally, the routine quality surveying of roads is of profound significance to the development of digital cities. A crucial step in the 3D reconstruction of roads is stereo matching of the images.

[0003] Fully supervised learning algorithms for stereo matching require labeled data to analyze and predict results from unknown data. For large, labelable training datasets, manual data labeling is time-consuming and inefficient. Furthermore, for datasets where ground truth values ​​are difficult to obtain, such as in stereo matching of road surface information, datasets with absolute ground truth values ​​are extremely rare. Virtual datasets are often used for training, but models trained on virtual data suffer from low accuracy when applied to real-world scenarios. Summary of the Invention

[0004] The purpose of this invention is to provide a self-supervised stereo matching method based on reconstruction loss, which improves the accuracy of stereo matching and reduces the dependence on the disparity ground truth, thus having practicality.

[0005] The objective of this invention can be achieved through the following technical solutions:

[0006] A self-supervised stereo matching method based on reconstruction loss includes the following steps:

[0007] Acquire stereo image pair information, which includes a left image and a right image, as well as image depth information;

[0008] Establish an end-to-end stereo matching network PSMNet;

[0009] The stereo image information is input into the end-to-end stereo matching network PSMNet to obtain the predicted disparity value;

[0010] The right image is reconstructed based on the predicted disparity value, and the reconstruction loss between the reconstructed right image and the left image is calculated.

[0011] Iterative training of the end-to-end stereo matching network PSMNet was performed based on backpropagation of reconstruction loss.

[0012] The trained end-to-end stereo matching network PSMNet is used for stereo image matching.

[0013] The end-to-end stereo matching network PSMNet performs the following steps:

[0014] The left and right images are respectively input into the CNN convolutional neural network module for preliminary feature extraction;

[0015] The preliminary feature extraction results from the left and right images are sequentially input into the spatial pyramid pooling module and the convolution module to extract information near the feature points and obtain feature maps.

[0016] The left and right feature maps are concatenated into a cost loss matrix;

[0017] The cost loss value is obtained by fusing the cost loss matrix based on 3D CNN and upsampling;

[0018] The predicted disparity value is obtained by disparity regression calculation based on the cost loss value and depth information.

[0019] The first three layers of the CNN convolutional neural network module are standard convolutional layers, and the last four layers are standard convolutional layers and dilated convolutional layers using residual structures.

[0020] The spatial pyramid pooling module consists of four fixed-size average pooling layers, which shift the focus of information from global to local, and concatenates it with some results from the CNN convolutional neural network module to obtain feature information.

[0021] The step of concatenating the left and right feature maps into a cost loss matrix is ​​as follows: by connecting the left feature map and the corresponding right feature map at each disparity level, a four-dimensional matrix of height × width × disparity × channel is formed, and the cost loss matrix is ​​obtained.

[0022] The cost loss value is obtained by fusing the cost loss matrix based on the stacked hourglass convolutional network, which includes three main hourglass networks. The cost loss value is obtained by bilinear interpolation of the output of each hourglass network.

[0023] Assuming there are D disparity levels, what is the error probability c of the left and right feature maps for each disparity d? d The softmax operation σ(·) is performed to obtain the weights, c. d Greater than 0 and less than 1, where disparity is obtained based on depth information, and the error probability c of the left and right feature maps. d That is, the cost or loss value.

[0024] Then, predict the disparity value The sum of each disparity d weighted by probability:

[0025]

[0026] The larger the matching error of the current feature point under disparity d, the greater c d The larger -c d The smaller, σ(-cd The closer the function is to 0, the better. The smaller the impact.

[0027] The reconstruction of the right image based on the predicted disparity value involves shifting each pixel in the right image to the right by the size of the predicted disparity value.

[0028] When reconstructing the image on the right, the movement of its pixels is on the decimal level.

[0029] The reconstruction loss is calculated using smoothed L1 loss:

[0030]

[0031]

[0032] Where p is the pixel intensity of the point in the left image. The pixel intensity of the corresponding point in the reconstructed right image. This is the smoothed L1 loss function.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] (1) This invention proposes a more general stereo matching method and introduces the concept of reconstruction loss. If the disparity prediction is accurate, the RGB values ​​of the reconstructed pixels are similar to those of the left image; if the disparity prediction is wrong, the reconstructed image is significantly different from the original image, resulting in a large reconstruction loss. The model is then iteratively modified to optimize network performance and reduce the dependence of the stereo matching network on the true disparity value.

[0035] (2) The introduction of self-supervised networks enables the use of real datasets that cannot be used by fully supervised networks for training and testing, improving the model’s stereo matching performance in real scenarios and breaking the bottleneck of no benchmark true value in real datasets.

[0036] (3) The present invention selects the end-to-end network PSMNet, which does not require post-processing, and this network is highly applicable to road surface datasets.

[0037] (4) The use of dilated convolution in the CNN convolutional neural network module of this invention expands the receptive field without increasing the amount of computation, and does not sacrifice resolution metrics compared to the same pooling operation.

[0038] (5) This invention uses Smooth L1 Loss to calculate the reconstruction loss, which combines the advantages of mean absolute error and mean square error. It solves the disadvantage of the mean square error being discontinuous near the zero point, and also draws on the advantage of the mean absolute error having a fast convergence speed. It can ensure that the value is stable at a certain value in the early stage of training and can also converge well to a certain value at the end of training, thus ensuring gradient descent. Attached Figure Description

[0039] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0040] Figure 2 A flowchart illustrating the workflow of the end-to-end stereo matching network PSMNet;

[0041] Figure 3 This is a schematic diagram of the structure of a CNN convolutional neural network module;

[0042] Figure 4 This is a schematic diagram of the spatial pyramid pooling module;

[0043] Figure 5 Schematic diagrams of two network architectures for fusing cost loss matrices. Detailed Implementation

[0044] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0045] This embodiment provides a self-supervised stereo matching method based on reconstruction loss, such as... Figure 1 As shown, it includes the following steps:

[0046] 1) Obtain stereo image pair information, which includes a left image and a right image, as well as image depth information.

[0047] 2) Establish an end-to-end stereo matching network PSMNet.

[0048] The workflow diagram of the end-to-end stereo matching network PSMNet is as follows: Figure 2 As shown, it performs the following steps:

[0049] 21) Input the left and right images into the CNN convolutional neural network module for preliminary feature extraction.

[0050] The structure of a CNN convolutional neural network module is as follows: Figure 3As shown, the first three layers are standard convolutional layers with a kernel size of 3*3, a channel size of 32, and a stride of 2, downsampling the original image size to half. The last four layers are standard convolutional layers and dilated convolutional layers using residual structures. `conv1_x` is repeated 3 times, `conv2_x` is repeated 16 times, downsampling the original image size to half, and has 64 channels; `conv3_x` and `conv4_x` are repeated 3 times using dilated convolution with a dilation rate of 2. The use of dilated convolution expands the receptive field without increasing computational cost, and compared to the same pooling operation, it does not sacrifice resolution.

[0051] 22) Input the preliminary feature extraction results of the left and right images into the spatial pyramid pooling module and the convolution module respectively to extract information near the feature points and obtain feature maps.

[0052] The structure of the spatial pyramid pooling module is as follows: Figure 4 As shown, it consists of four fixed-size average pooling layers: 64×64, 32×32, 16×16 and 8×8, with the pooling size decreasing from large to small. This shifts the focus of information from global to local, and it is concatenated with some results from the CNN convolutional neural network module. The final output includes features of information around pixels at different levels of the average pooling layer, as well as information from different receptive fields of the CNN module.

[0053] 23) Concatenate the left and right feature maps into a cost loss matrix: By connecting the left feature map and the corresponding right feature map at each disparity level, a four-dimensional matrix of height × width × disparity × channel is formed, and the cost loss matrix is ​​obtained.

[0054] 24) The cost loss value is obtained by fusing the cost loss matrix based on 3D CNN and upsampling.

[0055] When fusing the cost loss matrix, there are two optional architectures, such as Figure 5 As shown, one is a basic architecture, and the other is a stacked hourglass convolutional network. In this network, the broken line information represents copying and adding the information of the current layer to the neural network layer corresponding to the arrow.

[0056] The infrastructure consists of 12 3×3×3 convolutional layers stacked by residual structures, and then the feature maps are sampled to size H×W×D by bilinear interpolation.

[0057] The stacked hourglass convolutional network consists of three main hourglass networks. The outputs of each hourglass network are subjected to bilinear interpolation to obtain the cost loss value. Each hourglass network outputs a disparity map. The stacked hourglass network can utilize and fuse feature information at different scales.

[0058] 25) Based on the cost loss value and depth information, disparity regression is performed to obtain the predicted disparity value.

[0059] Assuming there are D disparity levels, what is the error probability c of the left and right feature maps for each disparity d? d The softmax operation σ(·) is performed to obtain the weights, c. d Greater than 0 and less than 1, where disparity is obtained based on depth information, and the error probability c of the left and right feature maps. d That is, the cost or loss value.

[0060] Then, predict the disparity value The sum of each disparity d weighted by probability:

[0061]

[0062] The larger the matching error of the current feature point under disparity d, the greater c d The larger -c d The smaller, σ(-c d The closer the function is to 0, the better. The smaller the impact.

[0063] This disparity regression calculation method is more robust than classification-based stereo matching methods.

[0064] 3) Input the stereo image pair information into the end-to-end stereo matching network PSMNet to obtain the predicted disparity value.

[0065] 4) Move each pixel in the right image to the right by the pixel whose disparity value is predicted, complete the reconstruction of the right image, and calculate the reconstruction loss between the reconstructed right image and the left image.

[0066] In the reconstruction of the right image, the movement of pixels is at the decimal level, and there is no process of converting the model calculation result to an integer before moving them.

[0067] The reconstruction loss is calculated using the smooth L1 loss:

[0068]

[0069]

[0070] Where p is the pixel intensity of the point in the left image. The pixel intensity of the corresponding point in the reconstructed right image. This is the smoothed L1 loss function.

[0071] 5) Iteratively train the end-to-end stereo matching network PSMNet based on backpropagation of reconstruction loss.

[0072] 6) Use the trained end-to-end stereo matching network PSMNet to match stereo images.

[0073] Based on the above method, this embodiment conducted training and testing on data with and without perspective transformation, as well as on existing fully supervised models and the model described in this invention.

[0074] A. Evaluation Indicators

[0075] The following parameters were used to test and evaluate the stereo matching results in this embodiment.

[0076] (1) Average pixel error

[0077] The average pixel error represents the pixel difference between the network's predicted disparity and the baseline ground truth. It represents the average level of the network's predictions, and its specific calculation method is shown in the following formula:

[0078]

[0079] Where m is the number of pixels in the image. For example, in a specific experiment, the input PSMNet image was randomly cropped for data augmentation, and the final cropped size was 640×480, then m would be 307200. And predisp... i This represents the disparity prediction result of PSMNet for the i-th pixel, gt i This represents the true disparity value for this pixel. The absolute difference between the predicted and true disparity values ​​represents the difference between the predicted and true disparity values ​​for each pixel. Summing these values ​​and dividing by the number of pixels gives the average pixel error. For example, if the average pixel error is 1.488, then the network error accuracy can be said to be 1.488.

[0080] (2) 1 pixel error

[0081] Unlike the average pixel error, the 1-pixel error represents the percentage of pixels in the network's prediction results whose prediction disparity is within 1 pixel. Its specific calculation method is shown in the following formula:

[0082]

[0083] In the field of image processing, the original error is often multiplied by 100, and the percentage sign is omitted. For example, if the 1-pixel error result of stereo matching is 5.314, it means that in the entire image, 5.314% of the pixels have an absolute value greater than 1 pixel for the difference between the prediction error and the true disparity value.

[0084] (3) 3-pixel error

[0085] A 1-pixel error condition can sometimes be too stringent. In the field of stereo matching, a 3-pixel error is often used for error testing. The specific calculation method is shown in the following formula:

[0086]

[0087] The tolerance for a 3-pixel error is much higher than that for a 1-pixel error, so this evaluation metric is more commonly used.

[0088] B. Test Results

[0089] Table 1 shows the test results of existing fully supervised models trained and validated based on ground truth, and Table 2 shows the test results of the method described in this invention. In the table, the numbers in parentheses represent the corresponding epochs, the numbers represent the performance at this epoch, and ▲ represents the global optimum. The last row of the table shows the test results with the reconstruction loss introduced.

[0090] Table 1

[0091]

[0092] Table 2

[0093]

[0094] As shown in the table above, the method described in this invention can effectively reduce various system errors, but it is still slightly inferior to existing fully supervised networks that introduce reconstruction loss. The essence of reconstruction loss error is calculating the similarity between pixels in the left image and pixels in the right image corresponding to the disparity distance. However, for single-color regions or repetitive texture regions, the loss costs calculated by reconstruction loss may be extremely similar. For example, pixel block A in the left image and pixel block B in the right image should have a one-to-one correspondence, but due to the high similarity of the surrounding scene or the indistinct texture region, the stereo matching network matches pixel block A with pixel block C in the right image and calculates the predicted disparity D. If it is a fully supervised stereo matching training, the loss cost between the predicted disparity D and the ground truth disparity D' is calculated, indicating a high loss cost. The model prediction is inaccurate, and the disparity result is transmitted in reverse to allow the model to iterate in the correct direction. However, if the self-supervised stereo matching method described in this invention is used, the high similarity between pixel block B and pixel block C may result in a low reconstruction loss, causing the network model to believe that the match is correct, and the network architecture parameters will not be effectively updated. The drawback of reconstruction loss is not only present in this method, but also in all unsupervised model networks. Therefore, the performance of the self-supervised network described in this invention, as shown in the test results, is slightly inferior to the fully supervised network that introduces reconstruction loss. However, the method described in this invention can be well applied to test scenarios that do not require ground truth, breaking the bottleneck of real datasets lacking benchmark ground truth.

[0095] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A self-supervised stereo matching method based on reconstruction loss, characterized in that, Includes the following steps: Acquire stereo image pair information, which includes a left image and a right image, as well as image depth information; Establish an end-to-end stereo matching network PSMNet; The stereo image information is input into the end-to-end stereo matching network PSMNet to obtain the predicted disparity value; The right image is reconstructed based on the predicted disparity value, and the reconstruction loss between the reconstructed right image and the left image is calculated. Iterative training of the end-to-end stereo matching network PSMNet was performed based on backpropagation of reconstruction loss. Stereo image matching is performed using the trained end-to-end stereo matching network PSMNet. The end-to-end stereo matching network PSMNet performs the following steps: The left and right images are respectively input into the CNN convolutional neural network module for preliminary feature extraction; The preliminary feature extraction results from the left and right images are sequentially input into the spatial pyramid pooling module and the convolution module to extract information near the feature points and obtain feature maps. The left and right feature maps are concatenated into a cost loss matrix; The cost loss value is obtained by fusing the cost loss matrix based on 3D CNN and upsampling; The predicted disparity value is obtained by disparity regression calculation based on the cost loss value and depth information.

2. The self-supervised stereo matching method based on reconstruction loss according to claim 1, characterized in that, The first three layers of the CNN convolutional neural network module are standard convolutional layers, and the last four layers are standard convolutional layers and dilated convolutional layers using residual structures.

3. The self-supervised stereo matching method based on reconstruction loss according to claim 1, characterized in that, The spatial pyramid pooling module consists of four fixed-size average pooling layers, each outputting pooling features at different scales. The pooling features output by each average pooling layer are concatenated with some output features from the CNN convolutional neural network module to obtain spatial pyramid pooling feature information.

4. The self-supervised stereo matching method based on reconstruction loss according to claim 1, characterized in that, The step of concatenating the left and right feature maps into a cost loss matrix is ​​as follows: by connecting the left feature map and the corresponding right feature map at each disparity level, a four-dimensional matrix of height × width × disparity × channel is formed, and the cost loss matrix is ​​obtained.

5. The self-supervised stereo matching method based on reconstruction loss according to claim 1, characterized in that, The cost loss value is obtained by fusing the cost loss matrix based on the stacked hourglass convolutional network, wherein the stacked hourglass convolutional network includes three hourglass networks stacked in series, and the cost loss value is obtained by bilinear interpolation of the output of the hourglass networks.

6. The self-supervised stereo matching method based on reconstruction loss according to claim 1, characterized in that, Assuming there are parallax levels For each parallax Error probabilities of left and right feature maps Perform softmax operation Obtain the weights. Greater than 0 and less than 1, where disparity is obtained based on depth information, and the error probability of the left and right feature maps. That is, the cost or loss value. Then, predict the disparity value For each disparity weighted by probability The sum: Current feature point in disparity The larger the matching error, the better. The larger, The smaller, The closer the function is to 0, the better. The smaller the impact.

7. The self-supervised stereo matching method based on reconstruction loss according to claim 1, characterized in that, The reconstruction of the right image based on the predicted disparity value specifically involves shifting each pixel in the right image horizontally to the right by a distance equal to the predicted disparity value.

8. The self-supervised stereo matching method based on reconstruction loss according to claim 7, characterized in that, When reconstructing the image on the right, the movement of its pixels is on the decimal level.

9. The self-supervised stereo matching method based on reconstruction loss according to claim 1, characterized in that, The reconstruction loss is calculated using smoothed L1 loss: in, The pixel intensity of the pixel in the left image. The pixel intensity of the corresponding pixel in the reconstructed right image. The function is the smoothing L1 loss function, where N is the total number of pixels in the image.

Citation Information

Patent Citations

  • End-to-end stereo matching method based on convolutional neural network

    CN111696148A

  • Binocular deep learning method based on adaptive single-peak stereo matching cost filtering

    CN111709977A

  • Semi-supervised learning three-dimensional reconstruction method based on relative depth training

    CN113762358A

  • End-to-end binocular stereo matching method based on convolutional neural network

    CN114972822A