A Real-time Stereo Matching Method Based on ZNCC Algorithm and Neural Network

By combining ZNCC algorithm and neural network, real-time stereoscopic visual matching on low-power embedded devices is achieved, solving the problems of insufficient accuracy of traditional algorithms and limited computing capabilities of neural network algorithms, and achieving a balance between speed and accuracy.

CN114972474BActive Publication Date: 2025-06-10NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210595064.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-28
Publication Date
2025-06-10
Estimated Expiration
2042-05-28

AI Technical Summary

Technical Problem

The prior art is difficult to achieve both fast and accurate stereoscopic visual matching on low-power embedded devices, and traditional algorithms are insufficient in accuracy and limited in computing capabilities of neural network algorithms.

Method used

The real-time stereo matching method based on ZNCC algorithm and neural network is adopted to quickly calculate the disparity map through the ZNCC algorithm, and the disparity map is enlarged and optimized by neural network to achieve a balance of speed and accuracy.

Benefits of technology

It realizes fast and accurate stereoscopic visual matching on low-power embedded devices, and can run in real time on embedded GPUs to meet practical application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972474B_ABST
    Figure CN114972474B_ABST
Patent Text Reader

Abstract

The present invention discloses a real-time stereo matching method based on the ZNCC algorithm and neural network, which can quickly and accurately calculate a depth image from a target image and a reference image. First, the input RGB image is downsampled and converted into a grayscale image to obtain a low-resolution grayscale image. Then, the ZNCC algorithm is used to calculate the matching cost of the low-resolution image, and an initial low-resolution disparity map is generated through the SGM algorithm, WTA algorithm, and SMP algorithm. Then, a high-resolution disparity map is obtained through the neural network algorithm. Finally, it is converted into a depth image. The method of the present invention is a hybrid system based on a traditional scheme and a neural network algorithm, wherein the traditional scheme is used to quickly generate a disparity map, and the neural network focuses on magnifying and optimizing the disparity map. Through the technical solution of the present invention, a real-time stereo vision matching system is realized, and the system has high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and particularly relates to a real-time stereo matching method based on the ZNCC algorithm and neural network. Background Art

[0002] Stereo vision matching is an important step in the 3D reconstruction process. Generally, through two or more image collectors, stereo matching is performed on a pair of images collected by the two collectors at the same moment to obtain the disparity values of each pixel point in the images, and the disparity values can be further converted into depth values. The process of stereo vision matching can be divided into two steps: depth image calculation and depth image optimization. The current stereo vision matching algorithms can be divided into traditional algorithms and neural network algorithms. Generally speaking, the stereo matching system built based on the neural network algorithm has higher accuracy, while the stereo matching system based on the traditional algorithm runs faster.

[0003] In the actual application process, limited by the power supply device, only low-power embedded devices can be used for stereo vision matching. The computing power of low-power embedded devices is limited, resulting in that the stereo vision matching system built based on the neural network algorithm is difficult to meet the actual application requirements. And directly using the traditional algorithm for stereo vision matching, its accuracy is difficult to meet the actual application requirements. Summary of the Invention

[0004] Object of the Invention: Aiming at the above-mentioned existing technologies, a real-time stereo matching method based on the ZNCC algorithm and neural network is proposed, which can quickly and accurately calculate the depth image through the target image and the reference image, and achieve the balance between speed and accuracy.

[0005] Technical Solution: A real-time stereo matching method based on the ZNCC algorithm and neural network includes:

[0006] Step 1: Downsample the input RGB image and convert it into a grayscale image to obtain a low-resolution grayscale image;

[0007] Step 2: Use the ZNCC algorithm, semi-global algorithm, winner-takes-all algorithm, and single-piece matching optimization for fast stereo vision matching to obtain a disparity map with large errors and low resolution;

[0008] Step 3: Use the neural network to magnify and optimize the disparity map obtained in Step 2 to obtain a disparity map with the original resolution;

[0009] Step 4: Convert the disparity map obtained in Step 3 into a depth image.

[0010] Further, the Step 2 includes the following sub-steps:

[0011] Step 2-1: Calculate the matching cost for the grayscale image obtained in Step 1 using the ZNCC algorithm, specifically including: Define a matching window W, and calculate the difference between the pixel points in the reference image and the target image according to the following formula:

[0012]

[0013] where,

[0014]

[0015]

[0016] and:

[0017]

[0018]

[0019] In the formula, C(x,y,d) represents the similarity between the two pixel points I R (x,y) and I T (x-d,y), which is the matching cost. (x,y) and d represent the coordinates and disparity respectively. The value of C(x,y,d) is between 0 and 1. The smaller the value, the more similar the two pixel points are; ΔI R (x,y) represents the result after subtracting the mean value of its neighborhood from the pixel point I R (x,y) in the reference image; ΔI T (x-d,y) represents the result after subtracting the mean value of its neighborhood from the pixel point I T (x-d,y) in the target image; σ R (x,y) and σ T (x-d,y) are the standard deviations of the pixel values in the matching window W; and are the average values of the pixel values in the matching window W;

[0020] Step 2-2: Aggregate the matching cost obtained in Step 2-1 through the semi-global algorithm. The specific formula is as follows:

[0021]

[0022] where, C r (x,y,d) represents the aggregation of the matching cost in the r direction, and r represents any direction; C r (x-r,y,d) represents the aggregation of the matching costs of all pixels with disparity d in front of the point (x,y) in the r direction; C r (x-t,y,d+1) represents the aggregation of the matching costs of all pixels with disparity (d+1) in front of the point (x,y) in the r direction;r (x - r, y, d - 1) represents the aggregation of the matching costs of all pixels with a disparity of (d - 1) in front of the point (x, y) in the r direction; P1 and P2 are penalty coefficients; C r (x - r, y, i) represents the aggregation of the matching costs of all pixels with a disparity of i in front of the point (x, y) in the r direction, where i is any disparity value other than d and (d ± 1); k is the disparity value that minimizes the aggregation of the matching costs of all pixels in front of the coordinate (x, y) in the r direction;

[0023] The final cost value C SGM (x, y, d) is aggregated from different directions as follows:

[0024]

[0025] Finally, an initial low - resolution disparity map is obtained through the Winner - Takes - All (WTA) algorithm. The specific formula is as follows:

[0026] D map (x, y) = argmin d (C SGM (x, y, d))

[0027] where D map (x, y) represents the value of the pixel at the coordinate (x, y) in the low - resolution disparity map;

[0028] Step 2 - 3: Optimize the low - resolution disparity map obtained in Step 2 - 2 through the unilateral matching phase algorithm. The specific constraint equation is as follows:

[0029]

[0030] where D is the maximum disparity value in the disparity map, D map (x - l, y) represents the value of the pixel at the coordinate (x - l, y) in the disparity map, and l represents any positive integer less than D; points that do not satisfy this equation will be marked as invalid points, and their disparity values will be replaced by the nearest and smallest valid values in the left - right direction.

[0031] Furthermore, Step 3 includes the following sub - steps:

[0032] Step 3 - 1: Extract feature maps from the disparity map obtained in Step 2 - 3 and the grayscale image obtained in Step 1 through a neural network. The number of channels in the convolutional layer is 32;

[0033] Step 3-2: Downsample the feature information obtained in Step 3-1 twice. The input for the first downsampling is the feature map obtained in Step 3-1, and the number of channels in the convolutional layer is 16. The input for the second downsampling is the feature information extracted by the residual block from the output of the first downsampling, and the number of channels in the convolutional layer is 32, thereby obtaining a compressed feature map;

[0034] Step 3-3: Convert the feature map obtained in Step 3-2 into a latent distribution G, and the specific formula is as follows:

[0035] G = gaussin * exp sigma-1 +(mu - 1)

[0036] where gaussin represents a Gaussian random distribution; sigma and mu are the variance and mean of the latent distribution respectively, and are obtained by reducing the dimension of the feature map obtained in Step 3-2 using a neural network;

[0037] Step 3-4: Convert the latent distribution obtained in Step 3-3 into a feature map, and the number of channels in the convolutional layer is 32;

[0038] Step 3-5: Upsample the feature map obtained in Step 3-4 twice, and the number of channels in the convolutional layer is 16. The input for the first upsampling is the concatenation of the feature map obtained in Step 3-4 and the feature information extracted by the residual block from the output of the second downsampling in Step 3-2. The input for the second upsampling is the concatenation of the feature information extracted by the residual block from the output of the first upsampling and the feature information extracted by the residual block from the output of Step 3-2. The number of channels in the convolutional layer is 16, and the output of Step 3-5 is the reconstructed feature map;

[0039] Step 3-6: Upsample the reconstructed feature map obtained in Step 3-5, the feature map obtained in Step 3-1, and the low-resolution disparity obtained in Step 2-2 Figure 1 to obtain a disparity map at the original resolution.

[0040] Furthermore, the residual block structure is a single-layer neural network, and the number of channels is half of that of the previous convolutional layer. The output of the residual block is the concatenation of the input information and the deep feature information extracted by the single-layer neural network.

[0041] Furthermore, for all neural networks, the convolutional kernel size is 5×5, the activation function is the Leaky_ReLU function, and the loss function of the network is:

[0042]

[0043] In the formula, gt represents the labeled data, Y represents the output result of the neural network, and L 1 is the loss function calculated from the output image of the real-time stereo matching method and the labeled data, and L2 The loss function is calculated based on the mean and variance of the latent distribution of the feature information obtained through neural network calculation and the standard Gaussian distribution.

[0044] Furthermore, for the neural network used, the optimizer for training is the Adam optimizer.

[0045] Beneficial effects: The present invention uses the ZNCC algorithm to calculate the disparity map and uses a neural network to optimize and amplify the disparity map. Compared with most current stereo matching systems that use neural networks to extract feature points, the present invention uses a neural network to amplify and optimize the disparity map obtained by the ZNCC algorithm. The neural network is only used in the optimization step after obtaining the disparity map. Therefore, it can better achieve the balance between speed and accuracy. In addition, the present invention uses low-resolution images for stereo matching calculation, and through the deconvolution layer, the neural network directly learns the relationship between the low-resolution disparity map and the original-resolution disparity map, which can better save memory and accelerate the operation, enabling this method to run in real time on an embedded GPU without any additional optimization steps. Description of the Drawings

[0046] Figure 1 is the flowchart of the method of the present invention;

[0047] Figure 2 are the low-resolution black-and-white image and the disparity map in the embodiment of the present invention;

[0048] Figure 3 is the network structure diagram of the method of the present invention;

[0049] Figure 4 are the original-resolution image and the depth image output by the system in the embodiment of the present invention. Detailed Embodiments

[0050] The present invention will be further described in detail below in conjunction with the embodiments and the drawings.

[0051] As Figure 1 shown, a real-time stereo matching method based on the ZNCC algorithm and a neural network includes:

[0052] Step 1: Downsample the input RGB image and convert it to a grayscale image to obtain a low-resolution grayscale image.

[0053] Step 2: Use the ZNCC algorithm for fast stereo vision matching to obtain a disparity map with large errors and low resolution, which specifically includes the following sub-steps:

[0054] Step 2-1: Calculate the matching cost of the grayscale image obtained in Step 1 using the ZNCC algorithm, which specifically includes: Define a matching window W with a size of 3×3, and calculate the difference between the pixel points in the reference image and the target image according to the following formula:

[0055]

[0056] Among them,

[0057]

[0058]

[0059] And:

[0060]

[0061]

[0062] In the formula, C(x, y, d) represents the similarity between two pixel points I R (x, y) and I T (x - d, y), which is the matching cost. (x, y) and d represent the coordinates and disparity respectively. The value of C(x, y, d) is between 0 and 1. The smaller the value, the more similar the two pixel points are; multiply the calculated matching cost by 21 and then keep the integer part; ΔI R (x, y) represents the result after subtracting the mean of its neighborhood from the pixel point I R (x, y) in the reference image; ΔI T (x - d, y) represents the result after subtracting the mean of its neighborhood from the pixel point I T (x - d, y) in the target image; σ R (x, y) and σ T (x - d, y) are the standard deviations of the pixel values in the matching window W, which are used to normalize the correlation coefficient between them; represents the average value of the surrounding pixel points in the reference image; represents the average value of the surrounding pixel points in the target image.

[0063] Step 2 - 2: Aggregate the matching cost obtained in Step 2 - 1 through the semi - global (SGM) algorithm. The specific formula is as follows:

[0064]

[0065] Among them, C r (x, y, d) represents the aggregation of the matching cost in the r direction, where r represents any direction; C r (x - r, y, d) represents the aggregation of the matching costs of all pixels with disparity d in front of the point (x, y) in the r direction; C r (x - r, y, d + 1) represents the aggregation of the matching costs of all pixels with disparity (d + 1) in front of the point (x, y) in the r direction; C r(x-r,y,d-1) represents the aggregation of the matching costs of all pixels with a disparity of (d-1) in front of the point (x,y) in the r direction; P1 and P2 are penalty coefficients. The cost of the current pixel is added to the adjacent pixels, and according to the different disparity values of the adjacent pixels, the penalty coefficient P1 or P2 is added respectively. This can fully consider the constraints of adjacent pixels, and subtracting the minimum value of the previous pixel can ensure that the cost value does not overflow; C r (x-r,y,i) represents the aggregation of the matching costs of all pixels with a disparity of i in front of the point (x,y) in the r direction, where i is any disparity value other than d and (d±1); k is the disparity value that minimizes the aggregation of the matching costs of all pixels in front of the coordinate (x,y) in the r direction. In this embodiment, the selected penalty parameter P1 is 18 and P2 is 185.

[0066] The final cost value C SGM (x,y,d) is summarized from different directions as follows:

[0067]

[0068] Finally, the initial low-resolution disparity map is obtained through the WTA algorithm. The specific formula is as follows:

[0069] D map (x,y) = argmin d (C SGM (x,y,d))

[0070] where D map (x,y) represents the value of the pixel at the coordinate (x,y) in the low-resolution disparity map.

[0071] Step 2-3: Optimize the low-resolution disparity map obtained in Step 2-2 through the Single-sided Matching Phase (SMP) algorithm to obtain an optimized disparity image. The idea of the single-sided matching phase algorithm is that any pixel in the reference image cannot be matched by two or more pixels on the target image at the same time. The specific constraint equation is as follows:

[0072]

[0073] where D is the maximum disparity value in the disparity map, and D map (x-l,y) represents the value of the pixel at the coordinate (x-l,y) in the disparity map, and l represents any positive integer less than D; points that do not satisfy this equation will be marked as invalid points, and their disparity values will be replaced by the nearest and smallest valid values in the left and right directions.

[0074] Step 3: Use a neural network to magnify and optimize the disparity map obtained in Step 2 to obtain the original-resolution disparity map, specifically including:

[0075] Step 3-1: Extract feature maps from the disparity map obtained in Step 2-3 and the grayscale image obtained in Step 1 through a neural network. The number of channels in the convolutional layer is 32.

[0076] Step 3-2: Perform 2 times of downsampling on the feature information obtained in Step 3-1. The input of the first downsampling is the feature map obtained in Step 3-1, and the number of channels in the convolutional layer is 16. The input of the second downsampling is the feature information extracted by the residual block from the output of the first downsampling, and the number of channels in the convolutional layer is 32, so as to obtain a compressed feature map.

[0077] Step 3-3: Convert the feature map obtained in Step 3-2 into a latent distribution G. The specific formula is as follows:

[0078] G = gaussin * exp sigma-1 +(mu - 1)

[0079] where, gaussin represents a Gaussian random distribution; sigma and mu are the variance and mean of the latent distribution respectively, and are obtained by reducing the dimension of the feature map obtained in Step 3-2 using a neural network.

[0080] Step 3-4: Convert the latent distribution obtained in Step 3-3 into a feature map, and the specific number of channels is 32.

[0081] Step 3-5: Perform 2 times of upsampling on the feature map obtained in Step 3-4. The specific number of channels is 16. The input of the first upsampling is the concatenation of the feature map obtained in Step 3-4 and the feature information extracted by the residual block from the output of the second downsampling in Step 3-2. The input of the second upsampling is the concatenation of the feature information extracted by the residual block from the output of the first upsampling and the feature information extracted by the residual block from the output of Step 3-2. The specific number of channels is 16. The output of Step 3-5 is a reconstructed feature map.

[0082] Step 3-6: Upsample the reconstructed feature map obtained in Step 3-5, the feature map obtained in Step 3-1, and the low-resolution disparity Figure 1 obtained in Step 2-2 to obtain a disparity map with the original resolution.

[0083] Among them, the residual block structure is a single-layer neural network, and the number of channels is half of that of the previous convolutional layer. The output of the residual block is the concatenation of the input information and the deep feature information extracted by the single-layer neural network.

[0084] In this embodiment, the training parameters of the network are as follows:

[0085] Parameter Name Parameter Value Image Size 1242*376 Batch Size 1 Initial Learning Rate 0.0005 Training Epoch 1000 Learning Rate Decay Decay by 0.03 every 10 epochs

[0086] For all neural networks, the convolution kernel size is 5×5, the activation function is the Leaky_ReLU function, and the loss function of the network is:

[0087]

[0088] wherein, gt represents the labeled data, Y represents the output result of the neural network, and L 1 is the loss function calculated by the output image of the real-time stereo matching method and the labeled data, and L 2 is the loss function calculated by the mean and variance of the hidden distribution of the feature information obtained by the neural network and the standard Gaussian distribution. The final loss function is the sum of the two.

[0089] In this embodiment, for all neural networks, the optimizer for training is the Adam optimizer, and the neural network model is established based on the weights obtained by minimizing the loss function.

[0090] Step 4: Convert the disparity map obtained in Step 3 into a depth image.

[0091] Figure 1 is the flow chart of the method of the present invention, Figure 2 (a) and (b) of which are respectively the low-resolution black and white image in the method of the present invention and the disparity map obtained in Step 2. Figure 3 is the structural diagram of the neural network in the embodiment of the present invention. Figure 4 is the original-resolution image and the depth image output by the system in the embodiment of the present invention. Figure 4 (a) of which is an image containing a transparent area, Figure 4 (b) of which is an image containing a weak texture area. The miscomputed disparity values are represented by white dots. The proposed method performs well in these two areas, and the obtained disparity map has a high correct rate.

[0092] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A real-time stereo matching method based on the ZNCC algorithm and neural network, characterized in that, it includes: Step 1: Downsample the input RGB image and convert it to a grayscale image to obtain a low-resolution grayscale image; Step 2: Use the ZNCC algorithm, semi-global algorithm, winner-takes-all algorithm, and single-piece matching phase optimization for fast stereo vision matching to obtain a disparity map with large errors and low resolution; Step 3: Use a neural network to amplify and optimize the disparity map obtained in Step 2 to obtain a disparity map with the original resolution; Step 4: Convert the disparity map obtained in Step 3 into a depth image; The said Step 2 includes the following sub-steps: Step 2-1: Calculate the matching cost of the grayscale image obtained in Step 1 using the ZNCC algorithm, specifically including: Define the matching window W, and calculate the difference between the pixel points in the reference image and the target image according to the following formula: Where, And: Where, C(x, y, d) represents I R (x, y) and I T (x - d, y). The similarity between these two pixel points is the matching cost. (x, y) and d represent the coordinates and disparity respectively. The value of C(x, y, d) ranges from 0 to 1. The smaller the value, the more similar the two pixel points are; ΔI R (x, y) represents the result after subtracting the mean value of its neighborhood from the pixel point I R (x, y) in the reference image; ΔI T (x - d, y) represents the result after subtracting the mean value of its neighborhood from the pixel point I T (x - d, y) in the target image; σ R (x, y) and σ T (x - d, y) are the standard deviations of the pixel values in the matching window W; and are the average values of the pixel values in the matching window W; Step 2-2: Aggregate the matching cost obtained in Step 2-1 through the semi-global algorithm, and the specific formula is as follows: Among them, C r (x, y, d) represents the aggregation of the matching cost in the r direction, where r represents any direction; C r (x - r, y, d) represents the aggregation of the matching cost with a disparity of d for all pixels in front of the point (x, y) in the r direction; C r (x - r, y, d + 1) represents the aggregation of the matching cost with a disparity of (d + 1) for all pixels in front of the point (x, y) in the r direction; C r (x - r, y, d - 1) represents the aggregation of the matching cost with a disparity of (d - 1) for all pixels in front of the point (x, y) in the r direction; P1 and P2 are penalty coefficients; C r (x - r, y, i) represents the aggregation of the matching cost with a disparity of i for all pixels in front of the point (x, y) in the r direction, where i is any disparity value other than d and (d ± 1); k is the disparity value that minimizes the aggregation of the matching cost for all pixels in front of the coordinate (x, y) in the r direction; Final cost value C SGM (x, y, d) are summarized from different directions as follows: Finally, obtain the initial low-resolution disparity map through the winner-takes-all (WTA) algorithm, and the specific formula is as follows: D map (x,y) = argmin d (C SGM (x,y,d)) Among them, D map (x, y) represents the value of the pixel at coordinates (x, y) in the low-resolution disparity map; Step 2-3: Optimize the low-resolution disparity map obtained in Step 2-2 through the unilateral matching phase algorithm, and the specific constraint equation is as follows: where D is the maximum disparity value in the disparity map, D map (x - l, y) represents the value of the pixel at coordinates (x - l, y) in the disparity map, and l represents any positive integer less than D; points that do not satisfy this equation will be marked as invalid points, and their disparity values will be replaced by the nearest and smallest valid values in the left - right direction; The said Step 3 includes the following sub-steps: Step 3-1: Extract the feature map of the disparity map obtained in Step 2-3 and the grayscale image obtained in Step 1 through a neural network, and the number of channels of the convolutional layer is 32; Step 3-2: Perform 2 times of downsampling on the feature information obtained in Step 3-1. The input of the first downsampling is the feature map obtained in Step 3-1, and the number of channels of the convolutional layer is 16. The input of the second downsampling is the feature information extracted by the residual block from the output of the first downsampling, and the number of channels of the convolutional layer is 32, so as to obtain a compressed feature map; Step 3-3: Convert the feature map obtained in Step 3-2 into a latent distribution G, and the specific formula is as follows: G = gaussin * exp sigma-1 +(mu - 1) Where, gaussin represents the Gaussian random distribution; sigma and mu are respectively the variance and mean of the latent distribution, and are obtained by reducing the dimension of the feature map obtained in Step 3-2 using a neural network; Step 3-4: Convert the latent distribution obtained in Step 3-3 into a feature map, and the number of channels of the convolutional layer is 32; Step 3-5: Perform 2 times of upsampling on the feature map obtained in Step 3-4, and the number of channels of the convolutional layer is 16. The input of the first upsampling is the splicing of the feature map obtained in Step 3-4 and the feature information extracted by the residual block from the output of the second downsampling in Step 3-2. The input of the second upsampling is the splicing of the feature information extracted by the residual block from the output of the first upsampling and the feature information extracted by the residual block from the output of Step 3-2, and the number of channels of the convolutional layer is 16. The output of Step 3-5 is the reconstructed feature map; Step 3-6: Upsample the reconstructed feature map obtained in Step 3-5, the feature map obtained in Step 3-1, and the low-resolution disparity map obtained in Step 2-2 together to obtain a disparity map with the original resolution.

2. The real-time stereo matching method according to claim 1, characterized in that, The residual block structure is a single-layer neural network with the number of channels being half of that of the previous convolutional layer. The output of the residual block is the concatenation of the input information and the deep feature information extracted by the single-layer neural network.

3. The real-time stereo matching method according to claim 2, characterized in that for all neural networks, the convolutional kernel size is 5×5, the activation function is the Leaky_ReLU function, and the loss function of the network is: where \(gt\) represents the labeled data, \(Y\) represents the output result of the neural network, and \(L\) 1 is the loss function calculated by the output image of the real-time stereo matching method and the labeled data, and \(L\) 2 is the loss function calculated by the mean and variance of the implicit distribution of the feature information obtained by the neural network and the standard Gaussian distribution.

4. The real-time stereo matching method according to claim 3, characterized in that for the used neural network, the optimizer for training is the Adam optimizer.

Citation Information

Patent Citations

  • Real-time stereo matching system and method based on ZNCC algorithm

    CN104794703A

  • Disparity map super-resolution method based on neural network

    CN114494016A