A 3D Visual Perception Method and System Based on the Fusion of Stereo Vision and TOF

By combining stereo vision and TOF technology, using FPGA and GPU boards to fusion in depth maps, the limitations of stereo vision and TOF are solved, and high-speed, high-precision and robust three-dimensional visual perception are achieved.

CN115714855BActive Publication Date: 2025-07-29HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211240792.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-11
Publication Date
2025-07-29
Estimated Expiration
2042-10-11

AI Technical Summary

Technical Problem

The existing stereo vision technology has problems such as ambient lighting changes, weak texture scenes and high computational complexity. The TOF technology has limited resolution and range, resulting in the visual perception method being not high-speed and robust enough.

Method used

Two-way vision imaging units are used to combine TOF imaging equipment, and the fusion of stereo vision and TOF depth map is performed through FPGA and GPU board, and the stereo vision calculation is performed using the Fast-SGM algorithm, and the depth map is fusion combined with binocular and TOF confidence.

Benefits of technology

It realizes stable three-dimensional visual perception in low light, weak texture and other scenarios, outputs a fusion depth map with high resolution and high frame rate, and has high speed, robustness and high precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115714855B_ABST
    Figure CN115714855B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of digital image acquisition and processing of robots, and specifically relates to a three-dimensional visual perception method and system based on the fusion of stereo vision and TOF, including: using two visual imaging units to acquire two RGB images; using a TOF imaging device to acquire TOF depth images; using an FPGA board and a high-speed image acquisition and processing module to perform stereo vision calculations on any one of the RGB images to obtain the stereo vision depth map corresponding to the RGB image of this path; using a GPU board and a depth fusion module based on CUDA to perform fusion processing on the stereo vision depth map and the TOF depth map to obtain a fused depth map; using the two RGB images, the stereo vision depth map, and the fused depth map as the results of three-dimensional visual perception to complete three-dimensional visual perception. The present invention has developed a high-speed and high-precision visual system based on the fusion of stereo vision and TOF, which gives full play to the advantages of the two sensors, improves the quality of the output depth map, and has the characteristics of high speed, high robustness, and high precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of robot digital image acquisition and processing, and more specifically, relates to a three-dimensional visual perception method and system based on the fusion of stereo vision and TOF. Background Art

[0002] With the rapid development of information age technology and the increasingly important role played by various robot devices in today's social production and life, the research on robots has become increasingly important. Machine vision is an important part of the development of robots.

[0003] Stereo vision technology is an important form of machine vision. Its principle is to use an imaging device to obtain multiple images of the object to be measured from different directions and angles based on parallax, and to restore the three-dimensional geometric information of the object by calculating the position deviation between corresponding pixel points and using mapping or three-dimensional reconstruction technology. The stereo vision system has the characteristics of low cost and good adaptability, and can adapt to image acquisition applications in indoor and outdoor environments. However, stereo vision also has its disadvantages. The reason is that its parallax estimation is based on the corresponding feature relationship between two images, and the specific limitations are as follows:

[0004] 1) Sensitive to environmental light. Changes in the angle and intensity of environmental light will cause a sharp decline in the effect of the stereo matching algorithm.

[0005] 2) Not applicable to monotonous scenes lacking texture information. Since the binocular stereo vision system performs image matching based on visual features, scenes lacking visual features will cause problems in matching.

[0006] 3) High algorithm complexity. Stereo matching needs to be calculated pixel by pixel, and is affected by factors such as baseline, measurement range, and shooting environment, resulting in a large amount of calculation and a long calculation time.

[0007] TOF is an active ranging method that can directly obtain the three-dimensional coordinate information of the target. Compared with stereo matching, TOF has higher accuracy, and can still obtain correct measurement results in cases such as weak texture regions and depth mutations, but there are problems of insufficient resolution and limited range.

[0008] Therefore, there is an urgent need for a high-speed and highly robust visual perception method at present. Summary of the Invention

[0009] In view of the defects and improvement requirements of the prior art, the present invention provides a three-dimensional visual perception method and system based on the fusion of stereo vision and TOF, aiming to give full play to the advantages of stereo vision sensors and TOF sensors, improve the quality of the output depth map, and achieve high-speed, robust, and high-precision three-dimensional visual perception.

[0010] To achieve the above object, according to one aspect of the present invention, a three-dimensional visual perception method based on the fusion of stereo vision and TOF is provided, including:

[0011] Two visual imaging units are used to collect two RGB images;

[0012] A TOF imaging device is used to collect TOF depth images;

[0013] An FPGA board and a high-speed image acquisition and processing module are used to perform stereo vision calculations on any one of the two RGB images based on the two RGB images to obtain a stereo vision depth map corresponding to the RGB image of this path;

[0014] A GPU board and a depth fusion module based on CUDA are used to perform fusion processing on the stereo vision depth map and the TOF depth map to obtain a fused depth map;

[0015] The two RGB images, the stereo vision depth map, and the fused depth map are used as the results of three-dimensional visual perception to complete three-dimensional visual perception.

[0016] Furthermore, the Fast-SGM algorithm suitable for FPGA implementation is used to calculate the stereo vision depth map for any one of the RGB images. The implementation method is as follows:

[0017] (a) Synchronize the data of the two RGB images;

[0018] (b) Calculate the initial cost of each image respectively;

[0019] (c) Use the initial cost to calculate the aggregated cost of each image respectively;

[0020] (d) Use the aggregated cost to perform sub-pixel interpolation calculation to obtain the final stereo vision disparity map;

[0021] (e) Based on the final disparity and the calibration parameters of the two visual imaging units, obtain the stereo vision depth map.

[0022] Furthermore, to efficiently implement step (b) in the stereo vision calculation, the firmware design method of the high-speed image acquisition and processing module on the FPGA board is as follows:

[0023] A 5×5 sliding window is adopted to calculate the grayscale value and gradient value for each pixel point used in the initial cost calculation, and 4 row caches and a 5×5 window cache are configured; for the data in the middle 3×3 window, the sobel gradient template is used to obtain the gradient values in the x and y directions of the current pixel point. At the same time, the 5×5 window is sampled in a 16-point fixed pattern. By summing up all the sampled points and shifting four bits to the right, the average grayscale value is obtained as the grayscale value of the current pixel point; since the initial cost calculation needs to calculate the initial cost within the 0-d max parallax range for each pixel point, d max shift registers are used to buffer the grayscale values and gradient values of the left and right images, and calculate all the initial costs corresponding to each pixel point within the parallax range.

[0024] Furthermore, to efficiently implement step (c) in the stereo vision calculation, the firmware design method of the high-speed image acquisition and processing module on the FPGA board is as follows:

[0025] The RAM resource is adopted to cache the path costs of the previous pixel on each path of the current pixel for calculating the path costs of each path of the current pixel. Specifically: for the calculation of the path costs of the upper left path, upper path, and upper right path of the current pixel, cache all the path costs of the previous row of the current pixel; for the calculation of the path cost of the left path of the current pixel, no row caching is required, and only the path cost calculated by the previous pixel of the current pixel needs to be cached; for the calculation of the path cost of the right path of the current pixel, use a ping-pong cache to reverse the path cost of the right path; finally, accumulate the path costs of the five paths to obtain the final aggregated cost.

[0026] Furthermore, to efficiently implement step (d) in the stereo vision calculation, the firmware design method of the high-speed image acquisition and processing module on the FPGA board is as follows:

[0027] Shift the numerator of the sub-pixel interpolation calculation formula 3 bits to the left, that is, expand it by 8 times, and use a divider to calculate its division quotient; expand the integer parallax corresponding to the minimum aggregated cost of the current pixel by 8 times; add or subtract the quotient obtained by the divider and the integer parallax expanded by the same 8 times. Add when the numerator is positive and subtract when the numerator is negative to obtain the parallax map expanded by 8 times; perform division by the same multiple on the parallax map expanded by 8 times transmitted from the FPGA at the GPU end to obtain the stereo vision parallax map after approximate sub-pixel interpolation.

[0028] Furthermore, for any RGB image for stereo vision calculation, the stereo vision calculation further includes:

[0029] Calculate the disparity map of the other RGB image, and mark the "bad pixels" by comparing whether the disparities of the corresponding points in the two stereo matching disparity maps are consistent.

[0030] Furthermore, the implementation method for fusing the stereo vision depth map and the TOF depth image is as follows:

[0031] (1) Project the TOF depth image into the RGB left camera coordinate system according to the calibration parameters and convert it into a TOF disparity map; upsample the TOF disparity map to obtain a TOF disparity map with the same resolution as the two RGB images;

[0032] (2) Calculate the binocular vision confidence and the TOF confidence. Among them, the binocular stereo vision confidence at each pixel in the stereo vision depth map is:

[0033]

[0034]

[0035] In the formula, V(p) represents the local variance at each pixel point p, N r (p) represents the rectangular neighborhood region with a radius of r of the pixel point p, |N r (p)| is the number of pixel points in the neighborhood region, d j represents the disparity value of the pixel point p, represents the neighborhood N of point p r (p), and S P

[0036] (3) The TOF confidence at each pixel in the TOF depth image is:

[0037] P T =P TA P TD ;

[0038]

[0039]

[0040] In the formula, P TA represents the TOF amplitude confidence, P TD represents the TOF disparity confidence, A represents the TOF amplitude map, A min and A max are the minimum and maximum thresholds respectively, D(p) represents the local variance of each pixel in the TOF disparity map, N(p) represents the neighborhood region of the pixel point p, |N(p)| is the number of pixel points in the neighborhood region, d iIndicates the disparity value of pixel point p, d j Indicates the disparity values of pixel points within the neighborhood N(p) of point p, and T represents the set maximum valid threshold;

[0041] (4) Adopt a local consistency disparity algorithm to fuse the stereoscopic vision disparity map and the TOF disparity map based on binocular stereoscopic vision confidence and TOF confidence.

[0042] The present invention also provides a three-dimensional visual perception system based on the fusion of stereoscopic vision and TOF, including: two-way stereoscopic vision imaging devices, a TOF imaging device, a heterogeneous high-speed image processing system, and a heterogeneous depth fusion system;

[0043] Among them, the heterogeneous high-speed image processing system includes an FPGA board and a high-speed image acquisition and processing module provided on the FPGA board. The firmware design and algorithm steps of the high-speed image acquisition and processing module are determined according to a three-dimensional visual perception method based on the fusion of stereoscopic vision and TOF as described above; the heterogeneous depth fusion system includes a GPU board and a CUDA-based depth fusion module provided on the GPU board. The depth fusion module is used to execute the fusion processing steps in a three-dimensional visual perception method based on the fusion of stereoscopic vision and TOF as described above.

[0044] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0045] The present invention combines the binocular depth map and the TOF depth map, giving full play to the advantages of the two depth estimation methods, and can output a high-speed and high-precision fused depth map. The fused depth map has high robustness because three-dimensional visual perception can be stably performed in scenes where binocular stereoscopic vision fails such as weak light and weak texture, as well as in scenes where TOF fails such as strong light and long distance; under the calculation resource limitations of the FPGA board and the high-speed image acquisition and processing module, high-speed two-way RGB images and stereoscopic vision depth maps with a resolution of 1024×1024 and a frame rate of 120 frames / s can be output; under the calculation resource limitations of the GPU board and the CUDA-based depth fusion module, a robust fused depth map with a resolution of 1024×1024 and a frame rate of 60 frames / s can be output. Description of the Drawings

[0046] Figure 1 It is a flowchart of a three-dimensional visual perception method based on the fusion of stereoscopic vision and TOF provided by an embodiment of the present invention;

[0047] Figure 2 It is a schematic diagram of the cost aggregation path of the Fast-SGM algorithm provided by an embodiment of the present invention;

[0048] Figure 3 Schematic diagram of the initial cost calculation firmware implementation of the Fast - SGM algorithm provided by the embodiments of the present invention;

[0049] Figure 4 Framework diagram of the firmware implementation of the Fast - SGM algorithm for FPGA platform development provided by the embodiments of the present invention;

[0050] Figure 5 Schematic diagram of the sub - pixel interpolation firmware implementation of the Fast - SGM algorithm provided by the embodiments of the present invention;

[0051] Figure 6 Flowchart of the depth fusion algorithm implemented based on CUDA provided by the embodiments of the present invention;

[0052] Figure 7 Physical prototype diagram of the micro - small high - speed and high - precision vision system based on the fusion of stereo vision and TOF provided by the embodiments of the present invention;

[0053] Figure 8 Schematic diagram of the micro - small high - speed and high - precision vision system based on the fusion of stereo vision and TOF provided by the embodiments of the present invention;

[0054] Figure 9 Schematic diagram of the FPGA board hardware of the heterogeneous high - speed image processing system provided by the embodiments of the present invention. Specific embodiments

[0055] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0056] Embodiment 1

[0057] A three - dimensional vision perception method based on the fusion of stereo vision and TOF, as Figure 1 shown, includes:

[0058] Using two visual imaging units to collect two RGB images;

[0059] Using a TOF imaging device to collect TOF depth images;

[0060] Using an FPGA board and a high - speed image acquisition and processing module, based on the two RGB images, performing stereo vision calculation on any one of the RGB images to obtain the stereo vision depth map corresponding to the RGB image;

[0061] Using a GPU board and a CUDA-based deep fusion module, the stereo vision depth map and the TOF depth map are fused to obtain a fused depth map;

[0062] Taking the two-channel RGB images, the stereo vision depth map, and the fused depth map as the results of three-dimensional visual perception, three-dimensional visual perception is completed.

[0063] It should be noted that stereo vision is a technology that uses two-channel RGB images and calculates the stereo vision depth map according to the stereo matching algorithm.

[0064] The fused depth map in this embodiment has high robustness because three-dimensional visual perception can be stably performed in scenarios where binocular stereo vision fails such as weak light and weak texture, and in scenarios where TOF fails such as strong light and long distance; under the computing resource constraints of the FPGA board and the high-speed image acquisition and processing module, high-speed two-channel RGB images and stereo vision depth maps with a resolution of 1024×1024 and a frame rate of 120 frames / s can be output; under the computing resource constraints of the GPU board and the CUDA-based deep fusion module, a robust fused depth map with a resolution of 1024×1024 and a frame rate of 60 frames / s can be output.

[0065] Based on the deficiencies of the existing stereo vision system and combined with the actual requirements, this embodiment proposes a micro-miniature high-speed and high-precision vision system based on the fusion of stereo vision and TOF, which gives full play to the advantages of the two sensors, improves the quality of the output depth map, and has the characteristics of high speed, robustness, and high precision.

[0066] Preferably, the Fast-SGM algorithm is used to perform stereo vision calculation on any one of the RGB images. As Figure 4 shown, the implementation method is:

[0067] (a) Synchronize the data of the two-channel RGB images;

[0068] During the image acquisition and preprocessing process, the left and right images are calculated separately. Even if the left and right cameras are synchronized in exposure, there may be a problem that the two images do not arrive completely synchronously. Therefore, data alignment is required for subsequent accurate calculation. Specifically, after preprocessing the data synchronization of the two-channel RGB images, two FIFOs with a depth of m and a state machine used to represent whether each FIFO is in a waiting data state or a data output state are used to align the two-channel data. When the frame start signals of the two images both arrive, the two-channel data of the FIFO is synchronously output. If only the frame start signal of the left image or the frame start signal of the right image arrives, wait for the frame start signal of the other image to arrive. Among them, m is taken as a value greater than the pixel delay between the two images. For example, m is taken as 512.

[0069] (b) Calculate the initial cost of each image separately;

[0070] Specifically, the initial cost uses the threshold-truncated AD cost and Sobel gradient cost based on pixel points. The costs of the two at the disparity d at the point (x, y) can be written as the following equations respectively:

[0071] Cost(x, y, d) = min{M(x, y, d), T gray}+ min{G(x, y, d), T grad};

[0072] In the formula, Cost(x, y, d) represents the initial cost at the disparity d at the point (x, y), which is a matrix, and M(x, y, d) represents the grayscale cost at the disparity d at the point (x, y), which is a matrix.

[0073] M(x, y, d) = |I L (x, y) - I R (x - d, y)|;

[0074] In the formula, I L (x, y) represents the grayscale value of the left image at the point (x, y), and I R (x - d, y) represents the grayscale value of the corresponding point (x - d, y) in the right image at the point (x, y) of the left image; G(x, y, d) represents the gradient cost at the disparity d at the point (x, y), which is a matrix.

[0075]

[0076] In the formula, represents the horizontal gradient value of the left image at the point (x, y), represents the horizontal gradient value of the right image at the point (x, y), represents the vertical gradient value of the left image at the point (x, y), represents the vertical gradient value of the right image at the point (x, y); T gray and T grad are the pixel intensity threshold and gradient threshold respectively.

[0077] The grayscale cost based on pixel points can greatly reduce the algorithm complexity and is convenient to be implemented using basic logic units in FPGA to achieve accelerated operation. And texture features have been proven to be one of the most important features in the field of image matching. The Sobel gradient cost can characterize the texture information in the image and can better describe the discontinuous regions such as object boundaries and regions with rich surface textures at a relatively small computational cost, which helps to overcome the problem of poor boundary performance of stereo matching algorithms.

[0078] (c) Calculate the aggregation cost of each image using the initial cost;

[0079] Fast - SGM generally follows the semi - global cost aggregation framework, converting the two - dimensional aggregation problem of global matching into a one - dimensional aggregation problem along fixed paths, as shown in the following formula:

[0080]

[0081] To reduce the computational complexity, and due to the inherent characteristics of the FPGA pipeline making some paths difficult to implement or requiring a large amount of storage resources for caching, the Fast - SGM algorithm changes the eight search paths to five, and the aggregation paths are as Figure 2 shown. The final total cost is shown in the following formula, which is the aggregation of the costs of the five paths:

[0082]

[0083] In the formula, L r (p, d) represents the cost of the pixel point p at the disparity d on the path r; S(p, d) represents the aggregation cost of the pixel point p at the disparity d; Cost(p, d) represents the initial cost at the pixel point p at the disparity d, P1 and P2 represent penalty coefficients, and L r (p - r, k) represents the path cost of the previous pixel point on the path r at the disparity k.

[0084] (d) Use the aggregation cost to perform sub - pixel interpolation calculation to obtain the final stereo matching disparity map;

[0085] Generally speaking, the SGM algorithm will use the disparity with the minimum cost of each pixel as the final disparity. However, since the obtained disparity is discrete and quantized, there is a loss of accuracy. Therefore, sub - pixel interpolation is needed to improve the original pixel - level accuracy. By minimizing the cost of the surrounding disparities for curve fitting, the disparity at the sub - pixel level is calculated, and the interpolation formula for the disparity at each pixel is as shown below:

[0086]

[0087]

[0088] Among them, S d represents the minimum cost of the current pixel, and d represents the disparity corresponding to the minimum cost.

[0089] (e) Based on the final disparity and the calibration parameters of the two vision imaging units, obtain the stereo vision depth map to improve the original pixel - level accuracy.

[0090] Preferably, since the gray - level cost based on pixel points is vulnerable to noise and it is difficult for FPGA to handle division operations, the Fast - SGM algorithm designs a 16 - point sampling pattern to smooth the 5×5 area around the pixel points to obtain the gray - level values for calculating the initial cost. The sum of the 16 sampled pixel values can be added through a register and then right - shifted by 4 bits to achieve accelerated operation. The sampling pattern is as Figure 3 shown. Meanwhile, the Fast - SGM algorithm adopts a threshold truncation operation to uniformly truncate the cost exceeding the threshold to the maximum threshold, thereby reducing the influence of noise on the algorithm accuracy.

[0091] Specifically, as Figure 3 shown, to implement step (b) in the above - mentioned stereo vision calculation, the firmware design method of the high - speed image acquisition and processing module on the FPGA board is as follows:

[0092] Adopt a 5×5 sliding window to calculate the gray - level value and gradient value of each pixel point for the initial cost calculation, and configure 4 row caches and a 5×5 window cache; use the sobel gradient template for the data in the middle 3×3 window to obtain the x and y direction gradient values of the current pixel point. At the same time, perform 16 - point sampling on the 5×5 window, and obtain the average gray - level value by adding all the sampled points and right - shifting by four bits as the gray - level value of the current pixel point; since the initial cost calculation needs to calculate the initial cost within the 0 - d max disparity range for each pixel point, use d max shift registers to buffer the gray - level values and gradient values of the left and right two - way images, and calculate all the initial costs corresponding to each pixel point within the disparity range.

[0093] Preferably, to implement step (c) in the stereo vision calculation, the firmware design method of the high - speed image acquisition and processing module on the FPGA board is as follows:

[0094] Adopt RAM resources to cache the path costs of the previous pixel on each path of the current pixel for calculating the path costs of each path of the current pixel. Specifically: for calculating the path costs of the upper - left path, upper path, and upper - right path (as Figure 2 shown) of the current pixel, cache all the path costs of the previous row of the current pixel; for calculating the path cost of the left path (as Figure 2 shown) of the current pixel, no row cache is required, and only the path cost calculated by the previous pixel of the current pixel needs to be cached; for calculating the path cost of the right path (as Figure 2 shown) of the current pixel, since the calculation direction is opposite to the direction of the image data stream input, a ping - pong cache is required to reverse its path result; finally, accumulate the path costs of the five paths to obtain the final aggregated cost.

[0095] Preferably, to implement step (d) in stereo vision calculation, the firmware design method of the high-speed image acquisition and processing module on the FPGA board is as follows:

[0096] In FPGA implementation, since the FPGA cannot process floating-point numbers, and the designed divider can only obtain the integer quotient and remainder, the Fast-SGM algorithm adopts the method of shifting the numerator of the sub-pixel interpolation calculation formula 3 bits to the left, that is, expanding it by 8 times, and uses the divider to calculate the division quotient for it; expands the integer disparity corresponding to the minimum aggregation cost of the current pixel by 8 times; adds or subtracts the quotient obtained by the divider from the integer disparity that is also expanded by 8 times. When the numerator is positive, add; when the numerator is negative, subtract, to obtain a disparity map expanded by 8 times; perform division (divide by 8) of the same multiple on the disparity map expanded by 8 times transmitted from the FPGA at the GPU end to obtain the disparity map after approximate sub-pixel interpolation. The block diagram of the sub-pixel interpolation firmware implementation is as Figure 5 shown.

[0097] Preferably, for any RGB image for stereo vision calculation, the stereo vision calculation further includes:

[0098] Calculating the disparity map of the other RGB image, and comparing whether the disparities of the corresponding points in the two stereo matching disparity maps are consistent to find out the "bad points" to verify the correctness of the disparity calculation.

[0099] Some occluded and mismatched areas can be excluded. Therefore, it is also necessary to use a shift register to cache the disparities within the range of 0-d max If the disparity of the right pixel corresponding to the disparity of the left pixel is inconsistent, the current pixel is considered invalid and assigned a value of 0.

[0100] Preferably, as Figure 6 shown, the depth fusion algorithm based on CUDA implementation proposed in this embodiment, as a preferred embodiment, mainly includes three parts: TOF projection and upsampling, TOF and binocular stereo matching depth map confidence calculation, and depth fusion. Specifically, it mainly includes the following steps:

[0101] (1) Since the original TOF depth map and the binocular depth map are not in the same coordinate system and cannot be fused in the subsequent fusion operation, it is necessary to first project the TOF depth data into the RGB left camera coordinate system according to the calibration parameters and convert it into a disparity map. At the same time, since the TOF depth map has a relatively low resolution compared to the RGB image, the TOF data projected into the RGB left camera coordinate system shows sparse attributes, which is not conducive to subsequent depth map fusion. It is necessary to obtain high-resolution TOF depth maps and disparity maps through an upsampling algorithm.

[0102] For the pixel point (x, y) in the image I(x, y), the spatial weight kernel and gray weight are shown as follows:

[0103]

[0104]

[0105] where i and j are the displacements relative to the central pixel (x, y) in the window, and σ s and σ c are the standard deviations of the Gaussian kernels for the spatial weight and the gray - level weight respectively. Therefore, for each pixel point (x, y), the weight of (i, j) within its window is: w(i, j) = w s (i, j)w c (i, j).

[0106] In actual up - sampling usage, since there are invalid disparities for some pixel points, the depth of the current pixel (x, y) is taken as the weighted average of all valid pixels, as shown in the following formula:

[0107]

[0108] where Z(x, y) represents the depth image and W represents the window area.

[0109] After up - sampling to obtain a high - resolution TOF depth map, the fusion algorithm converts the depth data into disparity data according to the following conversion formula: where d and Z represent the disparity value and the depth value respectively, B is the baseline of the binocular stereo system, and f is the focal length.

[0110] (2) Calculate the binocular vision confidence and the TOF confidence. Among them, the binocular stereo vision confidence at each pixel in the stereo vision depth map is:

[0111]

[0112]

[0113] In the formula, V(p) represents the local variance at each pixel point p, N r (p) represents the rectangular neighborhood area with a radius of r for the pixel point p, |N r (p)| is the number of pixel points in the neighborhood area, d j represents the disparity value of the pixel point p, represents the average of the disparity values of all pixel points within the neighborhood N r (p) of the point p, and P S (p) represents the binocular stereo vision confidence at each pixel;

[0114] Specifically, since binocular vision mainly relies on visual information and the prior information of depth continuity in the scene for calculation, the binocular vision confidence calculation is mainly based on the visual feature matching characteristics and the assumption of disparity continuity. The algorithm uses the local variance of the disparity map as the confidence of binocular vision. For the stereo matching disparity map calculated by the Fast-SGM algorithm, the local variance at each pixel is shown as follows: Among them, N r (p) represents the rectangular neighborhood area with a radius of r of pixel point p, and |N r (p)| is the number of pixel points in the neighborhood area, and d j represents the disparity value of pixel point p, represents the average of the disparity values of all pixel points in the neighborhood N r (p) of point p. Since the greater the variance, the greater the degree of local disparity change and the lower the confidence, it is necessary to take the opposite of the variance as an index and normalize it to [0,1].

[0115] The final binocular confidence is shown as follows:

[0116] In addition, the basis of the TOF confidence mainly comes from two parts: the error generation model of TOF and the prior of depth continuity in the scene. The error source of TOF is mainly related to the intensity of the reflected light in the scene received by the sensor, and is subject to uncertain factors in the scene, such as black areas with low reflectivity, specular reflection surfaces at reflection angles, etc. Therefore, the proposed fusion algorithm uses the TOF reflected light amplitude and local disparity to jointly form the TOF confidence.

[0117] For the TOF amplitude map A, a piecewise linear function is used to model the confidence:

[0118]

[0119] Among them, A min and A max are the minimum and maximum thresholds respectively.

[0120] When the disparity values of a pixel point and its neighborhood change greatly, such as approaching a discontinuous point, the obtained estimated depth measurement value is a convex combination of different depth values. Therefore, the algorithm assigns a lower confidence to pixel points with large neighborhood disparity changes. The formula for calculating the neighborhood disparity change is as follows:

[0121]

[0122] Among them, N(p) represents the neighborhood area of pixel point p, |N(p)| is the number of pixel points in the neighborhood area, and d i represents the disparity value of pixel point p, and d jRepresents the disparity of pixel points within the neighborhood N(p) of point p. To further calculate the confidence, the fusion algorithm normalizes the neighborhood change D to [0, 1] by setting a maximum effective absolute difference threshold T = 0.3, and the TOF data confidence P TD The calculation formula is as follows:

[0123]

[0124] Finally, the TOF confidence consists of two parts, as shown in the following formula: P T = P TA P TD .

[0125] (3) Adopt the locally consistent disparity technique (Locally consistent, LC) to fuse the two types of disparity data. The core idea is: Given a Fast-SGM disparity map D S and a TOF disparity map D T , then each pixel point p contains a stereo vision disparity and a TOF disparity That is, there are two assumed disparity values for each pixel point. For each assumed disparity value, the fusion algorithm measures the reliability of the corresponding assumed disparity value by considering spatial consistency and color consistency, and selects the more reliable assumed disparity value as the final estimated disparity value d of this pixel point p . The reliability measurement formulas for the assumed disparity values of stereo vision and TOF at pixel point p are as follows:

[0126]

[0127] Among them, M ∈ {S, T} represents the data source, S is the stereo vision data, and T is the TOF data, is the confidence of pixel point p, N(p) represents the critical region of pixel point p, m is the neighborhood point within the neighborhood region of point p, is the pixel point coordinate corresponding to pixel point p in the right camera pixel coordinate system when its disparity is , is the neighborhood point within the neighborhood region of point q, and r(p, q, m, n) is the neighborhood reliability function:

[0128]

[0129] Among them, Δ p,m is the spatial distance between point p and its neighborhood point m, is the color distance between point p and its neighborhood point m, is the color distance between point q and its neighborhood point n, is the color distance between neighborhood point m and its corresponding neighborhood point n.

[0130] For each pixel p, the fusion algorithm calculates the corresponding reliability based on its stereo disparity and a TOF disparity respectively and selects the disparity with higher reliability as the fused disparity d of this point p :

[0131] Embodiment 2

[0132] A three-dimensional visual perception system based on the fusion of stereo vision and TOF, as Figure 7 shown, includes: two stereo vision imaging devices, a TOF imaging device, a heterogeneous high-speed image processing system, and a heterogeneous depth fusion system;

[0133] Among them, the heterogeneous high-speed image processing system includes an FPGA board and a high-speed image acquisition and processing module arranged on the FPGA board. The algorithm steps and firmware design of the high-speed image acquisition and processing module are obtained according to a three-dimensional visual perception method based on the fusion of stereo vision and TOF described in Embodiment 1 above; the heterogeneous depth fusion system includes a GPU board and a CUDA-based depth fusion module arranged on the GPU board. The depth fusion module is used to execute the fusion processing steps in a three-dimensional visual perception method based on the fusion of stereo vision and TOF described in Embodiment 1 above.

[0134] As Figure 8 and Figure 9 shown, the above two stereo vision imaging devices can be designed as follows: the stereo vision imaging device includes two CMOS cameras for photoelectric conversion, lenses, filters, and an IMU attitude sensor for measuring three-axis attitude. The synchronization signals of the two CMOS cameras and the IMU sensor in this imaging device are connected, and the CMOS and IMU are synchronously triggered through an external trigger signal to respectively collect the left and right images and the IMU attitude data at the same time, and are transmitted into the FPGA image acquisition module on the FPGA board through LVDS signal lines.

[0135] That is, the stereo vision imaging device includes two lenses and two image acquisition boards. The image acquisition board includes a CMOS image sensor module, an IMU three-axis attitude sensor module, and a power supply module. The CMOS image sensor generates high-quality and high-frame-rate images, and together with the angle and acceleration data generated by the IMU three-axis attitude sensor, is transmitted to the heterogeneous high-speed image processing system.

[0136] The stereo vision imaging device provides an image acquisition system with strict hardware synchronization between the left and right CMOS cameras and the IMU. Among them, the CMOS sensor uses a global shutter, which can reduce the motion distortion generated during camera shooting and ensure the synchronization and fidelity of the image and sensor signal during high-speed acquisition.

[0137] Specifically, in this embodiment, the CMOS sensor of the stereo vision imaging device selects the PYTHON5000 high-performance CMOS high-speed vision sensor of ON Semiconductor Corporation in the United States. The shutter mode is a global shutter, and it can output a low-noise, high-resolution, high-frame-rate massive color image data stream of up to 2592 pixels × 2048 pixels @ 100 fps at most. The highest frame rate can reach 255 fps at a resolution of 1920 pixels × 1080 pixels. The IMU attitude sensor of the stereo vision imaging device selects the MTi1 series of XSENS Corporation in the Netherlands as the inertial measurement unit of the micro-miniature vision unit. It has a small volume and low power consumption. The core components adopt high-precision MEMS accelerometers, gyroscopes, and magnetometers, and can stably output high-precision roll, pitch, and heading and other information. The stereo vision imaging device communicates with the FPGA board through an FPC cable. The synchronization signals of the CMOS sensor and the IMU attitude sensor are connected, and the FPGA board synchronously controls the sensor image acquisition and IMU data acquisition to ensure the strict synchronization of image exposure and three-axis attitude data.

[0138] The above TOF imaging device measures the depth of the scene through the time-of-flight ranging principle and can output high-precision depth data in real time. The TOF imaging device can adopt the following design: The TOF imaging device includes multiple TOF modules and a time-sharing control module. The TOF module is mainly used to collect the scene depth data from the TOF sensor. The time-sharing control module triggers different TOF modules at different times in the form of external triggers to ensure that there is a TOF module working at each moment, and then integrates the data of multiple TOF modules to achieve an increase in the frame rate of the overall TOF imaging device relative to a single TOF module.

[0139] For example, the TOF imaging device includes four TOF sensors. Each TOF sensor can stably output depth images at 30 fps to the heterogeneous high-speed image processing system, and is collected by the heterogeneous depth fusion system in a time-sharing manner, and a total of 120 fps of TOF depth images can be generated.

[0140] The TOF imaging device provides a method to improve the frame rate of the TOF module, solves the problem of too low frame rate of a single-module TOF in a high-speed imaging system, and enables low-frame-rate TOF sensors to meet the requirements of high-speed systems.

[0141] Specifically, in this embodiment, the TOF imaging device uses the domestically produced OPNOUS OPNM8808C TOF module, which includes a laser transmitter, a laser receiver, and a data processing board. It can stably output TOF depth images with a resolution of 320×240 pixels and a frame rate of 30 fps. The TOF imaging device has a reserved interface connected to the GPU board, which time-shares the four imaging devices to achieve 120 fps high-speed depth map acquisition.

[0142] The aforementioned heterogeneous high-speed image processing system can be designed as follows: The heterogeneous high-speed image processing system includes an FPGA board and a high-speed image acquisition and processing module, consisting of an image input unit, an image processing unit, and an image output unit. The image input unit can acquire two RGB images output by a stereoscopic imaging device. The image processing unit performs high-speed, real-time preprocessing and depth calculation on these two RGB images. The image output unit outputs high-frame-rate RGB and depth images and communicates with the GPU board. In other words, the heterogeneous high-speed image processing system performs high-speed, high-frame-rate image acquisition, preprocessing, and binocular depth estimation, and then sends the processed two-channel image and depth map data to the GPU via PCIe.

[0143] The high-speed image acquisition and processing module includes an image processing ISP sub-module, a depth estimation algorithm sub-module and an image output sub-module. The image processing ISP sub-module includes image acquisition firmware, image preprocessing firmware, automatic exposure, white balance firmware, image correction firmware and image output firmware, which are used for high-speed image acquisition and preprocessing; the depth estimation algorithm sub-module includes image correction firmware and Fast-SGM algorithm firmware developed for FPGA platform, which are used to estimate depth, perform high-speed stereo matching through binocular images, and generate a depth map of the scene; the image output sub-module is used to distribute high-speed images and depth maps to the USB and PCIe output interfaces, and send them to the host computer and GPU board respectively.

[0144] That is, the heterogeneous high-speed image processing system is specifically an FPGA board and high-speed image processing firmware, which is used to collect two RGB images for pre-processing and stereo matching depth image calculation, send them to the host computer to display the video stream, and at the same time transmit them to the GPU board for deep fusion based on CUDA.

[0145] The Fast-SGM algorithm developed for the FPGA platform optimizes multiple steps of the SGM algorithm to meet the requirements of high speed and high precision. A firmware implementation framework is designed, and various parallel acceleration strategies are adopted to improve the algorithm operation efficiency. Finally, real-time high-precision depth map calculation of 1024 pixels × 1024 pixels @ 120 fps can be achieved. The algorithm steps include four parts: preprocessing, cost calculation, cost aggregation, and postprocessing. Among them, in the cost calculation part, the threshold truncation AD cost and Sobel gradient cost based on pixel points are used, which can enhance the accuracy and robustness of the algorithm; in the cost aggregation part, a variable penalty coefficient and an adaptive cost weight are added on the basis of the SGBM framework, which can improve the accuracy of discontinuous regions. At the same time, the global 8 search paths are changed to 4, which is convenient for the FPGA platform to adopt the pipeline design idea to accelerate the calculation. That is, the Fast-SGBM stereo matching algorithm developed for the FPGA platform completely uses FPGA logic to perform the calculation and processing of the stereo matching algorithm, improving the calculation speed of the stereo matching algorithm and the real-time performance of the entire binocular vision system.

[0146] Specifically, as Figure 9 shown, the FPGA board in the heterogeneous high-speed image processing system takes the FPGA chip as the core, and also includes two FPC interfaces, four USB3.0 interfaces, a PCIE interface, a clock module, a power module, and a storage module. Among them, the main chip selects the xczu19eg-ffvc1760-2-i chip of Xilinx company as the processing chip. This chip belongs to a heterogeneous SoC and is an innovative +FPGA architecture, including functional units and modules such as a video codec unit, an advanced dynamic power management unit, and DDR4 memory interface support, which can meet the processing requirements of real-time high-speed images while ensuring low power consumption, miniaturization, and high processing efficiency of the system; the FPC interface is used to receive images captured by the CMOS camera, with a maximum transmission rate of 5.4 Gbps; the USB3.0 interface is used to transmit two preprocessed RGB images and the depth image calculated by the stereo matching algorithm to the host computer, with a maximum transmission rate of 5 Gbps; the PCIE interface is used to communicate and transfer data with the GPU board, with a maximum transmission rate of 4 GB / s; the clock module is used to generate the clocks required by each device to ensure normal operation; the power module is used to provide different voltages for different devices to provide stable power for the entire board; the storage module consists of two 512 MB DDR4s, which are used to store or buffer images when the DDR storage space in the FPGA chip is insufficient.

[0147] The above heterogeneous depth fusion system can be designed as follows: The heterogeneous depth fusion system includes a GPU board and a depth fusion algorithm based on CUDA. The GPU board consists of an image input unit, an image processing unit, and an image output unit. The image input unit collects the depth images output by the FPGA board and the TOF imaging device. The image processing unit fuses the depth images, making the obtained final depth map have higher measurement accuracy and stronger robustness. The image output unit can wirelessly transmit the fused depth map. The depth fusion algorithm based on CUDA realizes the fusion of high-frame-rate and high-precision binocular stereo vision depth maps and TOF depth maps, and can output high-speed and high-precision depth maps. That is, the heterogeneous depth fusion system is used to collect data from stereo vision imaging devices and perform high-speed image preprocessing and calculate stereo vision depth maps in real time through the image processing unit.

[0148] According to the first embodiment, a parallel depth fusion algorithm architecture is designed for the depth fusion algorithm implemented based on CUDA, which greatly improves the running efficiency of the depth fusion algorithm on the GPU. At the same time, by combining binocular depth maps and TOF depth maps, the advantages of the two depth estimation methods are fully utilized, and a high-speed and high-precision fused depth map can be output. The specific algorithm steps are as described in the first embodiment, including TOF projection and upsampling, confidence calculation of TOF and binocular stereo matching depth maps, and depth fusion. Among them, TOF projection and upsampling are used to project TOF depth map data onto the binocular depth map to align the two types of data in the image, laying a foundation for subsequent calculations. The confidence calculation of TOF and binocular stereo matching depth maps is used to calculate the depth estimation confidence of each point based on the depth maps of binocular images and TOF depth maps, serving as a guide for the subsequent fusion algorithm. Depth fusion is used to perform depth fusion through the depth map confidence of the two types of data to finally obtain the fused depth map.

[0149] That is to say, the heterogeneous depth fusion system specifically refers to a GPU board and a depth fusion algorithm implemented based on CUDA, which is used to collect the depth map output by the TOF imaging device and fuse it with the stereo matching depth map output by the heterogeneous high-speed image processing system, and output a high-speed and high-precision depth map.

[0150] Specifically, in this embodiment, the GPU board in the heterogeneous high-speed image fusion system selects the NVIDIA Jetson-AGX-Xavier development kit, which is equipped with a 512-core Volta GPU with Tensor Core and an 8-core ARM v8.2 64-bit CPU, having sufficient computing resources and peripheral interfaces, and can ensure the real-time and high-precision fusion of depth data of binocular vision and TOF data. In addition, this kit has sufficient module interfaces, including PCIE interface, Ethernet interface, USB interface, M.2 interface, HDMI interface, etc., which are convenient for data transmission with the FPGA board or the host computer.

[0151] This embodiment uses the above-mentioned miniature high-speed and high-precision vision system for image acquisition and processing. Combined with the method of Example 1, after the FPGA board receives the two-way Bayer color image of the stereoscopic vision imaging device, it is stored in the DDR4 of the FPGA chip. The image is then pre-processed, including Bayer array interpolation, Gaussian filtering, automatic white balance, automatic gain, and image correction (the purpose is to make the image effect better in a variety of environments). After pre-processing, the two-way RGB image calculates the depth map according to the stereo matching algorithm. In order to achieve the purpose of high speed and high precision, the system proposes the Fast-SGM algorithm developed for the FPGA platform, optimizes the cost calculation and aggregation of the SGM algorithm, further improves the algorithm performance, and adopts a variety of parallel acceleration strategies to improve the algorithm operation efficiency, and ultimately can achieve 1024 pixels × 1024 pixels @ 120fps real-time high-precision depth map calculation.

[0152] The FPGA board then transmits the two RGB images and the stereo matching depth image to the host computer via four USB 3.0 interfaces for video display. Simultaneously, the stereo matching depth map is transmitted to the GPU board via the PCIE interface for depth map fusion. The GPU board then acquires the depth maps from the four TOF imaging devices through time-sharing control and fuses them using a CUDA-based depth fusion algorithm. The resulting highly accurate and robust fused depth map is then transmitted to the host computer via the USB 3.0 interface for display.

[0153] It should be noted that the descriptions of the first embodiment and the second embodiment are mutually effective and will not be repeated here.

[0154] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A three-dimensional visual perception method based on the fusion of stereo vision and TOF, characterized in that, Including: Two visual imaging units are adopted to collect two paths of RGB images; A TOF imaging device is adopted to collect TOF depth images; An FPGA board and a high-speed image acquisition and processing module are adopted. Based on the two paths of RGB images, stereo vision calculation is performed on any one path of RGB image to obtain a stereo vision depth map corresponding to this path of RGB image; A GPU board and a depth fusion module based on CUDA are adopted to perform fusion processing on the stereo vision depth map and the TOF depth map to obtain a fused depth map; The two paths of RGB images, the stereo vision depth map and the fused depth map are used as the results of three-dimensional visual perception to complete three-dimensional visual perception; Among them, the implementation method of performing fusion processing on the stereo vision depth map and the TOF depth image is: (1) Project the TOF depth image into the RGB left camera coordinate system according to the calibration parameters and convert it into a TOF disparity map; Upsample the TOF disparity map to obtain a TOF disparity map with the same resolution as the two paths of RGB images; (2) Calculate the binocular vision confidence and the TOF confidence. Among them, the binocular stereo vision confidence at each pixel in the stereo vision depth map is: ; ; In the formula, represents the local variance at each pixel point , represents the rectangular neighborhood area with a radius of for the pixel point , is the number of pixel points in the neighborhood area, represents the disparity value of the pixel point , represents the average of the disparity values of all pixel points in the neighborhood of the point ; represents the binocular stereo vision confidence at each pixel; (3) The TOF confidence at each pixel in the TOF depth image is: ; ; ; In the formula, represents the TOF amplitude confidence, represents the TOF parallax confidence, represents the TOF amplitude map, and are the minimum and maximum thresholds respectively, represents the local variance of each pixel in the TOF parallax map, represents the pixel point 's neighborhood area, is the number of pixel points in the neighborhood area, represents the pixel point 's parallax value, represents the point 's neighborhood the parallax values of the pixel points within, represents the set maximum effective threshold; (4) Adopt a local consistency disparity algorithm to fuse the stereo vision disparity map and the TOF disparity map based on the binocular stereo vision confidence and the TOF confidence.

2. The three-dimensional visual perception method according to claim 1, characterized in that, Adopt the Fast-SGM algorithm suitable for FPGA implementation to calculate the stereo vision depth map for any one path of RGB image. The implementation method is: (a) Synchronize the data of the two paths of RGB images; (b) Calculate the initial cost of each path of image respectively; (c) Adopt the initial cost to calculate the aggregated cost of each path of image respectively; (d) Adopt the aggregated cost to perform sub-pixel interpolation calculation to obtain the final stereo vision disparity map; (e) Based on the final stereo vision disparity map and the calibration parameters of the two visual imaging units, obtain the stereo vision depth map.

3. The three-dimensional visual perception method according to claim 2, characterized in that, To efficiently implement step (b) in the stereo vision calculation, the firmware design method of the high-speed image acquisition and processing module on the FPGA board is: Adopt a 5×5 sliding window to calculate the gray value and gradient value for the initial cost calculation of each pixel point, and configure 4 row caches and a 5×5 window cache; Use the Sobel gradient template for the data in the middle 3×3 window to obtain the gradient values in the x and y directions of the current pixel. At the same time, perform 16-point fixed-mode sampling on the 5×5 window, and obtain the average gray value by summing all the sampled points and shifting four bits to the right as the gray value of the current pixel; since the initial cost calculation requires calculating the initial cost within the disparity range, use shift registers to buffer the gray values and gradient values of the left and right images, and calculate all the initial costs corresponding to each pixel within the disparity range.

4. The three-dimensional visual perception method according to claim 2, wherein To efficiently implement step (c) in the stereo vision calculation, the firmware design method of the high-speed image acquisition and processing module on the FPGA board is: Adopt RAM resources to cache the path costs of the previous pixel on each path of the current pixel for calculating the path costs of each path of the current pixel. Specifically: for calculating the path costs of the upper left path, upper path, and upper right path of the current pixel, cache all the path costs of the previous row of the current pixel; for calculating the path cost of the left path of the current pixel, no row caching is required, and only the path cost calculated for the previous pixel of the current pixel needs to be cached; for calculating the path cost of the right path of the current pixel, use ping-pong caching to reverse the path cost of the right path; finally, accumulate the path costs of the five paths to obtain the final aggregation cost.

5. The three-dimensional visual perception method according to claim 2, wherein To efficiently implement step (d) in the stereo vision calculation, the firmware design method of the high-speed image acquisition and processing module on the FPGA board is as follows: Shift the numerator of the sub-pixel interpolation calculation formula 3 bits to the left, that is, expand it by 8 times, and use a divider to calculate its division quotient; expand the integer disparity corresponding to the minimum aggregation cost of the current pixel by 8 times; add or subtract the quotient obtained by the divider from the integer disparity that is also expanded by 8 times. When the numerator is positive, add; when the numerator is negative, subtract, to obtain a disparity map expanded by 8 times; on the GPU side, perform division by the same multiple on the disparity map expanded by 8 times transmitted from the FPGA to obtain a stereo vision disparity map after approximate sub-pixel interpolation.

6. The three-dimensional visual perception method according to claim 2, wherein For performing stereo vision calculation on any one of the RGB images, the stereo vision calculation further includes: Calculate the disparity map of the other RGB image, and mark "bad points" by comparing whether the disparities of the corresponding points in the two stereo matching disparity maps are consistent.

7. A three-dimensional visual perception system based on the fusion of stereo vision and TOF, characterized in that, It includes: Two stereo vision imaging devices, a TOF imaging device, a heterogeneous high-speed image processing system, and a heterogeneous depth fusion system; Among them, the heterogeneous high-speed image processing system includes an FPGA board and a high-speed image acquisition and processing module arranged on the FPGA board. The firmware design and algorithm steps of the high-speed image acquisition and processing module are determined according to a three-dimensional vision perception method based on the fusion of stereo vision and TOF as described in any one of claims 1 to 6; the heterogeneous depth fusion system includes a GPU board and a CUDA-based depth fusion module arranged on the GPU board. The depth fusion module is used to execute the fusion processing steps in a three-dimensional vision perception method based on the fusion of stereo vision and TOF as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Robot microminiature high-speed binocular stereoscopic vision system

    CN111524177A

  • Depth Information Acquisition Method and Device

    US20200128225A1