Line scanning depth calculation method and device, calibration system and TOF camera

By using a binocular vision-based line scan depth calculation method and a TOF camera calibration system, the problem of low depth ground truth accuracy was solved, achieving high-precision depth ground truth acquisition and training sample provision, thus improving the accuracy of deep learning.

CN115908528BActive Publication Date: 2026-05-15SHENZHEN ORBBEC CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN ORBBEC CO LTD
Filing Date
2022-10-27
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Among existing depth ground truth acquisition technologies, radar-based methods have a large measurement range but insufficient accuracy, while structured light solutions are not accurate enough in specific scenarios, resulting in low depth ground truth accuracy.

Method used

A binocular vision-based line scan depth calculation method is adopted. By acquiring the left and right images of the target object containing multi-line patterns, performing clustering and binarization operations, calculating the center point and performing stereo matching, and combining the calibration system of TOF camera and binocular camera, high-precision depth values ​​are obtained.

Benefits of technology

It improves the accuracy of ground truth depth data, ensuring the accuracy of binocular depth images in different scenarios, and provides high-precision ground truth depth training samples for TOF cameras, supporting the accurate application of deep learning algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115908528B_ABST
    Figure CN115908528B_ABST
Patent Text Reader

Abstract

The application discloses a line scanning depth calculation method applied to binocular vision, and comprises the following steps: acquiring left and right images of a target object containing a multi-line pattern; performing clustering binarization operation on the left and right images to obtain binary left and right images and left and right clustering pixel groups corresponding to the binary left and right images; acquiring a standard line region width according to the left and right clustering pixel groups; taking the standard line region width as a mask width to perform mask operation on the binary left and right images to obtain binary left and right line region images; calculating center points of the binary left and right line region images to obtain left and right center point arrays; performing interpolation calculation on the left and right center point arrays to obtain corresponding left and right matching points; performing stereo matching calculation by using the left and right matching points to obtain a disparity, and converting the disparity into a depth to obtain a binocular depth image. The application can improve the precision of line scanning depth, thereby improving the accuracy of depth true value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and camera calibration technology, and in particular to a method, apparatus, calibration system, and TOF camera for calculating line scan depth in binocular vision. Background Technology

[0002] Depth ground truth refers to real depth data. It can be used to reconstruct 3D information in real scenes and can solve the problem of insufficient samples in deep learning schemes that require a large number of high-precision data samples. The accuracy of depth ground truth determines the accuracy of subsequent deep learning calculations. Therefore, the acquisition of depth ground truth plays an important role in 3D information calculation.

[0003] Existing depth ground truth acquisition technologies are mostly based on radar or structured light. In practical applications, although radar-based ground truth acquisition technologies have a large measurable range, the measurement accuracy cannot be guaranteed. Structured light solutions, on the other hand, rely on speckle and fringe measurements, which are not accurate enough in certain scenarios, resulting in low precision of the acquired depth ground truth. Summary of the Invention

[0004] This invention provides a method, apparatus, calibration system, and TOF camera for line scanning depth calculation in binocular vision, with the main objective of solving the problem of low accuracy of the acquired depth true value.

[0005] To achieve the above objectives, the present invention provides a line scanning depth calculation method for binocular vision, comprising: acquiring left and right images of a target object containing a multi-line pattern; performing clustering binarization on the left and right images to obtain binary left and right images and left and right clustered pixel groups corresponding to the binary left and right images; obtaining the width of a standard line region based on the left and right clustered pixel groups; using the width of the standard line region as a mask width to perform a masking operation on the binary left and right images to obtain binary left and right line region images; calculating the center points of the binary left and right line region images to obtain left and right center point arrays; performing interpolation calculation on the left and right center point arrays to obtain corresponding left and right matching points; using the left and right matching points to perform stereo matching calculation to obtain disparity; and converting the disparity into depth to obtain a binocular depth image.

[0006] To address the aforementioned problems, this invention also provides a line-scanning depth calculation device for binocular vision, comprising: a binary clustering module for acquiring left and right images of a target object containing a multi-line pattern, performing clustering binarization on the left and right images to obtain binary left and right images and corresponding left and right clustered pixel groups; wherein the multi-line pattern is formed by multiple line beams projected onto the target object; a region extraction module for obtaining the width of a standard line region based on the left and right clustered pixel groups, using the standard line region width as a mask width to perform a masking operation on the binary left and right images to obtain binary left and right line region images; wherein the line region refers to the region composed of pixels in the binocular camera that respond to the line beams reflected back from the target; a matching point calculation module for calculating the center points of the binary left and right line region images to obtain left and right center point arrays, performing interpolation calculation on the left and right center point arrays to obtain corresponding left and right matching points; and a depth calculation module for performing stereo matching calculation using the left and right matching points to obtain disparity, and converting the disparity into depth to obtain a binocular depth image.

[0007] To address the aforementioned issues, this invention also provides a calibration system, including a TOF camera comprising a first transmitter and a receiver. The first transmitter emits at least one line beam onto a calibration plate, which is then received by the receiver and used to generate a corresponding TOF depth image, which is then transmitted to a processor. A binocular camera comprises a second transmitter, a left camera, and a right camera. The second transmitter emits at least one line beam onto the calibration plate, which is then captured by the left and right cameras to generate a left image and a right image. The image is then further processed using the aforementioned line scan depth calculation method to obtain a binocular depth image including the calibration plate, which is then transmitted to the processor. The processor controls the activation of the binocular camera and the TOF camera, and also converts the binocular depth image into a point cloud and projects the point cloud onto the image plane of the TOF camera to obtain a TOF depth ground truth image. The TOF depth image and the TOF depth ground truth image are used to train a preset neural network model to obtain a TOF depth ground truth network model.

[0008] To address the aforementioned problems, the present invention also provides a computer-readable storage medium storing at least one computer program, which is executed by a processor in a system to implement the line scan depth calculation method described above.

[0009] To address the aforementioned problems, the present invention also provides a Time-of-Flight (TOF) camera, comprising a transmitter, a receiver, and a processor; wherein the transmitter is used to emit a line beam toward a target, the receiver is used to acquire the line beam reflected by the target and generate a raw-phase image which is then transmitted to the processor; the processor is used to process the raw-phase image to obtain a TOF depth image, and input the TOF depth image into a preset TOF depth ground truth network model to obtain the target's depth ground truth; wherein the preset TOF depth ground truth network model is a neural network model pre-trained using the aforementioned calibration system.

[0010] Compared with existing technologies, this application proposes a line-scan depth calculation method and calibration system for binocular vision. On the one hand, by using a binocular camera with a variable baseline to lengthen or shorten the baseline and the line-scan depth calculation method in different scenarios, the accuracy of the binocular depth image is guaranteed. On the other hand, by converting the high-precision binocular depth image to the TOF camera coordinate system to obtain the true depth value of the TOF camera, the problem of low accuracy of the true depth value obtained by the TOF camera can be solved. Moreover, in practical applications, it can conveniently and accurately provide true depth training samples for deep learning algorithms. Attached Figure Description

[0011] Figure 1 This is a system architecture diagram of a calibration system for obtaining a TOF deep ground truth network model according to an embodiment of the present invention;

[0012] Figure 2 This is a schematic diagram of a multi-line pattern provided in an embodiment of the present invention;

[0013] Figure 3 This is a flowchart illustrating a line scan depth calculation method for binocular vision provided in an embodiment of the present invention.

[0014] Figure 4 This is a schematic diagram of the process of clustering left and right clustered pixel groups according to an embodiment of the present invention;

[0015] Figure 5 This is a schematic diagram of the clustered line region provided in an embodiment of the present invention;

[0016] Figure 6 This is a schematic diagram of the process for calculating matching points according to an embodiment of the present invention;

[0017] Figure 7 This is a functional block diagram of a line scan depth calculation device for binocular vision provided in an embodiment of the present invention.

[0018] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0019] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0020] Figure 1 This is a system architecture diagram of a calibration system for acquiring a TOF depth ground truth network model according to an embodiment of the present invention. The calibration system includes a calibration board 10, a TOF camera 11, a binocular camera with adjustable baseline length 12, and a processor 13, wherein:

[0021] The TOF camera 11 includes a first transmitter 110 and a receiver 111. The first transmitter 110 is used to emit at least one line beam onto the calibration plate 10 and receive it from the receiver 111 to further generate a TOF depth image containing the calibration plate and transmit it to the processor 13.

[0022] The binocular camera 12 includes a second transmitter 120, a left camera 121, a right camera 122, and an adjustable baseline mounting base 123. The second transmitter 120 is disposed between the left camera 121 and the right camera 122. The left camera 121 and the right camera 122 have the same focal length and are mounted on the adjustable length mounting base 123 to form a binocular camera with a variable baseline. The second transmitter 120 is used to emit at least one line beam to the calibration plate 10, which is then captured by the left camera 121 and the right camera 122 to generate a left image and a right image. The image is then further processed according to the line scan depth calculation method for binocular vision provided in one or more embodiments of this application to obtain a binocular depth image including the calibration plate and transmitted to the processor 13.

[0023] The processor 13 is used to control the simultaneous start-up of the TOF camera 11 and the binocular camera 12, and is also used to convert the binocular depth image into a point cloud and project the point cloud onto the image plane of the TOF camera to obtain the TOF depth ground truth image, thereby using the TOF depth image and the TOF depth ground truth image to train a preset neural network model to obtain a TOF depth ground truth network model.

[0024] In one embodiment, the binocular camera 12 includes an adjustable baseline mounting base 123 for adjusting the baseline length between the binocular cameras 12. When acquiring TOF depth images and TOF ground truth images of calibration plates at different distances, different baselines are used to generate TOF ground truth images. This not only improves the flexibility of the binocular camera in generating ground truth depth images and enables the same binocular camera to continuously calibrate calibration plates at different distances, but also improves measurement accuracy. When the main subject in the left image of the binocular camera is a near-field object, both cameras need to move towards the center simultaneously. When the main subject in the left image of the binocular camera is a far-field object, both cameras need to move to the sides simultaneously. The specific moving distance can be designed according to the actual situation.

[0025] In one embodiment, the processor 13 controls the first transmitter 110 of the TOF camera 11 or the second transmitter 120 of the binocular camera 12 to emit a line beam to a preset beam scanner. The preset beam scanner deflects the emitted line beam to form a multi-line pattern 30 to scan the calibration plate. The multi-line pattern 30 is composed of multiple line beams 301, such as... Figure 2 As shown; the processor 13 is also used to control the left and right cameras and the TOF camera to acquire the multi-line pattern reflected back by the calibration plate and generate left and right images and raw phase images. The raw phase image is the raw data from the receiver in the TOF camera, which converts the acquired light signal into a digital signal, and is used to generate the TOF depth image corresponding to the left and right images. It should be noted that the preset beam scanner in this embodiment can be a MEMS, a rotating prism, a galvanometer, etc. It should also be noted that the first transmitter 110 of the TOF camera 11 and the second transmitter 120 of the binocular camera 12 can be integrated into a single transmitter attached to the TOF camera 11 or the binocular camera 12; this application does not impose any limitations on this.

[0026] In one embodiment, after receiving the binocular depth image and the TOF depth image, the processor 13 is further configured to calculate the transformation relationship between the TOF camera and either of the binocular cameras using the pixels of the calibration plate contained in both the TOF depth image and the binocular depth image. This allows the binocular depth image acquired by the binocular camera 12 to be converted into a point cloud, and the transformation relationship is then projected onto the image plane of the TOF camera 11 to obtain the TOF depth ground truth image. Thus, during training using the TOF depth image and the TOF depth ground truth image, each pixel in the TOF depth image can directly learn the pixel depth ground truth at its corresponding coordinates.

[0027] In one embodiment, to obtain the transformation relationship between the TOF camera and any one of the binocular cameras 12, such as... Figure 1 As shown, in this embodiment, the right camera 122 is used as the reference camera. The TOF camera 11 is placed close to the right camera 122, and the receiver 111 of the TOF camera is placed adjacent to or attached to the right camera 122. This maximizes the shared field of view between the receiver 111 of the TOF depth camera and the right camera 122, ensuring that the TOF depth image and the binocular depth image have as many corresponding pixels as possible. This not only ensures the accuracy of the transformation relationship between the cameras but also ensures that each pixel of the receiver 111 of the TOF depth camera has a corresponding ground truth depth value. It should be noted that in this embodiment, the left camera 121 can also be used as the reference camera to calculate the transformation relationship between the TOF camera 11 and the left camera in the binocular camera 12, thereby ensuring a one-to-one pixel correspondence between the TOF depth image and the TOF ground truth depth image. The calibration board 10 can be a calibration whiteboard, a white wall, a checkerboard, etc.

[0028] Furthermore, the processor 13 transforms the binocular depth image to obtain a point cloud, and uses the transformation relationship between the TOF camera and either the binocular camera to project the point cloud onto the image plane of the TOF camera to obtain a true TOF depth image.

[0029] In one embodiment, converting the binocular depth image acquired by the binocular camera 12 into a point cloud and projecting it onto the image plane of the TOF camera 11 using a transformation relationship to obtain a TOF depth ground truth image includes: converting the binocular depth image acquired by the binocular camera into a point cloud; projecting the point cloud through perspective onto the image plane of the TOF camera receiver according to the transformation relationship between the cameras to obtain a TOF planar point cloud; and performing depth transformation based on the coordinate information of the TOF planar point cloud to obtain the TOF depth ground truth image of the TOF camera. In another embodiment, since the coordinates of each point in the projection plane point cloud are not necessarily integers on the image plane, the image plane triangle can be generated using the two-dimensional Delaunay method, and the depth ground truth corresponding to each pixel on the TOF camera receiver can be solved by determining the position of the pixel in the TOF camera receiver.

[0030] In practical applications, TOF depth images and ground truth TOF depth images acquired at the same time can be used as a set of training samples and fed into a neural network model. Multiple sets of training samples from different scenarios can be obtained, and the preset neural network model can be trained using these samples to obtain a neural network model with optimal weight parameters; this is the TOF depth ground truth network model. It should be understood that the baseline length of the stereo camera can be adjusted when acquiring different scenes to obtain high-precision stereo depth images, thereby obtaining high-precision TOF depth ground truth images. By updating and iterating the weight parameters of the neural network model using multiple sets of TOF depth images and ground truth TOF depth images from different scenarios to obtain the TOF depth ground truth network model with optimal weight parameters, it is ensured that the TOF depth images acquired by the TOF camera, when input into the network model, output high-precision and accurate depth images.

[0031] In one embodiment, the calibration system further includes a memory for storing computer programs that the processor can execute, such as a computer program for a line scan depth calculation method. The memory can be an internal storage unit, such as a hard disk or RAM; it can also be an external storage device, such as an external hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, or a Flash Card. Furthermore, the memory can include both internal and external storage units, and can also be used to temporarily store data that has been output or will be output; the memory can be integrated with the processor or set up independently of the processor.

[0032] In some embodiments, the processor 13 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processing units, neural network chips, and various control chips. The processor is the control unit of the calibration system, connecting various components of the entire calibration system through various interfaces and lines. It executes various functions of the calibration system and processes data by running or executing programs or modules stored in memory (e.g., executing depth truth acquisition programs) and calling data stored in memory. It should be understood that when the processor is a neural network chip, the calibration system does not include memory. Whether the calibration system needs to use memory to store the corresponding computer program depends on the type of processor. Furthermore, in this embodiment, the TOF camera 11 and the binocular camera 12 implicitly include a depth calculation engine for performing depth calculations on the acquired images. However, this depth calculation engine may be a separate processing chip from the processor 13, or it may be integrated with the processor 13; no limitation is made here.

[0033] It should be noted that the figure only shows the components of the system. Those skilled in the art will understand that the structure shown in the figure does not constitute a limitation on the system and may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0034] In one embodiment, the present invention also provides a Time-of-Flight (TOF) camera, which includes a transmitter, a receiver, and a processor. The transmitter emits a line beam towards a target; the receiver acquires the line beam reflected from the target and generates a raw-phase image, which is then transmitted to the processor. The processor processes the raw-phase image to obtain a TOF depth image and inputs the TOF depth image into a preset TOF depth network model to obtain the true depth value of the target. The preset TOF depth network model is... Figure 1 The calibration system shown was trained.

[0035] Figure 3 The diagram shown is a flowchart of a line scan depth calculation method for binocular vision provided by an embodiment of the present invention. The method includes:

[0036] S1. Obtain the left and right images of the target object containing the multi-line pattern, perform clustering binarization on the left and right images to obtain binary left and right images and the left and right clustered pixel groups corresponding to the binary left and right images.

[0037] Among them, the multi-line pattern is formed by projecting multiple beams of light onto the target object.

[0038] In one embodiment, acquiring left and right images of a target object containing a multi-line pattern specifically includes: taking pictures of the target object containing the multi-line pattern using a left camera and a right camera to obtain initial left and right images; and performing stereo correction on the initial left and right images to obtain the final left and right images. Wherein, when this method is applied to calibrate the intrinsic and extrinsic parameters of a binocular camera or to calibrate within a calibration system, the target object is the calibration board.

[0039] Specifically, stereo correction is performed on the initial left and right images to obtain the left and right images, including: randomly selecting one of the initial left and right images as the image to be corrected, and using the image other than the image to be corrected as the target reference image; obtaining the camera distortion intrinsic parameters corresponding to the image to be corrected, and using the camera distortion intrinsic parameters to perform distortion correction on the image to be corrected to obtain a standard image to be corrected; calculating the camera distortion extrinsic parameters corresponding to the image to be corrected based on the pixel mapping relationship between the image to be corrected and the target reference image, both of which contain the target object's pixels; performing a planar transformation on the standard image to be corrected based on the camera distortion extrinsic parameters to obtain a target transformed image, and combining the target transformed image and the target reference image to form the left and right images. Here, the camera distortion intrinsic parameters refer to the intrinsic parameters corresponding to the camera lens.

[0040] In one embodiment, using camera distortion intrinsic parameters to correct the distortion of the image to be corrected, and obtaining a standard image to be corrected, refers to calculating the standard image to be corrected using a tangential distortion correction algorithm based on the camera distortion intrinsic parameters and the image to be corrected. Specifically, the Fusiello epipolar correction algorithm can be used to calculate the camera distortion extrinsic parameters corresponding to the image to be corrected based on the mapping relationship between the image to be corrected and the target reference image. The Fusiello epipolar correction algorithm is an algorithm used to correct camera epipolar lines.

[0041] In this embodiment of the invention, by performing stereo correction on the initial left and right images, the left and right images are obtained. This includes correcting image distortion in the left and right images caused by differences in angles or distortion parameters between the left and right cameras, so that subsequent matching point searches only need to be performed in the horizontal direction, thus accelerating the calculation speed of the depth truth.

[0042] In one embodiment, performing clustering binarization on the left and right images to obtain binary left and right images and corresponding left and right clustered pixel groups includes: performing pixel clustering on the left and right images to obtain left and right clustered pixel groups; and performing binarization on the left and right images based on the left and right clustered pixel groups to obtain binary left and right images.

[0043] Specifically, refer to Figure 4As shown, pixel clustering is performed on the left and right images to obtain left and right clustered pixel groups, including:

[0044] S41. Divide the pixels in the left and right images into several pixel groups respectively, randomly select the initial center point of each pixel group, and calculate the distance from each pixel in the left and right images to the initial center point of each pixel group.

[0045] S42. Based on the distance of each pixel to the initial center point and the principle of proximity, group the pixels in the left and right images to obtain multiple standard pixel groups in the left and right images respectively.

[0046] S43. Calculate the secondary center point of each standard pixel group using the coordinate information of each pixel in each standard pixel group in the left and right images, and calculate the distance between each pixel in each standard pixel group and its corresponding secondary center point.

[0047] S44. Based on the distance of each pixel to the secondary center point and the principle of proximity, regroup each standard pixel group. Then repeat step S43 to iteratively calculate the center point of each group of pixels until the difference between the center points obtained by adjacent iterations is less than the distance threshold, thereby obtaining the left and right clustered pixel groups corresponding to the left and right images.

[0048] Specifically, binarizing the left and right images based on the left and right clustered pixel groups to obtain binary left and right images involves setting the pixel values ​​of pixels located inside the left and right clustered pixel groups to the first pixel value, and setting the pixel values ​​of pixels located outside the left and right clustered pixel groups to the second pixel value, thus obtaining binary left and right images. It should be noted that the first and second pixel values ​​are preferably 1 and 0, but other values ​​are also possible, and this application does not impose any restrictions on them.

[0049] In this embodiment of the invention, by acquiring the left and right images of the target object containing multi-line patterns, image distortion caused by the difference in angle or distortion parameters between the left and right cameras can be corrected. This allows subsequent searches for matching points to be performed only in the horizontal direction, accelerating the calculation of the true depth value. Performing clustering binarization on the left and right images to obtain binary left and right images and the corresponding left and right clustered pixel groups enables the left and right clustered pixel groups to correspond with the emitted beam of the TOF camera, thereby facilitating the formation of subsequent line regions and the subsequent calculation of the line region width.

[0050] S2. Obtain the width of the standard line region based on the left and right clustered pixel groups, and use the width of the standard line region as the mask width to perform a masking operation on the binary left and right images to obtain the binary left and right line region images.

[0051] The line region refers to the area in a binocular camera that is composed of pixels that respond to a line beam of light reflected back from the target object.

[0052] In one embodiment, a masking operation is performed on binary left and right images based on left and right clustered pixel groups to obtain binary left and right line region images. This includes: counting the number of pixels in each pixel set of the left and right clustered pixel groups to obtain left and right pixel arrays; and calculating the standard line region width corresponding to the left and right images based on the left and right pixel arrays. The standard line region width refers to the standard width of the line region, such as... Figure 5 As shown, line region 51 is the region composed of pixels 50 in the binocular camera that respond to each pair of beams reflected back from the target object (i.e., multi-line pattern); the width of the standard line region is used as the mask width, and the corresponding interest mask image is obtained according to the mask width; the interest mask image is used to perform masking operations on the binary left and right images to obtain binary left and right line region images; where the interest mask image is the mask image used to extract the region of interest, which is the line region; the masking operation refers to multiplying the interest mask image with the binary left and right images.

[0053] It should be noted that the left and right pixel arrays include a left pixel array and a right pixel array. Each element in the left pixel array refers to the number of pixels in the corresponding left cluster pixel group, and each element in the right pixel array refers to the number of pixels in the corresponding right cluster pixel group.

[0054] In one embodiment, obtaining the standard line region width based on left and right clustered pixel groups includes: calculating the left and right line region widths corresponding to the left and right images based on the left and right pixel arrays; calculating the average width of the left and right line region widths, and using the average width as the standard line region width. Further, the left and right images can be subdivided into subpixel segments based on the left and right pixel arrays using grayscale centroid methods or interpolation methods to obtain the left and right line region widths. In this embodiment of the invention, obtaining the standard line region width based on left and right clustered pixel groups improves the flexibility of depth ground truth calculation while increasing the accuracy to the subpixel level, thus improving the precision of depth ground truth calculation.

[0055] In this embodiment of the invention, by performing a masking operation on the binary left and right images based on the left and right clustered pixel groups, binary left and right line region images are obtained, which can further remove redundant pixels, improve the accuracy of depth calculation, and thus improve the accuracy of depth truth calculation.

[0056] S3. Calculate the center points of the binary left and right line region images to obtain the left and right center point arrays, and perform interpolation calculation on the left and right center point arrays to obtain the corresponding left and right matching points.

[0057] In one embodiment, step S3 includes: segmenting the binary left and right line region images according to pixel brightness to obtain corresponding left and right sub-region groups, and extracting left and right ROI region images from the left and right sub-region groups; calculating the center points of the left and right ROI region images to obtain left and right center point arrays, and interpolating the left and right center point arrays to obtain the left and right matching points corresponding to the left and right ROI region images; wherein, the ROI region in the left and right ROI region images refers to the region of interest, that is, the region of interest.

[0058] Specifically, the binary left and right line region images are segmented based on pixel brightness to obtain corresponding left and right sub-region groups. This involves traversing the line regions of each image in the binary left and right line region images using a sliding window approach, and selecting the line regions with the highest average pixel brightness to form the corresponding left and right sub-region images of the binary left and right line region images. By extracting the corresponding left and right sub-region images from the binary left and right line region images, the most obvious regions can be selected as samples for calculation, improving the accuracy of depth ground truth calculation.

[0059] In one embodiment, extracting left and right ROI region images from left and right sub-region images includes: using a bounding box algorithm to extract left and right ROI region images from the left and right sub-region images, i.e., the left and right ROI region image set is the minimum bounding rectangle of the left and right sub-region images. It should be noted that the bounding box algorithm is an algorithm for solving the optimal bounding space of a discrete point set. The basic idea is to approximate complex geometric objects with a slightly larger and simpler geometric shape.

[0060] In one embodiment, such as Figure 6 As shown, the center points of the left and right ROI regions are calculated to obtain left and right center point arrays. Interpolation is then performed on these left and right center point arrays to obtain the corresponding left and right matching points for the left and right ROI regions, including:

[0061] S61. Perform Gaussian smoothing on the left and right ROI region images based on the width of the standard line region to obtain smoothed left and right ROI images.

[0062] In one embodiment, Gaussian smoothing is performed on the left and right ROI region images based on the width of the standard line region to obtain smoothed left and right ROI images, wherein: pixels in the left and right ROI region images are selected one by one as target ROI pixels; Gaussian smoothing is performed on the target ROI pixels based on the line width to obtain smoothed ROI pixels; and the target ROI image is updated based on all the smoothed ROI pixels to obtain smoothed left and right ROI images.

[0063] In one embodiment, Gaussian smoothing includes:

[0064]

[0065] Where g(j) is the distance from the smoothed ROI pixel to the pixel center, g() is the Gaussian function symbol, j is the distance from the target ROI pixel to the pixel center, σ is the standard deviation, the specific value of σ is determined by the width of the standard line region, and exp is the exponential function symbol.

[0066] In this embodiment of the invention, by performing Gaussian smoothing on the target ROI pixels, smoothed ROI pixels are obtained, which can further remove noise in the image, thereby improving the accuracy of depth ground truth extraction.

[0067] S62. Use the preset smoothing feature algorithm to extract features from the left and right smoothed ROI images to obtain the left and right feature pixel groups.

[0068] Specifically, feature extraction is performed on the left and right smoothed ROI images using a preset smoothing feature algorithm to obtain feature pixel groups, including: feature extraction is performed on the left and right smoothed ROI images using the following smoothing feature algorithm to obtain left and right feature pixel groups:

[0069]

[0070] Here, H(x,y) represents the Hessian matrix, indicating the pixel feature corresponding to the pixel with coordinates (x,y) in the smoothed left and right ROI images after feature extraction. x is the x-coordinate of the pixel in the smoothed left and right ROI images, y is the y-coordinate of the pixel, and g() is the Gaussian function mentioned above. It should be noted that since the eigenvalues ​​of the Hessian matrix represent the concavity / convexity of its eigenvectors near a point, the larger the eigenvalue, the stronger the convexity, and the eigenvalue of the Hessian matrix at the center point is the largest.

[0071] In this embodiment, by using a smoothing feature algorithm to extract features from the left and right smoothed ROI images, feature pixel groups are obtained. The maximum feature value of the feature pixel group can be determined, which facilitates the subsequent extraction of the ground truth depth value.

[0072] S63. Calculate the center points of the left and right feature pixel groups using the preset sub-pixel coordinate formula to obtain the left and right center point arrays, and calculate the left and right matching points based on the left and right center point arrays by interpolation.

[0073] In one embodiment, left and right center point arrays are obtained by centering the left and right feature pixel groups using a preset sub-pixel coordinate formula. This includes: determining the center points of the left and right feature pixel groups and the corresponding horizontal and vertical normal vectors based on the maximum pixel feature values ​​of the left and right feature pixel groups; and calculating the sub-pixel coordinates of each center point in the left and right smoothed ROI images based on the horizontal and vertical normal vectors using the following sub-pixel coordinate formula:

[0074]

[0075] (p x ,p y )=(x+tn x ,y+tn y )

[0076] Where t is the subpixel coefficient of the subpixel coordinate formula, and n x It refers to the horizontal normal vector, n y This refers to the vertical normal vector, where n is the symbol for the normal vector, x is the x-coordinate of a pixel in the left and right smoothed ROI images, y is the y-coordinate of a pixel in the left and right smoothed ROI images, g() is the symbol for the Gaussian function, and p x This refers to the x-coordinate of the sub-pixel coordinates of the center point, p y It refers to the ordinate of the sub-pixel coordinates of the center point;

[0077] The coordinates of the center points of the left and right feature pixel groups are updated using the coordinates of each sub-pixel, resulting in the left and right feature sub-pixel groups. Feature value filtering is then performed on the left and right feature sub-pixel groups to obtain the left and right center point arrays. Here, the feature value filtering of the feature sub-pixel groups to obtain the center point arrays refers to filtering out pixels in the feature sub-pixel groups whose brightness difference is greater than a preset brightness threshold and whose feature value ratio is greater than a preset feature threshold.

[0078] In this embodiment of the invention, by using the subpixel coordinate formula to calculate the subpixel coordinates of each center point in the left and right smooth ROI images based on the horizontal and vertical normal vectors, subpixel-level values ​​can be obtained, thereby improving the accuracy of subsequent depth ground truth calculation.

[0079] In one embodiment, interpolating the left and right center point arrays to obtain the left and right matching points corresponding to the left and right ROI region images includes: selecting the side corresponding to the array with a smaller number of pixels in the left and right center point arrays as the starting point and sending a ray to the side corresponding to the array with a larger number of pixels; then performing linear interpolation on the two nearest coordinate points above and below the ray in the vertical direction to obtain the matching point. This is because the coordinates and number of the left and right center points obtained from the left and right ROI region images may be different, making one-to-one matching impossible. Therefore, interpolation can be performed based on each center point to obtain the matching point first.

[0080] Specifically, the matching point is obtained by linear interpolation of the two nearest coordinate points above and below the ray along the vertical axis using the following formula:

[0081]

[0082] Where (x0,y0) and (x1,y1) are the coordinates of the two points, b is the coordinate value of the matching point in the vertical direction, and a is the coordinate value of the matching point in the horizontal direction.

[0083] In this embodiment of the invention, by using the left and right sub-region groups to calculate the left and right ROI region image groups, most of the black areas can be removed, and the minimum bounding rectangle of the collected line beam portion can be retained. By using a preset Gaussian feature algorithm to extract the center point array from the left and right ROI region image groups, and interpolating the center point array to calculate the matching point, the accuracy of calculating the matching point can be improved, thereby improving the accuracy of subsequent depth ground truth calculation.

[0084] S4. Use the left and right matching points to perform stereo matching calculations to obtain the disparity, and convert the disparity into depth to obtain a binocular depth image.

[0085] Furthermore, after acquiring the stereo depth image, the stereo depth image is converted into a point cloud, and according to the transformation relationship between the TOF camera and either the stereo camera, the point cloud is transformed onto the image plane of the TOF depth camera, thereby obtaining the true TOF depth value of the TOF camera.

[0086] Figure 7 The diagram shows a functional block diagram of a line scan depth calculation device applied in a binocular camera according to an embodiment of the present invention. The line scan depth calculation device 700 of the present invention can be applied in a system. Depending on the functions implemented, the line scan depth calculation device 700 may include a binary clustering module 701, a region extraction module 702, a matching point calculation module 703, and a depth calculation module 704. A module of the present invention can also be referred to as a unit, which refers to a series of computer program segments that can be executed by a processor in a system and can perform a fixed function.

[0087] In this embodiment, the functions of each module / unit are as follows:

[0088] Binary clustering module 701 is used to acquire left and right images of a calibration board containing a multi-line pattern, perform clustering binarization on the left and right images to obtain binary left and right images and left and right clustered pixel groups corresponding to the binary left and right images; wherein, the multi-line pattern is formed by multiple line beams projected onto the target object;

[0089] The region extraction module 702 is used to obtain the standard line region width based on the left and right clustered pixel groups, use the standard line region width as the mask width to perform a mask operation on the binary left and right images to obtain binary left and right line region images, and extract the corresponding left and right sub-region groups from the binary left and right line region images based on the pixel brightness; wherein, the line region refers to the region composed of pixels in the binocular camera that respond to the line beam reflected back by the target object;

[0090] The matching point calculation module 703 is used to calculate the center points of the binary left and right line region images to obtain left and right center point arrays, and to perform interpolation calculations on the left and right center point arrays to obtain the corresponding left and right matching points.

[0091] The depth calculation module 704 is used to perform stereo matching calculations using the left and right matching points to obtain disparity, and convert the disparity into depth to obtain a binocular depth image.

[0092] It should be noted that, in this embodiment of the invention, each module in the line scan depth calculation device 700 adopts the same method as described above. Figures 3 to 6 The same technical means are used to calculate the line scan depth, and it can produce the same technical effect, so it will not be elaborated here.

[0093] Furthermore, if the integrated modules / units of the device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, a computer-readable medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, and read-only memory (ROM).

[0094] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a system processor, can implement the line scan depth calculation method of one or more embodiments provided in this application.

[0095] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.

[0096] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0097] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0098] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0099] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.

[0100] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0101] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a system claim may also be implemented by a single unit or device through software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any specific order.

[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for calculating line scan depth applied to binocular vision, characterized in that, The method includes: The left and right images of a target object containing a multi-line pattern are acquired, and the left and right images are subjected to clustering binarization to obtain binary left and right images and left and right clustered pixel groups corresponding to the binary left and right images; wherein, the multi-line pattern is formed by multiple line beams projected onto the target object; The standard line region width is obtained based on the left and right clustered pixel groups. The standard line region width is used as the mask width to perform a masking operation on the binary left and right images to obtain binary left and right line region images. Here, the line region refers to the region composed of pixels in the binocular camera that respond to the line beam reflected back by the target object. The center points of the binary left and right line region images are calculated to obtain left and right center point arrays. Interpolation calculations are performed on the left and right center point arrays to obtain the corresponding left and right matching points. The disparity is obtained by performing stereo matching calculations using the left and right matching points, and then the disparity is converted into depth to obtain a binocular depth image.

2. The line scan depth calculation method as described in claim 1, characterized in that, The acquisition of left and right images of the target object containing multi-line patterns includes: The target object containing the multi-line pattern is photographed using the left and right cameras to obtain initial left and right images; The initial left and right images are stereoscopically corrected to obtain the left and right images.

3. The line scan depth calculation method as described in claim 2, characterized in that, The step of performing stereoscopic correction on the initial left and right images to obtain the left and right images includes: One of the initial left and right images is randomly selected as the image to be corrected, and the images in the initial left and right images other than the image to be corrected are used as the target reference images; Obtain the camera distortion intrinsic parameters corresponding to the image to be corrected, and use the camera distortion intrinsic parameters to perform distortion correction on the image to be corrected to obtain a standard image to be corrected; The camera distortion extrinsic parameters corresponding to the image to be corrected are calculated based on the pixel mapping relationship between the image to be corrected and the target reference image, both of which contain the pixels of the target object. The standard image to be corrected is transformed using the camera distortion extrinsic parameters to obtain the target transformed image, and the target transformed image and the target reference image are combined to form left and right images.

4. The line scan depth calculation method as described in claim 1, characterized in that, The step of performing clustering binarization on the left and right images to obtain binary left and right images and corresponding left and right clustered pixel groups includes: Perform pixel clustering on the left and right images to obtain left and right clustered pixel groups; Binarize the left and right images based on the left and right clustered pixel groups to obtain binary left and right images.

5. The line scan depth calculation method as described in claim 4, characterized in that, The step of performing pixel clustering on the left and right images to obtain left and right clustered pixel groups includes: The pixels in the left and right images are divided into several pixel groups respectively. The initial center point of each pixel group is randomly selected, and the distance from each pixel in the left and right images to the initial center point of each pixel group is calculated. Based on the distance of each pixel to the initial center point and the principle of proximity, the pixels in the left and right images are grouped to obtain multiple standard pixel groups in the left and right images respectively. The secondary center point of each standard pixel group is calculated using the pixel coordinate information in each standard pixel group in the left and right images, and the distance between each pixel in each standard pixel group and its corresponding secondary center point is calculated. Based on the distance of each pixel to the secondary center point and the principle of proximity, the standard pixel groups are regrouped and the center point of each group of pixels is repeatedly calculated iteratively to obtain the left and right clustered pixel groups corresponding to the left and right images.

6. The line scan depth calculation method as described in claim 1, characterized in that, The step of obtaining the standard line region width based on the left and right clustered pixel groups, and using the standard line region width as the mask width to perform a masking operation on the binary left and right images to obtain binary left and right line region images includes: The number of pixels in each pixel set in the left and right clustered pixel groups is counted to obtain the left and right pixel arrays. The width of the standard line region corresponding to the left and right images is calculated based on the left and right pixel arrays. The width of the standard line region is used as the mask width, and the corresponding interest mask image is obtained based on the mask width; wherein, the interest mask image is a mask image used to extract the region of interest, and the region of interest is the line region; The interest mask image is used to perform masking operations on the binary left and right images to obtain the binary left and right line region images.

7. The line scan depth calculation method as described in claim 1, characterized in that, The calculation of the center points of the binary left and right line region images yields left and right center point arrays. Interpolation calculations are then performed on the left and right center point arrays to obtain corresponding left and right matching points, including: The binary left and right line region images are segmented according to pixel brightness to obtain corresponding left and right sub-region groups, and the left and right ROI region images are extracted from the left and right sub-region groups; The center points of the left and right ROI region images are calculated to obtain left and right center point arrays. The left and right center point arrays are interpolated to obtain the left and right matching points corresponding to the left and right ROI region images.

8. The line scan depth calculation method as described in claim 7, characterized in that, The calculation of the center points of the left and right ROI region images yields left and right center point arrays. Interpolation is then performed on these left and right center point arrays to obtain the left and right matching points corresponding to the left and right ROI region images, including: Gaussian smoothing is performed on the left and right ROI region images based on the width of the standard line region to obtain smoothed left and right ROI images; The left and right smoothed ROI images are used to extract features using a preset smoothing feature algorithm to obtain left and right feature pixel groups; The center points of the left and right feature pixel groups are calculated using a preset sub-pixel coordinate formula to obtain left and right center point arrays, and the left and right matching points are calculated by interpolation based on the left and right center point arrays.

9. A line-scan depth calculation device for binocular vision, characterized in that, include: The binary clustering module is used to acquire left and right images of a target object containing a multi-line pattern, perform clustering binarization on the left and right images to obtain binary left and right images and left and right clustered pixel groups corresponding to the binary left and right images; wherein, the multi-line pattern is formed by multiple line beams projected onto the target object; The region extraction module is used to obtain the standard line region width based on the left and right clustered pixel groups, and to perform a masking operation on the binary left and right images using the standard line region width as the mask width to obtain binary left and right line region images; wherein, the line region refers to the region composed of pixels in the binocular camera that respond to the line beam reflected back by the target object; The matching point calculation module is used to calculate the center points of the binary left and right line region images to obtain left and right center point arrays, and to perform interpolation calculations on the left and right center point arrays to obtain the corresponding left and right matching points. The depth calculation module is used to perform stereo matching calculations using the left and right matching points to obtain disparity, and then convert the disparity into depth to obtain a binocular depth image.

10. A calibration system, characterized in that, Includes a calibration board, a stereo camera, a TOF camera, and a processor, among which: The TOF camera includes a first transmitter and a receiver. The first transmitter is used to emit at least one line beam onto the calibration plate and receive it from the receiver to further generate a corresponding TOF depth image, which is then transmitted to the processor. The binocular camera includes a second transmitter, a left camera, and a right camera. The second transmitter is used to emit at least one line beam onto the calibration plate, and after being acquired by the left camera and the right camera to generate a left image and a right image, the depth is further calculated according to the line scan depth calculation method according to any one of claims 1-8 to obtain a binocular depth image including the calibration plate and transmitted to the processor. The processor is used to control the startup of the binocular camera and the TOF camera, and is also used to convert the binocular depth image into a point cloud and project the point cloud onto the image plane of the TOF camera to obtain a TOF depth ground truth image, and use the TOF depth image and the TOF depth ground truth image to train a preset neural network model to obtain a TOF depth ground truth network model.

11. The calibration system as described in claim 10, characterized in that, The binocular camera also includes an adjustable baseline mounting base. The second transmitter is located between the left camera and the right camera. The left camera and the right camera have the same focal length and are mounted on the adjustable length mounting base to form an adjustable baseline binocular camera.

12. A TOF camera, characterized in that, Includes the transmitter, receiver, and processor; among which: The transmitting end is used to transmit a line beam to the target; The receiving end is used to acquire the line beam reflected by the target and generate a raw phase image, which is then transmitted to the processor. The processor is configured to process the rawphase image to obtain a TOF depth image, and input the TOF depth image into a preset TOF depth ground truth network model to obtain the depth ground truth of the target; wherein, the preset TOF depth ground truth network model is a neural network model trained in advance using the calibration system of claim 10 or 11.

13. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the line scan depth calculation method as described in any one of claims 1 to 8.