Binocular structured light 3D reconstruction method, reconstruction system and reconstruction equipment

Through the TOF device-assisted binocular structured light three-dimensional reconstruction method, a set of structured light images and TOF device data are used to match points of the same name, which solves the problems of long image acquisition time and large data processing volume, and achieves efficient three-dimensional reconstruction.

CN114219866BActive Publication Date: 2025-08-19SUZHOU INST OF NANO TECH & NANO BIONICS CHINESE ACEDEMY OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111555932.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-17
Publication Date
2025-08-19
Estimated Expiration
2041-12-17

AI Technical Summary

Technical Problem

In the existing three-dimensional reconstruction method of binocular structured light, the image acquisition time is long, the number of images is large, resulting in slow reconstruction speed and large data processing volume.

Method used

The TOF device assisted in the three-dimensional reconstruction method of binocular structured light is used to calibrate the camera and project the structured light image, and a set of structured light images are used to match the same name point with the TOF device data, reducing the image acquisition time and data processing amount.

Benefits of technology

Significantly improve the speed of three-dimensional reconstruction, reduce image acquisition time by 2/3, greatly reduce data processing volume, and maintain high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114219866B_ABST
    Figure CN114219866B_ABST
Patent Text Reader

Abstract

The present invention discloses a binocular structured light 3D reconstruction method, reconstruction system, and reconstruction equipment. The reconstruction method includes a point cloud construction unit that constructs a 3D point cloud for a target; a first video capture unit and a second video capture unit each capture structured light image information of the target, wherein the structured light image information includes a set of multiple structured light images with different phases; the structured light image information acquired by the first video capture unit is matched with the 3D point cloud information by the same name to obtain first matching information; the first matching information is matched with the structured light image information acquired by the second video capture unit by the same name to obtain second matching information; and 3D reconstruction is performed based on the second matching information. The present invention significantly reduces image acquisition time and the amount of 3D modeling data processing, significantly improving the speed of 3D reconstruction while achieving good accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to three-dimensional modeling technology, and in particular to a binocular structured light three-dimensional reconstruction method, reconstruction system and reconstruction equipment. Background Art

[0002] As production and daily life demands increase, the use of 3D data is becoming increasingly widespread. For example, in production, 3D data of finished workpieces is collected for conformity testing, and robotic arms use 3D data to handle objects. In daily life, 3D facial recognition is commonplace, and autonomous driving is rapidly developing. Structured light-based 3D reconstruction has been widely used in various scenarios in recent years due to its high precision, high speed, low cost, and simple structure. Structured light 3D reconstruction can be performed using either monocular or binocular systems. Binocular systems offer higher 3D reconstruction accuracy due to their greater calibration precision.

[0003] Currently, binocular-based structured light 3D reconstruction methods typically use structured light with three frequencies and multiple phase shifts. To achieve higher accuracy, the images are divided into three groups: frequency 1, frequency 2, and frequency 3. A truncated phase is generated for each group. These three truncated phases are then used to create an unwrapped phase for matching homonymous points. The frequencies of the three image groups increase in sequence. The high frequency allows for accurate marking of homonymous points, but matching homonymous points can be challenging and requires combining the other two image groups, which requires the acquisition of a larger number of images.

[0004] Due to the camera exposure time, it takes a long time to collect images, and a large number of images has a significant impact on the reconstruction speed. The present invention improves this method, reduces the number of images required, and increases the speed of 3D reconstruction.

[0005] The information disclosed in this background technology section is only intended to enhance understanding of the overall background of the invention and should not be regarded as an admission or any form of suggestion that the information constitutes the prior art already known to a person skilled in the art. Summary of the Invention

[0006] The purpose of the present invention is to provide a binocular structured light 3D reconstruction method, reconstruction system and reconstruction equipment. Combined with the data of the TOF device, the matching of homonymous points can be achieved, which significantly reduces the time of image acquisition and greatly reduces the data that needs to be processed. It can significantly improve the speed of 3D reconstruction while maintaining accuracy.

[0007] To achieve the above-mentioned objectives, an embodiment of the present invention provides a binocular structured light three-dimensional reconstruction method, including: a point cloud construction unit, which constructs a three-dimensional point cloud for a target; a first video capture unit and a second video capture unit respectively collect structured light image information of the target, and the structured light image information includes a group of multiple structured light images with different phases; the structured light image information obtained by the first video capture unit is matched with the information of the three-dimensional point cloud by the same name to obtain first matching information; the first matching information is matched with the structured light image information obtained by the second video capture unit by the same name to obtain second matching information; and three-dimensional reconstruction is performed based on the second matching information.

[0008] In one or more embodiments of the present invention, a binocular structured light three-dimensional reconstruction method includes: constructing a coordinate system, calibrating a first video capture unit, a second video capture unit, and a point cloud construction unit; a projection unit projects structured light including phase information onto a target; a point cloud construction unit constructs a three-dimensional point cloud for the target; the first video capture unit and the second video capture unit respectively collect structured light image information of pixel coordinates (x, y) in the target, where the structured light image information includes a group of multiple structured light images with different phases; the structured light image information obtained by the first video capture unit is matched with the information of the three-dimensional point cloud with the same name to obtain first matching information; the first matching information is matched with the structured light image information obtained by the second video capture unit with the same name to obtain second matching information; and three-dimensional reconstruction is performed based on the second matching information.

[0009] In one or more embodiments of the present invention, the first matching information includes at least any one of the following: a first truncation phase, a first depth value, first disparity information, first disparity upsampling information, and first truncation phase correction information.

[0010] In one or more embodiments of the present invention, the first disparity upsampling information is obtained by upsampling the first disparity information.

[0011] In one or more embodiments of the present invention, the first truncation phase correction information is obtained by correcting the first truncation phase.

[0012] In one or more embodiments of the present invention, the second matching information includes at least any one of the following: second disparity information and second disparity upsampling information. Preferably, the second disparity upsampling information is obtained by upsampling the second disparity information.

[0013] In one or more embodiments of the present invention, the second disparity information is obtained by matching the first disparity information with the same name, and the premise of the matching with the same name satisfies the first disparity information: d err <p T , d err is the error and p of the first disparity information T is the phase period of the first disparity information.

[0014] In one or more embodiments of the present invention, a binocular structured light 3D reconstruction system includes a point cloud unit for collecting 3D information of a target object, constructing a 3D point cloud, and obtaining point cloud information; a plurality of video capture units for collecting structured light images of the target and obtaining structural image information; and a matching unit for pairing point cloud information with structural image information of the same name to generate 3D reconstruction information.

[0015] In one or more embodiments of the present invention, a binocular structured light 3D reconstruction system includes a point cloud unit for collecting 3D information of a target object, constructing a 3D point cloud, and obtaining point cloud information; a plurality of video capture units for collecting structured light images of the target and obtaining at least one set of structured light images with the same frequency for coordinates (x, y) to obtain structural image information; and a matching unit for pairing point cloud information with structural image information of the same name to generate 3D reconstruction information.

[0016] In one or more embodiments of the present invention, a binocular structured light three-dimensional reconstruction system includes at least two video capture units and a structured light projection unit, and also includes: a control unit, at least for calibrating the two video capture devices; a point cloud unit, for collecting three-dimensional information of the target object, constructing a three-dimensional point cloud, and obtaining point cloud information; the video capture unit, for collecting a structured light image of the target object, and obtaining at least one set of structured light images with the same frequency for the coordinates (x, y) to obtain structural image information; a matching unit, for pairing point cloud information and structural image information with the same name to generate three-dimensional reconstruction information.

[0017] In one or more embodiments of the present invention, the structured light image includes a group of first structured light images with the same frequency.

[0018] In one or more embodiments of the present invention, a set of first-page structured light images with the same frequency includes four images with a phase difference of π / 2.

[0019] In one or more embodiments of the present invention, the reconstruction system further comprises a projection unit for projecting information including different phases onto the target.

[0020] In one or more embodiments of the present invention, the binocular structured light 3D reconstruction device includes the binocular structured light 3D reconstruction system as described above.

[0021] Compared with the prior art, the binocular structured light 3D reconstruction method, reconstruction system and reconstruction device according to the embodiment of the present invention first calibrate the binocular camera, then project structured light stripes with phase information onto the object, and then use the camera to collect the deformed structured light image on the object, and then use the image to solve the phase information, and then use the phase information to match the same-name points to obtain the disparity, and finally obtain the three-dimensional coordinates through the disparity to achieve 3D reconstruction. In the implementation of the present invention, only one set of images in the three frequencies needs to be taken, and then combined with the data of the TOF device to achieve the matching of the same-name points, which reduces the image acquisition time by 2 / 3 and greatly reduces the data that needs to be processed. It can significantly improve the speed of 3D reconstruction while maintaining accuracy. The coarse disparity is obtained by using the TOF device point cloud, which greatly shortens the time spent on image acquisition and calculation; the coarse disparity narrows the search range of the same-name points and improves the speed of 3D reconstruction.

[0022] This method uses TOF device data to assist structured light to achieve the same accuracy while using fewer images and achieving 3D reconstruction at a faster speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a schematic diagram of a reconstruction process according to an embodiment of the present invention;

[0024] Figure 2 is a schematic diagram of reconstruction according to one embodiment of the present invention;

[0025] Figure 3 FIG. 1 is a schematic diagram of φ(x, y) performing homonymous point matching according to an embodiment of the present invention. DETAILED DESCRIPTION

[0026] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings, but it should be understood that the protection scope of the present invention is not limited by the specific embodiments.

[0027] Unless expressly stated otherwise, throughout the specification and claims, the term "comprise" or variations such as "include" or "comprising", etc., will be understood to include the stated elements or components but not to exclude other elements or other components.

[0028] like Figures 1 to 3 As shown, according to the preferred embodiment of the present invention:

[0029] Please note that including but not limited to the first video capture unit and the second video capture unit shown in the embodiment of the present invention is only for distinguishing video capture units with different lens directions, rather than just limiting the number of video capture units.

[0030] The low-precision point cloud constructed by the point cloud construction unit for 3D reconstruction provides a basis for homonymous matching. The first video capture unit and the second video capture unit set in different directions can be as follows: Figure 2 The two cameras shown have different lens directions. Of course, more cameras can be provided, and these cameras can all have different lens directions, so as to repeat the matching of the same name of the present invention to further improve the reconstruction accuracy.

[0031] In the implementation of the present invention, the above-mentioned video capture units can be used to capture the structural images of the target, and these structural images can be images with different frequency information as markers or distinctions. The capture can be performed under a structured light source.

[0032] At the same time, it should be noted that during the implementation of the present invention, Figure 1 In the same execution process shown, two pairs of the aforementioned video capture units are selected to participate. In multiple execution processes for the same coordinate, multiple groups can be used for reconstruction, and then downsampling, approximation, or interpolation operations can be performed on the multiple reconstruction results to determine the optimal reconstructed coordinates.

[0033] like Figure 1-3 As shown, the present invention is described below with a specific embodiment, but it is not intended to limit the scope of protection of the present invention:

[0034] Figure 2 The following is a diagram of the system setup, which includes two cameras, a TOF device, and a projector. First, the binocular system is calibrated to obtain the intrinsic parameters K and extrinsic parameters M of the two cameras:

[0035]

[0036] Among them, f x is the camera focal length f relative to the pixel width h x The ratio of f y is the camera focal length f relative to the pixel height h y The ratio of u0 and v0 is the coordinates of the center point of the pixel plane.

[0037]

[0038] Among them, r ij Represents the rotation relationship between the two camera coordinate systems, t i Represents the translation relationship between the origins of the two camera coordinate systems.

[0039] Also for TOF device and left camera (including but not limited to the "left and right" here, is for Figure 2The displayed paper image is calibrated with the observer’s viewing direction to obtain the external reference M T , which is the same as formula (2). TOF equipment can directly obtain the three-dimensional information of the object, but the accuracy is very low.

[0040] Next, the binocular system is used to collect structured light images, and the four projected images are expressed by formula (3)-( 6 ) display, you can choose the set of images with the highest frequency. Images with other frequencies can also be used. The selection of images with different frequencies depends at least on the accuracy of the TOF device:

[0041] I1=I′+I″cos[φ(x,y)] (3)

[0042]

[0043] I3=I′+I″cos[φ(x,y)+π] (5)

[0044]

[0045] Where (x, y) represents the pixel coordinates, φ(x, y) is the phase corresponding to the pixel, I′ is the background brightness of the projected structured light, and I″ is the modulated brightness of the projected structured light.

[0046] The truncated phase is solved according to formula (7). The truncated phase (which can be regarded as the first truncated phase) is the basis for matching the same-name points:

[0047] φ(x,y)=arctan((I4-I2) / (I1-I3)) (7)

[0048] The truncated phase can be epipolar corrected (which can be considered as obtaining the first truncated phase correction information) so that the homonymous points in the two cameras are located in the same row of pixels. Then, homonymous points are matched in corresponding rows. For example, after epipolar correction, the homonymous point in row 10 of the left camera is in row 10 of the right camera.

[0049] Epipolar correction can be exemplified as follows:

[0050] The Bouguet algorithm is used here to implement epipolar correction. This algorithm manipulates the left and right camera coordinate systems of the binocular camera to be coplanar and parallel by solving a rotation matrix, ensuring that points with the same name are in the same row of pixels. This process is divided into two steps: using the camera calibration parameters to generate a rotation and translation matrix that makes the two camera coordinate systems coplanar and parallel; then determining the corresponding pixel coordinates in the pre- and post-correction images to obtain the corrected image.

[0051] First, rotate the camera coordinate systems of the left and right cameras by half an angle each to make the image planes parallel. The matrix that can achieve this is the composite rotation matrix R of the left and right cameras. l 、R r According to the calibration, the rotation matrix R between the two camera coordinate systems can be obtained:

[0052]

[0053] The synthetic rotation matrix of the left camera and the right camera is obtained by formula (9):

[0054]

[0055] Then, the camera image plane is transferred to the coplanar plane. This step requires constructing the matrix R rect According to the relationship between the camera coordinate system and the rotated coordinate system:

[0056]

[0057] in, e3=e1×e2,Tc=[T x , T y , T z ], Tc is obtained through calibration, e1, e2, and e3 are obtained through Tc. Then the rotation matrix that makes the two camera coordinate systems coplanar and parallel can be obtained:

[0058] R′ l =R rect *R l , R′ r =R rect *R r (11)

[0059] Next, determine the correspondence between the pixel coordinates in the image before and after correction to obtain the corrected image. Let the pixel coordinates of the corrected image be (u, v), and the pixel coordinates are obtained according to equations (12) and (13):

[0060] x=(u-u0) / f x (12)

[0061] y=(v-v0) / f y (13)

[0062] Then the rotation matrix R′ in formula (11) l , R′ r Substitute into equation (14) and normalize using equations (15) and (16) to obtain the undistorted image coordinates (x′, y′) in the original image coordinate system:

[0063]

[0064] x′=X / W (15)

[0065] y′=Y / W (16)

[0066] Then, by adding the camera distortion through equations (17)-(19), we can obtain its coordinates (x″, y″) in the original image coordinate system:

[0067] r = x′ 2 +y′ 2 (17)

[0068] x″=x′(1+k1r 2 +k2r 4 )+2p1x′y′+p2(r 2 +2x′ 2 ) (18)

[0069] y″=y′(1+k1r 2 +k2r 4 )+p1(r 2 +2y′ 2 )+2p2x′y′ (19)

[0070] Among them, k1 and k2 are radial distortion coefficients, and p1 and p2 are tangential distortion coefficients.

[0071] Then, the image coordinates are converted to pixel coordinates using equations (20) and (21), and the pixel coordinates (u, v) of the corrected image corresponding to the pixel coordinates (u′, v′) in the original image are obtained:

[0072] u′=x″fx+u0 (20)

[0073] v′=y″f y +v0 (21)

[0074] Finally, the corrected image is obtained and the epipolar correction is completed.

[0075] However, since a row of pixels in the truncated phase contains multiple phases with a period of [0, 2π], there will be multiple points with the same phase when matching homonymous points based on phase, but only one is a homonymous point, such as Figure 3 , a truncated phase in the left camera is φ(x, y), and the corresponding pixel row in the right camera may have corresponding points φ(x1, y), φ(x2, y), and φ(x3, y). The next step is to use TOF to obtain the coarse disparity, narrow the matching range, and achieve matching of same-name points.

[0076] During the implementation of this solution, considering that when the truncated phase is used for matching homonymous points, due to the discreteness of pixels, the homonymous points may be located in the middle of the truncated phases of two adjacent pixels, so interpolation is used to achieve sub-pixel precision matching to obtain higher accuracy.

[0077] Use the TOF device to achieve the matching of homonymous points. First, obtain the TOF low-precision point cloud, project the point cloud of the TOF device into the left camera coordinate system, and perform this coordinate transformation using formula (22):

[0078]

[0079] in, is the coordinate in the left camera coordinate system, is the coordinate in the TOF device coordinate system, M T is the external parameter obtained by calibration. Then the point cloud in the left camera coordinate system is projected onto the image plane of the left camera, and the conversion is achieved through formula (23):

[0080]

[0081] Among them, (u, v) is the pixel coordinate corresponding to the three-dimensional coordinate, and K is the internal parameter obtained by calibration. At this time, the pixel (u, v) of the left camera can obtain a depth value z that may have lower accuracy. l (That is, it can be regarded as the first depth value), and then use it to calculate the disparity. Since there is a large error in the depth, this is a low-precision disparity, called coarse disparity or low-precision disparity (that is, it can be regarded as the first disparity information).

[0082] The process of obtaining coarse disparity is as follows, according to the relationship between disparity d and depth z (24):

[0083]

[0084] Among them, f x Same as (1), T is the camera system baseline, obtained by calibration. According to the low-precision depth value z l And (24), we can get the coarse parallax.

[0085] During the implementation process, since the resolution of the TOF device and the camera may be inconsistent, the parallax can be upsampled after the three-dimensional coordinates are projected onto the image plane to obtain dense parallax (that is, it can be regarded as the first parallax upsampling information) to overcome the impact of the inconsistent resolution of the TOF device and the camera.

[0086] On this basis, coarse disparity can be used to match homonymous points to obtain high-precision disparity (which can be regarded as second disparity information). Similarly, in the process of correcting the differences in high-precision disparity brought about by different coarse disparities, upsampling can also be applied to overcome the consistency problem, and then the second disparity upsampling information can be obtained. The premise of using coarse disparity for homonymous point matching is that the error d of the coarse disparity is err and phase period p T Satisfaction relationship: d err <p T , their units are the number of pixels. The reason is that Figure 3 As shown, there are multiple phase cycles in the search range when matching homonymous points. Therefore, to find homonymous points through coarse parallax, the coarse parallax must be able to narrow the search range to one cycle, that is, the error of coarse parallax should be within one phase cycle, otherwise more than one identical phase will be matched.

[0087] Taking the TOF device Azure Kinect DK as an example, its system error E systematic <11mm+0.1% distance; the binocular system is calibrated using a 130W pixel camera and an 8mm lens to obtain f x =1640.72205, baseline T = 400mm. When the working distance of the binocular system is 700mm, that is, the minimum working depth is 700mm, according to formula (24), the actual parallax is about 937 pixels, the coarse parallax obtained by the low-precision depth of the TOF device is about 922 pixels, and the parallax error is 15 pixels, ensuring that the phase period p in the image is T >15 pixels is sufficient. The parallax error varies under different circumstances. As can be seen from Equation (24), it is related to the parameters of the selected equipment and the system built.

[0088] Matching homonymous points. Based on the obtained coarse disparity, the approximate location of the homonymous point can be determined. Then, based on the truncated phase, points with the same phase within one phase cycle nearby are searched and identified as homonymous points. Specifically, if homonymous point matching is performed on the pixels in the i-th row, and the coarse disparity d is calculated for the pixels in the j-th column of the left camera, the homonymous point in the right camera is located near the jd-th column. Then, based on the truncated phase φ(i,j) of that pixel in the left camera, the pixel (m,n) with the same truncated phase is matched near the jd-th column of the right camera, that is, within the range (jd-pT, j-d+pT). This is the homonymous point.

[0089] When the phase is truncated for matching homonymous points, due to the discreteness of pixels, the homonymous points may be located in the middle of two adjacent pixels. Interpolation can be used to achieve sub-pixel precision matching to obtain higher accuracy.

[0090] The high-precision parallax is obtained and a point cloud is generated to realize three-dimensional reconstruction. That is, the TOF point cloud data in the scheme of the present invention is converted through formula (22), and the point cloud constructed by the final three-dimensional reconstruction is in the left camera coordinate system of the binocular system.

[0091] Compared to existing technologies, this solution significantly reduces the search range for homonymous points by using coarse parallax, shortening matching time and improving reconstruction speed. Existing homonymous point matching begins by searching the entire row, starting from the first pixel in the corresponding row. This solution uses coarse parallax to limit the matching range to within a single phase period, thus reducing the search range at the beginning of a row.

[0092] The foregoing descriptions of specific exemplary embodiments of the present invention are for purposes of illustration and description. These descriptions are not intended to limit the invention to the precise forms disclosed, and it is apparent that many variations and modifications are possible in light of the foregoing teachings. The exemplary embodiments have been selected and described for the purpose of explaining the specific principles of the invention and their practical application, thereby enabling those skilled in the art to realize and utilize a variety of exemplary embodiments of the invention and various options and modifications. The scope of the invention is intended to be defined by the claims and their equivalents.

Claims

1. A binocular structured light 3D reconstruction method, characterized in that: include: Point cloud construction unit, constructing a three-dimensional point cloud for the target; The first video capture unit and the second video capture unit respectively capture structured light image information of the target, the structured light image information including a set of multiple structured light images with the same frequency and different phases, solving a first truncated phase based on the multiple structured light images, performing epipolar correction on the truncated phase to obtain first truncated phase correction information, so that the same-name points in the first video capture unit and the second video capture unit are located in the same row of pixels; The structured light image information obtained by the first video capture unit is matched with the information of the three-dimensional point cloud to obtain first matching information, which includes a first depth value and a coarse disparity, the coarse disparity being first disparity information; matching the first matching information with the structured light image information obtained by the second video capture unit to obtain second matching information; The homonymous matching is to obtain the approximate location of the homonymous point based on the obtained coarse parallax, and then find the point with the same phase within a phase cycle nearby according to the first truncated phase, and determine it as the homonymous point; The three-dimensional reconstruction is performed based on the second matching information, the second matching information includes high-precision disparity, and the high-precision disparity is the second disparity information. The second disparity information is obtained by matching the first disparity information with the same name, and the premise of the same name matching satisfies the first disparity information: d err <p T , d err is the error and p of the first disparity information T is the phase period of the first disparity information.

2. The binocular structured light 3D reconstruction method according to claim 1, wherein: The first matching information includes: first disparity upsampling information.

3. The binocular structured light 3D reconstruction method according to claim 2, wherein: The first disparity upsampling information is obtained by upsampling the first disparity information.

4. The binocular structured light 3D reconstruction method according to claim 1, wherein: The second matching information includes: second disparity upsampling information.

5. A binocular structured light 3D reconstruction system using the method according to any one of claims 1 to 4, characterized in that: include Point cloud unit, used to collect three-dimensional information of the target object, construct three-dimensional point cloud, and obtain point cloud information; Several video capture units, used for collecting structured light images of the target and obtaining structural image information; The matching unit is used to pair point cloud information and structural image information with the same name to generate three-dimensional reconstruction information.

6. The binocular structured light 3D reconstruction system according to claim 5, wherein: The structured light image includes a group of first structured light images with the same frequency.

7. The binocular structured light 3D reconstruction system according to claim 6, wherein: The reconstruction system further comprises a projection unit for projecting information comprising different phases onto a target.

8. A binocular structured light 3D reconstruction device, comprising the binocular structured light 3D reconstruction system according to any one of claims 5 to 7.