Binocular structured light three-dimensional reconstruction method, device and system

By dividing the area to be tested into sub-regions and flexibly allocating the number of bits and offsets in the encoding, the problems of slow speed and low accuracy of 3D reconstruction in large scenes are solved, and efficient and accurate 3D reconstruction is achieved.

CN122115710APending Publication Date: 2026-05-29BINZHOU WEIQIAO NATIONAL SCIENCE & TECHNOLOGY ADVANCED TECHNOLOGY RESEARCH INSTITUTE +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610074434.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-20
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Traditional binocular structured light 3D reconstruction technology is slow and inaccurate in large scenes, and has high coding redundancy and mismatch rate, making it difficult to maintain global consistency.

Method used

The test area is divided into multiple sub-regions. The number of bits and offsets are flexibly allocated according to the complexity of the sub-regions. The global encoding value is determined by the local encoding value and the encoding offset, thus achieving global splicing.

Benefits of technology

It significantly reduces the amount of encoded data, shortens the acquisition time, improves the speed and accuracy of 3D reconstruction, avoids encoding conflicts, and ensures global consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115710A_ABST
    Figure CN122115710A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of three-dimensional reconstruction, and discloses a binocular structured light three-dimensional reconstruction method, device and system, which comprises the following steps: determining the complexity of each to-be-measured subregion in a to-be-measured region, determining the encoding bit number and the encoding offset of each to-be-measured subregion according to the complexity, sequentially projecting a structured light image sequence to each to-be-measured subregion according to different encoding bit numbers, and obtaining corresponding multiple frames of binocular views, determining the global encoding value of each to-be-measured subregion according to the multiple frames of binocular views and the encoding offset of each to-be-measured subregion, globally splicing each to-be-measured subregion based on the global encoding value, and obtaining a panoramic three-dimensional point cloud of the to-be-measured region. The application can realize flexible encoding of each to-be-measured subregion, significantly reduce the encoding data amount of simple regions, make the encoding values of different subregions form natural differentiation, effectively avoid partition encoding conflicts, and improve the speed and accuracy of three-dimensional reconstruction of the overall scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of three-dimensional reconstruction technology, specifically to a binocular structured light three-dimensional reconstruction method, apparatus, and system. Background Technology

[0002] Binocular structured light 3D reconstruction technology uses a projector to project structured light encoded patterns (such as Gray code) onto the surface of the object being measured. Binocular cameras simultaneously acquire deformed images, decode them to obtain the spatial position of each pixel, and perform triangulation through binocular parallax to achieve 3D point cloud reconstruction.

[0003] Taking Gray coding as an example, traditional Gray coding requires a uniform number of bits and resolution for the entire area under test. When the area under test exceeds the field of view of a single projection or camera, it needs to be scanned multiple times and the point cloud stitched together. If multiple areas are independently and fixedly coded, not only will a large amount of redundant coded data be generated in the low-detail areas of the area under test, increasing the acquisition time and data volume, but the same coded values ​​may also appear, resulting in slow speed, high mismatch rate, and difficulty in maintaining global consistency during 3D reconstruction. Summary of the Invention

[0004] This invention provides a binocular structured light 3D reconstruction method, device, and system to solve the problems of slow speed and low accuracy in 3D reconstruction in large scenes.

[0005] In a first aspect, the present invention provides a binocular structured light 3D reconstruction method, the method comprising: dividing the region to be tested into multiple sub-regions to be tested, determining the complexity of each sub-region to be tested, and determining the number of encoding bits and the encoding offset of the sub-region to be tested based on the complexity; sequentially projecting a sequence of structured light images onto the corresponding sub-regions to be tested according to different encoding bits, and obtaining multi-frame binocular views corresponding to each sub-region to be tested; for each sub-region to be tested, determining the corresponding local encoding value based on the multi-frame binocular views of the sub-region to be tested, and determining the global encoding value of the sub-region to be tested based on the local encoding value and the encoding offset; and globally stitching the sub-regions to be tested based on the global encoding value to obtain a panoramic 3D point cloud of the region to be tested.

[0006] The binocular structured light 3D reconstruction method provided by this invention determines the complexity of each sub-region in the test area, and determines the encoding bit depth and encoding offset of each sub-region based on the complexity. Then, structured light image sequences are projected onto each sub-region sequentially according to different encoding bit depths, acquiring corresponding multi-frame binocular views. A global encoding value is determined based on the multi-frame binocular views and encoding offset of each sub-region. Based on the global encoding value, each sub-region is globally stitched together to obtain a panoramic 3D point cloud of the test area. This invention assigns appropriate encoding bit depths to different sub-regions based on complexity determination, achieving flexible encoding of each sub-region. This not only significantly reduces the amount of encoded data in simple regions and shortens the structured light projection and image acquisition time, but also achieves a balance between accuracy and efficiency. Simultaneously, based on flexible encoding, a unique encoding offset is assigned to each sub-region. The global encoding value can be determined by the local encoding value and the encoding offset, naturally distinguishing the encoding values ​​of different sub-regions and effectively avoiding partition encoding conflicts. Finally, through flexible and precisely correlated regional encoding, the speed and accuracy of overall scene 3D reconstruction are improved.

[0007] In one optional implementation, the region to be tested is divided into multiple sub-regions, the complexity of each sub-region is determined, and the encoding bit length and encoding offset of each sub-region are determined based on the complexity. This includes: controlling a binocular imaging device to scan the region to be tested at a first resolution to obtain a first image; dividing the first image into multiple second images according to a preset division method to obtain multiple sub-regions, and determining the region index of each sub-region; scoring the complexity of each second image to obtain the complexity of each second image; determining the encoding bit length of each sub-region based on the complexity; and determining the corresponding encoding offset based on the encoding bit length and region index of each sub-region.

[0008] This invention performs a rapid, low-resolution pre-scan of the entire scene of the area to be tested. It can obtain scene information such as the approximate texture and detail distribution of the entire scene using only a small number of low-pixel images. Based on the scene information, the complexity of different sub-regions can be determined, thus providing a basis for determining the number of bits for encoding. This reduces the amount of encoded data in simple regions, ensures the amount of encoding in complex regions, and enables precise on-demand allocation of encoding resources, avoiding resource waste and precision loss.

[0009] In one optional implementation, determining the number of encoding bits for the sub-region under test based on complexity includes: for each sub-region under test, determining whether the sub-region under test is a complex region or a simple region based on complexity and a preset threshold; if it is a complex region, determining the number of encoding bits as a first number of encoding bits; if it is a simple region, determining the number of encoding bits as a second number of encoding bits, wherein the first number of encoding bits is greater than the second number of encoding bits; or, determining the maximum number of encoding bits and the minimum number of encoding bits, and determining the number of encoding bits for each sub-region under test based on the maximum number of encoding bits, the minimum number of encoding bits, and the complexity of each sub-region under test.

[0010] This invention provides two flexible and adaptable coding bit determination strategies, allowing for optimal selection based on the needs of different application scenarios. The first strategy, which classifies the complexity of a sub-region through a single comparison with a preset threshold, ensures both a minimum accuracy and a maximum efficiency, rapidly outputting coding bit determination results and adapting to rapid partitioning scans of large-scale test regions. The second strategy, which allocates an appropriate coding bit within the range of maximum and minimum coding bit based on the actual complexity of the sub-region, enables fine-grained matching of coding resources, covering scenarios with diverse complexity gradients.

[0011] In one optional implementation, the process of determining the preset threshold includes: setting the preset threshold as a fixed threshold; or, determining the preset threshold according to a preset ratio based on the complexity distribution of all sub-regions to be tested; or, constructing a comprehensive loss function based on minimizing the expected reconstruction error and sampling time cost, and determining the preset threshold based on the comprehensive loss function.

[0012] This invention provides three flexible and adjustable preset threshold determination methods to adapt to the accuracy, efficiency, and operational requirements of different application scenarios. The fixed threshold method has simple logic and low operation threshold, making it suitable for rapid partitioning scenarios of large-scale test areas. The preset proportional threshold method based on complexity distribution has strong scenario adaptability, can avoid the judgment bias of the fixed threshold, and balance global reconstruction resources. The cost balance method based on comprehensive loss function takes into account both accuracy and efficiency, and achieves the global optimal solution of the threshold, making it suitable for complex scenarios with high requirements for both accuracy and efficiency.

[0013] In one optional implementation, the binocular view is obtained by scanning the sub-region under test by a binocular imaging device at a second resolution, where the second resolution is higher than the first resolution. Local coding values ​​are determined based on multiple frames of binocular views of the sub-region under test, and global coding values ​​are determined based on the local coding values ​​and coding offsets. This includes: calculating a first difference image between the left view and the reference image in each frame of the binocular view, and calculating a second difference image between the right view and the reference image in the binocular view; calculating the first sine component and the first cosine component of each first pixel in the left view based on the multiple first difference images corresponding to the multiple frames of the binocular view, and calculating the first sine component and the first cosine component of each first pixel in the left view based on the multiple second difference images corresponding to the multiple frames of the binocular view. For the difference image, calculate the second sine component and the second cosine component of each second pixel in the right view; for any first pixel in the left view, calculate the first wrapping phase of the first pixel based on the first sine component and the first cosine component, and determine the first absolute phase of the first pixel based on the first wrapping phase; for any second pixel in the right view, calculate the second wrapping phase of the second pixel based on the second sine component and the second cosine component, and determine the second absolute phase of the second pixel based on the second wrapping phase; decode based on the first absolute phase or the second absolute phase to obtain the local encoded value of the sub-region to be tested, and superimpose the local encoded value with the encoded offset to obtain the global encoded value.

[0014] This invention enables the encoding values ​​of different sub-regions to form natural numerical ranges by locally encoding each sub-region and combining the encoding offset of the sub-region relative to the whole region. This avoids the problem of repeated encoding values ​​that easily occurs in traditional partitioned independent encoding, and provides a globally unique encoding identifier for the splicing of adjacent sub-regions. Initial alignment can be achieved without relying on additional feature matching algorithms, which greatly improves splicing efficiency and robustness.

[0015] In one optional implementation, global stitching is performed on each sub-region to be tested based on global encoding values ​​to obtain a panoramic 3D point cloud of the region to be tested. This includes: within each sub-region to be tested, determining pixels with the same phase in the left and right views based on the first and second absolute phases, and determining the disparity of the binocular imaging device based on the pixels with the same phase; converting the two-dimensional planar coordinates of the first pixel in the left view into local point cloud coordinates based on the focal length, baseline distance, principal point coordinates, and disparity of the binocular imaging device; determining the overlapping area of ​​each pair of adjacent sub-regions to be tested according to the preset division method and scanning order of the sub-regions to be tested, and performing local stitching on each pair of adjacent sub-regions to be tested according to the overlapping area, converting the local point cloud coordinates within the sub-regions to be tested into global point cloud coordinates; and performing global stitching based on the global point cloud coordinates of all sub-regions to be tested to obtain a panoramic 3D point cloud of the region to be tested.

[0016] This invention transforms the two-dimensional planar coordinates of pixels in a binocular view into local point cloud coordinates, and after stitching together the various sub-regions to be measured, transforms the local point cloud coordinates into global point cloud coordinates. This reduces the computational pressure of single-batch data processing by calculating independently for each sub-region, avoiding the risk of memory overflow during large-scene reconstruction. Furthermore, it extracts overlapping areas according to a preset division method and scanning order, providing a clear spatial correlation basis for local stitching, effectively controlling the cumulative error during the stitching process, and solving the problem of traditional large-field-of-view partitioned scanning becoming increasingly off-center as the stitching progresses.

[0017] In one optional implementation, local stitching is performed on each pair of adjacent test sub-regions according to the overlapping region, including: obtaining a first overlapping point cloud fragment of the first test sub-region and a second overlapping point cloud fragment of the second test sub-region in each pair of adjacent test sub-regions based on the overlapping region; for each pair of adjacent test sub-regions, the first test sub-region and the second test sub-region are initially aligned according to the first global encoding value of each first spatial point in the first overlapping point cloud region and the second global encoding value of each second spatial point in the second overlapping point cloud fragment; after the initial alignment, each first spatial point and each second spatial point are re-aligned based on the iterative nearest point algorithm, and the three-dimensional rigid body transformation matrix of the iterative nearest point algorithm is adjusted during the re-alignment process; an error function is constructed to calculate the reprojection error after adjustment, and the optimal transformation matrix is ​​determined when the reprojection error reaches the minimum value; the local point cloud coordinates in each first test sub-region and the second test sub-region are converted into global point cloud coordinates according to the optimal transformation matrix.

[0018] This invention identifies overlapping point cloud fragments in adjacent test sub-regions, thereby locking in the spatial correlation data required for stitching. Based on the unique and continuous characteristics of globally encoded sub-regions, it aligns overlapping sub-regions, accurately mapping the local point cloud coordinates of each sub-region to the global coordinate system, achieving an orderly progression from local stitching to global uniformity.

[0019] In an optional implementation, when performing global stitching based on the global point cloud coordinates of all sub-regions to be tested, the method further includes: weighting and summing the first global point cloud coordinates of spatial points in the first overlapping point cloud segment and the second global point cloud coordinates of spatial points in the second overlapping point cloud segment of each pair of adjacent sub-regions to be tested according to a preset weight coefficient, so as to obtain the global point cloud coordinates of spatial points within the overlapping region.

[0020] This invention achieves a balanced correction of two sets of deviation data by weighted summation of the global point cloud coordinates of corresponding spatial points in the overlapping point cloud segments of adjacent test sub-regions. This avoids the problem of deviation amplification caused by directly selecting the coordinates of a single sub-region, realizes a smooth transition of the point cloud in the overlapping area, eliminates splicing discontinuities, and makes the final global point cloud coordinates of the overlapping area closer to the real spatial position of the object.

[0021] Secondly, the present invention provides a binocular structured light 3D reconstruction device, comprising: an encoding bit depth determination module, used to divide the region to be tested into multiple sub-regions to be tested, determine the complexity of each sub-region to be tested, and determine the encoding bit depth and encoding offset of the sub-region to be tested based on the complexity; an encoding pattern projection module, used to sequentially project a sequence of structured light images onto the corresponding sub-regions to be tested according to different encoding bit depths, and obtain multi-frame binocular views corresponding to each sub-region to be tested; a global encoding determination module, used to determine the corresponding local encoding value for each sub-region to be tested based on the multi-frame binocular views of the sub-region to be tested, and determine the global encoding value of the sub-region to be tested based on the local encoding value and the encoding offset; and a 3D point cloud reconstruction module, used to globally stitch together each sub-region to be tested based on the global encoding value to obtain a panoramic 3D point cloud of the region to be tested.

[0022] Thirdly, the present invention provides a binocular structured light three-dimensional reconstruction system, comprising: a controller, a structured light projection device, and a binocular imaging device, wherein the structured light projection device and the binocular imaging device are both connected to the controller; the controller comprises: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the binocular structured light three-dimensional reconstruction method of the first aspect or any corresponding embodiment described above. Attached Figure Description

[0023] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0024] Figure 1 This is a schematic diagram of the structure of a binocular structured light three-dimensional reconstruction system according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the first process of the binocular structured light three-dimensional reconstruction method according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the second process of the binocular structured light three-dimensional reconstruction method according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the third process of the binocular structured light three-dimensional reconstruction method according to an embodiment of the present invention; Figure 5 This is a structural block diagram of a binocular structured light three-dimensional reconstruction device according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the hardware structure of the controller according to an embodiment of the present invention. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.

[0027] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0028] As an optional application scenario of this invention, such as Figure 1 As shown, the binocular structured light 3D reconstruction system 100 includes: a controller 101, a structured light projection device 102, and a binocular imaging device 103. Both the structured light projection device 102 and the binocular imaging device 103 are connected to the controller 101.

[0029] Among them, binocular structured light 3D reconstruction technology is a high-precision non-contact optical 3D measurement method that integrates binocular stereo vision and active structured light projection. By projecting a known coded structured light pattern onto the object being measured, the binocular camera synchronously acquires images of the deformed light pattern from different perspectives. Combining the triangulation principle and phase calculation technology, the 3D coordinate information of the object's surface is accurately restored, and finally a complete 3D point cloud model or mesh model is generated.

[0030] In this embodiment of the invention, the core of the structured light projection device 102 is a digital micromirror device projector or a laser projector, which is used to generate and project structured light patterns with specific coding rules. Common types include multi-frequency phase-shift sinusoidal fringes, Gray codes, random speckle patterns, etc.

[0031] The binocular imaging device 103 consists of two industrial cameras with identical parameters. Each camera has a built-in high-resolution CMOS / CCD sensor and a fixed-focus lens. The two cameras are placed parallel or symmetrically. The baseline distance is designed according to the measurement range and accuracy requirements. Generally, a larger baseline results in a larger measurement range, while a smaller baseline results in higher measurement accuracy.

[0032] The controller 101 consists of an industrial computer and a dedicated image processing card. It is responsible for the operation of core algorithms such as synchronous control of the system (coordinating the timing of projector pattern switching and camera image acquisition), image preprocessing, phase calculation, corresponding point matching, three-dimensional coordinate calculation, and point cloud post-processing (denoising, stitching, and meshing).

[0033] In the process of 3D reconstruction of the area under test, if the area exceeds the field of view of a single projection or camera, multiple scans and point cloud stitching are required. However, related technologies typically use a uniform encoding bit depth and resolution for the entire area under test during encoding. For example, a projection resolution of 1024×768 requires at least 10 bits of binary encoding. Globally fixed encoding generates a large amount of redundant information in low-detail areas, increasing acquisition time and data volume. Furthermore, if the area under test is scanned multiple times and point cloud stitching is performed, traditional stitching algorithms rely on the target board or external sensors, which is cumbersome and costly. Moreover, when multiple areas are encoded independently, the same encoding value may occur, making it difficult to maintain global consistency during stitching. In addition, when aligning point clouds for stitching, the commonly used Iterative Closest Point (ICP) algorithm depends on the point cloud density and initial position, which is prone to mismatches due to noise and occlusion, making it difficult to guarantee stitching accuracy. Therefore, this invention provides a binocular structured light 3D reconstruction method. By flexibly encoding sub-regions of the area to be measured and assigning a unique encoding offset to each sub-region, the encoding values ​​of different sub-regions can be distinguished, thereby achieving flexible and accurate association encoding of sub-regions and improving the speed and accuracy of 3D reconstruction.

[0034] According to an embodiment of the present invention, a binocular structured light three-dimensional reconstruction method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0035] This embodiment provides a binocular structured light 3D reconstruction method, which can be used in the aforementioned binocular structured light 3D reconstruction system. Figure 2 This is a flowchart of a binocular structured light 3D reconstruction method according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps: Step S201: Divide the region to be tested into multiple sub-regions, determine the complexity of each sub-region, and determine the number of bits and the encoding offset of the sub-region based on the complexity.

[0036] Specifically, in this embodiment of the invention, before the formal scanning of the area to be tested, a rapid pre-scanning step is first performed. The pre-scanning resolution is lower than that of the formal projection resolution, thereby obtaining a low-pixel image. Based on the low-pixel image, the area to be tested is divided into multiple sub-images, each sub-image corresponding to a sub-region to be tested, and the complexity of the sub-region to be tested is determined according to the sub-image. The complexity can comprehensively quantify the richness of surface details, geometric features, and difficulty of structured light stripe deformation of each sub-region to be tested. Therefore, this embodiment of the invention determines the number of encoding bits for each sub-region to be tested in the subsequent 3D reconstruction process based on the complexity of each sub-region to be tested, and assigns a unique encoding offset to each sub-region to be tested based on the number of encoding bits.

[0037] In step S202, structured light image sequences are projected sequentially onto the corresponding test sub-regions according to different encoding bit lengths, and multi-frame binocular views corresponding to each test sub-region are obtained.

[0038] Specifically, in this embodiment of the invention, the controller generates a structured light image sequence matching the number of encoding bits for each sub-region to be tested, determined by a pre-scan. Taking Gray code as an example, the number of encoding bits directly determines the number of frames in the structured light image sequence. For regions with high encoding bit bits, more frames or higher frequency stripes are projected; for regions with low encoding bit bits, fewer frames are projected. For example, if the encoding bit bit for the sub-region to be tested is 8, then 8 frames of structured light images with different encoding patterns are generated. The generated structured light image sequence has strict encoding logic, with only one bit difference in the encoding values ​​of adjacent frame patterns, ensuring the uniqueness of subsequent decoding.

[0039] In some optional implementations, the structured light projection device, under the coordinated control of the controller, performs directional focusing projection on each sub-region under test according to the scanning sequence of the sub-regions to be tested, such as a zone-by-zone scanning sequence from left to right or from top to bottom. During the projection process, the projection range of the structured light can be limited by means of lens clipping or mechanical blocking, covering only the currently traversed sub-region under test, thus avoiding interference from the projected light to adjacent sub-regions under test. The projection timing is strictly controlled, and the projection duration of each frame of structured light image is precisely synchronized with the exposure duration of the binocular imaging device.

[0040] Furthermore, while the structured light projection device projects each frame of the coded pattern, the binocular imaging device synchronously triggers exposure at a resolution corresponding to the projection sequence (higher than the low resolution of the pre-scan), acquiring a structured light deformation image of the surface of the sub-region under test. For each frame of structured light image projected, the binocular imaging device outputs one frame of binocular view; after the projection of all coded patterns for a sub-region under test is completed, multiple frames of binocular views consistent with the number of frames for the coded bit depth can be obtained, such as 8 frames of binocular views for 8-bit coding, which is only an example and not a limitation.

[0041] Step S203: For each sub-region to be tested, determine the corresponding local coding value based on the multi-frame binocular view of the sub-region to be tested, and determine the global coding value of the sub-region to be tested based on the local coding value and the coding offset.

[0042] Specifically, in this embodiment of the invention, the encoding of each sub-region under test is consistent with normal encoding. It involves projecting a structured light image sequence matching the number of encoding bits, acquiring a binocular deformed view, and then performing steps such as difference image denoising, phase calculation, and phase decoding to obtain the local encoding value of each pixel within the sub-region under test. However, before encoding, this embodiment assigns a unique encoding offset to each sub-region under test. This encoding offset is a non-negative integer strongly correlated with the number of encoding bits and the region index of the sub-region. Therefore, after determining the corresponding local encoding value of any sub-region under test based on multiple frames of binocular views, the local encoding value is numerically superimposed with the unique encoding offset assigned to that sub-region to generate a globally unique global encoding value. This fundamentally avoids encoding value conflicts arising from independent encoding of different sub-regions under test, providing a precise and unified encoding identifier for subsequent global stitching.

[0043] Step S204: Based on the global encoding value, perform global stitching on each sub-region to be tested to obtain a panoramic 3D point cloud of the region to be tested.

[0044] Specifically, in this embodiment of the invention, the controller uses the global encoding value of each sub-region under test as the core association basis. First, for each sub-region under test, the controller combines the disparity calculated from its multi-frame binocular view and the intrinsic and extrinsic parameters of the binocular imaging device to complete the 3D coordinate transformation of the pixels inside the sub-region under test, thereby obtaining the local 3D point cloud of each sub-region under test. Then, relying on the uniqueness and continuity of the global encoding value, the controller establishes the spatial correspondence between adjacent sub-regions under test, and integrates the local 3D point clouds of all sub-regions under test into the same global coordinate system in sequence. Finally, by uniformly standardizing the integrated point cloud data, a panoramic 3D point cloud covering the entire sub-region under test is formed, thereby realizing the complete 3D reconstruction of the sub-region under test.

[0045] The binocular structured light 3D reconstruction method provided by this invention determines the complexity of each sub-region in the test area, and determines the encoding bit depth and encoding offset of each sub-region based on the complexity. Then, structured light image sequences are projected onto each sub-region sequentially according to different encoding bit depths, acquiring corresponding multi-frame binocular views. A global encoding value is determined based on the multi-frame binocular views and encoding offset of each sub-region. Based on the global encoding value, each sub-region is globally stitched together to obtain a panoramic 3D point cloud of the test area. This invention assigns appropriate encoding bit depths to different sub-regions based on complexity determination, achieving flexible encoding of each sub-region. This not only significantly reduces the amount of encoded data in simple regions and shortens the structured light projection and image acquisition time, but also achieves a balance between accuracy and efficiency. Simultaneously, based on flexible encoding, a unique encoding offset is assigned to each sub-region. The global encoding value can be determined by the local encoding value and the encoding offset, naturally distinguishing the encoding values ​​of different sub-regions and effectively avoiding partition encoding conflicts. Finally, through flexible and precisely correlated regional encoding, the speed and accuracy of overall scene 3D reconstruction are improved.

[0046] This embodiment provides a binocular structured light 3D reconstruction method, which can be used in the aforementioned binocular structured light 3D reconstruction system. Figure 3 This is a flowchart of a binocular structured light 3D reconstruction method according to an embodiment of the present invention, such as... Figure 3 As shown, the process includes the following steps: Step S301: Divide the region to be tested into multiple sub-regions, determine the complexity of each sub-region, and determine the number of bits and the encoding offset of the sub-region based on the complexity.

[0047] Specifically, step S301 includes: Step S3011: Control the binocular imaging device to scan the area to be measured at a first resolution to obtain a first image.

[0048] Specifically, in this embodiment of the invention, a first resolution of 1 / 4 or 1 / 8 of the normal horizontal resolution is preferred. This allows the binocular imaging device to perform a low-resolution rapid pre-scan of the entire area to be tested according to the first resolution, obtaining a low-pixel first image. This minimizes time costs while ensuring the complexity of the detection requirements. For example, if the normal horizontal resolution is 1024 pixels, the first resolution for the rapid pre-scan is 128 to 256 pixels. This is merely an example and not a limitation.

[0049] Step S3012: Divide the first image into multiple second images according to a preset division method to obtain multiple sub-regions to be tested, and determine the region index of each sub-region to be tested.

[0050] Specifically, in this embodiment of the invention, the region to be tested is divided based on the low-resolution image obtained from the pre-scanning according to a preset division method, resulting in multiple second images. Each second image corresponds to a sub-region to be tested, thereby dividing the region to be tested into multiple sub-regions. The preset division method can be fixed or determined according to the range of the region to be tested, and is not limited here. After obtaining multiple sub-regions to be tested, an index is assigned to each sub-region to determine the region index of each sub-region.

[0051] Step S3013: Score the complexity of each second image to obtain the complexity of each second image.

[0052] Specifically, in this embodiment of the invention, a convolutional neural network (CNN) or texture analysis algorithm is used to score the complexity of each sub-region to be tested. In one approach, if a deep learning (CNN)-based method is used, the second image corresponding to each test sub-region is input into a pre-trained lightweight convolutional neural network. The network automatically extracts the edge and texture features of the image through convolutional layers and maps them to a value between 0 and 1 through fully connected layers. This value is the complexity score. If a texture gradient-based method is used, a gradient operator (such as the Sobel operator) is used to calculate the pixel grayscale gradient of the second image corresponding to each test sub-region. The average gradient magnitude within the region is calculated (a larger gradient indicates richer texture), and this magnitude is normalized to a value between 0 and 1, which is then used as the complexity score.

[0053] Step S3014: Determine the number of bits in the encoding of the sub-region to be tested based on the complexity.

[0054] Specifically, in this embodiment of the invention, two flexible and adaptable coding bit determination strategies are provided, including a fixed hierarchical coding method and a dynamic continuous coding method, and the optimal one is selected according to the needs of different application scenarios.

[0055] The fixed hierarchical coding method involves determining whether each sub-region to be tested is complex or simple based on its complexity and a preset threshold. If it is a complex region, the first coding bit length is set to a certain value; if it is a simple region, the second coding bit length is set to a certain value. The first coding bit length is greater than the second coding bit length. In other words, the complexity of each sub-region to be tested is... With preset threshold To make a comparison, if If the sub-region to be tested is determined to be a complex region, a high number of coding bits are allocated; if If the sub-region to be tested is determined to be a simple region, the low coding bit length, high coding bit length, and low coding bit length are allocated in advance, and must be higher than the minimum coding bit length of the projector resolution. For example, for a projection with a width of 1024, the minimum number of bits required for encoding. for:

[0056] In one optional implementation, a preset threshold is used in the fixed hierarchical coding method. There are three ways to determine the threshold: ① Set the preset threshold to a fixed threshold, for example... ② Based on the complexity distribution of all sub-regions to be tested, determine a preset threshold according to a preset ratio. For example, after sorting the complexity of all sub-regions to be tested from low to high, an ordered complexity data sequence is formed, and the complexity at the 70th percentile of the entire sequence is taken. As a preset threshold ③ A comprehensive loss function is constructed based on minimizing the expected reconstruction error and the sampling time cost. A preset threshold is determined based on the comprehensive loss function. The expected reconstruction error is the accuracy deviation of the 3D reconstruction of the corresponding sub-region when different thresholds are used to divide complex / simple regions. The sampling time cost is the total time taken to complete structured light projection and image acquisition for all sub-regions under the corresponding threshold. The two indicators are weighted according to actual needs to construct the comprehensive loss function. The threshold that minimizes the loss function value is selected as the preset threshold for dividing complex / simple regions. .

[0057] Furthermore, the dynamic continuous encoding method involves determining the maximum and minimum encoding bit lengths, and then determining the encoding bit length for each sub-region under test based on these two factors and the complexity of each sub-region. In other words, the encoding bit length of the sub-region under test is dynamically mapped according to its complexity, and the mapping rule is as follows:

[0058] in, The maximum number of bits allowed by the system. The minimum number of bits for encoding is as described above. These are adjustable parameters used to control the rate at which the number of bits in the encoding changes with complexity. Optimal sampling configurations for regions with different texture richness can be achieved through dynamic continuous encoding.

[0059] Step S3015: Determine the corresponding encoding offset based on the encoding bit length and region index of each sub-region to be tested.

[0060] Specifically, in this embodiment of the invention, the encoding offset of the sub-region to be tested is set. for:

[0061] in, The region index for the sub-region to be tested. The number of bits used to encode the sub-region under test. This represents the maximum upper limit of the local encoded value for this sub-region, such as the number of bits used in encoding. hour, The local encoding value ranges from 0 to 255. The encoding offset calculated using the above formula ensures that the offset value of each sub-region is an integer multiple of its local encoding maximum value. The global encoding value generated by superimposing the local encoding values ​​will fall within non-overlapping numerical ranges, such as the index. , The sub-region has a global encoding range of 256~511; index , The sub-regions have a global encoding range of 512 to 767, which can fundamentally avoid encoding value conflicts between different sub-regions and provide reliable encoding identifiers for subsequent global splicing.

[0062] Step S302: Structured light image sequences are sequentially projected onto the corresponding sub-regions under test according to different encoding bit depths, and multi-frame binocular views corresponding to each sub-region under test are obtained. For details, please refer to [link to details]. Figure 2 Step S202 of the illustrated embodiment will not be described again here.

[0063] Step S303: For each sub-region under test, determine the corresponding local coding value based on the multi-frame stereo view of the sub-region under test, and determine the global coding value of the sub-region under test based on the local coding value and the coding offset. For details, please refer to [link to relevant documentation]. Figure 2 Step S203 of the illustrated embodiment will not be described again here.

[0064] Step S304: Based on the global encoding values, perform global stitching on each sub-region to be tested to obtain a panoramic 3D point cloud of the region to be tested. For details, please refer to [link to relevant documentation]. Figure 2 Step S204 of the illustrated embodiment will not be described again here.

[0065] The binocular structured light 3D reconstruction method provided by this invention determines the complexity of each sub-region in the test area, and determines the encoding bit depth and encoding offset of each sub-region based on the complexity. Then, structured light image sequences are projected onto each sub-region sequentially according to different encoding bit depths, acquiring corresponding multi-frame binocular views. A global encoding value is determined based on the multi-frame binocular views and encoding offset of each sub-region. Based on the global encoding value, each sub-region is globally stitched together to obtain a panoramic 3D point cloud of the test area. This invention assigns appropriate encoding bit depths to different sub-regions based on complexity determination, achieving flexible encoding of each sub-region. This not only significantly reduces the amount of encoded data in simple regions and shortens the structured light projection and image acquisition time, but also achieves a balance between accuracy and efficiency. Simultaneously, based on flexible encoding, a unique encoding offset is assigned to each sub-region. The global encoding value can be determined by the local encoding value and the encoding offset, naturally distinguishing the encoding values ​​of different sub-regions and effectively avoiding partition encoding conflicts. Finally, through flexible and precisely correlated regional encoding, the speed and accuracy of overall scene 3D reconstruction are improved.

[0066] This embodiment provides a binocular structured light 3D reconstruction method, which can be used in the aforementioned binocular structured light 3D reconstruction system. Figure 4 This is a flowchart of a binocular structured light 3D reconstruction method according to an embodiment of the present invention, such as... Figure 4 As shown, the process includes the following steps: Step S401: Divide the region to be tested into multiple sub-regions, determine the complexity of each sub-region, and determine the number of bits and the encoding offset of each sub-region based on the complexity. For details, please refer to [link to relevant documentation]. Figure 3 Step S301 of the illustrated embodiment will not be described again here.

[0067] Step S402: Structured light image sequences are sequentially projected onto the corresponding sub-regions under test according to different encoding bit depths, and multi-frame binocular views corresponding to each sub-region under test are obtained. For details, please refer to [link to details]. Figure 3 Step S302 of the illustrated embodiment will not be described again here.

[0068] Step S403: For each sub-region to be tested, determine the corresponding local coding value based on the multi-frame binocular view of the sub-region to be tested, and determine the global coding value of the sub-region to be tested based on the local coding value and the coding offset.

[0069] Specifically, step S403 includes: Step S4031: Calculate the first difference image between the left view and the reference image in each frame of the binocular view, and calculate the second difference image between the right view and the reference image in the binocular view.

[0070] Specifically, in this embodiment of the invention, each sub-region to be tested is decoded according to normal encoding logic. When the projector projects the corresponding number of encoded bits onto any sub-region to be tested... After obtaining the structured light image sequence, acquire multiple frames of binocular views of the sub-region under test, including left and right views.

[0071] For the left or right view, to eliminate interference from ambient light and uneven surface reflectivity of the object, a difference image is calculated. Let... For the first The pixel intensity of the left or right view during frame structured light projection. As the reference image (the image without structured light or with a full white field projected), then the difference image The calculation is as follows:

[0072] in, These are the coordinates of each pixel in the image. Subsequent decoding and phase calculations are based on this difference image. This process significantly improves the signal-to-noise ratio.

[0073] Step S4032: Based on the multiple first difference images corresponding to the multiple frames of binocular views, calculate the first sine component and the first cosine component of each first pixel in the left view, and based on the multiple second difference images corresponding to the multiple frames of binocular views, calculate the second sine component and the second cosine component of each second pixel in the right view.

[0074] Specifically, in this embodiment of the invention, for the left view or the right view, the projection is assumed to be... The step phase-shifting method utilizes the difference image. Calculate the sinusoidal components Sum and cosine components :

[0075]

[0076] in, This indicates the number of phase-shifted images acquired.

[0077] Step S4033: For any first pixel in the left view, calculate the first wrapping phase of the first pixel based on the first sine component and the first cosine component, and determine the first absolute phase of the first pixel based on the first wrapping phase.

[0078] Specifically, in this embodiment of the invention, for the left or right view, the arctangent function is used to calculate the wrapped phase. :

[0079] in, The range of values ​​is .

[0080] Step S4034: For any second pixel in the right view, calculate the second wrapping phase of the second pixel based on the second sine component and the second cosine component, and determine the second absolute phase of the second pixel based on the second wrapping phase.

[0081] Specifically, in this embodiment of the invention, for the left or right view, the wrapped phase is unwrapped by combining the order information provided by Gray code or multi-frequency heterodyne method to obtain a continuous absolute phase. Φ(u, v) The unfolding process is a conventional technique in this field and will not be described in detail here.

[0082] Step S4035: Decode based on the first absolute phase or the second absolute phase to obtain the local coding value of the sub-region to be tested, and superimpose the local coding value with the coding offset to obtain the global coding value.

[0083] Specifically, in this embodiment of the invention, when decoding is performed based on the absolute phase, the local encoded value of the sub-region to be tested is obtained. Subsequently, to achieve multi-region stitching, it is necessary to combine the region index of the sub-region to be tested. i Introduce offset This will enable the local encoded value With encoding offset By superimposing the values, we obtain the global encoded value. The formula is shown below:

[0084] Step S404: Based on the global encoding value, perform global stitching on each sub-region to be tested to obtain a panoramic 3D point cloud of the region to be tested.

[0085] Specifically, step S404 includes: Step S4041: In each sub-region to be tested, determine the same phase pixels in the left and right views based on the first and second absolute phases, and determine the parallax of the binocular imaging device based on the same phase pixels.

[0086] Specifically, in this embodiment of the invention, the calculated absolute phase is used Φ ( u , v This establishes a sub-pixel-level correspondence between the left and right camera images. In a stereo system with epipolar correction, for any pixel in the left view... Find pixels with the same phase on the epipolar line in the right view. Calculate parallax .

[0087] Step S4042: Based on the focal length, baseline distance, principal point coordinates, and parallax of the binocular imaging device, the two-dimensional planar coordinates of the first pixel in the left view are converted into local point cloud coordinates.

[0088] Specifically, in this embodiment of the invention, based on the principle of binocular triangulation and the focal length determined during camera calibration, Baseline distance and principal point coordinates (The point where the camera's optical axis falls on the image sensor), reconstructing the three-dimensional coordinates of each pixel in the camera coordinate system of the current sub-region to be measured. :

[0089]

[0090]

[0091] This yields the current sub-region. i Local high-precision point cloud .

[0092] Step S4043: According to the preset division method and scanning order of the sub-region to be tested, determine the overlapping area of ​​each pair of adjacent sub-regions to be tested, and perform local stitching of each pair of adjacent sub-regions to be tested according to the overlapping area, so as to convert the local point cloud coordinates in the sub-region to be tested into global point cloud coordinates.

[0093] Specifically, in this embodiment of the invention, automatic registration can be achieved without external markers or sensors by utilizing the parallax consistency constraint of adjacent test sub-regions in the overlapping field of view.

[0094] In some optional implementations, step S4043 above includes: Step a1: Based on the overlapping region, obtain the first overlapping point cloud fragment of the first sub-region under test and the second overlapping point cloud fragment of the second sub-region under test in each pair of adjacent sub-regions under test.

[0095] Step a2: For each pair of adjacent sub-regions to be tested, the first sub-region to be tested and the second sub-region to be tested are initially aligned based on the first global encoding value of each first spatial point in the first overlapping point cloud region and the second global encoding value of each second spatial point in the second overlapping point cloud segment.

[0096] Step a3: After the initial alignment, the first spatial points and the second spatial points are re-aligned based on the iterative nearest point algorithm, and the three-dimensional rigid body transformation matrix of the iterative nearest point algorithm is adjusted during the re-alignment process.

[0097] Step a4: Construct an error function to calculate the adjusted reprojection error. Once the reprojection error reaches its minimum value, determine the optimal transformation matrix.

[0098] Step a5: Based on the optimal transformation matrix, transform the local point cloud coordinates of each of the first and second sub-regions to be tested into global point cloud coordinates.

[0099] Specifically, in this embodiment of the invention, for adjacent sub-regions i and j Extract overlapping parts (e.g., regions) of the images according to the scanning order. i 10% on the right and the region j (10% from the left side) to obtain the corresponding 3D point cloud fragment. and Coarse registration is performed using the continuity between the globally encoded values ​​determined above, that is, checking the globally encoded values ​​of overlapping regions to ensure that the regions are properly matched. i Maximum encoding and region j Minimum encoding alignment. If the scan path is known, the translation vector can be estimated. For example, when using a horizontal sliding rail, the translation vector For along x The axis movement is approximately 500mm, which is only an example and is not a limitation.

[0100] In some alternative implementations, after initial alignment, a modified Iterative Closest Point (ICP) algorithm is used for fine registration. For a spatial point within the overlapping region... If the coordinate systems of the two sub-regions to be measured are not correctly aligned, then the points will be... The three-dimensional rigid body transformation matrix to be optimized After transformation, the predicted coordinates obtained by backprojection onto the camera plane will deviate from the actual measured image coordinates (i.e., disparity prediction residuals). Therefore, this embodiment of the invention constructs an error function. Optimize:

[0101] in, These are the pixel coordinates measured in the original image; The matrix represents the 3D rigid body transformation (rotation and translation) to be optimized. These are 3D points in a local point cloud. These are weighting coefficients. It is the error term in the ICP algorithm. This is the projection function for the camera.

[0102] Adjust the 3D rigid body transformation matrix and calculate the reprojection error based on the error function. Determine the optimal transformation matrix when the reprojection error reaches its minimum value. At this point, the reprojection error of the stitched point cloud in the binocular view is minimized, ensuring geometric consistency.

[0103] Therefore, by minimizing the aforementioned error function, the embodiments of the present invention calculate the precise pose of each sub-region under test relative to the global coordinate system. All local point clouds Transform to the global coordinate system.

[0104] Step S4044: Perform global stitching based on the global point cloud coordinates of all sub-regions to be tested to obtain a panoramic 3D point cloud of the region to be tested.

[0105] Specifically, in this embodiment of the invention, the controller aggregates the 3D point cloud coordinate data of all sub-regions to be measured that have been converted to the global coordinate system. Based on the preset position arrangement of each sub-region in the space to be measured and the spatial relationship between adjacent sub-regions, the scattered local point clouds are sequentially connected and integrated according to their actual spatial positions. At the same time, the spliced ​​point cloud is uniformly regulated to eliminate splicing gaps and data redundancy between regions, and finally a panoramic 3D point cloud covering the entire area to be measured with complete geometric shape and unified coordinates is formed.

[0106] In some optional implementations, during the integration process, for the overlapping areas between adjacent test sub-regions, the first global point cloud coordinates of spatial points within the first overlapping point cloud segment and the second global point cloud coordinates of spatial points within the second overlapping point cloud segment in each pair of adjacent test sub-regions are weighted and summed according to a preset weighting coefficient to obtain the global point cloud coordinates of spatial points within the overlapping area. The weighted fusion formula is as follows:

[0107] in, The first global point cloud coordinates are the spatial coordinates of the first overlapping point cloud segment within the adjacent sub-regions to be tested. For the corresponding weighting coefficients, The second global point cloud coordinates are the spatial coordinates of the spatial points within the second overlapping point cloud segment in the adjacent sub-regions to be tested. These are the corresponding weighting coefficients.

[0108] The binocular structured light 3D reconstruction method provided by this invention determines the complexity of each sub-region in the test area, and determines the encoding bit depth and encoding offset of each sub-region based on the complexity. Then, structured light image sequences are projected onto each sub-region sequentially according to different encoding bit depths, acquiring corresponding multi-frame binocular views. A global encoding value is determined based on the multi-frame binocular views and encoding offset of each sub-region. Based on the global encoding value, each sub-region is globally stitched together to obtain a panoramic 3D point cloud of the test area. This invention assigns appropriate encoding bit depths to different sub-regions based on complexity determination, achieving flexible encoding of each sub-region. This not only significantly reduces the amount of encoded data in simple regions and shortens the structured light projection and image acquisition time, but also achieves a balance between accuracy and efficiency. Simultaneously, based on flexible encoding, a unique encoding offset is assigned to each sub-region. The global encoding value can be determined by the local encoding value and the encoding offset, naturally distinguishing the encoding values ​​of different sub-regions and effectively avoiding partition encoding conflicts. Finally, through flexible and precisely correlated regional encoding, the speed and accuracy of overall scene 3D reconstruction are improved.

[0109] This embodiment also provides a binocular structured light 3D reconstruction device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0110] This embodiment provides a binocular structured light 3D reconstruction device, such as... Figure 5 As shown, it includes: The encoding bit determination module 501 is used to divide the region to be tested into multiple sub-regions to be tested, determine the complexity of each sub-region to be tested, and determine the encoding bit and encoding offset of the sub-region to be tested based on the complexity.

[0111] The coded pattern projection module 502 is used to project a sequence of structured light images onto the corresponding sub-regions to be tested in sequence according to different coded bit widths, and to obtain multi-frame binocular views corresponding to each sub-region to be tested.

[0112] The global encoding determination module 503 is used to determine the corresponding local encoding value for each sub-region under test based on the multi-frame stereo view of the sub-region under test, and to determine the global encoding value of the sub-region under test based on the local encoding value and the encoding offset.

[0113] The 3D point cloud reconstruction module 504 is used to globally stitch together each sub-region under test based on the global encoding value to obtain a panoramic 3D point cloud of the region under test.

[0114] In some optional implementations, the encoding bit determination module 501 includes: The pre-scanning unit is used to control the binocular imaging device to scan the area to be tested according to a first resolution to obtain a first image.

[0115] The region division unit is used to divide the first image into multiple second images according to a preset division method, thereby obtaining multiple sub-regions to be tested, and determining the region index of each sub-region to be tested.

[0116] The complexity calculation unit is used to score the complexity of each second image and obtain the complexity of each second image.

[0117] The encoding bit determination unit is used to determine the encoding bit of the sub-region to be tested based on the complexity.

[0118] The offset determination unit is used to determine the corresponding encoding offset based on the encoding bit length and region index of each sub-region to be tested.

[0119] In some optional implementations, the encoding bit depth determination unit includes: The first determining sub-unit is used to determine whether each sub-region to be tested is a complex region or a simple region based on its complexity and a preset threshold. If it is a complex region, the first encoding bit length is determined to be the first encoding bit length; if it is a simple region, the second encoding bit length is determined to be the second encoding bit length. The first encoding bit length is greater than the second encoding bit length.

[0120] The threshold determination subunit is used to set a preset threshold as a fixed threshold; or, based on the complexity distribution of all sub-regions to be tested, determine the preset threshold according to a preset ratio; or, construct a comprehensive loss function based on minimizing the expected reconstruction error and sampling time cost, and determine the preset threshold based on the comprehensive loss function.

[0121] The second determining subunit is used to determine the maximum and minimum coding bit lengths, and to determine the coding bit length of each sub-region under test based on the maximum coding bit length, the minimum coding bit length, and the complexity of each sub-region under test.

[0122] In some optional implementations, the global encoding determination module 503 includes: The difference image calculation unit is used to calculate the first difference image between the left view and the reference image in each frame of the binocular view, and to calculate the second difference image between the right view and the reference image in the binocular view.

[0123] The sine and cosine component calculation unit is used to calculate the first sine component and the first cosine component of each first pixel in the left view based on multiple first difference images corresponding to multiple frames of binocular views, and to calculate the second sine component and the second cosine component of each second pixel in the right view based on multiple second difference images corresponding to multiple frames of binocular views.

[0124] The first absolute phase determination unit is used to calculate the first wrapping phase of any first pixel in the left view based on the first sine component and the first cosine component, and to determine the first absolute phase of the first pixel based on the first wrapping phase.

[0125] The second absolute phase determination unit is used to calculate the second wrapping phase of any second pixel in the right view based on the second sine component and the second cosine component, and to determine the second absolute phase of the second pixel based on the second wrapping phase.

[0126] The global coding value determination unit is used to decode based on the first absolute phase or the second absolute phase to obtain the local coding value of the sub-region to be tested, and to superimpose the local coding value with the coding offset to obtain the global coding value.

[0127] In some alternative implementations, the 3D point cloud reconstruction module 504 includes: The disparity determination unit is used to determine the same phase pixels in the left and right views based on the first and second absolute phases within each sub-region to be measured, and to determine the disparity of the binocular imaging device based on the same phase pixels.

[0128] The local point cloud coordinate determination unit is used to convert the two-dimensional planar coordinates of the first pixel in the left view into local point cloud coordinates based on the focal length, baseline distance, principal point coordinates, and parallax of the binocular imaging device.

[0129] The global point cloud coordinate determination unit is used to determine the overlapping area of ​​each pair of adjacent sub-regions according to the preset division method and scanning order of the sub-region to be measured, and to perform local stitching of each pair of adjacent sub-regions according to the overlapping area, so as to convert the local point cloud coordinates in the sub-region to be measured into global point cloud coordinates.

[0130] The point cloud stitching unit is used to perform global stitching based on the global point cloud coordinates of all sub-regions to be tested, so as to obtain a panoramic 3D point cloud of the region to be tested.

[0131] In some optional implementations, the global point cloud coordinate determination unit includes: The overlapping point cloud determination sub-unit is used to obtain, based on the overlapping region, the first overlapping point cloud segment of the first sub-region under test and the second overlapping point cloud segment of the second sub-region under test in each pair of adjacent sub-regions under test.

[0132] The initial alignment sub-unit is used to initially align the first and second sub-regions under test for each pair of adjacent sub-regions under test based on the first global encoding value of each first spatial point in the first overlapping point cloud region and the second global encoding value of each second spatial point in the second overlapping point cloud segment.

[0133] The sub-units are re-aligned after the initial alignment. Based on the iterative nearest point algorithm, each first spatial point and each second spatial point are re-aligned, and the three-dimensional rigid body transformation matrix of the iterative nearest point algorithm is adjusted during the re-alignment process.

[0134] The optimal matrix determines the sub-unit, which is used to construct the error function to calculate the adjusted reprojection error. When the reprojection error reaches its minimum value, the optimal transformation matrix is ​​determined.

[0135] The point cloud coordinate transformation sub-unit is used to transform the local point cloud coordinates of each of the first and second sub-regions to be tested into global point cloud coordinates according to the optimal transformation matrix.

[0136] In some alternative implementations, the point cloud stitching unit includes: The point cloud fusion subunit is used to perform a weighted summation of the first global point cloud coordinates of spatial points in the first overlapping point cloud segment and the second global point cloud coordinates of spatial points in the second overlapping point cloud segment in each pair of adjacent test sub-regions according to a preset weight coefficient, so as to obtain the global point cloud coordinates of spatial points in the overlapping region.

[0137] The binocular structured light 3D reconstruction device provided in this embodiment of the invention can execute the binocular structured light 3D reconstruction method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.

[0138] Figure 6 This is a schematic diagram of the controller in a binocular structured light 3D reconstruction system provided in an embodiment of the present invention.

[0139] The following is a detailed reference. Figure 6 The diagram illustrates a structural schematic suitable for implementing a controller in an embodiment of the present invention. The controller may include a processor (e.g., a central processing unit, graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from memory 608 into random access memory (RAM) 603. The RAM 603 also stores various programs and data required for controller operation. The processor 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0140] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows the controller to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 A controller with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown, and may alternatively implement or have more or fewer devices.

[0141] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a memory 608, or installed from a ROM 602. When the computer program is executed by the processor 601, it performs the functions defined in the binocular structured light 3D reconstruction method of the embodiments of the present invention.

[0142] Figure 6 The controller shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0143] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the binocular structured light 3D reconstruction method shown in the above embodiments is implemented.

[0144] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0145] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A binocular structured light three-dimensional reconstruction method, characterized in that, The method includes: The region to be tested is divided into multiple sub-regions, the complexity of each sub-region is determined, and the encoding bit length and encoding offset of the sub-region are determined based on the complexity. Structured light image sequences are sequentially projected onto the corresponding test sub-regions according to different encoding bit lengths, and multi-frame binocular views corresponding to each test sub-region are obtained; For each sub-region under test, the corresponding local coding value is determined based on the multi-frame binocular view of the sub-region under test, and the global coding value of the sub-region under test is determined based on the local coding value and the coding offset. Based on the global encoding value, each of the sub-regions to be tested is globally stitched together to obtain a panoramic 3D point cloud of the region to be tested.

2. The method according to claim 1, characterized in that, The process of dividing the region to be tested into multiple sub-regions, determining the complexity of each sub-region, and determining the encoding bit length and encoding offset of the sub-region based on the complexity includes: The binocular imaging device is controlled to scan the area to be tested according to a first resolution to obtain a first image; The first image is divided into multiple second images according to a preset division method to obtain multiple sub-regions to be tested, and the region index of each sub-region to be tested is determined. The complexity of each second image is obtained by scoring its complexity. The number of bits used to encode the sub-region to be tested is determined based on the complexity. The corresponding encoding offset is determined based on the encoding bit length and the region index of each of the sub-regions to be tested.

3. The method according to claim 2, characterized in that, Determining the number of bits in the encoding of the sub-region to be tested based on the complexity includes: For each sub-region to be tested, the sub-region to be tested is determined to be a complex region or a simple region based on the complexity and a preset threshold. If it is a complex region, the encoding bit length is determined to be the first encoding bit length. If it is a simple region, the encoding bit length is determined to be the second encoding bit length. The first encoding bit length is greater than the second encoding bit length. Alternatively, determine the maximum and minimum number of encoding bits, and determine the number of encoding bits for each of the sub-regions under test based on the maximum number of encoding bits, the minimum number of encoding bits, and the complexity of each sub-region under test.

4. The method according to claim 3, characterized in that, The process of determining the preset threshold includes: Set the preset threshold to a fixed threshold; Alternatively, the preset threshold can be determined according to a preset ratio based on the complexity distribution of all the sub-regions to be tested. Alternatively, a comprehensive loss function can be constructed based on minimizing the expected reconstruction error and the sampling time cost, and the preset threshold can be determined according to the comprehensive loss function.

5. The method according to claim 2, characterized in that, The binocular view is obtained by scanning the sub-region under test by a binocular imaging device at a second resolution, wherein the second resolution is higher than the first resolution; The step of determining the corresponding local coding value based on the multi-frame stereo view of the sub-region under test, and determining the global coding value of the sub-region under test based on the local coding value and the coding offset, includes: Calculate the first difference image between the left view and the reference image in each frame of the binocular view, and calculate the second difference image between the right view and the reference image in the binocular view; Based on multiple first difference images corresponding to multiple frames of binocular views, calculate the first sine component and the first cosine component of each first pixel in the left view, and based on multiple second difference images corresponding to multiple frames of binocular views, calculate the second sine component and the second cosine component of each second pixel in the right view. For any first pixel in the left view, calculate the first wrapping phase of the first pixel based on the first sine component and the first cosine component, and determine the first absolute phase of the first pixel based on the first wrapping phase; For any second pixel in the right view, calculate the second wrapping phase of the second pixel based on the second sine component and the second cosine component, and determine the second absolute phase of the second pixel based on the second wrapping phase; Decoding is performed based on the first absolute phase or the second absolute phase to obtain the local encoded value of the sub-region under test, and the local encoded value is superimposed with the encoded offset to obtain the global encoded value.

6. The method according to claim 5, characterized in that, The step of globally stitching together each of the sub-regions to be tested based on the global encoding value to obtain a panoramic 3D point cloud of the region to be tested includes: Within each of the sub-regions to be tested, pixels with the same phase in the left and right views are determined based on the first and second absolute phases, and the parallax of the binocular imaging device is determined based on the pixels with the same phase. Based on the focal length, baseline distance, principal point coordinates, and parallax of the binocular imaging device, the two-dimensional planar coordinates of the first pixel in the left view are transformed into local point cloud coordinates; According to the preset division method and scanning order of the sub-region to be tested, the overlapping area of ​​each pair of adjacent sub-regions to be tested is determined, and each pair of adjacent sub-regions to be tested is locally stitched together according to the overlapping area, so as to convert the local point cloud coordinates in the sub-region to be tested into global point cloud coordinates. Global point cloud coordinates of all the sub-regions to be tested are used to stitch together the data to obtain a panoramic 3D point cloud of the region to be tested.

7. The method according to claim 6, characterized in that, The step of locally stitching together each pair of adjacent sub-regions to be tested according to the overlapping area includes: Based on the overlapping region, obtain the first overlapping point cloud fragment of the first sub-region under test and the second overlapping point cloud fragment of the second sub-region under test in each pair of adjacent sub-regions under test. For each pair of adjacent sub-regions to be tested, the first sub-region to be tested and the second sub-region to be tested are initially aligned according to the first global encoding value of each first spatial point in the first overlapping point cloud region and the second global encoding value of each second spatial point in the second overlapping point cloud segment. After the initial alignment, the first spatial points and the second spatial points are re-aligned based on the iterative nearest point algorithm, and the three-dimensional rigid body transformation matrix of the iterative nearest point algorithm is adjusted during the re-alignment process. An error function is constructed to calculate the adjusted reprojection error. Once the reprojection error reaches its minimum value, the optimal transformation matrix is ​​determined. The local point cloud coordinates within each of the first and second sub-regions to be tested are transformed into global point cloud coordinates based on the optimal transformation matrix.

8. The method according to claim 7, characterized in that, When performing global stitching based on the global point cloud coordinates of all the sub-regions to be tested, the method further includes: According to a preset weighting coefficient, the first global point cloud coordinates of the spatial points in the first overlapping point cloud segment and the second global point cloud coordinates of the spatial points in the second overlapping point cloud segment of each pair of adjacent test sub-regions are weighted and summed to obtain the global point cloud coordinates of the spatial points in the overlapping region.

9. A binocular structured light three-dimensional reconstruction device, characterized in that, The device includes: The encoding bit depth determination module is used to divide the region to be tested into multiple sub-regions to be tested, determine the complexity of each sub-region to be tested, and determine the encoding bit depth and encoding offset of the sub-region to be tested based on the complexity. The coded pattern projection module is used to sequentially project a sequence of structured light images onto the corresponding sub-regions to be tested according to different coded bit widths, and to obtain multi-frame binocular views corresponding to each sub-region to be tested. The global encoding determination module is used to determine the corresponding local encoding value for each sub-region under test based on the multi-frame binocular view of the sub-region under test, and to determine the global encoding value of the sub-region under test based on the local encoding value and the encoding offset. The 3D point cloud reconstruction module is used to globally stitch together each of the sub-regions to be tested based on the global encoding value to obtain a panoramic 3D point cloud of the region to be tested.

10. A binocular structured light 3D reconstruction system, characterized in that, include: The system includes a controller, a structured light projection device, and a binocular imaging device, wherein the structured light projection device and the binocular imaging device are both connected to the controller. The controller includes a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the binocular structured light three-dimensional reconstruction method according to any one of claims 1 to 8.