Structured light single-frame reconstruction method based on layered stereo matching

By treating the structured light system as a binocular system and constructing a simulation dataset, and using random speckle patterns and feature extraction networks for depth prediction, the error accumulation problem in single-frame reconstruction is solved, and high-precision and strong generalization 3D reconstruction effects are achieved.

CN120765835APending Publication Date: 2025-10-10SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510729478.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing structured light 3D reconstruction technology has a trade-off between accuracy and speed in single-frame reconstruction, and the generalization and accuracy of deep learning models are insufficient, especially in dynamic scenes.

Method used

The structured light active measurement system is regarded as a binocular passive measurement system. The measurement system's own parameters are used as constraints to build a model. A simulation dataset is constructed for training, and random speckle patterns and feature extraction networks are used for depth prediction.

Benefits of technology

It achieves high-precision 3D reconstruction in a single-frame image, has stronger dynamics and generalization performance, avoids error accumulation, and is suitable for dynamic scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765835A_ABST
    Figure CN120765835A_ABST
Patent Text Reader

Abstract

The invention discloses a structured light single-frame reconstruction method based on layered stereo matching. A structured light active measurement system is regarded as a binocular passive measurement system, random speckles are adopted as projection patterns, a single-frame depth estimation model is provided based on the stereo matching principle, and the single-frame depth estimation model mainly comprises a depth initialization module, a depth updating module and a depth refinement module. Compared with a traditional structured light three-dimensional reconstruction algorithm, the method has the advantages that three-dimensional reconstruction of an object can be completed only through a single-frame image, and the method has higher dynamics; compared with other deep learning based on a phase shift method, the method does not need to solve three-dimensional data through intermediate variables such as a wrapped phase and an absolute phase, and avoids the influence of an accumulation error on a final result. In addition, system calibration parameters are used as a priori construction model, and a simulation data set is adopted to assist training, so that the model has higher generalization performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of machine vision and three-dimensional reconstruction, and in particular to a structured light single-frame reconstruction method based on stereo matching. Background Art

[0002] Structured light 3D reconstruction is an important optical active measurement technology. It offers the advantages of non-contact, high precision, and immunity to surface texture. It also exhibits a certain degree of immunity to ambient light interference. With the rapid development of informatization and digitalization, 3D sensing technology has been widely applied in many fields, such as robotic positioning and grasping, facial recognition, product dimensional measurement, and cultural relic restoration.

[0003] Traditional monocular structured light 3D reconstruction technology is primarily based on phase shifting, which offers the advantages of high precision and robustness. Using more phase shift steps generally improves both accuracy and robustness. However, since a single measurement requires simultaneous projection and capture of multiple images, this increases measurement time and limits its application to static scenes. Many studies have had to balance accuracy and speed for different application scenarios.

[0004] In recent years, deep learning has brought new developments to structured light reconstruction technology, with many studies attempting to leverage deep learning techniques to achieve single-frame reconstruction. Most existing studies, still based on traditional phase-shifting methods, leverage deep learning to address single-frame wrapped phase resolution, phase unwrapping, and single-frame absolute phase resolution. Because these intermediate quantities inherently have certain errors, coupled with the influence of system calibration errors, the various accumulated errors can affect the final three-dimensional information. While some end-to-end research exists, most of these approaches are divorced from the measurement system itself, treating the model as a "black box" and using only camera-captured images and depth maps for training. This often results in poor generalization. Furthermore, due to the difficulty of acquiring real-world data, many studies have been forced to use a small number of data samples. Furthermore, the inherent deviation between real-world data acquisition and the true value further negatively impacts model training. Summary of the Invention

[0005] The purpose of the present invention is to overcome the above-mentioned shortcomings and deficiencies of the prior art and to provide a structured light single-frame reconstruction method based on hierarchical stereo matching.

[0006] The present invention regards the structured light active measurement system as a binocular passive measurement system, uses the measurement system's own parameters as constraints to build a model, and simultaneously constructs a simulation data set for model training.

[0007] The present invention is achieved through the following technical solutions:

[0008] S1: Generate a projection pattern, using random speckle as the optical projection pattern, projecting the pattern through the optical machine, and synchronously capturing the distorted image by the camera.

[0009] S2: Build a data set, construct simulation data with different system parameters, different lighting conditions, and different scenes, and combine it with the data collected from actual scenes for training.

[0010] S3: Data preprocessing: To make the two input images have a larger overlapping area, the projected pattern cropping coordinates are calculated according to the system calibration parameters and depth range, and the cropped pattern is upsampled.

[0011] S4: Feature extraction, extracting features of the camera captured image and projected pattern.

[0012] S5: Depth initialization, generating an initial depth map for each layer based on the assumed depth range and feature map.

[0013] S6: Iteratively update the predicted depth based on the initial depth of this layer, the predicted depth of the previous layer, and the feature map.

[0014] S7: Depth refinement, using camera captured patterns to further refine the predicted depth.

[0015] Furthermore, the step of generating the random speckle pattern in step S1 is as follows:

[0016] The image is divided into equally spaced grids. The grid size and speckle size determine the speckle density. Each speckle spot is randomly offset in the xy direction with the grid point as the center, thereby generating a random speckle distribution. The brightness value of the point marked as speckle is then set to 255.

[0017] Furthermore, step S2 screens suitable public models and uses the blender platform to construct simulation data samples. By adjusting parameters such as the light source, object material, camera focal length, and optical and mechanical light intensity around the scene, multiple different scenes that conform to real physics and are close to reality are rendered. At the same time, a virtual calibration plate is used in combination with the phase shift method to calibrate the system, calculate the system calibration parameters, and compare them with the set camera and optical and mechanical parameters to verify the accuracy of the simulation system.

[0018] Furthermore, step S3 crops and upsamples the projected pattern so that there is a larger overlap area between the two views. The transformation matrix from camera coordinates to optomechanical coordinates is:

[0019] proj = p_proj × inv(c_proj)

[0020] Among them, p_proj is the optical machine projection matrix, c_proj is the camera projection matrix, inv is the inverse operation, and × is matrix multiplication.

[0021] The homogeneous coordinates of the clipping position are calculated by the following formula:

[0022] [x,y,z] homo =(proj[:3,:3]×[0,0,1] T )*max_depth+proj[:3,3:4]

[0023] Where * is element-wise multiplication.

[0024] Clipping coordinates (x crop ,y crop )for:

[0025] [x crop ,y crop ,1]=[x,y,z] homo / [x,y,z] homo [2:3,:]

[0026] After cropping the image, bicubic interpolation is used to upsample the projected pattern to the same resolution as the camera image. The final input data consists of the processed projected pattern, camera image, system parameters, and cropping coordinates.

[0027] Furthermore, step S4 uses UNet as a feature extraction network, whose input is an image pair consisting of a camera image and a projection pattern, and whose output is feature maps of three different scales.

[0028] Furthermore, using the feature map obtained in step S4, step S5 obtains the rough initial depth at different scales according to the feature maps of different scales. The initialization steps are as follows:

[0029] cost=abs[warp(fea proj )-fea cam ]

[0030] depth init =sum(softmax(Conv3D(cost))*depth_range)

[0031] Among them, abs means taking the absolute value, sum means sum, cost means matching cost, Conv3D means 3D convolutional network, and depth_range means the assumed depth range.

[0032] In step S5, the projected pattern features are distorted according to the given depth range, the difference between the distorted map and the camera features is used to measure the matching cost, 3D convolution is used to aggregate the costs to obtain the confidence of each depth, the softmax operation is performed to obtain the weight of each depth value, and finally the weighted sum is performed to obtain the initial depth.

[0033] Furthermore, in step S6, the depth of this layer is initialized using Depth init And the predicted depth from the previous layer pre Iteratively update depth update , the final depth is:

[0034]

[0035] Indicates the depth value updated in the i-th iteration, feature is the feature map output by the feature extraction network, and initially Using Depth pre Up-sampled.

[0036] Furthermore, step S7 further refines the final depth map Depth obtained in step S6, using the img of the image captured by the camera cam The contextual information further refines the predicted depth.

[0037] Depth refine =Refine(img cam ,Depth).

[0038] Compared with the prior art, the present invention has the following advantages and effects:

[0039] Compared with traditional structured light 3D reconstruction algorithms, the present invention only requires a single frame of image to complete 3D reconstruction of an object, and has stronger dynamics.

[0040] Compared with other deep learning based on phase shift method, this invention regards the measurement system as a binocular passive measurement system and adopts an end-to-end model prediction model. It does not need to solve three-dimensional data through intermediate variables such as wrapped phase and absolute phase, thus avoiding the influence of cumulative error on the final result.

[0041] Compared with other end-to-end models that only use camera-captured images and depth values ​​for training, the present invention builds a model based on system calibration parameters as a priori and uses simulation data sets to assist in training, so the model has stronger generalization performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is the random speckle pattern projected by the present invention.

[0043] Figure 2 A scene graph built for rendering simulation data.

[0044] Figure 3 This is a rendering from the camera's perspective in the simulation environment.

[0045] Figure 4Set up output nodes for rendering images, depth images, and mask images in the simulation environment.

[0046] Figure 5 This is an example diagram of a sample in the data set of the present invention.

[0047] Figure 6 This is a diagram of the single-frame reconstruction network framework described in the present invention.

[0048] Figure 7 This is a schematic diagram of the initialization module of the present invention.

[0049] Figure 8 Schematic diagram of the update module in the network. DETAILED DESCRIPTION

[0050] The present invention will be further described in detail below with reference to examples and drawings, but the embodiments of the present invention are not limited thereto.

[0051] The present invention proposes a structured light single-frame reconstruction method based on hierarchical stereo matching, the main working process of which is as follows:

[0052] S1: Generate a random speckle pattern of height x width with a grid spacing of L, a speckle diameter of D, and an offset of a:

[0053] Initially, the speckle centers are distributed on a grid with an interval of L. All speckle center points are randomly offset by a length of 0 to a*L / 2 in the x and y directions, and all points within a circle with a diameter of D and the speckle center point as the center are marked as speckles. Finally, the brightness value of all speckle spots in the pattern is set to 255. The pseudo code is as follows:

[0054]

[0055] Attachment Figure 1 This is a random speckle pattern generated using this method.

[0056] S2: The simulation data set is rendered using Blender's Cycles engine. Figure 2 For a scene diagram built, attached Figure 3 Rendering from the camera perspective, attached Figure 4 For the corresponding output node settings, apply anti-aliasing operations to the depth and rendering output by the Cycles rendering engine, generate a mask map based on the depth value, and set the save format of each output node. The camera intrinsic parameters can be directly obtained from the set camera parameters, or they can be calibrated using the calibration method in the real environment for verification. Without considering camera distortion, the camera focal length f and image width width are used. img , the sensor size is width sensor For example, the corresponding internal parameters are:

[0057]

[0058] The method for calculating K(1,1) in the height direction is similar, and the bias parameter is:

[0059] K(0,2)=width img / 2

[0060] K(1,2)=height img / 2

[0061] The real data set is generated by the traditional phase shift method + multi-frequency method, and the phase shift + multi-frequency expansion method is used to assist calibration. Figure 5 is a sample example of the dataset.

[0062] S3: Since the image captured by the camera in the actual environment is part of the projection pattern, in order to make the two input images have a larger overlapping area, the projection pattern is cropped and up-sampled.

[0063] The transformation matrix from camera pixel coordinates to optical machine pixel coordinates is:

[0064] proj = p_proj × inv(c_proj)

[0065] Among them, p_proj is the optical machine projection matrix, c_proj is the camera projection matrix, inv is the inverse operation, and × is matrix multiplication.

[0066] The homogeneous coordinates of the clipping position are calculated by the following formula:

[0067] [x,y,z] homo =(proj[:3,:3]×[0,0,1] T )*max_depth+proj[:3,3:4]

[0068] Where * is element-wise multiplication.

[0069] Clipping coordinates (x crop ,y crop )for:

[0070] [x crop ,y crop ,1]=[x,y,z] homo / [x,y,z] homo [2:3,:]

[0071] S4: As attached Figure 6The network framework diagram of the present invention, wherein the feature extraction module is used to extract various effective features in the image. The present invention uses a three-layer UNet network for feature extraction, which generates three feature maps of different scales. If the resolution of the input image is (height, width), the resolution of the third layer feature map is the highest, then the feature map output by the i-th layer is (32 / 2 i ,height / 2 2-i ,width / 2 2-i ), where the first dimension is the number of channels, and the number of channels of the output feature map is 32, 16, and 8 respectively.

[0072] S5: As attached Figure 7 As shown in Figure 1, to prevent the errors of the previous layer from accumulating to the current layer, each layer contains an independent initial depth. This module distorts the feature map of the random speckle pattern according to the system parameters and the set depth range:

[0073] For a point p on the camera image c , whose pixel coordinate p on the projected pattern p :

[0074] proj = p_proj × inv(c_proj)

[0075] [x,y,z] homo =(proj[:3,:3]×p c )*depth+proj[:3,3:4]

[0076]

[0077] Among them, p_proj is the optical machine projection matrix, c_proj is the camera projection matrix, inv is the inverse operation, and × is matrix multiplication.

[0078] Since the projected pattern is clipped, the actual sampling coordinates are:

[0079] p sample =(p p -crop)*up_scale

[0080] crop is the projected pattern cropping coordinates, and up_scale is the upsampling multiple.

[0081] In the first layer, there is no depth update from the previous layer, so a relatively accurate initial depth needs to be generated for the update module. The depth range depth_range assumed by each layer is interval from coarse to fine, and the number of assumed depths of each layer is set to 192, 96, and 48 respectively. When the depth value is close to the true depth, the distorted feature map has a high degree of similarity with the camera feature map. The matching cost is constructed by the gap between the distorted features and the camera features, and the cost is aggregated using a 3D convolutional network. Finally, softmax is used to generate the corresponding weights at each depth, and each depth is weighted to generate the initial depth. The initialization steps are as follows:

[0082] cost=abs[warp(fea proj )-fea cam ]

[0083] depth init =sum(softmax(Conv3D(cost))*depth_range)

[0084] Among them, abs means taking the absolute value, sum means sum, cost means matching cost, Conv3D means 3D convolutional network, and depth_range means the assumed depth range.

[0085] S6: As attached Figure 8 As shown, step S6 uses the rough initial depth value Depth init , the output feature feature of the feature extraction network, and the depth map Depth predicted by the previous layer pre Iteratively update the depth value.

[0086]

[0087] Denotes the depth value updated in the i-th iteration, feature is the feature map output by the feature extraction network, and Depth is initially pre Bicubic upsampling makes its resolution consistent with the next layer resolution and uses the upsampled depth value as UpdateNet is a deep update network consisting of two connected residual blocks. The update network calculates Depth init The matching cost is calculated as follows:

[0088] cost=warp(fea proj )-fea cam

[0089] The sampling method of the warp operation is as described in S5 above. Finally, the matching cost, depth value, and feature map are used as the input of the residual network, and the output is the updated depth.

[0090] S7: Step S7 further refines the final depth map Depth obtained in step S6, using the img of the image taken by the camera cam The contextual information further refines the predicted depth.

[0091] Depth refine =Refine(img cam ,Depth)

[0092] The final refined depth map is:

[0093] Depth refine =Depth+ConvBlocks(cat([img cam ,Dept h])

[0094] ConvBlocks consists of three 3x3 convolutional blocks and one 1x1 convolutional block, and each convolutional block consists of a convolutional layer and a LeakyReLU activation function.

[0095] The present invention uses three different scales of output depth to calculate the loss. The final loss is composed of the l1 loss weighted by the predicted depth map and the true depth map output by the three stages:

[0096]

[0097] W 0 、W 1 、W 2 They are 0.5, 1.0, and 2.0 respectively.

[0098] The above examples are a preferred implementation of the present invention, but the present invention is not limited to these examples. Any changes, modifications, substitutions, combinations, or simplifications that do not depart from the core spirit and principles of the present invention shall be considered equivalent alternatives and shall also be included in the scope of protection of the present invention.

Claims

1. A structured light single-frame reconstruction method based on hierarchical stereo matching, characterized in that: The following steps are involved: S1: generation of speckle pattern; S2: Build dataset; S3: Data preprocessing: Calculate the projection pattern clipping coordinates based on the system calibration parameters and depth range, so that the projection pattern has a large overlapping area with the camera image; S4: Feature extraction, extracting features of the two input images for subsequent steps; S5: Depth initialization, generating an initial depth map for each layer based on the assumed depth range and feature map; S6: Iteratively update the predicted depth based on the initial depth of the current layer, the predicted depth of the previous layer, and the feature map; S7: Depth refinement, using camera captured patterns to further refine the predicted depth.

2. The structured light single-frame reconstruction method based on hierarchical stereo matching according to claim 1, characterized in that: In step S1, the generation of the speckle pattern specifically refers to dividing the image into equally spaced grids. The grid size and speckle size determine the speckle density. Each speckle spot is centered at a grid point and randomly offset in the xy direction, thereby generating a random speckle distribution. The brightness value of the point marked as speckle is then set to 255.

3. The structured light single-frame reconstruction method based on hierarchical stereo matching according to claim 1, characterized in that: In step S2, constructing a data set specifically refers to constructing simulation data with different system parameters, different lighting conditions, and different scenes, and performing training in combination with actual scene data sets.

4. The structured light single-frame reconstruction method based on hierarchical stereo matching according to claim 1, characterized in that: In step S3, data preprocessing specifically refers to cropping and upsampling the projection pattern so that there is a large overlap area between the two views. The transformation matrix from camera coordinates to optomechanical coordinates is: proj = p_proj × inv(c_proj); Where p_proj is the optical projection matrix, c_proj is the camera projection matrix, inv is the inverse operation, and × is the matrix multiplication; The homogeneous coordinates of the clipping position are calculated by the following formula: [x,y,z] homo =(proj[:3,:3]×[0,0,1] T )*max_depth+proj[:3,3:4]; Where * is element-by-element multiplication; Clipping coordinates (x crop ,y crop )for: [x crop ,y crop ,1]=[x,y,z] homo / [x,y,z] homo [2:3,:]。 5. The structured light single-frame reconstruction method based on hierarchical stereo matching according to claim 1, characterized in that: In step S4, feature extraction specifically refers to using UNet as a feature extraction network, whose input is an image pair consisting of a camera image and a random speckle pattern, and whose output is feature maps of three different scales.

6. The structured light single-frame reconstruction method based on hierarchical stereo matching according to claim 1, characterized in that: In step S5, depth initialization specifically refers to obtaining the initial depth at different scales based on the feature map. First, the feature map of the random speckle pattern is distorted according to the system parameters and the set depth range. The steps are as follows: For a point p on the camera image c , whose pixel coordinates p on the projected pattern p : proj = p_proj × inv(c_proj); [x,y,z] homo =(proj[:3,:3]×p c )*depth+proj[:3,3:4]; Where p_proj is the optical projection matrix, c_proj is the camera projection matrix, inv is the inverse operation, and × is the matrix multiplication; Since the projected pattern is clipped, the actual sampling coordinates are: p sample =(p p -crop)*up_scale; crop is the projected pattern cropping coordinate, up_scale is the upsampling multiple; The initialization steps are as follows: cost=abs[warp(fea proj )-fea cam ]; depth init =sum(softmax(Conv3D(cost))*depth_range); Where abs represents the absolute value, sum represents the sum, cost represents the matching cost, Conv3D represents the 3D convolutional network, and depth_range represents the assumed depth range. In the first layer, a relatively accurate initial depth needs to be generated for updating the module. The assumed depth range depth_range of each layer is from fine to coarse, and the number of assumed depths of each layer is set to 192, 96, and 48 respectively; when the depth value is close to the true depth, the distortion feature map has a large similarity with the camera feature map. The matching cost is constructed by the gap between the distortion feature and the camera feature, and the cost is aggregated using a 3D convolutional network. Finally, softmax is used to generate the corresponding weights at each depth, and each depth is weighted to generate the initial depth.

7. The structured light single-frame reconstruction method based on hierarchical stereo matching according to claim 1, characterized in that: In step S6, iteratively updating the predicted depth specifically refers to using the rough initial depth value Depth init , the output feature feature of the feature extraction network, and the depth map Depth predicted by the previous layer pre Iteratively update the depth value; Denotes the depth value updated in the i-th iteration, feature is the feature map output by the feature extraction network, and Depth is initially pre Bicubic upsampling makes its resolution consistent with the next layer resolution and uses the upsampled depth value as UpdateNet is a deep update network consisting of two connected residual blocks. The update network calculates Depth init The matching cost is calculated as follows: cost=warp(fea proj )-fea cam ; The sampling method of the warp operation is the depth initialization described in step S5. Finally, the matching cost, depth value, and feature map are used as the input of the residual network, and the final output is the updated depth 8. The structured light single-frame reconstruction method based on hierarchical stereo matching according to claim 1, characterized in that: The depth refinement in step S7 further refines the final depth map Depth obtained in step S6, using the img of the image captured by the camera cam The contextual information further refines the predicted depth; Depth refine =Refine(img cam ,Depth); The final refined depth map is: Depth refine =Depth+ConvBlocks(cat([img cam ,Depth])。 9. The structured light single-frame reconstruction method based on hierarchical stereo matching according to claim 8, characterized in that: ConvBlocks consists of three 3x3 convolution blocks and one 1x1 convolution block, and each convolution block consists of a convolution layer and a LeakyReLU activation function.

Citation Information

Cited By

  • Robust single-frame structured light three-dimensional imaging method and system based on neural feature decoding

    CN121505174A