A three-dimensional reconstruction method of dynamic structured light and its system and computing device
By introducing the RPSNet neural network model in structured light three-dimensional reconstruction, the problem that traditional methods cannot achieve high-quality three-dimensional reconstruction in dynamic scenarios is solved, and three-dimensional reconstruction adapted to random phase shifting length is realized, improving the reconstruction quality and real-timeness.
Patent Information
- Application Number
- CN202310643012.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-01
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-06-01
AI Technical Summary
Traditional structured light three-dimensional reconstruction methods cannot effectively achieve high-quality three-dimensional reconstruction in dynamic or real-time measurement scenarios, mainly because the acquisition of multiple phase shift images takes time and the errors caused by object movement are difficult to overcome.
A dynamic structured light three-dimensional reconstruction method based on the RPSNet neural network model is proposed. This model adopts the CycleGAN network framework, combines the AIR2U-net generator and multi-layer convolutional network discriminator, and adapts to the random phase shift step size through adversarial learning and attention mechanisms to realize three-dimensional reconstruction.
This method can generate high-quality depth maps in dynamic scenarios, overcomes the shortcomings of the traditional method in motion error and real-time, and is suitable for the three-dimensional reconstruction needs in factory environments.
Smart Images

Figure CN117132704B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine vision, and in particular relates to a three-dimensional reconstruction method of dynamic structured light and a system and computing equipment thereof. Background Art
[0002] As an optical non-contact 3D shape measurement technology, structured light has been widely used in intelligent manufacturing, reverse engineering, and heritage digitization. Structured light projection technology based on FPP (Fringe Projection Profilometry) is one of the most popular optical 3D imaging technologies due to its simple hardware structure, flexible implementation, and high measurement accuracy. With the improvement of the performance of imaging equipment and projection equipment, it is possible to achieve high-speed 3D shape measurement based on FPP structured light technology. At the same time, the importance of obtaining high-quality 3D information in high-speed scenes for online quality inspection, stress deformation analysis, rapid reverse forming and other applications is self-evident. In order to achieve 3D measurement in high-speed scenes, it is usually necessary to reduce the number of images required for each reconstruction to improve the measurement efficiency. The traditional structured light measurement method theoretically requires at least 3 phase-shifted images to complete a reconstruction. In the actual reconstruction process, 5 or more phase-shifted images are often required to obtain higher reconstruction quality, and the quality of 3D reconstruction is proportional to the number of phase-shifted images. This traditional method works well in static scenes and non-real-time measurement scenarios, but it cannot achieve the expected results in dynamic or real-time measurement scenarios. This is because the acquisition of multiple phase-shifted images takes time, which is unacceptable for applications that require high real-time performance. More importantly, since the object being measured is in motion, there are errors between the multiple phase-shifted images due to the movement of the object, which leads to unsatisfactory final 3D reconstruction results.
[0003] In recent years, with the improvement of neural network structure and the increase of computer computing power, deep learning has shown a strong fitting ability. A large number of studies have shown that deep learning is superior to traditional algorithms in terms of speed and robustness. In the research direction of structured light, deep learning can also be widely used, such as fringe denoising, fringe analysis, phase unwrapping, etc.
[0004] In a complex factory environment, factors such as external vibrations and electromagnetic interference on the motor will affect the uniform motion of the object being measured, resulting in uneven relative phase shift steps. Traditional three-step phase shift and twelve-step phase shift methods need to be used while ensuring that the phase shift step is uniform, which requires a method that can adapt to random phase shift steps. Summary of the invention
[0005] The purpose of the present invention is to address the deficiencies of the prior art and to propose a new dynamic structured light three-dimensional reconstruction method.
[0006] The present invention proposes a phase shifting method that adapts to the random phase shifting step size. The method is implemented based on the RPSNet (Random Phase-shifting Network) network model. The model uses CycleGAN (Cycle Generative Adversarial Network) as the network framework, including a generator and a discriminator.
[0007] The generator network adopts the AIR2U-net network, which is based on the U-net model and integrates the attention mechanism and IRR module. U-net is improved based on FCN (fully convolutional network), including encoder, bottleneck module, and decoder. The network obtains feature maps through the convolution and downsampling operations of the encoder, and then through the deconvolution and upsampling process, and adds the feature maps in the encoding process to this process, and finally obtains the output result. The addition of the attention mechanism greatly enhances the ability of the entire network to focus on target features and suppress irrelevant features. The present invention proposes an IRR module, which adds a recurrent convolutional layer (RCL, Recurrent Convolutional Layer) on the basis of the Inception-Res module, which increases the model's ability to recognize object features.
[0008] The discriminator uses a simple convolutional network structure. The input of the discriminator is the output result G(x) of the generator or the label y of the data set. After learning, it can determine whether the input data is real data. When the output of the generator is judged to be false, the generator will learn the characteristics of the input data, update the parameters through back propagation, and output the data to the discriminator again for true or false judgment, and repeat this process. This adversarial learning process can help the generator continuously learn the characteristic distribution of real data, thereby generating more realistic data. In addition, we can optimize the network structure and parameters of the discriminator to improve the discrimination accuracy of the discriminator, thereby improving the generation ability of the generator.
[0009] In a first aspect, the present invention provides a three-dimensional reconstruction method of dynamic structured light, comprising the following steps:
[0010] Step S1, obtaining multiple grating images of the object to be measured at consecutive moments under working conditions;
[0011] Step S2, using the dynamic structured light model RPSNet to perform three-dimensional reconstruction on the above-mentioned multiple grating images to obtain a depth map of the object to be measured;
[0012] The dynamic structured light model RPSNet is a generative adversarial network, including a generator AIR2U-net and a discriminator;
[0013] The generator AIR2U-net adopts the basic architecture of U-net network and adopts encoding-decoding structure; the feature map is obtained through the feature extraction and downsampling operation of the IRR module in the encoder, and then through the feature extraction and upsampling process of the IRR module in the decoder, and the feature map in the encoding process is added to the decoding process through the attention mechanism module RCAM through the jump connection, and finally the output result is obtained.
[0014] The discriminator adopts a multi-layer convolutional network.
[0015] In a second aspect, the present invention provides a dynamic structured light 3D reconstruction system, which is characterized by comprising the following steps:
[0016] The data acquisition module obtains multiple grating images of the object to be measured at consecutive moments under working conditions;
[0017] The 3D reconstruction module uses the dynamic structured light model RPSNet to perform 3D reconstruction on the above multiple grating images to obtain the depth map of the object to be measured.
[0018] In a third aspect, the present invention provides a computing device, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the described method is implemented.
[0019] The beneficial effects of the present invention are:
[0020] This paper proposes a new dynamic structured light 3D reconstruction method. It adopts reverse thinking, keeps the grating image of structured light projection unchanged, uses the movement of objects in a certain direction to form relative phase shift, and proposes a neural network model based on RPSNet to solve the problem of uncertainty in the step length of object movement. The model uses CycleGAN as the backbone network, in which the generator is based on U-net, and adds residual mechanism and attention mechanism to enable the generator to fully learn the features of the input feature map. After adversarial training between the generator and the discriminator, the output of the final generator can pass the judgment of the discriminator and output an ideal depth map. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is the RPSNet model training process and operation process of the present invention;
[0022] Figure 2 This is a network architecture diagram of the generator of the present invention;
[0023] Figure 3 This is a structural diagram of the IRR module of the present invention;
[0024] Figure 4 This is the structure diagram of the RCAM model;
[0025] Figure 5 This is the discriminator network architecture diagram;
[0026] Figure 6 Some models of the Thing10k 3D model dataset;
[0027] Figure 7 is the virtual FFP scene graph;
[0028] Figure 8 Set up for Blender shader nodes;
[0029] Fig. 9 Set up graphs for Blender compositing nodes;
[0030] Fig.10 Illustration of part of the generated data set;
[0031] Fig.11 The patterns to be projected in the 3Step&Gray scheme, where (a) are three phase-shifted images, (b) are four Gray code images, and (c) are two binary images.
[0032] Fig.12 The following are comparison charts of experimental results, including (a) 3Step&Gray point cloud image; (b) 12Step&MultiFreq point cloud image; (c) RPSNet point cloud image; (d) GT point cloud image. DETAILED DESCRIPTION
[0033] The present invention will be described in detail below.
[0034] A new dynamic structured light 3D reconstruction method includes the following steps:
[0035] Step S1, obtaining multiple grating images of the object to be measured at consecutive moments under working conditions;
[0036] Step S2, using the dynamic structured light model RPSNet to perform three-dimensional reconstruction on the above-mentioned multiple grating images to obtain a depth map of the object to be measured;
[0037] The dynamic structured light model RPSNet is a generative adversarial network, including a generator AIR2U-net and a discriminator;
[0038] The generator AIR2U-net adopts the basic architecture of U-net network, such as Figure 2The encoder-decoder structure is adopted; the feature map is obtained through feature extraction and downsampling operation of the residual network module IRR in the encoder, and then through the feature extraction and upsampling process of the IRR module in the decoder, and the feature map in the encoding process is added to the decoding process through the attention mechanism through the jump connection, and finally the output result is obtained;
[0039] In the encoder downsampling process, the 2-N layers are connected in series with an IRR module at the back end of each layer of the existing U-net network encoder; N represents the total number of layers of the encoder;
[0040] In the decoding layer upsampling process, the 2-N layers are connected in series with a Concatenation module at the front end of each layer of the existing U-net network decoding layer, and an IRR module at the back end;
[0041] The encoder and the decoder use skip connections in the 2nd to Nth layers, and an attention mechanism module RCAM is connected in series on each skip connection;
[0042] The Concatenation module is used to concatenate the output of the current layer of the encoder with the output of the previous layer of the decoder;
[0043] like Figure 3 The IRR module uses the Inception-ResNet module as the basic network framework, adds a recurrent convolution block (RCB) to enhance the features, specifically including parallel residual connections, 1×1 RCB, 3×3 RCB, 5×5 RCB; and 1×1 bottleneck layer and splicing layer;
[0044] The 1×1 bottleneck layer receives the outputs of 1×1RCB, 3×3RCB, and 5×5RCB, and performs dimension reduction and dimension increase on the channel dimension to effectively reduce the number of parameters in the network, avoid overfitting, and reduce the computational burden;
[0045] The concatenation layer receives the original input features output by the residual connection and the features output by the 1×1 bottleneck layer, and performs residual concatenation on them;
[0046] The specific process of RCB is as follows:
[0047]
[0048] in is the output of RCB with a convolution kernel of size s, and They are the input of the standard convolution layer and the RCB with a convolution kernel of size s. and is the weight of the standard convolutional layer and the kth RCB layer, and bk is the deviation;
[0049] This output is then fed into the standard ReLU activation function f, expressed as follows:
[0050]
[0051] Where F(x s ,w s ) represents the output of RCB with a convolution kernel of size s;
[0052] The specific process of the IRR module is as follows:
[0053] After multi-scale RCB, the input is effectively accumulated using the idea of circular convolution, and the result is input to the 1×1 bottleneck layer. Finally, the input X l The residual concatenation is performed with the output of the 1×1 bottleneck layer convolution operation, and the concatenated result is used as the output X of the entire IRR module. l+1 ; See formula (3):
[0054]
[0055] in are the outputs of the RCB modules with convolution kernel sizes of 1×1, 3×3, and 5×5, respectively. B(·) is the bottleneck layer function. Indicates that the matrix is spliced along the depth direction;
[0056] like Figure 4 The attention mechanism module RCAM includes a channel attention module and a spatial attention module;
[0057] The channel attention module first performs global maximum pooling and average pooling operations on each input channel, and then uses a fully connected layer with shared weights (i.e., multi-layer perceptron) to calculate the weights of different channels of the feature maps after the pooling operation, obtains the weights of the two feature maps in different channels, adds the weights together, and obtains the final weight result through the activation function;
[0058] The spatial attention module first uses maximum pooling and average pooling to perform pooling operations on the channel, concatenates the two feature maps, and uses a convolutional network to calculate the weights. Finally, the obtained weights are output through an activation function.
[0059] The whole process of the attention mechanism module RCAM is expressed using the following formula:
[0060]
[0061] where f d , f eDenote the feature maps of the decoder and encoder outputs respectively, Conv(·) denotes the convolution operation on the input, and M c (·), M s (·) perform channel attention and spatial attention operations on the input respectively, It represents the element-by-element multiplication of the left and right inputs, and O represents the final output of the module;
[0062] The channel attention module compresses the feature map in the spatial dimension and obtains a one-dimensional vector before operating. In this model, both average pooling and maximum pooling are used to aggregate the spatial information of the feature map, which is then sent to a multi-layer perceptron network with shared weights to compress the spatial dimension of the input feature map, and finally the channel attention map is generated by element-by-element summation. For a single image, the channel attention mechanism mainly learns what content in this image is important. Maximum pooling performs gradient back propagation calculations. The entire channel attention mechanism is described by the formula:
[0063] M c (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F))) (5)
[0064] AvgPool(·) and MaxPool(·) represent average pooling and maximum pooling of the input respectively, MLP(·) represents the weight-sharing multi-layer perceptron operation, and σ(·) is the activation function;
[0065] The spatial attention module compresses the channel and performs average pooling and maximum pooling on the channel dimension. The maximum pooling operation is to extract the maximum value on the channel, and the number of extractions is the height multiplied by the width. The average pooling operation is to extract the average value on the channel, and the number of extractions is also the height multiplied by the width. Then the previously extracted feature maps are merged to obtain the final output feature map. The process is expressed by the formula:
[0066] M s (F')=σ(Conv(AvgPool(F');MaxPool(F'))) (6)
[0067] The training and testing process of the dynamic structured light model RPSNet is as follows Figure 1 :
[0068] During training, the dynamic structured light model RPSNet includes two parts;
[0069] The first part is the grating image x of the object being measured passing through the generator G XY Converted into the depth map y' of the object being measured, and then generated by the generator G YX Generate a grating image x', and the two conversion results x' and y' of this process are respectively used by the discriminator DX , D Y To distinguish true from false, the generator calculates the adversarial loss based on the output of the discriminator and adjusts the output. In addition, in order to prevent the output of the generator from being too radical and losing the characteristics of the original input image, a consistency loss is added in the process of generating the image.
[0070] The second part is similar to the first part, converting the depth map y into a raster image x' and then into a depth map y". This process also requires the use of a discriminator and a consistency loss function for constraints. Through the alternating training of the above process, the entire network model can eventually learn key features and output high-quality depth images.
[0071] In the process of discriminator training, it is necessary to train the discriminator D of domain X and domain Y X and D Y , input the real X-domain or Y-domain image and the image generated by the generator respectively, calculate the loss function with the result obtained by the discriminator and the actual label, and then multiply the two loss functions by a coefficient of 0.5 and add them together; finally, update the model parameters through back propagation;
[0072] During training, the loss function is used to train the constructed generator and discriminator alternately. First, a discriminator is trained, and then a generator is trained alternately until the preset number of training rounds is reached.
[0073] During testing, the test data set is used only for the generator G XY Verify and test.
[0074] Generator G XY and the generator G YX The generator AIR2U-net is used in both cases.
[0075] The loss function of the dynamic structured light model RPSNet during training is composed of the loss function of the generator and the loss function of the discriminator;
[0076] The loss function of the generator includes adversarial loss, cycle consistency loss and identity loss. Since two generators are used in RPSNet training to convert the original domain and the target domain, the above three loss functions are composed of two parts, so the total loss function consists of six parts; as shown in formula (7):
[0077]
[0078] in and Denote the adversarial generation loss functions of domain X and domain Y respectively; and denote the cycle consistency loss function of domain X and domain Y respectively, λ X and λ Y Represent the weights of the X domain and the Y domain respectively; and Represent the identity loss function of the X domain and the Y domain, μ X and μ Y Then they represent the weights of the X domain and the Y domain respectively;
[0079] Since CycleGAN has two generators, the adversarial generation loss function of CycleGAN is expressed as follows:
[0080]
[0081] where f w (·) represents a set of functions that satisfy the K-Lipschitz condition; P gxy Denotes the generator G XY The resulting sample distribution; P gyx Denotes the generator G YX The sample distribution generated; E represents the mean;
[0082] The purpose of the cycle consistency loss function is to prevent the image generated by the generator from being too biased towards the target domain and losing the image information in the original domain. The input of the cycle consistency loss function is the image in the original domain and the image in the target domain output by the generator and the inverse generator. The two should be as similar as possible. Therefore, the cycle consistency loss function is defined by the following formula:
[0083]
[0084] Where X represents the original input raster image, Y represents the original input depth map, ||·|| 1 Represents the distance calculation function, such as G YX (G XY (X)) to X, G YX (G XY (X)) represents a raster image generated from a depth map generated from a raster image;
[0085] The above loss functions are only the input of the original domain, and do not consider the input of the target domain; for this reason, the identity loss function adds the input of the target domain, and its formula definition is as follows:
[0086]
[0087] The loss function of the discriminator also uses Wasserstein distance as a similarity measurement indicator. In order to correctly identify the real image and the generated image, the discriminator hopes that the distance between the real image and the generated image is as large as possible in the design of the loss function. Therefore, the loss function of the discriminator can be expressed as follows:
[0088]
[0089] like Figure 5 The discriminator uses a simple multi-layer convolutional network structure. The input of the discriminator is the output result G(x) of the generator or the label y of the data set. After learning, it can determine whether the input data is real data. When the output of the generator is judged to be false, the generator will learn the characteristics of the input data, update the parameters through back propagation, and output the data to the discriminator again for true and false judgment, and repeat this process. Specifically, there are two discriminators, one for distinguishing between real images and generated images, and the other for distinguishing the transformation results of real style images and generated images. The structures of these two discriminators are the same, both of which use deep neural networks composed of multiple convolutional layers and fully connected layers. During the training process, the goal of the discriminator is to distinguish between real samples and generated samples as much as possible, so its loss function usually uses a binary cross entropy loss function.
[0090] Experimental results and analysis
[0091] The training of the neural network model is inseparable from a large number of data sets. In order for the model to better adapt to the three-dimensional reconstruction task in the factory environment, the Thing10k 3D model library is used to establish a data set and the model is trained and tested on the data set. In order to verify the effectiveness of the model, the present invention also compares the model of the present invention with the measurement method using three-step phase shift with Gray code and the measurement method using twelve-step phase shift with multi-frequency heterodyne. Experiments show that the model proposed by the present invention has certain performance advantages and can better adapt to the measurement needs of workpieces in a factory environment.
[0092] 1) Experimental Dataset
[0093] Currently commonly used 3D model datasets include ModelNet, ShapeNet, ABC, Thingi10K, etc. When selecting a dataset, the present invention mainly considers two points. The first is the effective working distance of the FPP system under visible light, which is generally 1 to 2 meters, so the volume of the 3D model should not be too large; the second is that the application scenario of the model of the present invention is in the industrial production process, and the selected model should be similar to common artifacts in industrial scenarios. Based on the above two points, the present invention selects the Thingi10K dataset as the 3D model dataset used by the present invention. The dataset contains various 3D models of many common objects such as artifacts, sculptures, vases, etc., some of which are as follows Figure 6 The diversity and scale of these models help generate large and diverse data samples, thereby training more generalizable and generalizable models.
[0094] The effect diagram of the virtual FFP system is as follows: Figure 7 As shown. Blender is an open source 3D scene production software, and it can process images in batches through Python scripts. Using Blender simulation software, real-world scenes can be simulated in a virtual environment. In this virtual environment, by placing two virtual cameras and a virtual projector, and setting the projector to project sinusoidal stripes onto the object, the deformed stripes after the object height modulation are captured by the left and right cameras. The raster image captured by the virtual camera can be so realistic that the entire virtual system can simulate the real FPP system.
[0095] Using the virtual FFP system, the data set required for model training can be generated. In the virtual FPP system, a projector is first required to project a raster pattern, which can be achieved by setting up a shader node. The shader node is a module in Blender used to color and render the model, which can achieve different rendering effects. Figure 8 As shown, you first need to set up a node called "Image Texture" to select the source of the image projected by the projector. This node needs to select a sinusoidal stripe image as the input of the projection image. In this way, the projection of the grating pattern can be realized in the virtual FPP system, and the corresponding edge image and depth image can be generated.
[0096] Then, the raster image captured by the camera needs to be rendered. Specifically, the "image" or "depth" attribute in the "render layer" of the synthesis node is passed through the normalization node and output to the synthesis node, and then the raster image and depth image are rendered. The synthesis node settings are shown in the figure below. Fig. 9 shown.
[0097] In order to make the virtual FFP system closer to the real factory environment, the present invention uses a variety of methods to simulate the real environment. For example, in order to further enrich the data set, the present invention rotates the model in the three-dimensional model data set multiple times in all directions, which is also a simulation of the messy placement of real workpieces. In order to enhance the realism of the virtual FFP system, a background board is added to the scene, and photos of the real factory assembly line are rendered on the background board as maps. The data set obtained by some of the above-mentioned simulation methods can be as close to the data set collected in the real environment as possible. During training, the network model can better learn the input features and avoid overfitting.
[0098] Through the above-mentioned settings for building the FFP virtual system in Blender, a large number of simulation data sets can be collected. However, in practical applications, the above-mentioned graphical settings are extremely cumbersome and inefficient. For example, in order to train the model to learn the grating pattern of random phase shift steps, it is necessary to perform random displacement in a single direction on the object under test. Manual adjustment in the above-mentioned graphical interface is undoubtedly a huge workload. Blender provides a Python script method to build the entire simulation system, which is extremely convenient for users. Therefore, the construction of the above-mentioned FFP virtual system is implemented using Python scripts, and some of the generated data sets are as follows Fig.10 In order to correctly and effectively train and evaluate the model, the present invention divides the data set into a training data set and a test data set in a ratio of 3:1.
[0099] 2) Model implementation details
[0100] The present invention aims to illustrate the design and training details of the RPSNet model proposed in the present invention. When training RPSNet, the dataset used is based on the Thing10k dataset. The hardware environment used in the experiment is a 64-bit Windows system, the CPU is Intel Core i7-11700, the graphics card is NVIDIA RTX2080TI, the memory is 16GB, and all codes in the experiment are implemented using the Pytorch framework. Since RPSNet is a variant of CycleGAN, its training method is the same as CycleGAN, that is, the model is trained by alternately training the generator and the discriminator.
[0101] In the process of training the generator, it is necessary to simultaneously train the generator G that converts the X domain input into the Y domain input XY and a generator G that transforms Y domain input into X domain input YX Both generators need to use the adversarial loss function L GAN , cycle-consistent loss function L cyc And the identity loss function L idt For training, there are a total of six loss functions, among which the consistency loss function and the identity loss function both use L 1 loss, while the adversarial loss function uses Wasserstein distance to define the loss function. In order to control the weight of each loss function term, the coefficient of the cycle consistency loss term is set to 10, the coefficient of the identity loss term is set to 5, and the coefficient of the adversarial loss term is set to 1 when calculating the total loss. In the process of training the discriminator, it is necessary to train the discriminator D of the X domain and the Y domain X and D Y, input the real X-domain or Y-domain image and the image generated by the generator respectively, calculate the loss function with the result obtained by the discriminator and the actual label, and then add the two loss functions multiplied by a coefficient of 0.5. Finally, update the model parameters by back propagation. In order to ensure the training effect, the batch size (Batch size) of the training sample is set to 8 in the present invention, and the Adam optimization algorithm is selected for back propagation. The Adam optimizer is a commonly used optimization algorithm for solving the gradient descent problem. It updates the parameters by dynamically adjusting the first-order moment estimate and the second-order moment estimate of the gradient. In the present invention, the learning rate of the Adam optimizer is set to 0.002, and the momentum parameter β 1 is 0.5, β 2 The value of 0.999 is used to find the global optimal point and improve the training efficiency and model performance. Through such a training process, the generalization ability of the model can be improved, thus achieving better results in real application scenarios.
[0102] 3) Performance comparison experiment
[0103] The experiment of the present invention is intended to verify the effectiveness of the RPSNet model proposed in the present invention. To this end, the present invention uses a method of combining Gray code with a three-step phase shift, a method of combining multi-frequency heterodyne with a twelve-step phase shift, and the method proposed in the present invention for comparative experiments. Since there is currently no generally applicable method with high reconstruction quality for the measurement of dynamic objects, the present invention selects a reconstruction algorithm of multi-frequency heterodyne combined with a twelve-step phase shift to measure the measured object while keeping it static as the Ground Truth. Since the twelve-step phase shift can better suppress the errors introduced by factors such as the nonlinearity of the projection and the reflection of the object, it ensures that the three-dimensional object can be reconstructed with high quality.
[0104] The specific reconstruction algorithms used by each solution in the comparative experiment and the relevant settings of the corresponding experiments are described as follows:
[0105] ·3Step&Gray: This solution uses three-step phase shift calculation to wrap the phase and projects an additional Gray code pattern for phase marking for parsing the wrapped phase. In this solution, the number of Gray code patterns is set to 4, that is, it supports up to 16 cycles in the field of view, which is sufficient for an image of 512 pixels. The specific projected pattern is encoded according to the Gray code to reduce the influence of the object surface reflection on the grating pattern. The encoded pattern is as follows Fig.11 As shown. In actual applications, considering the reflected light on the surface of the object being measured and the uneven light in the environment, two additional patterns of full black and full white need to be projected to normalize the brightness, so as to facilitate the judgment of the brightness of the pixels in different light environments. In summary, this solution needs to project a total of 3+4+2=9 patterns.
[0106] ·12Step&MultiFreq: In this solution, twelve-step phase shift is first used to calculate the wrapped phase, and multi-frequency heterodyne method is used for phase unwrapping. In order to ensure that the total period obtained by multi-frequency heterodyne can cover the entire field of view, this solution adopts a three-frequency heterodyne algorithm, and the unwrapping calculation is performed on sinusoidal patterns with 25 pixels, 27 pixels, and 29 pixels as one period. Since the calculation of the wrapped phase uses twelve-step phase shift, a total of 12×3=36 grating patterns need to be projected.
[0107] RPSNet: This solution uses the dynamic structured light measurement model RPSNet proposed in the present invention that adapts to the random phase shift step. The model directly converts the input grating pattern into a depth image without the need for the calculation of the wrapped phase and the unwrapped phase process.
[0108] GT: In order to highlight the reconstruction effect of each scheme, this comparative experiment adopts a reconstruction algorithm of twelve-step phase shift combined with multi-frequency heterodyne, and measures the measured object while keeping it static. The relevant experimental settings of this reconstruction algorithm are consistent with the 12Step&MultiFreq reconstruction scheme.
[0109] Fig.12 The results of this comparative experiment are shown. Fig.12 As can be seen in (a), when the 3Step&Gray structured light reconstruction scheme is faced with the measurement of objects in dynamic scenes, the existence of motion errors leads to uneven planes in the point cloud and the entire point cloud is relatively sparse, and the reconstruction accuracy of the entire point cloud is also low. This is because the three-step phase shift algorithm cannot handle the errors introduced by factors such as the nonlinearity of the projector and the reflection of the object surface. Fig.12 (b) is the reconstruction result of the 12Step&MultiFreq scheme. Although this scheme can suppress the errors of the environment and equipment, the large number of phase shift steps leads to large motion errors. The RPSNet model proposed in this paper can better overcome the shortcomings of the above two reconstruction schemes. The final reconstructed point cloud is shown in Figure 2. Fig.12 As shown in (c), it can be seen that although the point cloud data reconstructed by this scheme is not as dense as that by the twelve-step phase shift, by comparison, the point cloud data reconstructed by this scheme can accurately reconstruct the basic features of the workpiece.
[0110] In order to conduct a specific quantitative evaluation of the reconstruction effects of the above methods, the present invention uses the above methods to conduct comparative experiments on standard parts. The results of the experiments are shown in the following table:
[0111] Table 1 Comparative experiment of each method measured on standard parts
[0112]
[0113] From the table above, it can be seen that the three-step phase shift combined with Gray code structured light measurement scheme has a large error when measuring moving objects, and the error fluctuates greatly. From the comparison between 12Step&MultiFreq and GT, the twelve-step phase shift combined with multi-frequency heterodyne measurement method has higher measurement accuracy in static scenes, but will have a large error when measuring objects in motion. The RPSNet proposed in the present invention can complete three-dimensional reconstruction in scenes where objects are moving, and has high accuracy, and can complete the grasping task.
[0114] 4) Ablation experiment
[0115] In order to prove the effectiveness of each module in the RPSNet model proposed in this paper, ablation experiments are designed for different modules, and standard parts are used as measurement objects. The reconstruction quality of different models is represented by comparing the error between the measured length of the standard parts and the actual length of the standard parts. The experimental settings are based on the relevant settings of the RPSNet items in the performance comparison experiment, and the experimental settings not mentioned are not changed.
[0116] (1) Effectiveness of the IRR model
[0117] To verify the effectiveness of the IRR model, the present invention designed three groups of experiments, and the experimental settings are as follows:
[0118] Conv: This group of experiments replaces the IRR module in the RPSNet model proposed in this paper with a common 3×3 convolution operation, with padding set to 1 and stride set to 1, while other parameters remain unchanged.
[0119] Recurrent Conv: This group of experiments replaces the IRR module with a recurrent convolution operation. The convolution settings are consistent with those in the Conv experiment. The specific implementation of recurrent convolution is to set the number of cycles to 3. Except for the first convolution operation, the result of the previous convolution operation is accumulated with the input of the recurrent convolution and then input into the convolution operation.
[0120] IRR: The relevant settings of this group of experiments are consistent with the RPSNet experimental settings in the performance comparison experiment.
[0121] Table 2 Comparative experimental results of ablation experiment on IRR module on standard parts
[0122]
[0123]
[0124] Table 2 shows the comparative experimental results of the ablation experiment of the IRR module on the standard parts. It can be seen that the IRR module proposed in the present invention reduces the measurement error to 0.36mm at the cost of increasing the time consumption by about 0.3 seconds, and obtains the best measurement results in the above-mentioned ablation experiment, which effectively proves the effectiveness of the module.
[0125] (2) Effectiveness of RCAM module
[0126] In the experiment of validating the RCAM module, the present invention designed three groups of comparative experiments, and the experimental settings are as follows:
[0127] None: No attention module is used in this set of experiments, and the results of each encoder level are directly input into the corresponding decoder.
[0128] AG: This group of experiments uses the Attention Gate module in the Attention U-net proposed by Oktay et al. instead of the RCAM module in RPSNet.
[0129] RCAM: The relevant settings of this group of experiments are consistent with those of the RPSNet experiment in the performance comparison experiment.
[0130] Table 3 Comparative experimental results of ablation experiment of RCAM module on standard parts
[0131]
[0132] The comparative experimental results of the ablation experiment of the RCAM module on the standard parts are shown in Table 3. From this table, it can be concluded that the RCAM module of the experiment of the present invention can improve the quality of reconstruction better than other modules in this ablation experiment, and the overall time consumption of the model is also within an acceptable range, which proves the effectiveness of the model.
[0133] In summary, the present invention is aimed at the problem of using structured light to perform three-dimensional reconstruction of the measured object in a dynamic scene. The so-called dynamic measurement refers to the object being measured being in motion, and the direction and speed of the movement are not fixed. For this dynamic three-dimensional measurement using structured light, there is currently no universal, perfect solution. The present invention designs a dynamic structured light three-dimensional reconstruction method based on RPSNet for the specific scenario of three-dimensional detection of objects on a conveyor belt. In this scenario, the measured object can be regarded as moving in one direction, and this movement just creates conditions for relative phase shift. Although this approach converts the original source of motion error into relative phase shift, it also introduces new errors, namely, uneven phase shifts, because the conveyor belt is a mechanical transmission device, which is inevitably subject to external interference and causes inconsistent speeds. To this end, the present invention is used to solve the structured light measurement method of random phase shift steps, but the problem that follows is that the biggest problem currently faced by this data-driven model is the lack of valid data. In order to establish a valid data set, the present invention creates a virtual FFP system in the software Blender and uses the 3D model data set Thing10k to make a data set. In order to verify the effectiveness of the method proposed in this paper, the method is compared with the three-step phase shift combined with Gray code, the twelve-step phase shift combined with multi-frequency heterodyne and other methods. The experiment shows that the reconstruction quality of the method proposed in this paper is close to the reconstruction result of the twelve-step phase shift combined with multi-frequency heterodyne method, and the accuracy error is 0.11mm. In general, it can meet the three-dimensional reconstruction needs of workpiece sorting in factory environments, and create conditions for subsequent three-dimensional positioning.
Claims
1. A 3D reconstruction method of dynamic structured light, Features The steps include: Step S1, obtaining multiple grating images of the object to be measured at consecutive moments under working conditions; Step S2, using the dynamic structured light model RPSNet to perform three-dimensional reconstruction on the above-mentioned multiple grating images to obtain a depth map of the object to be measured; The dynamic structured light model RPSNet is a generative adversarial network, including a generator AIR2U-net and a discriminator; The generator AIR2U-net adopts the basic architecture of the U-net network and an encoding-decoding structure; the feature map is obtained through the feature extraction and downsampling operation of the IRR module in the encoder, and then through the feature extraction and upsampling process of the IRR module in the decoder, and the feature map in the encoding process is added to the decoding process through the attention mechanism module RCAM through the jump connection, and finally the output result is obtained; The discriminator adopts a multi-layer convolutional network; In the encoder downsampling process, the 2-N layers are connected in series with an IRR module at the back end of each layer of the existing U-net network encoder; N represents the total number of layers of the encoder; In the decoder upsampling process, the 2-N layers are connected in series with a Concatenation module at the front end of each layer of the existing U-net network decoding layer, and an IRR module is connected in series at the back end; The encoder and the decoder use skip connections in the 2nd to Nth layers, and an attention mechanism module RCAM is connected in series on each skip connection; The Concatenation module is used to concatenate the output of the current layer of the encoder with the output of the previous layer of the decoder; The IRR module is based on the Inception-ResNet module as the basic network framework, and a recurrent convolution block RCB is added to enhance the features; The IRR module specifically includes parallel residual connections, 1×1 RCB, 3×3 RCB, 5×5 RCB; and 1×1 bottleneck layer and splicing layer; The 1×1 bottleneck layer receives the outputs of 1×1RCB, 3×3RCB, and 5×5RCB, and performs dimension reduction and dimension increase on the channel dimension to effectively reduce the number of parameters in the network, avoid overfitting, and reduce the computational burden; The concatenation layer receives the original input features output by the residual connection and the features output by the 1×1 bottleneck layer, and performs residual concatenation on them; The specific process of RCB is as follows: in is the output of RCB with a convolution kernel of size s, and They are the input of the standard convolution layer and the RCB with a convolution kernel of size s. and is the weight of the standard convolutional layer and the kth RCB layer, and b k is the deviation; This output is then fed into the standard ReLU activation function f, expressed as follows: Where F(x s ,w s ) represents the output of RCB with a convolution kernel of size s; The specific process of the IRR module is as follows: After multi-scale RCB, the input is effectively accumulated using the idea of circular convolution, and the result is input to the 1×1 bottleneck layer. Finally, the input X l The residual concatenation is performed with the output of the 1×1 bottleneck layer convolution operation, and the concatenated result is used as the output X of the entire IRR module. l+1 ; See formula (3): in are the outputs of the RCB modules with convolution kernel sizes of 1×1, 3×3, and 5×5, respectively. B(·) is the bottleneck layer function, and ° indicates the concatenation of matrices along the depth direction.
2. The method according to claim 1, Features The attention mechanism module RCAM includes a channel attention module and a spatial attention module; The channel attention module first performs global maximum pooling and average pooling operations on each input channel, and then uses a fully connected layer with shared weights to calculate the weights of different channels of the feature maps after the pooling operation, obtains the weights of the two feature maps in different channels, adds the weights together and obtains the final weight result through the activation function; The spatial attention module first uses maximum pooling and average pooling to perform channel pooling operations, concatenates the two feature maps, and uses a convolutional network to calculate the weights. Finally, the obtained weights are output through an activation function.
3. The method according to claim 2, Features The whole process of the attention mechanism module RCAM is expressed using the following formula: where f d , f e Denote the feature maps of the decoder and encoder outputs, respectively. Conv(·) denotes the convolution operation on the input. c (·), M s (·) perform channel attention and spatial attention operations on the input respectively, It represents the element-by-element multiplication of the left and right inputs, and P represents the final output of the module; M c (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F))) (5) Where AvgPool(·) and MaxPool(·) represent average pooling and maximum pooling of the input respectively, MLP(·) represents the weight-sharing multi-layer perceptron operation, and σ(·) is the activation function; M s (F')=σ(Conv(AvgPool(F');MaxPool(F'))) (6)。 4. The method according to claim 1, Features The training and testing process of the dynamic structured light model RPSNet is as follows: During training, the dynamic structured light model RPSNet includes two parts; The first part is the grating image x of the object being measured passing through the generator G XY Converted into the depth map y' of the object being measured, and then generated by the generator G YX Generate a grating image x', and the two conversion results x' and y' of this process are respectively used by the discriminator D X , D Y To distinguish true from false, the generator calculates the adversarial loss based on the output of the discriminator and adjusts the output. In addition, in order to prevent the output of the generator from being too radical and losing the characteristics of the original input image, a consistency loss is added in the process of generating the image. The second part is similar to the first part, converting the depth map y into a raster image x' and then into a depth map y". This process also requires the use of a discriminator and a consistency loss function for constraints. Through the alternating training of the above process, the entire network model can eventually learn key features and output high-quality depth images. In the process of discriminator training, it is necessary to train the discriminator D of domain X and domain Y X and D Y , input the real X domain or Y domain image and the image generated by the generator respectively, calculate the loss function with the result obtained by the discriminator and the actual label, and then multiply the two loss functions by a coefficient of 0.5 and add them together; Finally, the model parameters are updated through back propagation; During training, the loss function is used to train the constructed generator and discriminator alternately. First, a discriminator is trained, and then a generator is trained alternately until the preset number of training rounds is reached. During testing, the test data set is used only for the generator G XY Verify and test.
5. The method according to claim 4, Features The loss function of the dynamic structured light model RPSNet during training is composed of the loss function of the generator and the loss function of the discriminator; The loss function of the generator includes adversarial loss, cycle consistency loss and identity loss. Since two generators are used in RPSNet training to convert the original domain and the target domain, the above three loss functions are composed of two parts, so the total loss function is shown in the following formula: in and Denote the adversarial generation loss functions of domain X and domain Y respectively; and denote the cycle consistency loss function of domain X and domain Y respectively, λ X and λ Y Represent the weights of the X domain and the Y domain respectively; and Represent the identity loss function of the X domain and the Y domain, μ X and μ Y Represent the weights of the X domain and the Y domain respectively; The adversarial generation loss function is defined by the following formula: where f w (·) represents a set of functions that satisfy the K-Lipschitz condition; Denotes the generator G XY The resulting sample distribution; Denotes the generator G YX The sample distribution generated; E represents the mean; The cycle consistency loss function is defined as follows: Where X represents the original input raster image, Y represents the original input depth map, ||·|| 1 represents the distance calculation function; The identity loss function is defined by the following formula: The loss function of the discriminator is defined by the following formula; 6. A dynamic structured light 3D reconstruction system implementing the method according to any one of claims 1 to 5, Features include: The data acquisition module obtains multiple grating images of the object to be measured at consecutive moments under working conditions; The 3D reconstruction module uses the dynamic structured light model RPSNet to perform 3D reconstruction on the above-mentioned multiple grating images to obtain the depth map of the object to be measured.
7. A computing device, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Three-dimensional reconstruction method and device, point cloud fusion method and device, equipment and storage medium
CN114066960A
Three-dimensional face reconstruction method and device based on speckle structured light
CN116188701A