Binocular structured light transparent object three-dimensional reconstruction method and system based on diffusion model
By constructing a denoising U-Net network of diffusion generation model, the problem of difficulty in extracting aliased stripe images in transparent object reconstruction is solved, and high-precision three-dimensional reconstruction of transparent object surfaces is achieved.
Patent Information
- Application Number
- CN202510170461.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional phase measurement contour technique is difficult to accurately reconstruct the three-dimensional morphology of transparent objects, because aliased stripe images caused by refraction and reflection of light on the surface of transparent objects cannot accurately extract phase information.
Using a binocular structured light method based on the diffusion model, the denoised U-Net network in the diffusion generation model is constructed, and the binocular aliased stripe pattern and aliased stripe pattern sequence set of transparent objects are trained to restore the reflective stripe pattern on the front surface of the transparent object.
It realizes the extraction of reflected stripe information on the front surface of the transparent object from the aliased stripe pattern captured by the camera, accurately restores the three-dimensional morphology of the transparent object surface, and improves the reconstruction accuracy.
Smart Images

Figure CN120298569A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of three-dimensional reconstruction of objects, and particularly relates to a method and system for three-dimensional reconstruction of transparent objects based on a diffusion model using binocular structured light. Background Art
[0002] Phase measurement profilometry is a 3D measurement method that combines sinusoidal fringe projection and phase-shifting techniques. Its principle of implementation is to calculate the phase value containing the three-dimensional information of the surface of the object to be measured by capturing a series of sinusoidal fringe images with different frequencies and phases through a camera, and then to restore the height information of the object surface through the phase information. Currently, phase measurement profilometry is widely used to reconstruct the three-dimensional geometry of the surface of purely diffuse or slightly specularly reflecting objects. However, for transparent objects, this technology faces special challenges. First, when light passes through the surface of a transparent object, refraction occurs, causing the light to deviate from its original propagation path, and only a small amount of light is reflected from the surface of the transparent object and captured by the camera. In addition, the refracted light is reflected multiple times inside the object and then re-emitted, causing severe aliasing of the fringe images captured by the camera, resulting in errors in the process of unwrapping the phase and making it impossible to accurately restore the height information of the surface of the transparent object. Therefore, traditional phase measurement profilometry cannot accurately extract phase information from aliased fringes when reconstructing transparent objects. Summary of the Invention
[0003] One of the purposes of the present invention is to provide a method for three-dimensional reconstruction of transparent objects based on a diffusion model using binocular structured light, which can effectively extract the information of the reflected fringe pattern (i.e., the binocular non-aliased fringe pattern) on the front surface of the transparent object from the binocular aliased fringe pattern captured by the camera, accurately restore the three-dimensional topography of the surface of the transparent object, and achieve the reconstruction of the surface of the transparent object.
[0004] Another purpose of the present invention is to provide a system for three-dimensional reconstruction of transparent objects based on a diffusion model using binocular structured light.
[0005] To achieve the above-mentioned one purpose, the present invention is implemented by adopting the following technical solutions:
[0006] A method for three-dimensional reconstruction of transparent objects based on a diffusion model using binocular structured light, the method for three-dimensional reconstruction of transparent objects using binocular structured light comprising:
[0007] Step S1, obtaining a set of sequences of binocular aliased fringe patterns and corresponding sets of sequences of binocular non-aliased fringe patterns of transparent object samples with different shapes;
[0008] Step S2: Divide the binocular aliased fringe pattern sequence set and the corresponding binocular non-aliased fringe pattern sequence set into a binocular aliased fringe pattern sequence training set and the corresponding binocular non-aliased fringe pattern sequence training set, as well as a binocular aliased fringe pattern sequence test set and the corresponding binocular non-aliased fringe pattern sequence test set;
[0009] Step S3: Construct a diffusion generative model; and use the binocular aliased fringe pattern sequence training set and the corresponding binocular non-aliased fringe pattern sequence training set to train the denoising U-Net network in the diffusion generative model to obtain a trained denoising U-Net network;
[0010] Step S4: Use the binocular aliased fringe pattern sequence test set and the corresponding binocular non-aliased fringe pattern sequence test set to test the trained denoising U-Net network to obtain a tested denoising U-Net network;
[0011] Step S5: Obtain the binocular aliased fringe pattern of the transparent object to be restored; and use the binocular aliased fringe pattern of the transparent object to be restored and adopt the tested denoising U-Net network to perform three-dimensional reconstruction of the front surface of the transparent object.
[0012] Further, in the step S1, the specific process of obtaining the binocular aliased fringe pattern sequence set of different-shaped transparent object samples and the corresponding reflective object samples includes:
[0013] Step S11: Obtain different-shaped transparent object samples, and project a series of phase-shifted fringe patterns onto each shaped transparent object sample with different poses, and then use left and right binocular cameras to perform image acquisition to obtain the binocular aliased fringe pattern sequence set of each shaped transparent object sample in each pose;
[0014] In the step S11, at least one of the phase, frequency, and angle corresponding to the aliased fringe patterns in the binocular aliased fringe pattern sequence set is different;
[0015] Step S12: Spray a reflective layer on each shaped transparent object sample to obtain the corresponding reflective object sample of each shaped transparent object sample;
[0016] Step S13: Project a series of phase-shifted fringe patterns onto each shaped reflective object sample with different poses, and then use left and right binocular cameras to perform image acquisition to obtain the binocular non-aliased fringe pattern sequence set of each shaped transparent object sample in each pose;
[0017] In the step S13, at least one of the phase, frequency, and angle corresponding to the binocular non-aliased fringe pattern sequences in the binocular non-aliased fringe pattern sequence set is different.
[0018] Further, in the step S3, the specific process of the training includes:
[0019] Step S31, initialize the network parameters of the denoising U-Net network;
[0020] Step S32, set the serial number i = 1 in the binocular aliased fringe pattern sequence training set and the corresponding binocular non-aliased fringe pattern sequence training set;
[0021] Step S33, use the VAE encoder in the diffusion generation model to encode the i-th binocular aliased fringe pattern and the corresponding binocular non-aliased fringe pattern into a first latent space tensor and a second latent space tensor respectively;
[0022] Step S34, randomly generate Gaussian noise; and add the Gaussian noise to the second latent space tensor at time t to generate the second latent space tensor with Gaussian noise at time t;
[0023] Step S35, splice the first latent space tensor and the second latent space tensor with Gaussian noise, and then input them into the denoising U-Net network;
[0024] Step S36, the denoising U-Net network estimates the noise using the network parameters;
[0025] Step S37, determine whether the loss value between the noise estimation result and the Gaussian noise is less than the threshold. If so, save the network parameters and enter step S38; if not, update the network parameters and return to step S36;
[0026] Step S38, determine whether i is equal to I. If so, save the trained denoising U-Net network and end; if not, i = i + 1 and return to step S33;
[0027] Wherein, I is the number of binocular aliased fringe patterns in the binocular aliased fringe pattern sequence training set.
[0028] Further, in the step S4, the specific process of the test includes:
[0029] Step S41, set the serial number j = 1 in the binocular aliased fringe pattern sequence test set and the corresponding binocular non-aliased fringe pattern sequence test set;
[0030] Step S42, use the VAE encoder in the diffusion generation model to encode the j-th binocular aliased fringe pattern into a third latent space tensor;
[0031] Step S43, construct a fourth latent space tensor with noise at time t for the j-th binocular non-aliased fringe pattern;
[0032] Step S44: After splicing the third latent space tensor and the fourth latent space tensor, input them into the trained denoising U-Net network for noise prediction;
[0033] Step S45: Denoise the noise prediction result to obtain the noisy fifth latent space tensor of the j-th binocular aliasing-free fringe pattern at time t-1;
[0034] Step S46: Determine whether t-1 is 0. If so, use the VAE decoder in the diffusion generation model to decode the fifth latent space tensor to obtain the j-th binocular aliasing-free fringe pattern, and enter Step S47; if not, after fine-tuning the network parameters, use the fifth latent space tensor as the fourth latent space tensor, set t = t-1, and return to Step S44;
[0035] Step S47: Determine whether j is equal to J. If so, save the tested denoising U-Net network and end; if not, set j = j + 1 and return to Step S42;
[0036] Where J is the number of binocular aliasing fringe patterns in the binocular aliasing fringe pattern sequence test set.
[0037] Further, in the Step S5, the specific process of the three-dimensional reconstruction of the front surface of the transparent object includes:
[0038] Step S51: Use the tested denoising U-Net network to restore the binocular aliasing fringe pattern of the transparent object to be restored to obtain the binocular aliasing-free fringe pattern of the transparent object to be restored;
[0039] Step S52: Perform phase extraction, phase unwrapping, and binocular stereo matching on the sequence of binocular aliasing-free fringe patterns of the transparent object to be restored in sequence.
[0040] To achieve the second above object, the present invention adopts the following technical solution:
[0041] A three-dimensional reconstruction system for a binocular structured light transparent object based on a diffusion model, the three-dimensional reconstruction system for a binocular structured light transparent object includes:
[0042] An acquisition module, configured to acquire a set of binocular aliasing fringe pattern sequences and corresponding sets of binocular aliasing-free fringe pattern sequences of transparent object samples with different shapes;
[0043] A partitioning module, configured to partition the set of binocular aliasing fringe pattern sequences and the corresponding set of binocular aliasing-free fringe pattern sequences into a binocular aliasing fringe pattern sequence training set and a corresponding binocular aliasing-free fringe pattern sequence training set, and a binocular aliasing fringe pattern sequence test set and a corresponding binocular aliasing-free fringe pattern sequence test set;
[0044] A training module, configured to construct a diffusion generation model; and train the denoising U-Net network in the diffusion generation model by using the binocular aliased fringe pattern sequence training set and the corresponding binocular non-aliased fringe pattern sequence training set, to obtain a trained denoising U-Net network;
[0045] A testing module, configured to test the trained denoising U-Net network by using the binocular aliased fringe pattern sequence test set and the corresponding binocular non-aliased fringe pattern sequence test set, to obtain a tested denoising U-Net network;
[0046] A 3D reconstruction module, configured to obtain the binocular aliased fringe pattern of a transparent object to be restored; and perform 3D reconstruction of the front surface of the transparent object by using the binocular aliased fringe pattern of the transparent object to be restored and adopting the tested denoising U-Net network.
[0047] Further, the obtaining module includes:
[0048] A first projection and acquisition sub-module, configured to obtain transparent object samples of different shapes, and project a series of phase-shifted fringe patterns onto each transparent object sample in different poses, and then perform image acquisition by using left and right binocular cameras, so as to obtain a set of binocular aliased fringe pattern sequences of each transparent object sample in each pose;
[0049] At least one of the phase, frequency, and angle corresponding to the aliased fringe patterns in the set of binocular aliased fringe pattern sequences is different;
[0050] A spraying sub-module, configured to spray a reflective layer on each transparent object sample, so as to obtain a corresponding reflective object sample for each transparent object sample;
[0051] A second projection and acquisition sub-module, configured to project a series of phase-shifted fringe patterns onto each reflective object sample in different poses, and then perform image acquisition by using left and right binocular cameras, so as to obtain a set of binocular non-aliased fringe pattern sequences of each transparent object sample in each pose;
[0052] At least one of the phase, frequency, and angle corresponding to the binocular non-aliased fringe pattern sequences in the set of binocular non-aliased fringe pattern sequences is different.
[0053] Further, the training module includes:
[0054] An initialization sub-module, configured to initialize the network parameters of the denoising U-Net network;
[0055] The first setting sub-module is used to set the serial number i = 1 in the binocular aliased fringe pattern sequence training set and the corresponding binocular non-aliased fringe pattern sequence training set;
[0056] The first encoding sub-module is used to encode the i-th binocular aliased fringe pattern and the corresponding binocular non-aliased fringe pattern into a first latent space tensor and a second latent space tensor respectively by using the VAE encoder in the diffusion generation model;
[0057] The addition sub-module is used to randomly generate Gaussian noise; and add the Gaussian noise to the second latent space tensor at time t to generate the second latent space tensor with Gaussian noise at time t;
[0058] The first splicing sub-module is used to splice the first latent space tensor and the second latent space tensor with Gaussian noise and then input them into the denoising U-Net network;
[0059] The noise estimation sub-module is used for the denoising U-Net network to perform noise estimation by using the network parameters;
[0060] The first judgment sub-module is used to judge whether the loss value between the noise estimation result and the Gaussian noise is less than the threshold. If so, save the network parameters and transmit them to the second judgment sub-module; if not, update the network parameters and transmit them to the noise estimation sub-module;
[0061] The second judgment sub-module is used to judge whether i is equal to I. If so, save the trained denoising U-Net network and end; if not, i = i + 1 and transmit it to the first encoding sub-module;
[0062] Wherein, I is the number of binocular aliased fringe patterns in the binocular aliased fringe pattern sequence training set.
[0063] Further, the test module includes:
[0064] The first setting sub-module is used to set the serial number j = 1 in the binocular aliased fringe pattern sequence test set and the corresponding binocular non-aliased fringe pattern sequence test set;
[0065] The second encoding sub-module is used to encode the j-th binocular aliased fringe pattern into a third latent space tensor by using the VAE encoder in the diffusion generation model;
[0066] The construction sub-module is used to construct the fourth latent space tensor with noise at time t of the j-th binocular non-aliased fringe pattern;
[0067] The second splicing sub-module is used to splice the third latent space tensor and the fourth latent space tensor and then input them into the trained denoising U-Net network for noise prediction;
[0068] A denoising processing sub-module, configured to perform denoising processing on the noise prediction result to obtain a noisy fifth latent space tensor of the j-th binocular aliasing-free fringe pattern at time t-1;
[0069] A third judgment sub-module, configured to judge whether t-1 is 0. If so, use the VAE decoder in the diffusion generation model to decode the fifth latent space tensor to obtain the j-th binocular aliasing-free fringe pattern, and transmit it to the fourth judgment sub-module; if not, after fine-tuning the network parameters, use the fifth latent space tensor as the fourth latent space tensor, set t = t-1, and transmit it to the second splicing sub-module;
[0070] A fourth judgment sub-module, configured to judge whether j is equal to J. If so, save the tested denoising U-Net network and end; if not, set j = j + 1 and transmit it to the second encoding sub-module;
[0071] Where J is the number of binocular aliasing fringe patterns in the test set of the binocular aliasing fringe pattern sequence.
[0072] Further, the 3D reconstruction module includes:
[0073] A restoration sub-module, configured to use the tested denoising U-Net network to restore the binocular aliasing fringe pattern of the transparent object to be restored to obtain a sequence of binocular aliasing-free fringe patterns of the transparent object to be restored;
[0074] A processing sub-module, configured to sequentially perform phase extraction, phase unwrapping, and binocular stereo matching on the sequence of binocular aliasing-free fringe patterns of the transparent object to be restored.
[0075] In summary, the technical solution of the present invention has the following technical effects:
[0076] Based on the constructed set of binocular aliased fringe pattern sequences and the corresponding set of binocular non-aliased fringe pattern sequences for transparent object samples, the denoising U-Net network in the constructed diffusion generative model is trained to restore the aliased phase-shifted fringes (i.e., binocular aliased fringe patterns) captured by the camera for the transparent object samples under the fringe projection method to the non-aliased phase-shifted fringes (i.e., binocular non-aliased fringe patterns) reflected by the front surface of the transparent object samples; and by using the non-aliased phase-shifted fringes, high-precision surface three-dimensional reconstruction of the transparent object is realized, significantly improving the surface reconstruction accuracy of the transparent object; the present invention can effectively extract the reflected fringe pattern information (i.e., binocular non-aliased fringe patterns) of the front surface of the transparent object from the binocular aliased fringe patterns captured by the camera, accurately restore the three-dimensional surface topography of the transparent object, realize the reconstruction of the transparent object surface, solve the problem that phase profilometry cannot directly reconstruct transparent objects, and has broad application potential in related fields such as non-contact three-dimensional reconstruction and in-situ measurement of the transparent object surface. Description of the Drawings
[0077] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0078] Figure 1 Schematic flow chart of the binocular structured light three-dimensional reconstruction method for transparent objects based on the diffusion model in the embodiment of the present invention;
[0079] Figure 2 Schematic diagram of 10 transparent object samples with different shapes used in the construction of the fringe pattern data set in the embodiment of the present invention;
[0080] Figure 3 Schematic flow chart of obtaining the fringe pattern data set in the embodiment of the present invention;
[0081] Figure 4 Schematic diagram of the training of the denoising U-Net network in the embodiment of the present invention;
[0082] Figure 5 Schematic diagram of the network inference of the denoising U-Net network in the embodiment of the present invention;
[0083] Figure 6 Schematic flow chart of binocular structured light reconstruction using the inference network in the embodiment of the present invention. Detailed Embodiments
[0084] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0085] The basic principle of reconstructing an object with a diffuse reflection surface using the existing Phase Shift Profilometry (PSP) method: A projection device projects a series of phase shift fringe patterns onto the surface of the object to be measured. For the N-step phase measurement profilometry, the fringe patterns captured by the camera can be expressed as:
[0086]
[0087] where i is the serial number of the projected fringe pattern, i = 1, 2,..., N, I i (x, y) is the intensity distribution of the i-th fringe pattern captured by the camera at the coordinate (x, y); a is the ambient light intensity, b is the modulation amplitude of the fringe pattern intensity, and φ(x, y) is the phase distribution at the coordinate (x, y).
[0088] To obtain the phase value φ(x, y), it is calculated through the following formula:
[0089]
[0090] where the range of the phase value φ(x, y) of each pixel in the phase map is limited to -π to π. The wrapped phase between -π and π is unwrapped to obtain an unwrapped phase map with monotonic values, thereby eliminating the phase discontinuity. Finally, based on the obtained phase information and the camera calibration parameters, the precise reconstruction of the three-dimensional shape of the object is achieved. When reconstructing the surface of a transparent object using phase profilometry, the fringe patterns captured by the camera will be severely aliased, resulting in the failure of the phase extraction process and the inability to accurately recover the height information of the transparent object surface.
[0091] When reconstructing a transparent object using the phase measurement profilometry method, first, a sequence of phase shift fringes is projected onto the surface of the transparent object, and the light will be reflected and refracted on the front and back surfaces of the transparent object. Among them, a part of the light is reflected from the front surface and directly captured by the camera, recorded as F. When the light F is refracted through the front surface and reflected by the back surface, the light captured by the camera can be recorded as B. M is the light that is finally captured by the camera after multiple reflections between the front and back surfaces of the refracted light. Usually, the intensities of the light F and B are the strongest, while the intensity of the light M is the weakest.
[0092] By applying the above reflection model and refraction model to phase measurement profilometry, the fringe pattern captured by the camera at time t consists of four light sources, namely, the reflected fringe pattern on the front surface (i.e., the binocular aliasing-free fringe pattern) F(t), the fringe pattern B(t) reflected once by the back surface, the fringe pattern M(t) after multiple reflections, and the background light A(t) captured by the camera. According to the basic principle of imaging, at time t, the fringe pattern captured by the camera (i.e., the binocular aliased fringe pattern) I(t) can be expressed as the sum of the above four components:
[0093] I(t) = F(t) + B(t) + M(t) + A(t);
[0094] Among them, the front surface image fringe pattern (i.e., the binocular aliasing-free fringe pattern) F(t) can be used to reconstruct the three-dimensional shape of the front surface of the transparent object, and F(t) can be expressed as:
[0095] F(t) = I(t) - B(t) - M(t) - A(t)
[0096] Then, a mapping network (i.e., the denoising U-Net network in the diffusion generation model) D is constructed. Through the network parameter θ, the fringe pattern (i.e., the binocular aliased fringe pattern) I(t) at time t is mapped into an aliasing-free front surface fringe image (i.e., the binocular aliasing-free fringe pattern) F(t). Then, the obtained F(t) can be used in the traditional binocular structured light three-dimensional reconstruction process to restore the three-dimensional shape information of the front surface of the transparent object. The formula is described as follows:
[0097]
[0098] Among them, D is the mapping network (i.e., the denoising U-Net network in the diffusion generation model); I is the aliased fringe image of the transparent object captured by the camera (i.e., the binocular aliased fringe pattern); θ is the network parameter of the network model; is the aliasing-free fringe image of the front surface of the transparent object predicted by the network.
[0099] Based on the above principle analysis, this embodiment presents a three-dimensional reconstruction method for a binocular structured light transparent object based on a diffusion model. Refer to Figure 1 , this three-dimensional reconstruction method for a binocular structured light transparent object includes:
[0100] Step S1, obtain a set of sequences of binocular aliased fringe patterns and corresponding sets of sequences of binocular aliasing-free fringe patterns for transparent object samples of different shapes.
[0101] In this embodiment, a variety of transparent objects with different shapes are used to construct a set of sequences of binocular aliased fringe patterns and corresponding sets of sequences of binocular aliasing-free fringe patterns for transparent object samples in a real environment, and they are divided into a training set and a test set.
[0102] In this embodiment, the binocular camera method is adopted to capture the phase-shifted fringe sequences on the surface of the transparent object sample through the left and right cameras; the phase extraction and phase unwrapping processing are performed on each group of phase-shifted fringe sequences to obtain the unwrapped phase diagrams under the left and right cameras respectively. Then, the binocular stereo matching algorithm is used, combined with the calibration parameters of the binocular cameras, to achieve high-precision reconstruction of the three-dimensional shape of the object.
[0103] The shape of the transparent object sample adopted in this embodiment is as Figure 2 shown. There are 10 shapes in total, and the surface stripe data sets (i.e., the binocular aliased stripe pattern sequence set and the corresponding binocular non-aliased stripe pattern sequence set) are constructed to enrich the shapes of the transparent objects as much as possible.
[0104] For each transparent object sample, different projector and camera poses are constructed. Under each constructed projector and camera pose, different-angle stripe patterns are projected onto the transparent sample before spraying the DPT-5 white layer and after spraying the DPT-5 white layer (i.e., the reflection layer). Each group of patterns collected by the camera is used as a group of data pairs. Image data is collected under multiple poses to complete the construction of the data set. There are 10 transparent samples used to construct the data set. For each transparent sample, 20 sets of camera and projector poses are set, and 60 sets of stripe patterns with different phases, different frequencies, and different angles are projected under each pose. Then, the data set totals 12,000 groups of data pairs. Among them, the data size for network input is (12,000, 3, 512, 512), and the label data size is (12,000, 3, 512, 512). 12,000 is the total number of constructed data pairs, 3 is the number of image channels, and 512 is the length and width size of the collected images.
[0105] As Figure 3 shown, the specific process of obtaining the binocular aliased stripe pattern sequence set of transparent object samples with different shapes and the corresponding reflected object samples in this embodiment includes:
[0106] Step S11: Obtain transparent object samples with different shapes, and adopt different poses to project a series of phase-shifted stripe patterns onto each shape of the transparent object sample. Then, use the left and right binocular cameras to collect images to obtain the binocular aliased stripe pattern sequence set of each shape of the transparent object sample under each pose.
[0107] In the step S11, at least one of the phase, frequency, and angle corresponding to the aliased stripe patterns in the binocular aliased stripe pattern sequence set is different;
[0108] Step S12: Spray the reflection layer on each shape of the transparent object sample to obtain the corresponding reflected object sample of each shape of the transparent object sample;
[0109] Step S13: After projecting a series of phase-shifted fringe patterns onto the reflective object samples of each shape with different poses, use left and right binocular cameras to collect images, so as to obtain a set of binocular aliasing-free fringe pattern sequences of the transparent object samples of each shape in each pose;
[0110] In the step S13, at least one of the phase, frequency, and angle corresponding to the binocular aliasing-free fringe pattern sequences in the set of binocular aliasing-free fringe pattern sequences is different.
[0111] Step S2: Divide the set of binocular aliased fringe pattern sequences and the corresponding set of binocular aliasing-free fringe pattern sequences into a training set of binocular aliased fringe pattern sequences and the corresponding training set of binocular aliasing-free fringe pattern sequences, as well as a test set of binocular aliased fringe pattern sequences and the corresponding test set of binocular aliasing-free fringe pattern sequences.
[0112] Divide the input sample data and label data according to a ratio of 7:3 to construct a training set and a test set. Among them, the total network input size of the training set is (8400, 3, 512, 512), and the total network label size of the training set is (8400, 3, 512, 512). The total network input size of the test set is (3600, 3, 512, 512), and the total network label size of the test set is (3600, 3, 512, 512).
[0113] Step S3: Construct a diffusion generative model; and use the training set of binocular aliased fringe pattern sequences and the corresponding training set of binocular aliasing-free fringe pattern sequences to train the denoising U-Net network in the diffusion generative model to obtain a trained denoising U-Net network.
[0114] In this embodiment, a diffusion generative model is built using the Pytorch deep learning framework. The core structure of the diffusion generative model includes a VAE encoder, a VAE decoder, and a denoising U-Net network. Usually, the training of the diffusion generative model is very resource-consuming. To improve the training efficiency, the Stable Diffusion v2 model pre-trained on the LAION-5B dataset is modified to make it applicable to the prediction task of the front surface aliasing-free fringes of the transparent objects in this embodiment. The VAE encoder and decoder parts are the same as the VAE in Stable Diffiusion v2. The denoising U-Net network is the same as the U-Net in StableDiffiusionv2, and only its parameters are fine-tuned.
[0115] In this embodiment, the processes of adding noise and denoising in the training of the denoising U-Net network and network inference (network parameter fine-tuning) are as follows:
[0116] During the network forward propagation noise addition process, given the initial sample distribution F0 = F, input the sample F, and obtain the noise-free aliasing stripe sample F(t) by gradually adding Gaussian noise n at different times t of the non-aliasing stripe sample F0. The noise-free aliasing stripe sample F(t) is:
[0117]
[0118] where represents Gaussian noise obeys a multivariate normal distribution with a mean of 0 and a covariance matrix of the identity matrix I; {β1,...,β t} are covariance scheduling coefficients with t steps.
[0119] During the network backward propagation denoising process, gradually denoise the noise-free aliasing stripe image through a parameterized denoising model, and gradually obtain F t-1 、F t-2 、... and finally obtain the noise-free non-aliasing stripe sample F0. Refer to Figure 4 The specific process of the training process of the denoising U-Net network includes:
[0120] Step S31, initialize the network parameters of the denoising U-Net network;
[0121] Step S32, set the serial number i = 1 in the binocular aliasing stripe pattern sequence training set and the corresponding binocular non-aliasing stripe pattern sequence training set;
[0122] The stripe pattern in this embodiment is denoted as (I, F), where I is the aliasing stripe image (i.e., the binocular aliasing stripe pattern) and F is the non-aliasing front surface stripe image (i.e., the binocular non-aliasing stripe pattern).
[0123] Step S33, use the VAE encoder in the diffusion generation model to encode the i-th binocular aliasing stripe pattern and the corresponding binocular non-aliasing stripe pattern into a first latent space tensor and a second latent space tensor respectively.
[0124] The first latent space tensor and the second latent space tensor in this embodiment are denoted as Z (I) and Z (F) respectively.
[0125] Step S34, randomly generate Gaussian noise; and add the Gaussian noise to the second latent space tensor at time t to generate the second latent space tensor with Gaussian noise at time t.
[0126] The second latent space tensor with Gaussian noise at time t in this embodiment
[0127] Step S35: After concatenating the first latent space tensor and the second latent space tensor with Gaussian noise, input them into the denoising U-Net network;
[0128] Step S36: The denoising U-Net network estimates the noise using the network parameters;
[0129] Step S37: Determine whether the loss value between the noise estimation result and the Gaussian noise is less than the threshold. If so, save the network parameters and proceed to Step S38; if not, update the network parameters and return to Step S36;
[0130] Step S38: Determine whether i is equal to I. If so, save the trained denoising U-Net network and end; if not, set i = i + 1 and return to Step S33;
[0131] Where I is the number of binocular aliased fringe pattern images in the binocular aliased fringe pattern sequence training set.
[0132] In this embodiment, by minimizing a denoising diffusion objective function, the estimated Gaussian noise is made as close as possible to the added Gaussian noise n. By calculating the loss between the two and performing backpropagation to update the network parameters θ.
[0133] Step S4: Use the binocular aliased fringe pattern sequence test set and the corresponding binocular non-aliased fringe pattern sequence test set to test the trained denoising U-Net network to obtain the tested denoising U-Net network.
[0134] Load the network parameters obtained from the training in Step S3, perform network inference, and continuously denoise the pure noise image using the trained denoising U-Net network D to reconstruct the noise-free and non-aliased fringe pattern. And use this model for quantitative evaluation on the test set.
[0135] Network inference (network parameter fine-tuning) process: Use the VAE encoder to encode the noise-free aliased fringe image (i.e., the binocular aliased fringe pattern) into a latent space tensor. Then, construct the noisy sample Z at time t t (F) , and concatenate Z t (F) with Z (I) and send them into the trained denoising U-Net network D to predict the Gaussian noise for this iteration Then, obtain the denoised sample Z at time t - 1 through the denoising module t-1 (F) , and then repeat the noise addition and denoising process to continuously obtain the denoised samples Z at times t - 2, t - 3,... t-2 (F),Z t-3 (F) ,...,Z0 (F) , and then use the VAE decoder to process the finally obtained Z0 (F) Decode the aliasing-free latent space stripes to obtain an aliasing-free front surface stripe image with the original size of 512×512.
[0136] Based on the above network inference, refer to Figure 5 , the specific process of network inference includes:
[0137] Step S41: Set the serial number j = 1 in the binocular aliasing stripe pattern sequence test set and the corresponding binocular aliasing-free stripe pattern sequence test set;
[0138] Step S42: Use the VAE encoder in the diffusion generation model to encode the j-th binocular aliasing stripe pattern into a third latent space tensor;
[0139] Step S43: Construct a noisy fourth latent space tensor of the j-th binocular aliasing-free stripe pattern at time t;
[0140] Step S44: After splicing the third latent space tensor and the fourth latent space tensor, input them into the trained denoising U-Net network for noise prediction;
[0141] Step S45: Denoise the noise prediction result to obtain a noisy fifth latent space tensor of the j-th binocular aliasing-free stripe pattern at time t-1;
[0142] Step S46: Determine whether t-1 is 0. If so, use the VAE decoder in the diffusion generation model to decode the fifth latent space tensor to obtain the j-th binocular aliasing-free stripe pattern, and enter Step S47; if not, after fine-tuning the network parameters, use the fifth latent space tensor as the fourth latent space tensor, set t = t-1, and return to Step S44;
[0143] Step S47: Determine whether j is equal to J. If so, save the tested denoising U-Net network and end; if not, set j = j + 1 and return to Step S42;
[0144] where J is the number of binocular aliasing stripe patterns in the binocular aliasing stripe pattern sequence test set.
[0145] Step S5: Obtain the binocular aliasing stripe pattern of the transparent object to be restored; and use the binocular aliasing stripe pattern of the transparent object to be restored and adopt the tested denoising U-Net network to perform three-dimensional reconstruction of the front surface of the transparent object.
[0146] In a real scenario, using a projector, a sequence of phase fringes is projected onto the surface of a transparent object with an unknown surface shape, and Gaussian noise is added to the aliased fringe images captured by a binocular camera. Then, the obtained noisy image group is fed into the denoising U-Net network after network inference, and finally, a group of front-surface reflection fringes without noise and aliasing is restored.
[0147] Phase extraction, phase unwrapping, and binocular stereo matching are respectively performed on the sequence of front-surface reflection fringe patterns restored under the left camera and the right camera to achieve high-precision reconstruction of the front surface of the transparent object with an unknown surface shape. Refer to Figure 6 , the specific process of three-dimensional reconstruction of the front surface of the transparent object includes:
[0148] Step S51: Use the tested denoising U-Net network to restore the binocular aliased fringe patterns of the transparent object to be restored, and obtain a sequence of binocular non-aliased fringe patterns of the transparent object to be restored;
[0149] Step S52: Successively perform phase extraction, phase unwrapping, and binocular stereo matching on the sequence of binocular non-aliased fringe patterns of the transparent object to be restored.
[0150] Based on the constructed set of binocular aliased fringe pattern sequences and the corresponding set of binocular non-aliased fringe pattern sequences of the transparent object samples in this embodiment, the denoising U-Net network in the constructed diffusion generation model is trained, realizing the restoration of the aliased phase-shifted fringes (i.e., binocular aliased fringe patterns) captured by the camera under the fringe projection method for the transparent object samples to the non-aliased phase-shifted fringes (i.e., binocular non-aliased fringe patterns) reflected from the front surface of the transparent object samples; and using the non-aliased phase-shifted fringes to achieve high-precision three-dimensional reconstruction of the surface of the transparent object, significantly improving the surface reconstruction accuracy of the transparent object; this embodiment can effectively extract the reflected fringe pattern information (i.e., binocular non-aliased fringe patterns) of the front surface of the transparent object from the binocular aliased fringe patterns captured by the camera, accurately restore the three-dimensional surface topography of the transparent object, realize the reconstruction of the surface of the transparent object, solve the problem that the phase profilometry cannot directly reconstruct the transparent object, and has broad application potential in related fields such as non-contact three-dimensional reconstruction and in-situ measurement of the surface of the transparent object.
[0151] The above embodiment can be implemented by the technical solution given in the following embodiment:
[0152] Another embodiment provides a binocular structured light three-dimensional reconstruction system for transparent objects based on a diffusion model. The binocular structured light three-dimensional reconstruction system for transparent objects includes:
[0153] An acquisition module, configured to acquire a set of binocular aliased fringe pattern sequences and the corresponding set of binocular non-aliased fringe pattern sequences of transparent object samples with different shapes;
[0154] A partitioning module, configured to partition the binocular aliased fringe pattern sequence set and the corresponding binocular non-aliased fringe pattern sequence set into a binocular aliased fringe pattern sequence training set and the corresponding binocular non-aliased fringe pattern sequence training set, as well as a binocular aliased fringe pattern sequence test set and the corresponding binocular non-aliased fringe pattern sequence test set;
[0155] A training module, configured to construct a diffusion generative model; and use the binocular aliased fringe pattern sequence training set and the corresponding binocular non-aliased fringe pattern sequence training set to train the denoising U-Net network in the diffusion generative model to obtain a trained denoising U-Net network;
[0156] A testing module, configured to use the binocular aliased fringe pattern sequence test set and the corresponding binocular non-aliased fringe pattern sequence test set to test the trained denoising U-Net network to obtain a tested denoising U-Net network;
[0157] A three-dimensional reconstruction module, configured to obtain the binocular aliased fringe pattern of the transparent object to be restored; and use the binocular aliased fringe pattern of the transparent object to be restored and adopt the tested denoising U-Net network to perform three-dimensional reconstruction of the front surface of the transparent object.
[0158] Further, the obtaining module includes:
[0159] A first projection acquisition sub-module, configured to obtain transparent object samples of different shapes, and project a series of phase-shifted fringe patterns onto each shape of transparent object sample in different poses, and then use left and right binocular cameras to perform image acquisition to obtain a binocular aliased fringe pattern sequence set of each shape of transparent object sample in each pose;
[0160] At least one of the phase, frequency, and angle of the aliased fringe patterns corresponding to the aliased fringe pattern sequence set in the binocular aliased fringe pattern sequence set is different;
[0161] A spraying sub-module, configured to spray a reflective layer on each shape of transparent object sample to obtain a corresponding reflective object sample for each shape of transparent object sample;
[0162] A second projection acquisition sub-module, configured to project a series of phase-shifted fringe patterns onto each shape of reflective object sample in different poses, and then use left and right binocular cameras to perform image acquisition to obtain a binocular non-aliased fringe pattern sequence set of each shape of transparent object sample in each pose;
[0163] At least one of the phase, frequency, and angle of the binocular non-aliased fringe pattern sequences corresponding to the binocular non-aliased fringe pattern sequence set in the binocular non-aliased fringe pattern sequence set is different.
[0164] Further, the training module includes:
[0165] An initialization sub-module for initializing the network parameters of the denoising U-Net network;
[0166] A first setting sub-module for setting the serial number i = 1 in the binocular aliased fringe pattern sequence training set and the corresponding binocular non-aliased fringe pattern sequence training set;
[0167] A first encoding sub-module for using the VAE encoder in the diffusion generative model to encode the i-th binocular aliased fringe pattern and the corresponding binocular non-aliased fringe pattern into a first latent space tensor and a second latent space tensor respectively;
[0168] An addition sub-module for randomly generating Gaussian noise; and adding the Gaussian noise to the second latent space tensor at time t to generate the second latent space tensor with Gaussian noise at time t;
[0169] A first splicing sub-module for splicing the first latent space tensor and the second latent space tensor with Gaussian noise and then inputting them into the denoising U-Net network;
[0170] A noise estimation sub-module for the denoising U-Net network to perform noise estimation using the network parameters;
[0171] A first judgment sub-module for judging whether the loss value between the noise estimation result and the Gaussian noise is less than a threshold. If so, saving the network parameters and transmitting them to the second judgment sub-module; if not, updating the network parameters and transmitting them to the noise estimation sub-module;
[0172] A second judgment sub-module for judging whether i is equal to I. If so, saving the trained denoising U-Net network and ending; if not, setting i = i + 1 and transmitting it to the first encoding sub-module;
[0173] Wherein, I is the number of binocular aliased fringe patterns in the binocular aliased fringe pattern sequence training set.
[0174] Furthermore, the test module includes:
[0175] A first setting sub-module for setting the serial number j = 1 in the binocular aliased fringe pattern sequence test set and the corresponding binocular non-aliased fringe pattern sequence test set;
[0176] A second encoding sub-module for using the VAE encoder in the diffusion generative model to encode the j-th binocular aliased fringe pattern into a third latent space tensor;
[0177] A construction sub-module for constructing a fourth latent space tensor with noise at time t for the j-th binocular non-aliased fringe pattern;
[0178] The second splicing sub-module is used to splice the third latent space tensor and the fourth latent space tensor and then input them into the trained denoising U-Net network for noise prediction;
[0179] The denoising processing sub-module is used to perform denoising processing on the noise prediction result to obtain the noisy fifth latent space tensor of the j-th binocular aliasing-free fringe pattern at time t-1;
[0180] The third judgment sub-module is used to judge whether t-1 is 0. If so, the VAE decoder in the diffusion generation model is used to decode the fifth latent space tensor to obtain the j-th binocular aliasing-free fringe pattern and transmit it to the fourth judgment sub-module; if not, after fine-tuning the network parameters, the fifth latent space tensor is used as the fourth latent space tensor, let t = t-1, and transmit it to the second splicing sub-module;
[0181] The fourth judgment sub-module is used to judge whether j is equal to J. If so, save the tested denoising U-Net network and end; if not, j = j + 1 and transmit it to the second encoding sub-module;
[0182] Where J is the number of binocular aliasing fringe patterns in the binocular aliasing fringe pattern sequence test set.
[0183] Furthermore, the three-dimensional reconstruction module includes:
[0184] The restoration sub-module is used to use the tested denoising U-Net network to restore the binocular aliasing fringe pattern of the transparent object to be restored to obtain a sequence of binocular aliasing-free fringe patterns of the transparent object to be restored;
[0185] The processing sub-module is used to perform phase extraction, phase unwrapping, and binocular stereo matching on the sequence of binocular aliasing-free fringe patterns of the transparent object to be restored in sequence.
[0186] The principles, formulas, and their parameter definitions involved in the above embodiments are all applicable and will not be elaborated here one by one.
[0187] The above embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A three-dimensional reconstruction method for binocular structured light transparent objects based on a diffusion model, characterized in that, The binocular structured light three-dimensional reconstruction method for transparent objects includes: Step S1: Obtain a set of binocular aliased fringe pattern sequences and corresponding sets of binocular non-aliased fringe pattern sequences for transparent object samples of different shapes; Step S2: Divide the set of binocular aliased fringe pattern sequences and the corresponding set of binocular non-aliased fringe pattern sequences into a training set of binocular aliased fringe pattern sequences and the corresponding training set of binocular non-aliased fringe pattern sequences, as well as a test set of binocular aliased fringe pattern sequences and the corresponding test set of binocular non-aliased fringe pattern sequences; Step S3: Construct a diffusion generative model; and use the training set of binocular aliased fringe pattern sequences and the corresponding training set of binocular non-aliased fringe pattern sequences to train the denoising U-Net network in the diffusion generative model to obtain a trained denoising U-Net network; Step S4: Use the test set of binocular aliased fringe pattern sequences and the corresponding test set of binocular non-aliased fringe pattern sequences to test the trained denoising U-Net network to obtain a tested denoising U-Net network; Step S5: Obtain the binocular aliased fringe pattern of the transparent object to be restored; and use the binocular aliased fringe pattern of the transparent object to be restored and adopt the tested denoising U-Net network to perform three-dimensional reconstruction of the front surface of the transparent object.
2. The three-dimensional reconstruction method of a binocular structured light transparent object according to claim 1, characterized in that, In the step S1, the specific process of obtaining the set of binocular aliased fringe pattern sequences of the transparent object samples of different shapes and the corresponding reflective object samples includes: Step S11: Obtain transparent object samples of different shapes, and project a series of phase-shifted fringe patterns onto each shape of transparent object sample with different poses, and then use left and right binocular cameras to collect images to obtain a set of binocular aliased fringe pattern sequences of each shape of transparent object sample in each pose; In the step S11, at least one of the phase, frequency, and angle corresponding to the aliased fringe patterns in the set of binocular aliased fringe pattern sequences is different; Step S12: Spray a reflective layer on each shape of transparent object sample to obtain the corresponding reflective object sample of each shape of transparent object sample; Step S13: Project a series of phase-shifted fringe patterns onto each shape of reflective object sample with different poses, and then use left and right binocular cameras to collect images to obtain a set of binocular non-aliased fringe pattern sequences of each shape of transparent object sample in each pose; In the step S13, at least one of the phase, frequency, and angle corresponding to the binocular non-aliased fringe pattern sequences in the set of binocular non-aliased fringe pattern sequences is different.
3. The three-dimensional reconstruction method of a binocular structured light transparent object according to claim 2, characterized in that, In the step S3, the specific process of the training includes: Step S31: Initialize the network parameters of the denoising U-Net network; Step S32: Set the serial number i = 1 in the training set of binocular aliased fringe pattern sequences and the corresponding training set of binocular non-aliased fringe pattern sequences; Step S33: Use the VAE encoder in the diffusion generative model to encode the i-th binocular aliased fringe pattern and the corresponding binocular non-aliased fringe pattern into a first latent space tensor and a second latent space tensor respectively; Step S34: Randomly generate Gaussian noise; and at time t, add the Gaussian noise to the second latent space tensor to generate the second latent space tensor with Gaussian noise at time t; Step S35: After splicing the first latent space tensor and the second latent space tensor with Gaussian noise, input them into the denoising U-Net network; Step S36: The denoising U-Net network estimates the noise using the network parameters; Step S37: Determine whether the loss value between the noise estimation result and the Gaussian noise is less than the threshold. If so, save the network parameters and proceed to Step S38; if not, update the network parameters and return to Step S36; Step S38: Determine whether i is equal to I. If so, save the trained denoising U-Net network and end; if not, set i = i + 1 and return to Step S33; where I is the number of binocular aliased fringe pattern sequences in the binocular aliased fringe pattern sequence training set.
4. The three-dimensional reconstruction method of a binocular structured light transparent object according to claim 3, wherein, In Step S4, the specific process of the test includes: Step S41: Set the serial number j = 1 in the binocular aliased fringe pattern sequence test set and the corresponding binocular non-aliased fringe pattern sequence test set; Step S42: Use the VAE encoder in the diffusion generative model to encode the j-th binocular aliased fringe pattern into a third latent space tensor; Step S43: Construct a fourth latent space tensor with noise at time t for the j-th binocular non-aliased fringe pattern; Step S44: After splicing the third latent space tensor and the fourth latent space tensor, input them into the trained denoising U-Net network for noise prediction; Step S45: Denoise the noise prediction result to obtain a fifth latent space tensor with noise at time t - 1 for the j-th binocular non-aliased fringe pattern; Step S46: Determine whether t - 1 is 0. If so, use the VAE decoder in the diffusion generative model to decode the fifth latent space tensor to obtain the j-th binocular non-aliased fringe pattern and proceed to Step S47; if not, after fine-tuning the network parameters, use the fifth latent space tensor as the fourth latent space tensor, set t = t - 1, and return to Step S44; Step S47: Determine whether j is equal to J. If so, save the tested denoising U-Net network and end; if not, set j = j + 1 and return to Step S42; where J is the number of binocular aliased fringe patterns in the binocular aliased fringe pattern sequence test set.
5. The three-dimensional reconstruction method of a binocular structured light transparent object according to claim 4, characterized in that In Step S5, the specific process of the three-dimensional reconstruction of the front surface of the transparent object includes: Step S51: Use the tested denoising U-Net network to restore the binocular aliased fringe pattern of the transparent object to be restored to obtain the binocular non-aliased fringe pattern of the transparent object to be restored; Step S52: Successively perform phase extraction, phase unwrapping, and binocular stereo matching on the binocular non-aliased fringe pattern sequence of the transparent object to be restored.
6. A three-dimensional reconstruction system for binocular structured light transparent objects based on a diffusion model, characterized in that, The binocular structured light transparent object three-dimensional reconstruction system includes: An acquisition module for acquiring a set of binocular aliased fringe pattern sequences and corresponding sets of binocular non-aliased fringe pattern sequences of transparent object samples with different shapes; A partitioning module, configured to partition the binocular aliased fringe pattern sequence set and the corresponding binocular non-aliased fringe pattern sequence set into a binocular aliased fringe pattern sequence training set and the corresponding binocular non-aliased fringe pattern sequence training set, as well as a binocular aliased fringe pattern sequence test set and the corresponding binocular non-aliased fringe pattern sequence test set; A training module, configured to construct a diffusion generative model; and use the binocular aliased fringe pattern sequence training set and the corresponding binocular non-aliased fringe pattern sequence training set to train the denoising U-Net network in the diffusion generative model to obtain a trained denoising U-Net network; A testing module, configured to use the binocular aliased fringe pattern sequence test set and the corresponding binocular non-aliased fringe pattern sequence test set to test the trained denoising U-Net network to obtain a tested denoising U-Net network; A three-dimensional reconstruction module, configured to obtain the binocular aliased fringe pattern of the transparent object to be restored; and use the binocular aliased fringe pattern of the transparent object to be restored and adopt the tested denoising U-Net network to perform three-dimensional reconstruction of the front surface of the transparent object.
7. The three-dimensional reconstruction system for a binocular structured light transparent object according to claim 6, wherein The obtaining module includes: A first projection and acquisition sub-module, configured to obtain transparent object samples of different shapes, and project a series of phase-shifted fringe patterns onto each shape of transparent object sample with different poses, and then use left and right binocular cameras to perform image acquisition to obtain a binocular aliased fringe pattern sequence set of each shape of transparent object sample in each pose; At least one of the phase, frequency, and angle corresponding to the aliased fringe patterns in the binocular aliased fringe pattern sequence set is different; A spraying sub-module, configured to spray a reflective layer on each shape of transparent object sample to obtain a corresponding reflective object sample for each shape of transparent object sample; A second projection and acquisition sub-module, configured to project a series of phase-shifted fringe patterns onto each shape of reflective object sample with different poses, and then use left and right binocular cameras to perform image acquisition to obtain a binocular non-aliased fringe pattern sequence set of each shape of transparent object sample in each pose; At least one of the phase, frequency, and angle corresponding to the binocular non-aliased fringe pattern sequences in the binocular non-aliased fringe pattern sequence set is different.
8. The three-dimensional reconstruction system for a binocular structured light transparent object according to claim 7, wherein The training module includes: An initialization sub-module, configured to initialize the network parameters of the denoising U-Net network; A first setting sub-module, configured to set the serial number i = 1 in the binocular aliased fringe pattern sequence training set and the corresponding binocular non-aliased fringe pattern sequence training set; A first encoding sub-module, configured to use the VAE encoder in the diffusion generative model to encode the i-th binocular aliased fringe pattern and the corresponding binocular non-aliased fringe pattern into a first latent space tensor and a second latent space tensor respectively; An adding sub-module, configured to randomly generate Gaussian noise; and add the Gaussian noise to the second latent space tensor at time t to generate a second latent space tensor with Gaussian noise at time t; A first splicing sub-module, configured to splice the first latent space tensor and the second latent space tensor with Gaussian noise and then input them into the denoising U-Net network; A noise estimation sub-module, which is used to denoise the U-Net network and estimate noise using the network parameters; A first judgment sub-module, which is used to judge whether the loss value between the noise estimation result and the Gaussian noise is less than a threshold. If so, save the network parameters and transmit them to the second judgment sub-module; if not, update the network parameters and transmit them to the noise estimation sub-module; A second judgment sub-module, which is used to judge whether i is equal to I. If so, save the trained denoising U-Net network and end; if not, i = i + 1 and transmit it to the first encoding sub-module; Where I is the number of binocular aliased fringe pattern in the binocular aliased fringe pattern sequence training set.
9. The three-dimensional reconstruction system for a binocular structured light transparent object according to claim 8, characterized in that, The test module includes: A first setting sub-module, which is used to set the serial number j = 1 in the binocular aliased fringe pattern sequence test set and the corresponding binocular non-aliased fringe pattern sequence test set; A second encoding sub-module, which is used to encode the j-th binocular aliased fringe pattern into a third latent space tensor using the VAE encoder in the diffusion generation model; A construction sub-module, which is used to construct a noisy fourth latent space tensor of the j-th binocular non-aliased fringe pattern at time t; A second splicing sub-module, which is used to splice the third latent space tensor and the fourth latent space tensor and then input them into the trained denoising U-Net network for noise prediction; A denoising processing sub-module, which is used to perform denoising processing on the noise prediction result to obtain a noisy fifth latent space tensor of the j-th binocular non-aliased fringe pattern at time t - 1; A third judgment sub-module, which is used to judge whether t - 1 is 0. If so, use the VAE decoder in the diffusion generation model to decode the fifth latent space tensor to obtain the j-th binocular non-aliased fringe pattern and transmit it to the fourth judgment sub-module; if not, after fine-tuning the network parameters, use the fifth latent space tensor as the fourth latent space tensor, set t = t - 1, and transmit it to the second splicing sub-module; A fourth judgment sub-module, which is used to judge whether j is equal to J. If so, save the tested denoising U-Net network and end; if not, j = j + 1 and transmit it to the second encoding sub-module; Where J is the number of binocular aliased fringe pattern in the binocular aliased fringe pattern sequence test set.
10. The three-dimensional reconstruction system for a binocular structured light transparent object according to claim 9, characterized in that, The 3D reconstruction module includes: A restoration sub-module, which is used to use the tested denoising U-Net network to restore the binocular aliased fringe pattern of the transparent object to be restored to obtain a sequence of binocular non-aliased fringe patterns of the transparent object to be restored; A processing sub-module, which is used to perform phase extraction, phase unwrapping, and binocular stereo matching on the sequence of binocular non-aliased fringe patterns of the transparent object to be restored in sequence.