High-speed three-dimensional reconstruction system and method based on two-stage collaborative coding

The high-speed 3D reconstruction system using a two-level collaborative coding approach, combined with a liquid crystal spatial light modulator and a digital micromirror device, adjusts coding parameters in real time, solving the problem of 3D reconstruction under conditions of high reflectivity, low reflectivity, and high-speed motion in traditional systems, and achieving high signal-to-noise ratio and high-precision 3D reconstruction results.

CN121026012BActive Publication Date: 2025-12-30WUHAN YILUT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511553979.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2025-12-30
Estimated Expiration
2045-10-29

AI Technical Summary

Technical Problem

Existing structured light 3D measurement systems struggle to achieve coordinated control of phase modulation and amplitude modulation when dealing with highly reflective metals, low-reflection black materials, complex textures, and high-speed moving objects. This results in the inability to adjust the coded pattern in real time, leading to local overexposure, decreased signal-to-noise ratio, and motion artifacts. Furthermore, the large data volume and high transmission pressure make it difficult to meet the requirements of high-speed 3D reconstruction.

Method used

A high-speed 3D reconstruction system based on two-level collaborative coding is adopted. By combining a parallel light source module, a liquid crystal spatial light modulator, a 4f relay system, a digital micromirror device, a projection lens, a sparse light field sampling module, and an edge computing device, the coding parameters are dynamically adjusted in real time to achieve collaborative control of phase modulation and amplitude modulation. Furthermore, the amount of data is reduced through sparse light field sampling and compressed sensing technology.

Benefits of technology

It achieves high signal-to-noise ratio and high accuracy in 3D reconstruction under high-speed motion and complex surface conditions, eliminates motion artifacts, reduces data transmission volume, and overcomes the problem of not being able to achieve speed, accuracy and data volume simultaneously, thus meeting the needs of industrial inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121026012B_ABST
    Figure CN121026012B_ABST
Patent Text Reader

Abstract

The application provides a high-speed three-dimensional reconstruction system and method based on two-stage cooperative coding, which comprises a parallel light source module, a liquid crystal spatial light modulator, a 4f relay system, a digital micromirror device, a projection lens, a sparse light field sampling module and an edge computing device; parallel light beams output by the parallel light source module are transmitted to the digital micromirror device for high-speed binary processing through the 4f relay system after phase modulation, so that a structured light field modulated in phase and amplitude is formed; the projection lens projects the modulated light field to a measured surface of a workpiece, and the reflected light field of the workpiece is captured by the sparse sampling module and then compressed; the edge computing device dynamically adjusts parameters of the liquid crystal spatial light modulator, the digital micromirror device and the sparse light field sampling module according to real-time feedback, so as to optimize the adaptability of the structured light field to the surface characteristics; finally, the surface topography is accurately reconstructed through a three-dimensional reconstruction network, and the core problem that speed, precision and data volume cannot be simultaneously achieved is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of structured light 3D measurement, and specifically relates to a high-speed 3D reconstruction system and method based on two-level cooperative coding. Background Technology

[0002] In the field of structured light 3D measurement, traditional DLP or single-channel digital micromirror device (DMD) projection systems have long been limited by fixed stripes and temporal phase-shifting strategies. This prevents real-time adjustment of encoding based on the reflectivity, texture, and motion state of the measured surface, leading to overexposure in highly reflective metal areas, low signal-to-noise ratio in low-reflectivity black rubber, and phase ambiguity on brushed or textured surfaces. Furthermore, multi-frame acquisition in high-speed motion scenarios introduces significant motion artifacts. While liquid crystal spatial light modulators (LC-SLMs) can provide high-fidelity positional modulation, their refresh rates are typically less than 100 Hz, and high-speed DMDs, although reaching tens of kilohertz, can only achieve binary amplitude output. The two technologies have not yet achieved effective synergy at the hardware level. Existing cascaded solutions are mostly used for static holographic displays or optical encryption, lacking a real-time closed-loop control mechanism for 3D measurement and making it difficult to achieve adaptive encoding optimization in dynamic scenes.

[0003] Meanwhile, high-resolution structured light systems generate massive amounts of raw data. A 4K pattern at a frame rate of thousands of frames per second can have an instantaneous bandwidth of tens of gigabytes per second, far exceeding the transmission capacity of conventional interfaces. Existing systems are often forced to sacrifice resolution or frame rate to alleviate data transmission pressure. While compressed sensing theory has been applied to single-pixel cameras and Fourier single-lens imaging, it is rarely combined with high-speed structured light measurement. Current liquid crystal spatial light modulator (LC-SLM) architectures are limited to single-pixel global masks, making it difficult to simultaneously meet the demands of large field of view and dense point cloud reconstruction, thus failing to satisfy the dual standards of industrial inspection for 3D point cloud quality and data volume.

[0004] Therefore, how to provide a high-speed 3D reconstruction system and method based on two-level cooperative coding, realize the cooperative control of phase modulation and amplitude modulation, and overcome the core problem that speed, accuracy and data volume cannot be achieved simultaneously, is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] Existing structured light 3D measurement systems generally exhibit three fundamental bottlenecks when dealing with highly reflective metals, low-reflective black materials, complex textures, and high-speed moving objects: First, a single spatial light modulator cannot simultaneously achieve high-fidelity positional modulation and a high refresh rate, resulting in the coded pattern failing to adjust in real time according to the surface BRDF and motion state, leading to local overexposure, decreased signal-to-noise ratio, and motion artifacts; Second, fixed temporal phase shifts require multiple frame acquisitions, resulting in long measurement cycles and large data redundancy, making it difficult for gigabit-level interfaces to handle the instantaneous tens of gigabytes of bandwidth generated by 4K×4K, kilohertz frame rates; Third, the traditional sampling-then-compression link already generates massive amounts of raw data at the sampling end, further exacerbating the storage and transmission load.

[0006] This invention provides a high-speed 3D reconstruction system and method based on two-level cooperative coding to solve at least one of the above-mentioned technical problems.

[0007] To address the aforementioned technical problems, in a first aspect, the present invention provides a high-speed 3D reconstruction system based on two-level cooperative coding, the system comprising:

[0008] The system comprises a parallel light source module, a liquid crystal spatial light modulator, a 4f relay system, a digital micromirror device (DMM), a projection lens, a sparse light field sampling module, and an edge computing device. The parallel light source module outputs a parallel light beam. The liquid crystal spatial light modulator performs phase modulation on the parallel light beam. The 4f relay system transmits the phase-modulated light field to the DMM without scaling. The DMM performs binary pulse width modulation on the phase-modulated light field within a microsecond-level time window. The projection lens focuses the binary pulse width modulated light field onto the surface of the workpiece to be tested, forming an emitted light field. The sparse light field sampling module collects the reflected light field from the surface of the workpiece and calculates a compressed measurement vector based on the reflected light field. The edge computing device synchronously controls the liquid crystal spatial light modulator, the DMM, and the sparse light field sampling module.

[0009] The edge computing device is configured to perform the following operations: initialize and calculate the initial phase map of the liquid crystal spatial light modulator; dynamically adjust the phase order increment of the liquid crystal spatial light modulator, the duty cycle matrix of the digital micromirror device, and the camera exposure time of the sparse light field sampling module in real time; input the compressed measurement vector and the current phase map into the end-to-end 3D reconstruction network model to obtain a depth map, wherein the current phase map is generated by the liquid crystal spatial light modulator dynamically adjusting the initial phase map according to the phase order increment.

[0010] Preferably, the parallel light source module includes an LED light source and a collimating lens, wherein the LED light source is used to emit a 450nm high-brightness scattered beam; and the collimating lens is used to shape the scattered beam into the parallel beam.

[0011] Preferably, the sparse light field sampling module is composed of a microlens array and a compressed sensing sensor cascaded together. The microlens array divides the reflected light field into multi-view sub-aperture images. The compressed sensing sensor uses a random Toeplitz measurement matrix to perform linear observation of the sub-aperture images to obtain a compressed measurement vector. The compressed measurement vector is represented as y = Φx + ε, where Φ is the random Toeplitz measurement matrix, x is the vectorization of the multi-view sub-aperture images, and ε is the measurement noise.

[0012] Preferably, the real-time dynamic adjustment of the phase order increment of the liquid crystal spatial light modulator, the duty cycle matrix of the digital micromirror device, and the camera exposure time of the sparse light field sampling module includes:

[0013] The physical property parameters of the surface of the workpiece to be detected are acquired in real time by the edge computing device to form an 8-dimensional continuous state vector. The physical property parameters include BRDF parameters, surface normal vector and motion velocity vector. The BRDF parameters include diffuse reflectance, specular reflectance and roughness.

[0014] The normalized state vector is input into the trained reinforcement learning policy network, and the action space is output. The action space includes a duty cycle matrix, a phase order increment, and a camera exposure time. The edge computing device sends the duty cycle matrix to the digital micromirror device, the phase order increment to the liquid crystal spatial light modulator, and the camera exposure time to the compressed sensing sensor, thereby synchronously triggering dynamic parameter adjustments.

[0015] Preferably, the 4f relay system consists of a first lens and a second lens, both of which have a focal length of 100mm and are plano-convex lenses with coaxial optical axes.

[0016] Preferably, the step of inputting the compressed measurement vector and the current phase map into the end-to-end 3D reconstruction network model to obtain a depth map includes:

[0017] An end-to-end 3D reconstruction network model based on an encoder / decoder architecture is constructed. The encoder adopts the 3DSwin-Transformer structure and uses a multi-level hierarchical window attention mechanism to jointly extract features from the compressed measurement vector after channel dimension concatenation and the current phase map, generating a hybrid feature representation containing global context information and local detail features.

[0018] The hybrid feature representation output by the encoder is fused across scales through adaptive weight allocation, and then mapped to the decoder input space by the dimension transformation module;

[0019] The decoder is configured with a dual-branch output structure. The first branch regresses and predicts the depth map based on the fused hybrid features; the second branch simultaneously estimates the pixel-level confidence map.

[0020] The depth map is adaptively optimized based on the estimated confidence map, and the optimized depth map is combined with the pre-calibrated sensor intrinsic and extrinsic parameter matrices. The depth values ​​in the pixel coordinate system are converted into a 3D point cloud in the camera coordinate system through inverse perspective projection transformation.

[0021] Preferably, the loss function of the end-to-end 3D reconstruction network model is constructed according to the following formula:

[0022] L = L_depth + γ·L_photo + η·L_sparse;

[0023] Where L_depth is the L1 depth error; L_photo is the image consistency calculation based on differentiable rendering; L_sparse is the point cloud sparsity regularization; γ is the weight controlling photometric consistency; and η is the strength of the sparsity constraint.

[0024] Preferably, the initial phase diagram calculation of the liquid crystal spatial light modulator includes:

[0025] The edge computing device controls the sparse light field sampling module to perform a millisecond-level flash pre-scan on the surface of the workpiece to be tested, acquire multiple low-dose images, and obtain the prior information of the workpiece to be tested based on the low-dose images. The prior information is encoded into an 8-bit phase map to obtain an initial phase map. The prior information is the physical characteristic parameters of the workpiece to be tested in the initial state.

[0026] In a second aspect, the present invention also provides a three-dimensional reconstruction method, applied to a three-dimensional reconstruction system as described in any of the first aspects, the method comprising:

[0027] S1. Initialize the phase diagram of the liquid crystal spatial light modulator and load the initial phase diagram into the liquid crystal spatial light modulator;

[0028] S2. Real-time acquisition of physical property parameters of the current surface of the workpiece to be inspected to form a state vector. The physical property parameters include BRDF parameters, surface normal vector and motion velocity vector. The BRDF parameters include diffuse reflectance, specular reflectance and roughness.

[0029] S3. Input the normalized state vector into the trained reinforcement learning policy network and output the action space, which includes the duty cycle matrix, phase order increment and camera exposure time. Send the duty cycle matrix to the digital micromirror device, send the phase order increment to the liquid crystal spatial light modulator, and send the camera exposure time to the compressed sensing sensor to synchronously trigger dynamic parameter adjustment.

[0030] S4. Acquire the reflected light field of the workpiece under test after adjustment based on the current parameters through a microlens array, and divide it into sub-aperture images. The compressed sensing sensor uses a random Toeplitz measurement matrix to perform linear observation of the sub-aperture images to obtain a compressed measurement vector.

[0031] S5. Input the compressed measurement vector and the current phase map into the end-to-end 3D reconstruction network model to obtain the depth map. The current phase map is generated by the liquid crystal spatial light modulator dynamically adjusting the initial phase map according to the phase order increment.

[0032] Preferably, the step of inputting the compressed measurement vector and the current phase map into the end-to-end 3D reconstruction network model to obtain a depth map includes:

[0033] An end-to-end 3D reconstruction network model based on an encoder / decoder architecture is constructed. The encoder adopts the 3DSwin-Transformer structure and uses a multi-level hierarchical window attention mechanism to jointly extract features from the compressed measurement vector after channel dimension concatenation and the current phase map, generating a hybrid feature representation containing global context information and local detail features.

[0034] The hybrid feature representation output by the encoder is fused across scales through adaptive weight allocation, and then mapped to the decoder input space by the dimension transformation module;

[0035] The decoder is configured with a dual-branch output structure. The first branch regresses and predicts the depth map based on the fused hybrid features; the second branch simultaneously estimates the pixel-level confidence map.

[0036] The depth map is adaptively optimized based on the estimated confidence map, and the optimized depth map is combined with the pre-calibrated sensor intrinsic and extrinsic parameter matrices. The depth values ​​in the pixel coordinate system are converted into a 3D point cloud in the camera coordinate system through inverse perspective projection transformation.

[0037] Beneficial Effects: This invention proposes a high-speed 3D reconstruction system based on two-level cooperative coding. The high-speed 3D reconstruction system provided in this embodiment includes: a parallel light source module, a liquid crystal spatial light modulator, a 4f relay system, a digital micromirror device (DMM), a projection lens, a sparse light field sampling module, and an edge computing device. The parallel light source module outputs a parallel light beam, the liquid crystal spatial light modulator performs phase modulation on the parallel light beam, the 4f relay system transmits the phase-modulated light field to the DMM without scaling, the DMM performs binary pulse width modulation on the phase-modulated light field within a microsecond-level time window to form a structured light field modulated by both phase and amplitude, and the projection lens focuses the binary pulse width modulated light field onto the target object. The workpiece surface generates an emitted light field, and the sparse light field sampling module collects the reflected light field from the workpiece surface and calculates a compressed measurement vector based on the reflected light field. An edge computing device is used to synchronously control the liquid crystal spatial light modulator, digital micromirror device, and sparse light field sampling module. The edge computing device is configured to perform the following operations: calculate the initial phase map of the liquid crystal spatial light modulator; dynamically adjust the phase order increment of the liquid crystal spatial light modulator, the duty cycle matrix of the digital micromirror device, and the camera exposure time of the sparse light field sampling module in real time; input the compressed measurement vector and the current phase map into an end-to-end 3D reconstruction network model to obtain a depth map. The current phase map is generated by the liquid crystal spatial light modulator dynamically adjusting the initial phase map based on the phase order increment. This application dynamically adjusts encoding parameters based on real-time feedback using edge computing devices. Specifically, it balances projection brightness and modulation speed through duty cycle matrix adjustment, optimizes the adaptability of the structured light field to surface characteristics through phase order incremental adjustment, and dynamically matches surface reflection characteristics through exposure time to achieve coordinated control of phase modulation and amplitude modulation. The modulated structured light field can be adjusted in real time according to the physical characteristics and motion state of the workpiece surface, improving the signal-to-noise ratio while eliminating motion artifacts. The application also compresses the reflected light field projected onto the measured surface of the workpiece through a sparse sampling module to effectively reduce data transmission volume. Finally, it fuses compressed data with dynamic phase information through a 3D reconstruction network to achieve accurate reconstruction of the surface morphology, overcoming the core challenge of simultaneously achieving speed, accuracy, and data volume. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0039] Figure 1 A schematic diagram of the structure of the high-speed three-dimensional reconstruction system provided in this application;

[0040] Figure 2 A flowchart of the three-dimensional reconstruction method provided in this application;

[0041] Figure 3 A three-dimensional reconstruction image of a complex metal surface based on the traditional fixed stripe method provided for this application;

[0042] Figure 4 The complex metal surface provided in this application is based on the three-dimensional reconstruction image obtained in this application.

[0043] Attached image captions:

[0044] 11. LED light source; 12. Collimating lens; 2. Liquid crystal spatial light modulator; 3. 4f relay system; 4. Digital micromirror device; 5. Projection lens; 61. Microlens array; 62. Compressed sensing sensor; 7. Edge computing device; 8. Workpiece to be inspected. Detailed Implementation

[0045] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0046] Example 1

[0047] like Figure 1As shown, this embodiment provides a high-speed 3D reconstruction system based on two-level cooperative coding. The high-speed 3D reconstruction system provided in this embodiment includes: a parallel light source module, a liquid crystal spatial light modulator (LC-SLM), a 4f relay system, a digital micromirror device (DMD), a projection lens, a sparse light field sampling module, and an edge computing device. The parallel light source module is used to output a collimated parallel beam, providing a stable light source for subsequent modulation stages. The liquid crystal spatial light modulator is used to perform pixel-level phase modulation on the parallel beam. The 4f relay system transmits the phase-modulated light field to the digital micromirror device without scaling. The digital micromirror device performs binary pulse width modulation on the phase-modulated light field within a microsecond-level time window. The projection lens transmits the binary pulse width modulated light... The field is focused onto the surface of the workpiece to be inspected to form an emitted light field. The sparse light field sampling module collects the reflected light field from the surface of the workpiece and calculates the compressed measurement vector based on the reflected light field. The edge computing device is used to synchronously control the liquid crystal spatial light modulator, the digital micromirror device, and the sparse light field sampling module. The edge computing device is configured to perform the following operations: initialize and calculate the initial phase map of the liquid crystal spatial light modulator; dynamically adjust the phase order increment of the liquid crystal spatial light modulator, the duty cycle matrix of the digital micromirror device, and the camera exposure time of the sparse light field sampling module in real time; input the compressed measurement vector and the current phase map into the end-to-end 3D reconstruction network model to obtain the depth map. The current phase map is generated by the liquid crystal spatial light modulator dynamically adjusting the initial phase map according to the phase order increment.

[0048] Specifically, during system operation, a parallel beam is phase-modulated to form a structured light field, which is then losslessly transmitted to a digital micromirror device via a 4f relay system for high-speed binarization, resulting in a phase- and amplitude-modulated structured light field. The projection lens projects the modulated light field onto the surface under test, and the reflected light field is captured by a sparse sampling module and compressed to effectively reduce data transmission. The edge computing device dynamically adjusts the encoding parameters based on real-time feedback, that is, it balances the projection brightness and modulation speed through duty cycle matrix adjustment, optimizes the adaptability of the structured light field to surface characteristics through phase order increment adjustment, and dynamically matches the surface reflection characteristics through exposure time. Finally, the compressed data and dynamic phase information are fused through a 3D reconstruction network to achieve accurate reconstruction of the surface morphology.

[0049] In this application, a phase-based LC-SLM and a binary amplitude-based DMD are cascaded to form a two-stage cooperative coding module. The first stage uses Gerchberg–Saxton iteration to synthesize a continuous phase distribution φ(u,v) within the target depth z ∈ [z_min, z_max], resulting in uniform irradiance of the outgoing light field in both the lateral and axial directions, thus providing a low-distortion, high-signal-to-noise ratio "basis pattern" for the first frame projection. This phase map is transmitted to the second-stage DMD without scaling via a 4f relay system; the latter performs binary pulse width modulation (BPWM) on the basis pattern within a microsecond-level time window. Specifically, the DMD modulates the light field amplitude of each pixel to a duty cycle d_i ∈ [0,1], and uses the local duty cycle to compensate for the irradiance difference caused by the non-uniform reflection of the measured surface ρ(x,y) in a very short time, achieving pixel-level light intensity / phase fine-tuning, and finally outputting a spatiotemporally joint modulation pattern.

[0050] As one possible approach, the parallel light source module includes an LED light source and a collimating lens. The LED light source is used to emit a 450nm high-brightness scattered beam; the collimating lens is used to shape the scattered beam into a parallel beam, thereby providing a highly consistent optical field input for subsequent phase modulation.

[0051] Among them, the LED light source is a solid-state light source that uses light-emitting diodes as light-emitting elements. Specifically, a gallium nitride-based blue light chip combined with a fluorescence conversion layer can achieve high brightness output in the 450nm wavelength band. This wavelength selection takes into account both the need to resist ambient light interference in industrial environments and the transmission efficiency of optical devices. The collimating lens is a refractive optical element with a specific curved surface structure. Specifically, an aspherical lens design can be used to achieve beam collimation. By precisely calculating the lens curvature, the Lambertian divergence characteristics of the LED light source are eliminated, converting the scattered beam into a spatially uniform parallel light field.

[0052] As one feasible approach, the sparse light field sampling module consists of a cascaded microlens array and a compressed sensing sensor. The microlens array segments the reflected light field into multi-view sub-aperture images. The compressed sensing sensor uses a random Toeplitz measurement matrix to perform linear observations of the sub-aperture images, obtaining a compressed measurement vector. This compressed measurement vector is represented as y = Φx + ε, where Φ is the random Toeplitz measurement matrix, x is the vectorization of the multi-view sub-aperture images, and ε is the measurement noise. A 12 MB compressed measurement vector y can be obtained in a single frame.

[0053] The microlens array can be an optical element composed of multiple microlenses arranged in a regular pattern, specifically a hexagonal close-packed plano-convex lens array. Each microlens corresponds to a sub-aperture channel, used to divide the incident light field into multiple viewpoint-independent sub-light fields. The compressed sensing sensor is an imaging device capable of data compression through linear projection. It can be implemented using a single-pixel camera architecture based on digital micromirror devices, generating the measurement matrix by programming the flipping mode of the micromirror array. The random Toeplitz measurement matrix is ​​a structured random matrix with translational symmetry. It can be implemented by cyclically shifting binary pseudo-random sequences to generate row elements, with adjacent rows having a fixed offset correlation, reducing storage resource consumption and accelerating matrix operations. The compressed measurement vector is a low-dimensional data representation obtained through linear observation. It can be implemented by using a random Toeplitz matrix to compress and sample multi-view images, effectively reducing data transmission volume.

[0054] Specifically, the sparse light field sampling module first rearranges the light field of the projection lens imaging surface into a multi-view sub-aperture image using a microlens array (MLA). Where L is the number of angular channels, each sub-aperture corresponds to a local area of ​​the object's surface and a specific viewing angle, forming a multi-angle observation data set. Subsequently, the compressed sensing sensor utilizes a random Toeplitz measurement matrix. Linear observations are performed on each sub-aperture image to obtain a compressed measurement vector y = Φx + ε. The measurement vector y is concatenated with the current SLM coding pattern in the channel dimension, forming the input tensor of the 3D reconstruction network. Due to the cyclic shift property of the Toeplitz matrix, measurement vectors in different rows can be quickly generated by cyclic shifting a single random sequence, eliminating the need to store the complete matrix and reducing hardware implementation complexity. Parallel compressed observations of multi-view sub-aperture images can compress the original data volume proportionally to the number of rows in the measurement matrix while maintaining spatial resolution, alleviating data transmission bandwidth pressure.

[0055] One possible approach is to dynamically adjust the phase order increment of the liquid crystal spatial light modulator, the duty cycle matrix of the digital micromirror device, and the camera exposure time of the sparse light field sampling module in real time, including the following:

[0056] The physical property parameters of the workpiece surface to be inspected are acquired in real time by the edge computing device, forming an 8-dimensional continuous state vector. The physical property parameters include BRDF parameters, surface normal vector and motion velocity vector. The BRDF parameters include diffuse reflectance, specular reflectance and roughness.

[0057] The normalized state vector is input into the trained reinforcement learning policy network, and the output action space includes the duty cycle matrix, phase order increment and camera exposure time. The edge computing device sends the duty cycle matrix to the digital micromirror device, the phase order increment to the liquid crystal spatial light modulator, and the camera exposure time to the compressed sensing sensor, and synchronously triggers the dynamic adjustment of parameters.

[0058] The state vector is an 8-dimensional continuous state vector, a multi-dimensional data set composed of the BRDF parameters, surface normal vector, and motion velocity vector of the workpiece surface. It is used to characterize the optical properties, geometric morphology, and dynamic characteristics of the workpiece surface in real time. The 8-dimensional continuous state vector s can be represented as: s = (ρ_d, ρ_s, α, n_x, n_y, n_z, v_x, v_y). The reinforcement learning policy network is a policy model trained offline, specifically implemented using a deep deterministic policy gradient algorithm, used to map the state vector to the control parameter space. The duty cycle matrix is ​​the on-time ratio of each micromirror unit in the digital micromirror device, specifically generated by a pulse width modulation signal, used to adjust the spatiotemporal coding distribution of the projected light field. The phase order increment is the phase change between adjacent refresh cycles of the liquid crystal spatial light modulator, specifically achieved by adjusting the orientation of liquid crystal molecules through a voltage-driven circuit, used to dynamically optimize wavefront modulation accuracy. The camera exposure time refers to the integration time of the compressed sensing sensor, specifically adjusted by the electronic shutter control module, used to match the workpiece movement speed to avoid motion blur.

[0059] Specifically, during system operation, physical characteristic parameters are acquired in real time and converted into 8-dimensional state vectors, which are then normalized and input into the reinforcement learning policy network. This network generates combined parameters of duty cycle matrix, phase increment, and exposure time through online inference, and distributes them synchronously to each hardware module via edge computing devices. The phase increment controls the dynamic phase adjustment of the liquid crystal spatial light modulator, adapting the light field modulation mode to changes in the reflectivity of the workpiece surface; the duty cycle matrix adjusts the binary encoding distribution of the digital micromirror device, optimizing the spatiotemporal energy distribution of the projected pattern; and the camera exposure time is dynamically matched to the workpiece's motion speed, ensuring that the temporal resolution of the sampling process is consistent with the motion state. Through a closed-loop feedback mechanism, the coordinated adjustment of these three parameters can suppress overexposure in highly reflective areas in real time, enhance the signal strength in low-reflective areas, and eliminate image blurring caused by high-speed motion.

[0060] The dynamic parameter adjustment process can specifically employ a Markov decision process (S,A,R). The state space S consists of the BRDF parameters of the current surface region, the surface normal vector n, and the workpiece velocity v, forming an 8-dimensional continuous observation. The action space A includes the local duty cycle matrix of the DMD, the phase order increment of the LC-SLM, and the camera exposure time, forming a 3-dimensional discrete-continuous hybrid action. The reward function R is designed as a weighted difference between the negative logarithm of the reconstruction error and the projected luminous flux. λ is the energy penalty coefficient, and E_rec is the reconstruction error. The policy network π(a|s) is implemented using Proximal Policy Optimization (PPO) and updated online at a 1 ms cycle on the GPU of the edge computing device to ensure that the closed-loop bandwidth matches the production line cycle time.

[0061] As one feasible approach, a 4f relay system consists of a first lens and a second lens, both with a focal length of 100mm. Both lenses are coaxial plano-convex lenses to achieve non-scaling transmission of the light field, avoiding image distortion or resolution loss caused by scaling operations in traditional relay systems. Correspondingly, the collimating lens has a focal length of f=75mm, and the projection lens has a focal length of f=35mm.

[0062] As one possible approach, the depth map is obtained by inputting the compressed measurement vector into the current phase map into the end-to-end 3D reconstruction network model, specifically including the following steps (1)-(4):

[0063] (1) Construct an end-to-end 3D reconstruction network model based on encoder / decoder architecture, wherein the encoder adopts 3DSwin-Transformer structure, and performs joint feature extraction on the compressed measurement vector after channel dimension splicing and the current phase map through multi-level hierarchical window attention mechanism to generate a hybrid feature representation containing global context information and local detail features;

[0064] The 3D Swin-Transformer structure extends the two-dimensional window attention mechanism to three-dimensional space within a neural network module. Specifically, it uses a cubic window partitioning method to process spatiotemporal data and achieves cross-window interaction through shift window operations. This structure can capture the spatial correlation features between compressed measurement vectors and phase maps. The multi-level hierarchical window attention mechanism sets attention windows of different sizes at different network levels. Specifically, it uses progressively smaller window sizes to achieve feature focusing from global to local, effectively balancing the breadth and accuracy of feature extraction.

[0065] (2) The hybrid feature representation output by the encoder is fused across scales through adaptive weight allocation, and then mapped to the decoder input space through the dimension transformation module;

[0066] Specifically, adaptive weight allocation can generate weight coefficients by combining channel attention mechanism and spatial attention mechanism, thereby optimizing the integration effect of features at different resolutions.

[0067] (3) The decoder is set to a dual-branch output structure. The first branch predicts the depth map based on the fused hybrid features regression; the second branch simultaneously estimates the pixel-level confidence map; whereby the confidence map is used to quantify the confidence level of the depth prediction.

[0068] (4) Based on the estimated confidence map, the depth map is adaptively optimized, and the optimized depth map is combined with the pre-calibrated sensor intrinsic and extrinsic matrix. The depth values ​​in the pixel coordinate system are converted into three-dimensional point clouds in the camera coordinate system through perspective projection inverse transformation.

[0069] Specifically, the compressed measurement vector and the current phase map are concatenated along the channel dimension and then input into the encoder. The 3D Swin-Transformer structure divides the input data into multiple sub-cubes using cubic windows, performing self-attention computation within each sub-cube to establish cross-modal associations. The hierarchical window mechanism uses large windows to capture the global context in shallow layers and switches to small windows to focus on local details in deeper layers, achieving cross-regional information interaction through window shifting operations. The multi-scale features output by the encoder are processed by an adaptive weight module to calculate the contribution weights of each layer's features, preserving high-frequency details while suppressing redundant information. The decoder receives the fused features, and the first branch gradually restores the spatial resolution and regresses the depth value through deconvolution operations. The second branch calculates the confidence map for each pixel using the same features. Low-confidence regions are given higher smoothness constraints during the optimization phase. Finally, the 3D point cloud is mapped from 2D to 3D coordinates by combining the depth map with the camera's intrinsic and extrinsic parameter matrices and utilizing inverse perspective transformation. The network completes one forward inference cycle every 8 ms at FP16 accuracy, meeting real-time requirements.

[0070] As one possible approach, the loss function for an end-to-end 3D reconstruction network model can be constructed based on the following formula:

[0071] L = L_depth + γ·L_photo + η·L_sparse;

[0072] Where L_depth is the L1 depth error; L_photo is the image consistency calculation based on differentiable rendering; L_sparse is the point cloud sparsity regularization; γ is the weight that controls photometric consistency; and η is the strength of the sparsity constraint that adjusts the network to maintain geometric smoothness in textureless regions.

[0073] Specifically, this loss function achieves a comprehensive improvement in 3D reconstruction performance through a multi-objective joint optimization mechanism. The L1 depth error is calculated using the first-order norm to determine the absolute deviation between the predicted depth value and the true depth value. Its role is to suppress the interference of outliers on model training and enhance the robustness of depth estimation.

[0074] As one possible approach, the initial phase map of the computational liquid crystal spatial light modulator is initialized by the following steps: the edge computing device controls the sparse light field sampling module to perform a millisecond-level flash pre-scan on the surface of the workpiece to be inspected, acquires multiple low-dose images, and obtains the prior information of the workpiece to be inspected based on the low-dose images. The prior information is encoded into an 8-bit phase map to obtain the initial phase map. The prior information is the physical characteristic parameters of the workpiece to be inspected in the initial state.

[0075] Specifically, during system startup, a low-resolution image of the workpiece surface is rapidly acquired by triggering a millisecond-level flash pre-scan. Image processing algorithms are then used to extract the surface reflectivity distribution and roughness features, forming an initial set of physical property parameters. This parameter set is then input into the encoder, and an 8-bit phase map matching the resolution of the liquid crystal spatial light modulator is generated through grayscale mapping. This phase map serves as the reference for subsequent dynamic adjustments; its quantization level has a non-linear relationship with the surface reflectivity, ensuring that the initial modulation parameters can adapt to the differences in the optical properties of the workpiece surface.

[0076] Among them, millisecond-level flash pre-scanning refers to the process of rapidly acquiring optical information of the workpiece surface through short exposure. By controlling the exposure time to the millisecond level, low-latency data acquisition is achieved, avoiding the system response lag caused by traditional high-resolution scanning. Low-dose images refer to sparse image data acquired by reducing sampling resolution or integration time. Specifically, this can be achieved using the random sampling mode of a compressed sensing sensor. This reduces the transmission bandwidth requirement by reducing the amount of data per frame, while retaining key information about surface reflection characteristics. Physical characteristic parameters are quantitative indicators describing the optical properties of the workpiece surface. Specifically, they can be extracted from low-dose images using image inversion algorithms, providing initial conditions that match the surface reflection characteristics for phase modulation. The 8-bit phase map discretizes continuous physical parameters into a modulation spectrum of 256 phase levels, and generates an initialization input compatible with the liquid crystal spatial light modulator driving circuit through quantization processing.

[0077] This invention relates to computational imaging, structured light 3D measurement, light field manipulation and compressed sensing technologies, and is particularly suitable for industrial online inspection, topography measurement of high-speed moving objects, and high-precision 3D reconstruction of highly reflective / low-reflective / complex textured surfaces.

[0078] Taking an automated production line for automotive engine cylinder blocks as an example: The PLC sends a "workpiece in place" signal every 2 seconds, and the system immediately starts: a 50 ms pre-scan first uses low-power stripes to quickly establish a coarse BRDF model of the aluminum alloy-cast iron hybrid surface. Then, the PPO reinforcement learning network calculates the DMD local duty cycle matrix within 1 ms and simultaneously refreshes the LC-SLM phase. The 450 nm blue LED is exposed for only 1 ms, and the CS-CMOS outputs 12 MB compressed sampling values ​​at 1000 fps. The Jetson AGX Orin completes the Transformer network inference in 8 ms with FP16 precision, generating a dense point cloud of 3.3 M points. The edge RMS error is reduced from 85 µm in the traditional 3-step phase shift method to 21 µm, and the missing rate is reduced from 23% to <4%. The entire measurement-detection process is completed within 120 ms, and real-time transmission can be achieved via gigabit network. If a high-curvature black rubber sealing ring is used, the algorithm automatically reduces the DMD. The duty cycle is increased and the exposure time is extended to maintain the signal-to-noise ratio. For fan blades rotating at a tangential speed of 0.5 m / s, the external encoder triggers the DMD duty cycle to be adjusted in real time according to the angle, achieving 360° full-circumference measurement without blind spots.

[0079] Appendix Figure 3 This is a 3D reconstruction image of a complex metal surface based on the traditional fixed-fringe method. In this method, the exposure time is 2 ms, the phase shift is 3 steps, the data size is 72 MB, the reconstructed point cloud has 3.1 M points, and the missing rate in highly reflective areas is 23%. (Attached) Figure 4 For the 3D reconstruction image of a complex metal surface based on the 3D reconstruction system provided in this application, the exposure time is 1ms, the data volume is only 12 MB for single-frame encoding, the point cloud has 3.3 M points, the missing rate is <4%, and the edge RMS error is reduced from 85 μm to 21 μm.

[0080] Example 2

[0081] This invention also provides a three-dimensional reconstruction method, applied to the three-dimensional reconstruction system as described in any one of the embodiments, the method comprising the following steps S1-S5:

[0082] S1. Initialize the phase diagram of the liquid crystal spatial light modulator and load the initial phase diagram into the liquid crystal spatial light modulator;

[0083] The initial phase map is generated through pre-scanning. Specifically, it can be achieved by using millisecond-level flash pre-scanning to acquire low-dose images and encoding them into an 8-bit phase map, providing the system with an initial encoding reference adapted to the surface characteristics of the object under test. The specific generation process is as follows: the edge computing device controls the sparse light field sampling module to perform a millisecond-level flash pre-scan on the surface of the workpiece to be tested, acquiring multiple low-dose images. Based on the low-dose images, the prior information of the workpiece to be tested is obtained, and the prior information is encoded into an 8-bit phase map to obtain the initial phase map. The prior information consists of the physical characteristic parameters of the workpiece to be tested in its initial state.

[0084] S2. Real-time acquisition of the physical property parameters of the current surface of the workpiece to be inspected, forming a state vector. The physical property parameters include BRDF parameters, surface normal vector and motion velocity vector. The BRDF parameters include diffuse reflectance, specular reflectance and roughness.

[0085] Among them, the state vector is an 8-dimensional continuous state vector, which is a multi-dimensional data set composed of the BRDF parameters, surface normal vector and motion velocity vector of the workpiece surface to be detected. It is used to characterize the optical properties, geometric shape and dynamic characteristics of the workpiece surface in real time. The 8-dimensional continuous state vector s can be expressed as: s=(ρ_d, ρ_s, α, n_x, n_y, n_z, v_x, v_y).

[0086] Furthermore, the BRDF parameters and normal vectors can be obtained by combining the reflection distribution acquired by the camera with the known incident direction of the light source through network inversion, while the motion speed often depends on auxiliary sensors such as IMU or is estimated through inter-frame matching of video. Therefore, the state space is composed of optical acquisition results and auxiliary sensing information.

[0087] S3. Input the normalized state vector into the trained reinforcement learning policy network and output the action space, which includes the duty cycle matrix, phase order increment and camera exposure time. Send the duty cycle matrix to the digital micromirror device, the phase order increment to the liquid crystal spatial light modulator, and the camera exposure time to the compressed sensing sensor to synchronously trigger dynamic parameter adjustment.

[0088] The reinforcement learning policy network is a policy model trained offline, specifically implemented using a deep deterministic policy gradient algorithm, used to map the state vector to the control parameter space. The duty cycle matrix represents the on-time ratio of each micromirror unit in the digital micromirror device, specifically generated by a pulse width modulation signal, used to adjust the spatiotemporal coding distribution of the projected light field. The phase order increment is the phase change between adjacent refresh cycles of the liquid crystal spatial light modulator, specifically achieved by adjusting the orientation of liquid crystal molecules through a voltage-driven circuit, used to dynamically optimize wavefront modulation accuracy. The camera exposure time refers to the integration time of the compressed sensing sensor, specifically adjusted by the electronic shutter control module, used to match the workpiece movement speed to avoid motion blur.

[0089] Specifically, during system operation, physical characteristic parameters are acquired in real time and converted into 8-dimensional state vectors, which are then normalized and input into the reinforcement learning policy network. This network generates combined parameters of duty cycle matrix, phase increment, and exposure time through online inference, and distributes them synchronously to each hardware module via edge computing devices. The phase increment controls the dynamic phase adjustment of the liquid crystal spatial light modulator, adapting the light field modulation mode to changes in the reflectivity of the workpiece surface; the duty cycle matrix adjusts the binary encoding distribution of the digital micromirror device, optimizing the spatiotemporal energy distribution of the projected pattern; and the camera exposure time is dynamically matched to the workpiece's motion speed, ensuring that the temporal resolution of the sampling process is consistent with the motion state. Through a closed-loop feedback mechanism, the coordinated adjustment of these three parameters can suppress overexposure in highly reflective areas in real time, enhance the signal strength in low-reflective areas, and eliminate image blurring caused by high-speed motion.

[0090] S4. Acquire the reflected light field of the workpiece under test after adjustment based on the current parameters through a microlens array, and divide it into sub-aperture images. The compressed sensing sensor uses a random Toeplitz measurement matrix to perform linear observation of the sub-aperture images to obtain the compressed measurement vector.

[0091] The compressed measurement vector can be represented as y = Φx + ε, where Φ is the random Toeplitz measurement matrix, x is the vectorization of the multi-view sub-aperture image, and ε is the measurement noise.

[0092] S5. Input the compressed measurement vector and the current phase map into the end-to-end 3D reconstruction network model to obtain the depth map. The current phase map is generated by the liquid crystal spatial light modulator dynamically adjusting the initial phase map according to the phase order increment.

[0093] As one possible approach, step S5 above, which involves inputting the compressed measurement vector into the current phase map of the end-to-end 3D reconstruction network model to obtain a depth map, specifically includes the following:

[0094] An end-to-end 3D reconstruction network model based on an encoder / decoder architecture is constructed. The encoder adopts the 3DSwin-Transformer structure and uses a multi-level hierarchical window attention mechanism to jointly extract features from the compressed measurement vector after channel dimension concatenation and the current phase map, generating a hybrid feature representation containing global context information and local detail features.

[0095] The hybrid feature representation output by the encoder is fused across scales through adaptive weight allocation, and then mapped to the decoder input space by the dimension transformation module;

[0096] The decoder is configured with a dual-branch output structure. The first branch regresses and predicts the depth map based on the fused hybrid features; the second branch simultaneously estimates the pixel-level confidence map.

[0097] The depth map is adaptively optimized based on the estimated confidence map, and the optimized depth map is combined with the pre-calibrated sensor intrinsic and extrinsic parameter matrices. The depth values ​​in the pixel coordinate system are converted into a 3D point cloud in the camera coordinate system through inverse perspective projection transformation.

[0098] As one possible approach, after obtaining the depth map in step S5 above, this application performs the following operations based on the depth map obtained in the current frame:

[0099] The depth map is spatially registered and numerically compared with the preset offline CAD model and / or the reconstruction results of the preceding frame, and the root mean square error is calculated as a quantitative evaluation index.

[0100] The root mean square error is converted into an instant reward signal and input into the reinforcement learning policy network. The network parameters of the policy network are then updated and optimized online through the backpropagation algorithm.

[0101] The updated strategy network is directly deployed to the processing flow of the next frame of data, forming a closed-loop control loop with dynamic parameter adjustment.

[0102] The closed-loop control process runs synchronously at a clock frequency of 1000Hz, enabling the strategy network to have sub-millisecond real-time adaptive capabilities. This allows it to maintain an edge reconstruction accuracy of around 21μm under high-speed production line conditions and ensure that the data loss rate in the compressed sensing process remains stable below 4%.

[0103] The loss function for the end-to-end 3D reconstruction network model is constructed based on the following formula:

[0104] L = L_depth + γ·L_photo + η·L_sparse;

[0105] Where L_depth is the L1 depth error, L_photo is the image consistency calculated based on differentiable rendering, and L_sparse is the point cloud sparsity regularization term.

[0106] Where L_depth is the depth error; L_photo is the image consistency calculation based on differentiable rendering; L_sparse is the point cloud sparsity regularization; γ is the weight that controls the photometric consistency; and η is the force that adjusts the sparsity constraint to make the network maintain geometric smoothness in textureless regions.

[0107] Specifically, this loss function achieves a comprehensive improvement in 3D reconstruction performance through a multi-objective joint optimization mechanism. The L1 depth error is calculated using the first-order norm to determine the absolute deviation between the predicted depth value and the true depth value. Its role is to suppress the interference of outliers on model training and enhance the robustness of depth estimation.

[0108] In summary, this method achieves dynamic 3D reconstruction by constructing a closed-loop feedback control. During the initialization phase, a reference phase map adapted to the surface reflection characteristics of the measured object is generated through pre-scanning, establishing initial conditions for subsequent dynamic adjustments. During real-time operation, physical properties such as surface BRDF parameters, normal vectors, and motion velocity are continuously monitored to form a multi-dimensional state vector describing the current operating conditions. A reinforcement learning policy network performs online analysis of the state vector, simultaneously generating three control parameters: duty cycle matrix, phase order increment, and exposure time, achieving synergistic optimization of LC-SLM phase modulation and DMD high-speed modulation. The compressed sensing module uses a random Toeplitz matrix to compress and observe multi-view sub-aperture images, reducing the data volume to less than 4% of the original signal. The depth reconstruction network fuses the compressed measurement vector with the dynamically adjusted phase map, utilizes the multi-level attention mechanism of the 3D Swin-Transformer to extract global and local features, and synchronously outputs depth and confidence maps through a dual-branch decoding structure, ultimately generating high-precision 3D point cloud data.

[0109] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.

[0110] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0111] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0112] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0113] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A high-speed three-dimensional reconstruction system based on two-stage collaborative coding, characterized in that, The system comprises: a parallel light source module, a liquid crystal spatial light modulator, a 4f relay system, a digital micromirror device, a projection lens, a sparse light field sampling module and an edge computing device; The parallel light source module is used to output a parallel light beam, the liquid crystal spatial light modulator is used to perform phase modulation on the parallel light beam, the 4f relay system is used to scale-free transmit the phase-modulated light field to the digital micromirror device, the digital micromirror device is used to perform binary pulse width modulation on the phase-modulated light field within a microsecond time window, the projection lens is used to focus the binary pulse width modulated light field to the surface of a workpiece to be detected to form an emitted light field, and the sparse light field sampling module is used to collect a reflected light field of the surface of the workpiece to be detected and calculate a compressed measurement vector based on the reflected light field; and the edge computing device is used to synchronously control the liquid crystal spatial light modulator, the digital micromirror device and the sparse light field sampling module. The edge computing device is configured to perform the following operations: initializing an initial phase map of the liquid crystal spatial light modulator; dynamically adjusting a phase order increment of the liquid crystal spatial light modulator, a duty cycle matrix of the digital micromirror device and a camera exposure time of the sparse light field sampling module in real time; inputting the compressed measurement vector and a current phase map into an end-to-end three-dimensional reconstruction network model to obtain a depth map, wherein the current phase map is generated by dynamically adjusting the initial phase map according to the phase order increment.

2. The dual-stage collaborative coding based high-speed 3D reconstruction system according to claim 1, wherein, The parallel light source module comprises an LED light source and a collimating lens, the LED light source is used to emit a 450nm high-brightness scattered light beam, and the collimating lens is used to shape the scattered light beam into the parallel light beam.

3. The dual-stage collaborative coding based high-speed 3D reconstruction system according to claim 1, wherein, The sparse light field sampling module is composed of a microlens array and a compressive sensing sensor in cascade, the microlens array divides the reflected light field into multi-view sub-aperture images, the compressive sensing sensor adopts a random Toeplitz measurement matrix to linearly observe the sub-aperture images to obtain a compressed measurement vector, and the compressed measurement vector is expressed as y=Φx+ε, wherein Φ is a random Toeplitz measurement matrix, x is a vectorization of the multi-view sub-aperture images, and ε is a measurement noise.

4. The dual-stage collaborative coding based high-speed 3D reconstruction system according to claim 3, wherein, The real-time dynamic adjustment of the phase order increment of the liquid crystal spatial light modulator, the duty cycle matrix of the digital micromirror device and the camera exposure time of the sparse light field sampling module comprises: The edge computing device acquires physical characteristic parameters of the surface of the workpiece to be detected in real time to form a state vector, the physical characteristic parameters include BRDF parameters, a surface normal vector and a motion velocity vector, the BRDF parameters include a diffuse reflectance, a specular reflectance and a roughness; The normalized state vector is input into a trained reinforcement learning strategy network to output an action space, the action space includes a duty cycle matrix, a phase order increment and a camera exposure time, the edge computing device sends the duty cycle matrix to the digital micromirror device, sends the phase order increment to the liquid crystal spatial light modulator and sends the camera exposure time to the compressive sensing sensor, and synchronously triggers the dynamic adjustment of parameters.

5. The dual-stage collaborative coding based high-speed 3D reconstruction system according to claim 1, wherein, The 4f relay system is composed of a first lens and a second lens, the focal length of the first lens and the second lens is set to 100mm, and the first lens and the second lens are coaxial plano-convex lenses.

6. The dual-stage collaborative coding based high-speed 3D reconstruction system according to claim 4, wherein, The method comprises: An end-to-end three-dimensional reconstruction network model based on an encoder / decoder architecture is constructed, wherein the encoder adopts a 3D Swin-Transformer structure, and a multi-level hierarchical window attention mechanism is used to jointly extract features of the compressed measurement vector and the current phase map after channel dimension splicing, so as to generate a hybrid feature representation containing global context information and local detail features; The hybrid feature representation output by the encoder is fused across scales through adaptive weight distribution, and then mapped to the input space of the decoder through a dimension transformation module; The decoder is provided with a double-branch output structure, a first branch is used to regress and predict a depth map based on the fused hybrid feature, and a second branch is used to synchronously estimate a pixel-level confidence map; Based on the confidence map, the depth map is adaptively optimized, and the optimized depth map is combined with a pre-calibrated sensor intrinsic matrix and an extrinsic matrix, and the depth value in the pixel coordinate system is converted into a three-dimensional point cloud in the camera coordinate system through perspective projection inverse transformation.

7. The dual-stage collaborative coding-based high-speed three-dimensional reconstruction system according to claim 6, wherein: The loss function of the end-to-end three-dimensional reconstruction network model is constructed according to the following formula: L = L_depth + γ·L_photo + η·L_sparse; Wherein L_depth is L1 depth error; L_photo is image consistency based on differentiable rendering calculation; L_sparse is point cloud sparsity regularization, γ is the weight of controlling luminosity consistency, and η is the adjustment of sparsity constraint strength.

8. The dual-stage collaborative coding based high-speed 3D reconstruction system according to claim 7, wherein, The method comprises: The edge computing device controls the sparse light field sampling module to perform a millisecond-level flash pre-scan on the surface of the workpiece to be detected, collects multiple low-dose images, and obtains prior information of the workpiece to be detected based on the low-dose images, encodes the prior information into an 8-bit phase map to obtain an initial phase map, and the prior information is a physical characteristic parameter of the workpiece to be detected in an initial state.

9. A three-dimensional reconstruction method, characterized by, The method is applied to the three-dimensional reconstruction system according to any one of claims 1-8, and the method comprises: S1, initializing the phase map of the liquid crystal spatial light modulator and loading the initial phase map to the liquid crystal spatial light modulator; S2, real-time acquisition of the physical characteristic parameters of the surface of the workpiece to be detected to form a state vector, wherein the physical characteristic parameters include BRDF parameters, surface normal vectors and motion velocity vectors, and the BRDF parameters include diffuse reflectivity, specular reflectivity and roughness; S3, input the normalized state vector into the trained reinforcement learning strategy network, output the action space, the action space includes duty cycle matrix, phase order increment and camera exposure time, send the duty cycle matrix to the digital micromirror device, send the phase order increment to the liquid crystal spatial light modulator, send the camera exposure time to the compressed sensing sensor, and synchronously trigger parameter dynamic adjustment; S4, collect the reflected light field of the workpiece to be detected based on the current parameter adjustment through the microlens array, and segment it into sub-aperture images, and the compressed sensing sensor adopts a random Toeplitz measurement matrix to perform linear observation on the sub-aperture images to obtain a compressed measurement vector; S5, input the compressed measurement vector and the current phase map into an end-to-end three-dimensional reconstruction network model to obtain a depth map, wherein the current phase map is generated by dynamically adjusting an initial phase map according to the phase order increment by the liquid crystal spatial light modulator.

10. The three-dimensional reconstruction method of claim 9, wherein, The step of inputting the compressed measurement vector and the current phase map into the end-to-end three-dimensional reconstruction network model to obtain the depth map comprises: An end-to-end three-dimensional reconstruction network model based on an encoder / decoder architecture is constructed, wherein the encoder adopts a 3D Swin-Transformer structure, and a multi-level hierarchical window attention mechanism is used to perform joint feature extraction on the compressed measurement vector and the current phase map after channel dimension splicing, thereby generating a hybrid feature representation containing global context information and local detail features; The hybrid feature representation output by the encoder is cross-scale fused through adaptive weight distribution, and then mapped to the input space of the decoder through a dimension transformation module; The decoder is provided with a double-branch output structure, and the first branch regresses and predicts a depth map based on the fused hybrid feature; the second branch simultaneously estimates a pixel-level confidence map; The depth map is adaptively optimized based on the estimated confidence map, and the optimized depth map is combined with a pre-calibrated sensor intrinsic matrix and extrinsic matrix to convert the depth value in the pixel coordinate system into a three-dimensional point cloud in the camera coordinate system through perspective projection inverse transformation.

Citation Information

Patent Citations

  • Special environment laser radar system and method based on polarization modulation and laser coding

    CN120294771A

  • Processing method and system for sparse compression reconstruction of LDI sub-pixel image and application

    CN120447307A