A physically driven end-to-end 3D reconstruction method for low reflectivity targets

CN122820969APending Publication Date: 2026-09-25GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610909285.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,这类"黑盒"模型缺乏对光学成像物理机理的显式利用,割裂了物理干涉规律与极线几何约束之间的内在联系,导致模型泛化能力弱,难以保证亚像素级的测量精度,且在训练数据分布之外的场景中性能急剧下降

Benefits of technology

[0079]1、采用十字形均匀稀疏采样策略,仅需在频域正交轴上采集极少量的关键频率即可保留核心物理特征,突破了传统成像的速度瓶颈,实现了极低数据量下的高效三维测量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820969A_ABST
    Figure CN122820969A_ABST
Patent Text Reader

Abstract

The application discloses a physical driving end-to-end three-dimensional reconstruction method for a low reflectivity target, collects an under-sampling noise-containing fringe map of a low reflectivity target surface, and inputs the map into an end-to-end deep neural network, implements a two-stage reconstruction strategy based on physical priori and decoupling regression, in the first stage, extracts a pure light priori through the network, guides local adaptive feature fusion and self-attention modulation, and decouples a high-precision complex physical response of the target; in the second stage, a cross-interference cost volume is constructed based on the complex physical response, sub-pixel physical initial values are extracted, and parameter correction is performed by fusing geometric and frequency features by using a gated recurrent unit, and a normalized position along a polar line and a vertical polar line offset are decoupled and output; finally, the above parameters are directly reconstructed into accurate two-dimensional polar line projection coordinates and three-dimensional point cloud coordinates through a differentiable geometric projection layer. The application can effectively solve the problem of low signal-to-noise ratio of the low reflectivity target, and realize high-precision three-dimensional target reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision, optical measurement and deep learning, and in particular to a physical-driven end-to-end 3D reconstruction method for low reflectivity targets. Background Technology

[0002] 3D reconstruction technology has wide applications in fields such as industrial inspection, reverse engineering, cultural relic preservation, and medical imaging. Structured light 3D measurement, as a mainstream non-contact optical measurement method, works by projecting a specifically coded stripe pattern onto the surface of the object being measured using a projector, capturing the deformed stripe image modulated by the object's surface using a camera, and then reconstructing the object's 3D shape through steps such as phase unwrapping, stereo matching, and triangulation.

[0003] However, traditional structured light measurement techniques are highly dependent on the reflectivity of the object's surface. When dealing with black materials, light-absorbing coatings, distant low-light scenes, or targets with low reflectivity, the extremely low reflectivity of the object's surface results in a very low signal-to-noise ratio (SNR) for the fringe image captured by the camera, severely reducing fringe contrast and even causing it to be completely submerged by environmental noise. Under these conditions, traditional phase unwrapping algorithms (such as time-expansion-based phase-shifting methods and spatial-expansion-based Fourier transform profilometry) cannot accurately extract phase information, leading to phase unwrapping failure and consequently rendering subsequent 3D reconstruction completely ineffective.

[0004] In recent years, Parallel Single-Pixel Imaging (PSI) has emerged as a novel computational imaging method, demonstrating significant advantages in 3D measurement of low-reflectivity scenes. This technique achieves 3D reconstruction by independently recovering the light field transmission matrix for each pixel and locating the coordinates of directly reflected light. Essentially, it transforms the imaging problem into a frequency-domain sparse sampling and signal recovery problem. Compared to traditional methods, PSI exhibits inherent robustness to low-light-flux scenes, making it highly promising for 3D measurement of low-reflectivity targets.

[0005] However, parallel single-pixel imaging technology faces a fundamental contradiction in practical applications: the conflict between measurement efficiency and measurement accuracy. To improve measurement efficiency, the system often employs a sparse undersampling strategy, acquiring only a small number of frequency components in the frequency domain. However, under low-light conditions, drastically reducing the sampling rate leads to severe artifacts and distortions during the spectral reconstruction process. These frequency domain artifacts propagate directly to the epipolar matching stage, causing subsequent epipolar matching and 3D reconstruction to completely fail.

[0006] To address the above problems, existing technologies have proposed the following three types of solutions, but all of them have obvious limitations:

[0007] The first approach involves increasing the light source power or extending the exposure time. This improves the signal-to-noise ratio by increasing the incident light energy or extending the camera integration time. However, increasing the light source power leads to a significant increase in system power consumption, and for targets with extremely low reflectivity, such as light-absorbing materials, it is difficult to obtain an effective signal even with a substantial increase in light power. Extending the exposure time contradicts the original intention of high-speed measurement, making it unsuitable for dynamic scenes or online industrial inspection, and long exposure times can introduce problems such as motion blur.

[0008] The second approach employs compressed sensing and filtering algorithms. This method leverages the sparsity prior of the signal in a specific transform domain, using compressed sensing theory to recover the complete signal from a small number of samples, and supplements this with various filtering algorithms to suppress noise. However, compressed sensing algorithms have extremely high computational complexity, and the iterative optimization process is very time-consuming, making it difficult to meet real-time requirements. Furthermore, while filtering algorithms suppress noise, they often smooth high-frequency edge details of objects, leading to blurred edges and loss of detail in the reconstruction results.

[0009] The third approach is a purely data-driven deep learning method. This method leverages the powerful nonlinear fitting capabilities of deep neural networks to directly output 3D coordinates or fringe images from a low signal-to-noise ratio input. However, this type of "black box" model lacks explicit utilization of the physical mechanisms of optical imaging, severing the intrinsic connection between physical interference laws and epipolar geometric constraints. This results in weak model generalization ability, difficulty in guaranteeing sub-pixel level measurement accuracy, and a sharp decline in performance in scenarios outside the training data distribution.

[0010] Therefore, there is an urgent need for a solution that can achieve high-efficiency and high-precision three-dimensional topography reconstruction of low reflectivity targets under low signal-to-noise ratio conditions. Summary of the Invention

[0011] The purpose of this invention is to overcome the shortcomings of the prior art and provide a physical-driven end-to-end 3D reconstruction method for low reflectivity targets.

[0012] To achieve the above objectives, the technical solution provided by this invention is as follows:

[0013] A physical-driven end-to-end 3D reconstruction method for low-reflectivity targets includes the following steps:

[0014] S1. A sparse frequency sampling list is generated based on a cross-shaped uniform sparse sampling strategy, and a corresponding stripe pattern is generated by modulating the sparse frequency sampling list. The stripe pattern is then projected onto the surface of the target with low reflectivity, and the reflected light intensity signal of the scene is collected by a camera to obtain an undersampled noisy stripe pattern.

[0015] S2. The undersampled noisy fringe pattern obtained in step S1 is input into the first-stage deep neural network. The first-stage deep neural network first extracts pure illumination prior features and generates a pure mask through its illumination estimator. Then, it uses the illumination prior features as physical modulation conditions to perform spatial affine transformation and adaptive denoising on the multi-scale image features of the backbone network. At the same time, it performs adaptive soft fusion of the denoised features and the original features based on the pure mask. Finally, it reconstructs the complex physical response of the target through a decoupled dual-branch prediction head. The complex physical response includes a real part and an imaginary part.

[0016] S3. Based on the complex physical response obtained in step S2, construct the cross-interference cost volume to extract sub-pixel physical initial values; the sub-pixel physical initial values ​​are input to a coordinate regression network based on parameter decoupling and cyclic iteration; the coordinate regression network includes a complex physical feature extractor, a gated recurrent unit, and a differentiable geometric projection module; the complex physical feature extractor performs feature extraction and aggregation on the sub-pixel physical initial values ​​and the complex physical response to generate physical interference correlation features, local context features, and geometric query vectors; then, using the gated recurrent unit as the core updater, starting from the sub-pixel physical initial values, iteratively corrects the decoupled epipolar geometric parameters step by step; the epipolar geometric parameters include the normalized position along the epipolar line and the offset perpendicular to the epipolar line; after the iteration terminates, the epipolar geometric parameters are directly reconstructed into two-dimensional epipolar projection coordinates via the differentiable geometric projection module;

[0017] S4. Based on the system calibration parameters, the two-dimensional epipolar projection coordinates obtained in step S3 are mapped to the final three-dimensional point cloud coordinates through the second-stage deep neural network. The second-stage deep neural network is trained end-to-end with supervision using a dataset that pairs low signal-to-noise ratio noisy fringe images based on real physical scenes with the corresponding ground truth values ​​of two-dimensional epipolar projection coordinates, with the constraint of minimizing the geometric projection error on the two-dimensional image plane.

[0018] Further, in step S1, a sparse frequency sampling list is generated based on the cross-shaped uniform sparse sampling strategy, including:

[0019] Establish a frequency domain coordinate system: Set the image resolution to... Define the coordinates of the zero-frequency center as For any discrete frequency point in the two-dimensional frequency domain Their horizontal absolute physical frequencies are defined as follows: The vertical absolute physical frequency is ;

[0020] Determine the frequency band and number of sampling points: Set the low-frequency cutoff frequency for two-dimensional frequency domain sampling as... The high-frequency cutoff frequency is And set the number of single-arm target sampling points of the two-dimensional frequency domain cross coordinate axis as ;

[0021] Construct a continuous equidistant sequence and discretize the mapping: in the interval Internal calculation contains A linearly ordered sequence of theoretical feature points at equal intervals is obtained. Then, each floating-point element in the linearly ordered sequence is rounded down to the nearest integer, mapped to a discrete Fourier frequency domain grid index, and deduplication is performed to obtain the actual one-dimensional discrete target frequency set used for sampling. ;

[0022] Mapping a two-dimensional cross-shaped space: Traversing the two-dimensional frequency domain coefficient matrix, constructing screening conditions based on the cross-shaped orthogonal axis; if and only if any discrete frequency point A point is retained as a target feature point if any of the following conditions are met:

[0023] (1) Located on the horizontal frequency axis and belonging to the target set: i.e. and ;

[0024] (2) Located on the vertical frequency axis and belonging to the target set: that is and ;

[0025] Generate the final sampling sequence: Collect all extracted target feature points, merge them in index order, and perform stable deduplication to extract conjugate independent feature points, finally generating a sparse frequency sampling list for optical field modulation or inverse Fourier transform.

[0026] Further, in step S2, the pure illumination prior features are extracted using an illumination estimator and a pure mask is generated, including:

[0027] Input the original stripe image features as follows First, the mean value is calculated along the channel dimension, and then spatial low-pass filtering and pooling are applied to obtain the pure average reflectance. :

[0028] ;

[0029] For the original stripe image features in the first... Feature maps on each channel; The total number of channels for the original stripe image features; This is a space-average pooling operation;

[0030] Will and The enhanced tensor is obtained by splicing along the channel dimension. Illumination features are extracted using depthwise separable convolution. :

[0031] ;

[0032] This is the first convolutional layer, used to perform preliminary feature transformation on the augmentation tensor; Depthwise separable convolution is used to extract illumination features. ;

[0033] Through single-channel convolution and Activation function outputs a pure mask :

[0034] ;

[0035] This is the second convolutional layer, a single-channel convolution, used to map illumination features to single-channel mask predictions.

[0036] Furthermore, in step S2, the backbone network adopts a Swing Transformer structure and embeds spatial feature transformation layers in the shallow and mid-layer feature extraction, utilizing illumination features. As a physical modulation condition, spatial affine transformation is performed on image features. The process includes:

[0037] For input features Its spatial modulation formula is:

[0038] ;

[0039] ;

[0040] ;

[0041] Scaling convolutional layers are used to learn spatial scaling factors from illumination features. ; This is a translation convolutional layer used to learn the spatial translation factor from illumination features. ; For output features; This is an element-wise multiplication operation;

[0042] Scaling factor learned through space Spatial translation factor This enables pixel-by-pixel guidance of image features based on physical illumination priors.

[0043] Further, in step S2, the encoding part of the first-stage deep neural network includes a signal-to-noise ratio (SNR) sensing encoder; the SNR sensing encoder includes a denoising network, an original feature projection layer, and a gating network, used to adaptively soft-fuse the denoised features and the original features according to the clean mask; the specific calculation process is as follows:

[0044] The enhancement tensor is processed using a denoising network and the original feature projection layer, respectively. , to obtain denoising features With original features Processing pure masks using gating networks Generate spatial fusion weights restricted to a specific range. :

[0045] ;

[0046] For gating networks;

[0047] Through the fusion weight The two features are weighted and fused pixel by pixel to output high signal-to-noise ratio perceptual features. :

[0048] ;

[0049] in, This indicates an element-wise multiplication operation.

[0050] Furthermore, in step S2, the backbone network of the first-stage deep neural network employs an illumination-guided multi-head self-attention mechanism at the bottleneck layer; when calculating self-attention, the projected illumination features are incorporated into the calculation of the value matrix, specifically implemented as follows:

[0051] Map the input features to the query matrix Key matrix and initial value matrix ; for the query matrix Bond matrix The feature dimensions are subjected to L2 normalization; the downsampled illumination features are then utilized. Physical modulation of the initial value matrix yields the modulation value matrix. :

[0052] ;

[0053] The formula for calculating multi-head self-attention guided by illumination is:

[0054] ;

[0055] in, The learningable adaptive scaling parameters are used; the calculation results are linearly projected and then added to the positional encoding generated based on the modulation value matrix to output the final bottleneck features; finally, the deep decoding features are reconstructed into the complex physical response of the target through a shared feature mapping layer and a decoupled dual-branch prediction head.

[0056] The dual-branch prediction head includes a modulation branch and a phase direction branch; the modulation branch is transmitted through... The activation function outputs an absolutely normalized modulation distribution. The phase direction branch outputs the original phase features, which are then constrained to a unit vector through L2 normalization, and the cosine components are output. With sinusoidal components Finally, by combining the physical formulas, a high-precision complex physical response is reconstructed, which includes the real part. With the imaginary part :

[0057] ;

[0058] .

[0059] Further, in step S3, the cross-interference cost volume is constructed to extract sub-pixel physical initial values, including:

[0060] One-dimensional discrete sampling sequences are defined along the epipolar direction and perpendicular to the epipolar direction, respectively. Based on the epipolar start point, epipolar end point, and normal vector, the discrete sampling sequences are mapped to the absolute pixel physical coordinate system. Combined with the physical coordinates corresponding to each frequency point in the sparse frequency sampling list, the theoretical interference phase corresponding to each sampling point is calculated. ;

[0061] The real part of the complex physical response at each frequency extracted by the first-stage deep neural network. With the imaginary part Substituting the physical interference superposition equation, the total energy of each sampling point under multi-frequency interference is calculated, generating the cross cost volume corresponding to the epipolar direction and the perpendicular epipolar direction. :

[0062] ;

[0063] Cross cost volume A non-negative truncation operation is performed to eliminate negative interference responses; the position of the highest energy peak is searched within the truncated cross cost volume, and a local neighborhood window is constructed centered on this highest energy peak position; normalized weights are calculated using the energy values ​​of each sampling point within the local neighborhood window, and the sub-pixel physical initial values ​​of the relative positions in the epipolar direction are analyzed by weighted summation of local centroids. Subpixel physical initial value of pixel offset perpendicular to epipolar direction .

[0064] Furthermore, in step S3, the process of reconstructing the two-dimensional epipolar projection coordinates includes:

[0065] The complex physical feature extractor extracts and aggregates features from the sub-pixel physical initial value and the complex physical response to generate physical interference correlation features, local context features and geometric query vectors.

[0066] Using a gated loop unit as the core updater, in the first... In this iteration: the physical interference correlation features, local context features, geometric query vector, and latent state features from the previous time step are used as inputs; the gated recurrent unit outputs the position parameters at the current time step. and bias parameters The residual increment; using the residual increment and the built-in adaptive scaling gating mechanism, the parameters from the previous time step are updated:

[0067] ;

[0068] ;

[0069] in, This is a truncation function;

[0070] After the iteration terminates, based on the final output of the normalized position along the epipolar line and the offset perpendicular to the epipolar line, the two-dimensional epipolar projection coordinates are calculated using the vector composition formula through the differentiable geometric projection module. :

[0071] ;

[0072] The coordinates of the starting point of the polar line; is the unit basis vector along the polar direction; Let be the unit normal vector perpendicular to the polar direction. is the polar length.

[0073] Furthermore, during the training process in step S4, the geometric projection error on the two-dimensional image plane is minimized. To suppress physical noise interference;

[0074] The loss function for the second-stage deep neural network is defined as:

[0075] ;

[0076] in, This represents the total number of samples in the training batch. These are the two-dimensional epipolar projection coordinates output by the network forward propagation. The true values ​​of the corresponding two-dimensional epipolar projection coordinates. It is an L1 norm;

[0077] The training strategy employs the AdamW optimizer with a weight decay mechanism for gradient descent of network parameters and weight updates, and utilizes a single-cycle learning rate decay strategy to dynamically warm up and adjust the learning rate. By combining the optimization strategy with minimizing the geometric projection error, the network is forced to learn the mapping relationship from two-dimensional epipolar projection coordinates to three-dimensional point cloud coordinates.

[0078] Compared with existing technologies, the principles and advantages of this technical solution are as follows:

[0079] 1. By adopting a cross-shaped uniform sparse sampling strategy, only a very small number of key frequencies need to be collected on the orthogonal axis in the frequency domain to retain the core physical features, breaking through the speed bottleneck of traditional imaging and realizing efficient three-dimensional measurement with extremely low data volume.

[0080] 2. The system extracts prior physical illumination data to guide adaptive feature denoising, abandoning the traditional black-box image generation paradigm. Even under the dual harsh conditions of minimal sampling and low target reflectivity, it can still accurately reproduce high-precision complex physical responses, significantly improving the system's robustness in low-light scenarios.

[0081] 3. Constructing a cross-interference cost volume to extract sub-pixel physical initial values, and combining "epochal parameter decoupling + GRU cyclic iteration" for residual correction, can effectively suppress noise interference such as subsurface scattering, and still achieve high-fidelity sub-pixel level positioning accuracy with very little data input. Attached Figure Description

[0082] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0083] Figure 1 To realize the parallel single-pixel three-dimensional imaging system used in the embodiment of the present invention for a physically driven end-to-end three-dimensional reconstruction method for low reflectivity targets;

[0084] Figure 2 This is a flowchart illustrating the principle of a physical-driven end-to-end 3D reconstruction method for low-reflectivity targets according to an embodiment of the present invention.

[0085] Figure 3 This is a schematic diagram of a cross-shaped uniform sparse sampling strategy based on the two-dimensional discrete Fourier frequency domain space.

[0086] Figure 4 This is a schematic diagram of the structure of the first stage of a deep neural network;

[0087] Figure 5 This is a schematic diagram illustrating high-precision 3D point cloud reconstruction based on complex physical response and epipolar geometric parameters. Detailed Implementation

[0088] The present invention will be further described below with reference to specific embodiments:

[0089] like Figure 1 As shown, the parallel single-pixel 3D imaging system that implements the physical-driven end-to-end 3D reconstruction method for low-reflectivity targets described in this embodiment includes a computer 1, a projector 2, a color industrial camera 3, and a low-reflectivity target 4. Its working principle is as follows:

[0090] Computer 1 generates a sparse frequency sampling list based on a cross-shaped uniform sparse sampling strategy, and modulates several stripe patterns 5 according to this sparse frequency sampling list. The stripe patterns 5 are transmitted into projector 2 and projected onto the surface of the target 4. Next, a color industrial camera 3 acquires the corresponding color stripe pattern 6, and transmits the color stripe pattern 6 into computer 1 for processing. Computer 1 inputs these data into a pre-trained end-to-end deep neural network model (including a first-stage deep neural network and a second-stage deep neural network), and sequentially completes physical feature extraction, generative denoising, iterative coordinate regression, and low-reflectivity target point cloud reconstruction, obtaining result 7.

[0091] More specifically, such as Figure 2 As shown in this embodiment, a physical-driven end-to-end 3D reconstruction method for low-reflectivity targets includes the following steps:

[0092] S1. Computer 1 generates a sparse frequency sampling list based on a cross-shaped uniform sparse sampling strategy, and modulates and generates a corresponding stripe pattern according to the sparse frequency sampling list. The stripe pattern is then projected onto the surface of the low reflectivity target 4 through projector 2. The reflected light intensity signal of the scene is collected by color industrial camera 3 to obtain an undersampled noisy stripe pattern.

[0093] In this step, combined Figure 3 As shown, a sparse frequency sampling list is generated based on a cross-shaped uniform sparse sampling strategy, including:

[0094] Establish a frequency domain coordinate system: Set the image resolution to... Define the coordinates of the zero-frequency center as For any discrete frequency point in the two-dimensional frequency domain Their horizontal absolute physical frequencies are defined as follows: The vertical absolute physical frequency is ;

[0095] Determine the frequency band and number of sampling points: Set the low-frequency cutoff frequency for two-dimensional frequency domain sampling to be... The high-frequency cutoff frequency is And set the number of single-arm target sampling points of the two-dimensional frequency domain cross coordinate axis as ;

[0096] Construct a continuous equidistant sequence and discretize the mapping: in the interval Internal calculation contains A linearly spaced sequence of theoretical feature points is obtained; subsequently, each floating-point element in the linearly spaced sequence is rounded down to the nearest integer, mapped to a discrete Fourier frequency domain grid index, and deduplication is performed to obtain the actual one-dimensional discrete target frequency set used for sampling. ;

[0097] Mapping a two-dimensional cross-shaped space: Traversing the two-dimensional frequency domain coefficient matrix, constructing screening conditions based on the cross-shaped orthogonal axis; if and only if any discrete frequency point A point is retained as a target feature point if any of the following conditions are met:

[0098] (1) Located on the horizontal frequency axis and belonging to the target set: i.e. and ;

[0099] (2) Located on the vertical frequency axis and belonging to the target set: that is and ;

[0100] Generate the final sampling sequence: Collect all extracted target feature points, merge them in index order, and perform stable deduplication to extract conjugate independent feature points, finally generating a sparse frequency sampling list for optical field modulation or inverse Fourier transform.

[0101] S2. Input the undersampled noisy fringe pattern obtained in step S1 into the first-stage deep neural network, such as... Figure 4 As shown, the first-stage deep neural network first extracts pure illumination prior features and generates a pure mask through its internal illumination estimator. Then, it uses the illumination prior features as physical modulation conditions to perform spatial affine transformation and adaptive denoising on the multi-scale image features of the backbone network. At the same time, it performs adaptive soft fusion of the denoised features and the original features based on the pure mask. Finally, it reconstructs the high-precision complex physical response of the target through a decoupled dual-branch prediction head. The complex physical response includes a real part and an imaginary part.

[0102] Specifically, in this step, a clean illumination prior feature is extracted using an illumination estimator and a clean mask is generated, including:

[0103] Input the original stripe image features as follows First, the mean value is calculated along the channel dimension, and then spatial low-pass filtering and pooling are applied to obtain the pure average reflectance. :

[0104] ;

[0105] For the original stripe image features in the first... Feature maps on each channel; The total number of channels for the original stripe image features; This is a space-average pooling operation;

[0106] Will and The enhanced tensor is obtained by splicing along the channel dimension. Illumination features are extracted using depthwise separable convolution. :

[0107] ;

[0108] This is the first convolutional layer, used to perform preliminary feature transformation on the augmentation tensor; Depthwise separable convolution is used to extract illumination features. ;

[0109] Through single-channel convolution and Activation function outputs a pure mask :

[0110] ;

[0111] This is the second convolutional layer, a single-channel convolution, used to map illumination features to single-channel mask predictions.

[0112] Specifically, in this step, the backbone network adopts a Swin Transformer structure, and spatial feature transformation layers are embedded in the shallow and mid-level feature extraction layers to utilize illumination features. As a physical modulation condition, spatial affine transformation is performed on image features. The process includes:

[0113] For input features Its spatial modulation formula is:

[0114] ;

[0115] ;

[0116] ;

[0117] Scaling convolutional layers are used to learn spatial scaling factors from illumination features. ; This is a translation convolutional layer used to learn the spatial translation factor from illumination features. ; For output features; This is an element-wise multiplication operation;

[0118] Scaling factor learned through space Spatial translation factor This enables pixel-by-pixel guidance of image features based on physical illumination priors.

[0119] Specifically, in this step, the encoding part of the first-stage deep neural network includes a signal-to-noise ratio (SNR) sensing encoder; the SNR sensing encoder includes a denoising network, an original feature projection layer, and a gating network, used to adaptively soft-fuse the denoised features and the original features according to the clean mask; the specific calculation process is as follows:

[0120] The enhancement tensor is processed using a denoising network and the original feature projection layer, respectively. , to obtain denoising features With original features Processing pure masks using gating networks Generate spatial fusion weights restricted to a specific range. :

[0121] ;

[0122] For gating networks;

[0123] Through the fusion weight The two features are weighted and fused pixel by pixel to output high signal-to-noise ratio perceptual features. :

[0124] ;

[0125] in, This indicates an element-wise multiplication operation.

[0126] Specifically, in this step, the backbone network of the first-stage deep neural network adopts an illumination-guided multi-head self-attention mechanism at the bottleneck layer; when calculating self-attention, the projected illumination features are introduced into the calculation of the value matrix, specifically implemented as follows:

[0127] Map the input features to the query matrix Key matrix and initial value matrix ; for the query matrix Bond matrix The feature dimensions are subjected to L2 normalization; the downsampled illumination features are then utilized. Physical modulation of the initial value matrix yields the modulation value matrix. :

[0128] ;

[0129] The formula for calculating multi-head self-attention guided by illumination is:

[0130] ;

[0131] in, The learningable adaptive scaling parameters are used; the calculation results are linearly projected and then added to the positional encoding generated based on the modulation value matrix to output the final bottleneck features; finally, the deep decoding features are reconstructed into the complex physical response of the target through a shared feature mapping layer and a decoupled dual-branch prediction head.

[0132] The dual-branch prediction head includes a modulation branch and a phase direction branch; the modulation branch is transmitted through... The activation function outputs an absolutely normalized modulation distribution. The phase direction branch outputs the original phase features, which are then constrained to a unit vector through L2 normalization, and the cosine components are output. With sinusoidal components Finally, by combining the physical formulas, a high-precision complex physical response is reconstructed, which includes the real part. With the imaginary part :

[0133] ;

[0134] ;

[0135] This decoupled output design directly extracts the essential physical properties of the target under test, avoiding end-to-end black-box mapping errors in stripe image generation.

[0136] S3. Based on the complex physical response obtained in step S2, construct the cross-interference cost volume to extract sub-pixel physical initial values; the sub-pixel physical initial values ​​are input to a coordinate regression network based on parameter decoupling and cyclic iteration; the coordinate regression network includes a complex physical feature extractor, a gated recurrent unit, and a differentiable geometric projection module. The complex physical feature extractor performs feature extraction and aggregation on the sub-pixel physical initial values ​​and the complex physical response to generate physical interference correlation features, local context features, and geometric query vectors; then, using the gated recurrent unit as the core updater, starting from the sub-pixel physical initial values, iteratively corrects the decoupled epipolar geometric parameters step by step. The epipolar geometric parameters include the normalized position along the epipolar line and the offset perpendicular to the epipolar line. After the iteration terminates, the epipolar geometric parameters are directly reconstructed into two-dimensional epipolar projection coordinates via the differentiable geometric projection module.

[0137] In this step, the cross-interference cost volume is constructed to extract sub-pixel physical initial values, including:

[0138] One-dimensional discrete sampling sequences are defined along the epipolar direction and perpendicular to the epipolar direction, respectively. Based on the epipolar start point, epipolar end point, and normal vector, the discrete sampling sequences are mapped to the absolute pixel physical coordinate system. Combined with the physical coordinates corresponding to each frequency point in the sparse frequency sampling list, the theoretical interference phase corresponding to each sampling point is calculated. ;

[0139] The real part of the complex physical response at each frequency extracted by the first-stage deep neural network. With the imaginary part Substituting the physical interference superposition equation, the total energy of each sampling point under multi-frequency interference is calculated, generating the cross cost volume corresponding to the epipolar direction and the perpendicular epipolar direction. :

[0140] ;

[0141] Cross cost volume A non-negative truncation operation is performed to eliminate negative interference responses; the position of the highest energy peak is searched within the truncated cross cost volume, and a local neighborhood window is constructed centered on this highest energy peak position; normalized weights are calculated using the energy values ​​of each sampling point within the local neighborhood window, and the sub-pixel physical initial values ​​of the relative positions in the epipolar direction are analyzed by weighted summation of local centroids. Subpixel physical initial value of pixel offset perpendicular to epipolar direction .

[0142] In this step, the process of reconstructing the two-dimensional epipolar projection coordinates includes:

[0143] The complex physical feature extractor extracts and aggregates features from the sub-pixel physical initial value and the complex physical response to generate physical interference correlation features, local context features and geometric query vectors.

[0144] Using a gated loop unit as the core updater, in the first... In this iteration: the physical interference correlation features, local context features, geometric query vector, and latent state features from the previous time step are used as inputs; the gated recurrent unit outputs the position parameters at the current time step. and bias parameters The residual increment; using the residual increment and the built-in adaptive scaling gating mechanism, the parameters from the previous time step are updated:

[0145] ;

[0146] ;

[0147] in, This is a truncation function;

[0148] After the iteration terminates, based on the final output of the normalized position along the epipolar line and the offset perpendicular to the epipolar line, the two-dimensional epipolar projection coordinates are calculated using the vector composition formula through the differentiable geometric projection module. :

[0149] ;

[0150] The coordinates of the starting point of the polar line; is the unit basis vector along the polar direction; Let be the unit normal vector perpendicular to the polar direction. is the polar length.

[0151] S4. Based on the system calibration parameters, the two-dimensional epipolar projection coordinates obtained in step S3 are mapped to the final three-dimensional point cloud coordinates through the second-stage deep neural network. The second-stage deep neural network is trained end-to-end with supervision using a dataset that pairs low signal-to-noise ratio noisy fringe images based on real physical scenes with the corresponding ground truth values ​​of two-dimensional epipolar projection coordinates, with the constraint of minimizing the geometric projection error on the two-dimensional image plane.

[0152] During the training process in this step, the geometric projection error on the two-dimensional image plane is minimized. To suppress physical noise interference;

[0153] The loss function for the second-stage deep neural network is defined as:

[0154] ;

[0155] in, This represents the total number of samples in the training batch. These are the two-dimensional epipolar projection coordinates output by the network forward propagation. The true values ​​of the corresponding two-dimensional epipolar projection coordinates. It is an L1 norm;

[0156] The training strategy employs the AdamW optimizer with a weight decay mechanism for gradient descent of network parameters and weight updates, and utilizes a single-cycle learning rate decay strategy to dynamically warm up and adjust the learning rate. By combining the optimization strategy with minimizing the geometric projection error, the network is forced to learn the mapping relationship from two-dimensional epipolar projection coordinates to three-dimensional point cloud coordinates.

[0157] The above-described embodiments are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Therefore, any changes made in accordance with the shape and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. A physically driven end-to-end 3D reconstruction method for low reflectivity targets, characterized in that, Includes the following steps: S1. A sparse frequency sampling list is generated based on a cross-shaped uniform sparse sampling strategy, and a corresponding stripe pattern is generated by modulating the sparse frequency sampling list. The stripe pattern is then projected onto the surface of the target with low reflectivity, and the reflected light intensity signal of the scene is collected by a camera to obtain an undersampled noisy stripe pattern. S2. The undersampled noisy fringe pattern obtained in step S1 is input into the first-stage deep neural network. The first-stage deep neural network first extracts pure illumination prior features and generates a pure mask through its illumination estimator. Then, it uses the illumination prior features as physical modulation conditions to perform spatial affine transformation and adaptive denoising on the multi-scale image features of the backbone network. At the same time, it performs adaptive soft fusion of the denoised features and the original features based on the pure mask. Finally, it reconstructs the complex physical response of the target through a decoupled dual-branch prediction head. The complex physical response includes a real part and an imaginary part. S3. Based on the complex physical response obtained in step S2, construct the cross-interference cost volume to extract sub-pixel physical initial values; the sub-pixel physical initial values ​​are input to a coordinate regression network based on parameter decoupling and cyclic iteration; the coordinate regression network includes a complex physical feature extractor, a gated recurrent unit and a differentiable geometric projection module; the complex physical feature extractor performs feature extraction and aggregation on the sub-pixel physical initial values ​​and the complex physical response to generate physical interference correlation features, local context features and geometric query vectors; Then, using the gated loop unit as the core updater, starting from the sub-pixel physical initial value, the decoupled epipolar geometric parameters are iteratively corrected step by step. The epipolar geometric parameters include the normalized position along the epipolar line and the offset perpendicular to the epipolar line. After the iteration terminates, the epipolar geometric parameters are directly reconstructed into two-dimensional epipolar projection coordinates through the differentiable geometric projection module. S4. Based on the system calibration parameters, the two-dimensional epipolar projection coordinates obtained in step S3 are mapped to the final three-dimensional point cloud coordinates through the second-stage deep neural network. The second-stage deep neural network is trained end-to-end with supervision using a dataset that pairs low signal-to-noise ratio noisy fringe images based on real physical scenes with the corresponding ground truth values ​​of two-dimensional epipolar projection coordinates, with the constraint of minimizing the geometric projection error on the two-dimensional image plane.

2. The physical-driven end-to-end 3D reconstruction method for low-reflectivity targets according to claim 1, characterized in that, In step S1, a sparse frequency sampling list is generated based on the cross-shaped uniform sparse sampling strategy, including: Establish a frequency domain coordinate system: Set the image resolution to... Define the coordinates of the zero-frequency center as For any discrete frequency point in the two-dimensional frequency domain Their horizontal absolute physical frequencies are defined as follows: The vertical absolute physical frequency is ; Determine the frequency band and number of sampling points: Set the low-frequency cutoff frequency for two-dimensional frequency domain sampling to be... The high-frequency cutoff frequency is And set the number of single-arm target sampling points of the two-dimensional frequency domain cross coordinate axis as ; Construct a continuous equidistant sequence and discretize the mapping: in the interval Internal calculation contains A linearly spaced sequence of theoretical feature points is obtained; subsequently, each floating-point element in the linearly spaced sequence is rounded down to the nearest integer, mapped to a discrete Fourier frequency domain grid index, and deduplication is performed to obtain the actual one-dimensional discrete target frequency set used for sampling. ; Mapping a two-dimensional cross-shaped space: Traversing the two-dimensional frequency domain coefficient matrix, constructing screening conditions based on the cross-shaped orthogonal axis; if and only if any discrete frequency point A point is retained as a target feature point if any of the following conditions are met: (1) Located on the horizontal frequency axis and belonging to the target set: i.e. and ; (2) Located on the vertical frequency axis and belonging to the target set: that is and ; Generate the final sampling sequence: Collect all extracted target feature points, merge them in index order, and perform stable deduplication to extract conjugate independent feature points, finally generating a sparse frequency sampling list for optical field modulation or inverse Fourier transform.

3. The physically driven end-to-end 3D reconstruction method for low reflectivity targets according to claim 1, characterized in that, In step S2, the pure illumination prior features are extracted using an illumination estimator, and a pure mask is generated, including: Input the original stripe image features as follows First, the mean value is calculated along the channel dimension, and then spatial low-pass filtering and pooling are applied to obtain the pure average reflectance. : ; For the original stripe image features in the first... Feature maps on each channel; The total number of channels for the original stripe image features; This is a space-average pooling operation; Will and The enhanced tensor is obtained by splicing along the channel dimension. Illumination features are extracted using depthwise separable convolution. : ; This is the first convolutional layer, used to perform preliminary feature transformation on the augmentation tensor; Depthwise separable convolution is used to extract illumination features. ; Through single-channel convolution and Activation function outputs a pure mask : ; This is the second convolutional layer, a single-channel convolution, used to map illumination features to single-channel mask predictions.

4. The physical-driven end-to-end 3D reconstruction method for low-reflectivity targets according to claim 3, characterized in that, In step S2, the backbone network adopts a Swing Transformer structure and embeds spatial feature transformation layers in the shallow and mid-layer feature extraction, utilizing illumination features. As a physical modulation condition, spatial affine transformation is performed on image features. The process includes: For input features Its spatial modulation formula is: ; ; ; Scaling convolutional layers are used to learn spatial scaling factors from illumination features. ; This is a translation convolutional layer used to learn the spatial translation factor from illumination features. ; For output features; This is an element-wise multiplication operation; Scaling factor learned through space Spatial translation factor This enables pixel-by-pixel guidance of image features based on physical illumination priors.

5. The physically driven end-to-end 3D reconstruction method for low reflectivity targets according to claim 1, characterized in that, In step S2, the encoding part of the first-stage deep neural network includes a signal-to-noise ratio (SNR) sensing encoder; the SNR sensing encoder includes a denoising network, an original feature projection layer, and a gating network, used to adaptively soft-fuse the denoised features and the original features according to the clean mask; the specific calculation process is as follows: The enhancement tensor is processed using a denoising network and the original feature projection layer, respectively. , to obtain denoising features With original features ; Processing pure masks using gating networks Generate spatial fusion weights restricted to a specific range. : ; For gating networks; Through the fusion weight The two features are weighted and fused pixel by pixel to output the signal-to-noise ratio (SNR) perceived feature. : ; in, This indicates an element-wise multiplication operation.

6. The physically driven end-to-end 3D reconstruction method for low reflectivity targets according to claim 1, characterized in that, In step S2, the backbone network of the first-stage deep neural network employs an illumination-guided multi-head self-attention mechanism at the bottleneck layer; when calculating self-attention, the projected illumination features are incorporated into the calculation of the value matrix, specifically as follows: Map the input features to the query matrix Key matrix and initial value matrix ; for the query matrix Bond matrix The feature dimensions are subjected to L2 normalization; the downsampled illumination features are then utilized. Physical modulation of the initial value matrix yields the modulation value matrix. : ; The formula for calculating multi-head self-attention guided by illumination is: ; in, These are learnable adaptive scaling parameters; The calculation results are linearly projected and then added to the positional code generated based on the modulation value matrix to output the final bottleneck feature. Finally, the deep decoded features are reconstructed into the complex physical response of the target through a shared feature mapping layer and a decoupled dual-branch prediction head; The dual-branch prediction head includes a modulation branch and a phase direction branch; the modulation branch is transmitted through... The activation function outputs an absolutely normalized modulation distribution. The phase direction branch outputs the original phase features, which are then constrained to a unit vector through L2 normalization, and the cosine components are output. With sinusoidal components Finally, by combining the physical formulas, a high-precision complex physical response is reconstructed, which includes the real part. With the imaginary part : ; 。 7. The physically driven end-to-end 3D reconstruction method for low reflectivity targets according to claim 1, characterized in that, In step S3, the cross-interference cost volume is constructed to extract sub-pixel physical initial values, including: One-dimensional discrete sampling sequences are defined along the epipolar direction and perpendicular to the epipolar direction, respectively. Based on the epipolar start point, epipolar end point, and normal vector, the discrete sampling sequences are mapped to the absolute pixel physical coordinate system. Combined with the physical coordinates corresponding to each frequency point in the sparse frequency sampling list, the theoretical interference phase corresponding to each sampling point is calculated. ; The real part of the complex physical response at each frequency extracted by the first-stage deep neural network. With the imaginary part Substituting the physical interference superposition equation, the total energy of each sampling point under multi-frequency interference is calculated, generating the cross cost volume corresponding to the epipolar direction and the perpendicular epipolar direction. : ; Cross cost volume A non-negative truncation operation is performed to eliminate negative interference responses; the position of the highest energy peak is searched within the truncated cross cost volume, and a local neighborhood window is constructed centered on this highest energy peak position; normalized weights are calculated using the energy values ​​of each sampling point within the local neighborhood window, and the sub-pixel physical initial values ​​of the relative positions in the epipolar direction are analyzed by weighted summation of local centroids. Subpixel physical initial value of pixel offset perpendicular to the epipolar direction .

8. The physically driven end-to-end 3D reconstruction method for low reflectivity targets according to claim 1, characterized in that, In step S3, the process of reconstructing the two-dimensional epipolar projection coordinates includes: The complex physical feature extractor extracts and aggregates features from the sub-pixel physical initial value and the complex physical response to generate physical interference correlation features, local context features and geometric query vectors. Using a gated loop unit as the core updater, in the first... In this iteration: the physical interference correlation features, local context features, geometric query vector, and latent state features from the previous time step are used as inputs; the gated recurrent unit outputs the position parameters at the current time step. and bias parameters The residual increment; using the residual increment and the built-in adaptive scaling gating mechanism, the parameters from the previous time step are updated: ; ; in, This is a truncation function; After the iteration terminates, based on the final output of the normalized position along the epipolar line and the offset perpendicular to the epipolar line, the two-dimensional epipolar projection coordinates are calculated using the vector composition formula through the differentiable geometric projection module. : ; The coordinates of the starting point of the polar line; is the unit basis vector along the polar direction; Let be the unit normal vector perpendicular to the polar direction. is the polar length.

9. The physically driven end-to-end 3D reconstruction method for low reflectivity targets according to claim 1, characterized in that, During the training process in step S4, the geometric projection error on the two-dimensional image plane is minimized. To suppress physical noise interference; The loss function for the second-stage deep neural network is defined as: ; in, This represents the total number of samples in the training batch. These are the two-dimensional epipolar projection coordinates output by the network forward propagation. The true values ​​of the corresponding two-dimensional epipolar projection coordinates. It is an L1 norm; The training strategy employs the AdamW optimizer with a weight decay mechanism for gradient descent of network parameters and weight updates, and utilizes a single-cycle learning rate decay strategy to dynamically warm up and adjust the learning rate. By combining the optimization strategy with minimizing the geometric projection error, the network is forced to learn the mapping relationship from two-dimensional epipolar projection coordinates to three-dimensional point cloud coordinates.