A physically guided non-line-of-sight three-dimensional imaging method and system

By combining forward and reverse light transport models with the 3D Sobel convolution operator, a self-supervised training framework is constructed, which solves the problem of unstable reconstruction quality in non-view-of-view imaging, achieves efficient 3D reconstruction and generalization capabilities, and reduces data dependence.

CN122435162APending Publication Date: 2026-07-21NANCHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANCHANG UNIV
Filing Date
2026-06-22
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing non-view-of-sight imaging techniques are susceptible to noise, model assumptions, and acquisition parameters, resulting in reconstruction results with background noise, blurred edges, loss of details, and artifacts. Deep learning methods rely on a large number of paired 3D ground truth labels, which have insufficient generalization ability.

Method used

A physical-guided approach is adopted, which combines forward and reverse light transmission models with a 3D Sobel convolution operator to construct a self-supervised training framework. 3D voxel reconstruction is performed using unlabeled transient data, and light propagation law constraints and gradient regularization are introduced to reduce the dependence on labeled data.

Benefits of technology

It improves the stability and clarity of reconstruction results, enhances the model's generalization ability under different target categories and acquisition conditions, and reduces the cost of acquiring training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122435162A_ABST
    Figure CN122435162A_ABST
Patent Text Reader

Abstract

The application discloses a kind of physical guiding non-visual field three-dimensional imaging method and system, it is related to imaging technical field, the method includes obtaining unlabelled target transient data, input initial three-dimensional voxel reconstruction network obtains three-dimensional predicted voxel;Three-dimensional predicted voxel is input into forward light transmission model to generate simulated transient data, and calculate forward loss;Target transient data is input into reverse light transmission model to obtain physical reconstruction voxel, and calculate reverse loss;Three-dimensional Sobel convolution operator is used to calculate the voxel gradient of three-dimensional predicted voxel in three orthogonal directions, and the voxel space regularization loss is determined;Total loss function value is determined based on the above three losses, and network parameters are updated to convergence by back propagation, and target three-dimensional voxel reconstruction network is obtained;To-be-measured transient data is input into target network, and target three-dimensional voxel is output;The method reduces dependence on labeled data, improves the physical consistency, definition and generalization ability of reconstruction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of imaging technology, and in particular to a physically guided non-viewpoint three-dimensional imaging method and system. Background Technology

[0002] Non-line-of-sight imaging (NLOS imaging) is a computational optics technique that uses indirectly reflected light signals to image hidden targets outside the line of sight. This technique typically involves illuminating a relay wall or other diffuse reflective surface with a laser, acquiring the time-of-flight transient signals returned after multiple scatterings from the hidden target, and then using computational methods to reconstruct the target's three-dimensional shape, position, or reflectivity distribution. Existing research indicates that NLOS imaging can achieve the reconstruction, localization, and identification of hidden objects by capturing and analyzing the multiple scattered photons generated by them.

[0003] Existing non-line-of-sight (NLOS) imaging techniques are mainly divided into two categories: one is physical model-based reconstruction methods. These methods rely on pre-defined light transmission models and system geometric parameters, and are easily affected by model approximations, scanning parameters, noise, wall reflection characteristics, and the selection of filtering parameters. In complex hidden target or real experimental environments, this method often suffers from strong background noise, blurred edges, loss of details, or artifact residue, resulting in limited reconstruction fidelity; that is, the reconstruction quality of this method is limited by model assumptions and parameter sensitivity. The other category is data-driven methods based on deep learning. These methods require a large amount of paired training data, that is, they need to obtain transient measurement data and corresponding 3D ground truth labels simultaneously. However, the 3D ground truth of hidden targets in real NLOS scenarios is difficult to obtain directly, resulting in high training data construction costs and limited applicability. At the same time, pure data-driven models often do not fully embed the physical constraints of light transmission. When the test target category, scanning method, acquisition parameters, or noise distribution are inconsistent with the training data, the model is prone to insufficient generalization ability, reconstruction distortion, local artifacts, and noise sensitivity; that is, this method faces practical limitations such as computational complexity, noise sensitivity, and reliance on large-scale labeled data. Summary of the Invention

[0004] Based on this, this application provides a physically guided non-viewpoint 3D imaging method and system, aiming to solve the technical problems in related non-viewpoint 3D imaging methods, such as the physical model method being easily affected by noise, model assumptions and acquisition parameters, resulting in background noise, blurred edges, loss of details and artifacts in the reconstruction results, as well as the technical problems of existing deep learning methods relying on a large number of paired 3D ground truth labels and insufficient generalization ability under the condition of no target category or different scanning configurations.

[0005] In a first aspect, embodiments of this application provide a physically guided non-line-of-sight three-dimensional imaging method, including: Acquire unlabeled target transient data, input the target transient data into a preset initial three-dimensional voxel reconstruction network, and output three-dimensional predicted voxels; The three-dimensional predicted voxels are input into a preset forward light transmission model to remap the light propagation path, output simulated transient data, and the forward loss is determined based on the simulated transient data and the target transient data. The target transient data is input into a preset reverse optical transmission model to reverse map the transient data to the hidden space, thereby obtaining a physically reconstructed voxel. Based on the physically reconstructed voxel and the three-dimensional predicted voxel, the reverse loss is determined. The three-dimensional Sobel convolution operator is used to convolve the three-dimensional predicted voxels in three orthogonal directions to obtain multiple voxel gradients. Based on the multiple voxel gradients, the regularization loss of the voxel space is determined. Based on the forward loss, the inverse loss, and the regularization loss, the total loss function value is determined, and backpropagation is performed based on the total loss function value to update the network parameters of the initial 3D voxel reconstruction network until the loss converges, thus obtaining the target 3D voxel reconstruction network. The transient data to be tested is input into the target three-dimensional voxel reconstruction network, and the target three-dimensional voxel is output.

[0006] In some embodiments, in the steps of acquiring unlabeled target transient data, inputting the target transient data into a preset initial three-dimensional voxel reconstruction network, and outputting three-dimensional predicted voxels, the initial three-dimensional voxel reconstruction network adopts a hybrid U-Net architecture that integrates two-dimensional and three-dimensional convolutions. The hybrid U-Net architecture includes an encoder composed of multiple downsampling modules, a decoder composed of multiple upsampling modules, and skip connections set between corresponding levels of the encoder and the decoder.

[0007] In some embodiments, the step of acquiring unlabeled target transient data, inputting the target transient data into a preset initial 3D voxel reconstruction network, and outputting 3D predicted voxels includes the following data processing procedure for the initial 3D voxel reconstruction network: The target transient data is subjected to three-dimensional convolution and ReLU activation to multiply its feature channel number and obtain an initial feature map; The initial feature map is subjected to alternating downsampling processes of max pooling and three-dimensional convolutional blocks, which halves the spatiotemporal resolution of the feature map layer by layer, resulting in a deep semantic feature map with a preset number of channels. The deep semantic feature map is upsampled, and the features of each level during the upsampling process are spliced ​​and fused with the downsampled features of the corresponding level to restore spatial resolution and fuse multi-scale information to obtain three-dimensional voxel features. The three-dimensional voxel features are subjected to a three-dimensional to two-dimensional feature dimensionality reduction mapping to map the time-domain flight features into spatial-domain depth slice features, and the three-dimensional predicted voxel of the hidden target is output.

[0008] In some embodiments, the step of inputting the three-dimensional predicted voxel into a preset forward light transmission model to remap the light propagation path and output simulated transient data includes the following data processing procedure for the forward light transmission model: Based on the laser scanning point position, detection point position, voxel spatial coordinates, light speed, distance attenuation coefficient, and voxel albedo in the three-dimensional predicted voxel, the contribution value of each voxel to the transient signal is determined. The contribution values ​​are time-binned using a differentiable kernel density estimation method. The contribution values ​​of all voxels after time-binning are summed to obtain the simulated transient data for predicting flight time.

[0009] In some embodiments, in the step of inputting the three-dimensional predicted voxels into a preset forward light transmission model to remap the light propagation path and output simulated transient data, the calculation expression for the physical rendering process of the forward light transmission model is as follows:

[0010] In the formula, The transient light intensity probability distribution density. To hide the location in the space x Physical reconstruction voxels at the site, D For distance attenuation, n Where h is the number of samples, and h is the kernel density estimation bandwidth. For kernel function, c At the speed of light, x The three-dimensional spatial coordinates of the voxel to be calculated in the hidden space. For the first i The spatial coordinates of a laser scanning point or sampling point.

[0011] In some embodiments, the step of inputting the target transient data into a preset reverse optical transmission model to reverse-map the transient data to the hidden space and obtain the physically reconstructed voxel includes: The reverse optical transmission model employs a deconvolution algorithm and a wave algorithm to reverse-transmit the transient data of the target, thereby calculating the three-dimensional structure of the object in the hidden space and obtaining physically reconstructed voxels. The calculation expression for the physically reconstructed voxels is as follows:

[0012] In the formula, For scanning points or illumination points The transient light intensity signal collected at the location, The location of the scanning point or illumination point on the relay surface. t For photon flight time or time sampling variables, For the Dirac function, r The total optical path length from the relay surface to the hidden target and back to the detector. To integrate the positions of the scanning points on the relay surface, To integrate over the time dimension; For confocal mode:

[0013] For non-confocal modes:

[0014] In the formula, This indicates the location of the detection point.

[0015] In some embodiments, the step of performing convolution processing on the three-dimensional predicted voxels using a three-dimensional Sobel convolution operator in three orthogonal directions to obtain multiple voxel gradients, and determining the voxel space regularization loss based on the multiple voxel gradients, includes: The 3D prediction voxels are convolved using a 3×3×3 scale Sobel convolution operator in three orthogonal directions in space. The calculation expression is as follows:

[0016] In the formula, These are the three-dimensional Sobel convolution kernels in the x, y, and z directions, respectively. For convolution operations, For three-dimensional prediction voxels; Based on the voxel gradients in three directions, the regularization loss in the voxel space is calculated, where the regularization loss calculation expression is:

[0017] In the formula, The gradient energy aggregation function, Let x be the voxel gradient in the x-direction. Let be the voxel gradient in the y-direction. Let be the voxel gradient in the z-direction.

[0018] Compared with the prior art, the technical solution provided in the first aspect of this application includes at least the following beneficial effects or advantages: The imaging method provided in this application generates simulated transient data by inputting 3D predicted voxels into a forward light transmission model and calculating the forward loss, ensuring that the 3D reconstruction results output by the network conform to the actual light propagation laws and reducing the occurrence of false structures and non-physical reconstruction results. The target transient data is input into a reverse light transmission model to generate physically reconstructed voxels and the reverse loss is calculated, introducing physical prior constraints into the network training process. Combined with the feature extraction and nonlinear expression capabilities of neural networks, this suppresses background noise and artifacts in the reconstruction results, improving their stability. Furthermore, by calculating the forward and reverse light transmission consistency losses, the network can utilize unlabeled transient measurement data for self-supervised training, reducing reliance on manually labeled data or real 3D labeled data and lowering the cost of acquiring training data. Simultaneously, a 3D Sobel convolution operator is used to calculate the voxel gradients of the 3D predicted voxels in three orthogonal directions and construct a regularization loss, enhancing the contour structure and geometric edges of the hidden target, reducing edge blurring and detail loss, and improving the clarity of the 3D reconstruction results. By jointly driving the network parameter updates through forward loss, inverse loss, and regularization loss, the network learns physical features related to non-view imaging mechanisms, rather than simply memorizing the distribution of training samples. This improves the model's generalization ability under different target categories and acquisition conditions, making it suitable for various non-view imaging systems and different application scenarios.

[0019] Secondly, embodiments of this application provide a physically guided non-line-of-sight three-dimensional imaging system, comprising: The voxel reconstruction module is configured to acquire unlabeled target transient data, input the target transient data into a preset initial three-dimensional voxel reconstruction network, and output three-dimensional predicted voxels. The positive constraint module is configured to input the three-dimensional predicted voxel into a preset positive light transmission model to remap the light propagation path, output simulated transient data, and determine the positive loss based on the simulated transient data and the target transient data. The reverse processing module is configured to input the target transient data into a preset reverse optical transmission model to reverse map the transient data to the hidden space, obtain a physical reconstruction voxel, and determine the reverse loss based on the physical reconstruction voxel and the three-dimensional prediction voxel. The voxel regularization module is configured to perform convolution processing on the three-dimensional predicted voxels in three orthogonal directions using a three-dimensional Sobel convolution operator to obtain multiple voxel gradients, and determine the regularization loss of the voxel space based on the multiple voxel gradients. The network training module is configured to determine the total loss function value based on the forward loss, the inverse loss, and the regularization loss, and to perform backpropagation based on the total loss function value to update the network parameters of the initial 3D voxel reconstruction network until the loss converges, thereby obtaining the target 3D voxel reconstruction network. The reconstruction output module is configured to input the transient data to be tested into the target three-dimensional voxel reconstruction network and output the target three-dimensional voxel.

[0020] Thirdly, this application also provides an electronic device, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed, enable the at least one processor to perform the steps of the physically guided non-viewpoint 3D imaging method provided in the first aspect above.

[0021] Fourthly, this application also provides a computer-readable storage medium storing computer program instructions thereon, which, when executed by a processor, implement the steps of the physically guided non-viewpoint three-dimensional imaging method provided in the first aspect above.

[0022] It is understood that the beneficial effects of the technical solutions provided in the second, third and fourth aspects above can be found in the relevant descriptions in the first aspect above, and will not be repeated here.

[0023] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0024] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 A flowchart of a physically guided non-viewpoint 3D imaging method provided in an embodiment of this application; Figure 2 This is a schematic diagram of physical-guided self-supervised learning corresponding to the non-view imaging method provided in the embodiments of this application; Figure 3 This is a schematic diagram of the experimental scenario and experimental apparatus provided in the embodiments of this application; Figure 4 This is a schematic diagram of the reconstruction results of the experimental dataset of the physical-guided non-viewpoint 3D imaging method provided in the embodiments of this application; Figure 5 A structural block diagram of a physically guided non-viewpoint 3D imaging system provided in an embodiment of this application; Figure 6 A structural block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0026] To facilitate understanding of this application, a more complete description will be provided below with reference to the accompanying drawings, which illustrate several embodiments of the present application. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that the disclosure of this application will be thorough and complete.

[0027] It should be noted that, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0028] Please see Figures 1 to 2 , Figure 1 A flowchart of the physically guided non-viewpoint 3D imaging method provided in this embodiment; Figure 2 This is a schematic diagram of the physical-guided self-supervised learning method corresponding to the non-viewpoint imaging method provided in this embodiment, wherein, Figure 2 (a) shows the calculation process of the forward loss; (b) shows the process of reconstructing three-dimensional voxels from time-of-flight data through the U-net network; (c) shows the calculation process of the inverse loss; (d) shows the calculation process of the regularization loss. This embodiment provides a physically guided non-view-of-sight three-dimensional imaging method, specifically including steps S10 to S60.

[0029] Step S10: Obtain unlabeled target transient data, input the target transient data into a preset initial three-dimensional voxel reconstruction network, and output three-dimensional predicted voxels; In this step, the target transient data is Time of Flight (ToF) transient measurement data. For unlabeled target transient data, this method may include a pulsed laser, a scanning unit, a relay surface, a hidden target, a detector, and a time-correlated single-photon counting unit. The pulsed laser is used to emit short pulses of light onto the relay surface; the scanning unit controls the laser to scan the relay surface point by point; the hidden target is located in a position that the detector cannot directly observe; the detector receives the light signals returned by the hidden target after multiple scatterings from the relay surface; the time-correlated single-photon counting unit records the arrival time of photons at each scanning point, forming the Time of Flight transient measurement data. The ToF transient measurement data can be expressed as:

[0030] In the formula, H and W are the two-dimensional spatial resolution of the relay surface scanning points, and B is the number of time bins; For example, in terms of data dimensions, the spatial resolution of each set of transient measurement data is uniformly downsampled and normalized to... It is divided into 512 time histogram intervals along the time dimension, and the width of each time interval can be precisely set to 33 ps. The scanning resolution, time resolution, exposure time and detector type can also be changed according to the actual system.

[0031] Optionally, the acquired initial ToF transient measurement data can be preprocessed to adapt to the input of the preset initial 3D voxel reconstruction network. This preprocessing includes, but is not limited to, one or more of the following operations: 1. Background subtraction of transient data to remove the influence of ambient light and system dark count; 2. Calibration of the time zero point of different scan points; 3. Normalization of ToF data; 4. Clipping of the time dimension to retain only the time window containing valid echoes; 5. Filtering, smoothing, or resampling of low signal-to-noise ratio data; 6. Size adjustment of the input data to conform to the input format of the initial 3D voxel reconstruction network.

[0032] In some embodiments, the expression for the initial 3D voxel reconstruction network used to map input ToF transient measurement data to 3D voxel results of the hidden target can be:

[0033] In the formula, For parameters neural networks, For ToF transient measurement data, The result is a 3D voxel reconstruction output by the network.

[0034] Furthermore, the initial 3D voxel reconstruction network can be a hybrid U-Net architecture that integrates 2D and 3D convolutions. This hybrid U-Net architecture includes an encoder composed of multiple downsampling modules, a decoder composed of multiple upsampling modules, and skip connections between corresponding layers of the encoder and decoder. The goal of this network is to reconstruct the 3D voxel physical space representation of the hidden target end-to-end directly from the input time-of-flight transient measurement data. The data processing of the initial 3D voxel reconstruction network includes performing 3D convolution and ReLU activation on the transient data to enhance its features. The number of channels is multiplied to obtain an initial feature map. The initial feature map is then subjected to alternating downsampling processes of max pooling and 3D convolutional blocks, which halves the spatiotemporal resolution of the feature map layer by layer to obtain a deep semantic feature map with a preset number of channels. The deep semantic feature map is then upsampled, and the features of each level during the upsampling process are concatenated and fused with the downsampled features of the corresponding level to restore the spatial resolution and fuse multi-scale information to obtain 3D voxel features. The 3D voxel features are then subjected to 3D-to-2D feature dimensionality reduction mapping to map the temporal domain flight features to spatial domain depth slice features, and the 3D predicted voxels of the hidden target are output.

[0035] Combination Figure 2 As shown in one example, during the encoding stage of feature extraction, the network input is normalized transient time-of-flight measurement data, whose initial physical tensor dimension can contain 512 time histogram intervals and a spatial scan grid resolution of 64×64. The encoder progressively extracts hierarchical features through cascaded downsampling modules, compressing the high-dimensional photon time-of-flight data into a latent feature representation. The input data first passes through 3D convolutional and ReLU activation layers, increasing the number of feature channels to 64; subsequently, the network alternates between max pooling layers and 3D convolutional blocks, causing the spatiotemporal resolution of the feature map to be halved layer by layer (increased to 256×32). 2 128×16 2 The number of feature channels expands exponentially (increasing to 128 and 256 respectively) until it reaches the bottom layer of the network, forming a dimension of only 64×8. 2 However, the deep semantic feature representation has as many as 512 channels. Conversely, in the decoding stage of feature recovery, the decoder gradually recovers the spatiotemporal resolution of the feature map through upsampling operations and concatenates it with the corresponding feature map of the encoder, thereby fusing multi-scale information to generate a 3D voxel reconstruction result.

[0036] In this example, considering that the non-view reconstruction ultimately needs to output 3D spatial voxels represented by depth (Z-axis), a 3D to 2D feature dimensionality reduction transition mechanism is designed at the end of the network. When the number of output channels of the 3D decoder is compressed to a single channel (dimension 1×512×64),... 2Afterwards, the network eliminates redundant single-channel dimensions through a squeezing operation, directly transforming the 512 dimensions containing time-flight information into the feature channel number of subsequent 2D convolutional modules. This 3D-2D hybrid mechanism mathematically and architecturally achieves a smooth mapping from time-domain flight features to spatial-domain depth slices, ultimately outputting a 3D prediction voxel in the hidden space.

[0037] Step S20: Input the three-dimensional predicted voxel into a preset forward light transmission model to remap the light propagation path, output simulated transient data, and determine the forward loss based on the simulated transient data and the target transient data; In some embodiments, after the network deconstructs the three-dimensional predicted voxels of the hidden space in step S10, the complete physical process of photons propagating from the surface of the hidden target through a non-viewpoint path to the measurement plane is simulated based on a preset forward light transmission model. This allows the three-dimensional voxels predicted by the network to be rendered in reverse as simulated time-of-flight transient measurement data, i.e., simulated transient data. The model determines the contribution value of each voxel to the transient signal based on the laser scanning point position, detection point position, voxel spatial coordinates, light speed, distance attenuation coefficient, and voxel albedo in the three-dimensional predicted voxels. The contribution values ​​are time-binned using a differentiable kernel density estimation method, and all voxel contribution values ​​after time-binning are accumulated to obtain the simulated transient data of the predicted time of flight.

[0038] In one example, in the forward light propagation model, the model traces the complete light path from the pulsed laser source, through the hidden target voxel space where diffuse reflection occurs, and finally back to the detector, involving three diffuse reflections. During this process, the contribution of each discrete voxel in space to the light intensity received by the detector is influenced by the voxel's own albedo. The weighted modulation, along with the increase in the total optical path, produces a corresponding distance attenuation (the attenuation factor is denoted as 1 / D). The integral form of this physically rendered process can be expressed as:

[0039] in, Indicates spatial location x The albedo or scattering intensity of the voxel. D This represents the distance attenuation term related to the total optical path length. r Indicates the path length of light propagation. c Represents the speed of light. This is the Dirac function.

[0040] In this embodiment, since in general computer graphics rendering and physical modeling, the histogram statistics of photons on the detector are essentially based on the photon arrival time using an indicator function or Dirac delta function, The function performs a hard assignment. This operation is mathematically step-like, discrete, and non-differentiable (its gradient is zero at almost all non-zero points and undefined at the transition point). This non-differentiability directly blocks error backpropagation and gradient updates of depth coordinates in the 3D voxel reconstruction network throughout the closed-loop training process.

[0041] Therefore, to resolve this underlying mathematical conflict and establish a joint optimization path between the physical model and the deep neural network, this embodiment introduces the kernel density estimation (KDE) mechanism from nonparametric statistics. The physical meaning of kernel density estimation lies in using a continuous probability density function to softly assign the arrival time of photons. KDE expands each discrete photon arrival time point into a continuous Gaussian waveform with a specific variance. The complete differentiable forward light transmission model defined after introducing KDE defines the computational expression for the physical rendering process as follows:

[0042] In the formula, The transient light intensity probability distribution density. To hide the location in the space x Physical reconstruction voxels at the site, D For distance attenuation, n Where h is the number of samples, and h is the kernel density estimation bandwidth. For kernel function, c At the speed of light, x The three-dimensional spatial coordinates of the voxel to be calculated in the hidden space. For the first i The spatial coordinates of a laser scanning point or sampling point.

[0043] In some embodiments, relying on the aforementioned differentiable physical rendering engine, the self-supervised framework can generate predicted transient signals highly aligned with real physical processes. This is to quantize the synthesized forward-propagating Time-of-Flight (TOF) transient measurement data (T... pred ) and input TOF transient measurement data (T input The difference between the two is used to construct a positive consistency loss using the SSIM metric. Since the optimization objective of the deep learning optimizer is to minimize the loss function, this loss is defined by taking the negative of SSIM. The positive loss is calculated using the structural similarity loss expression: .

[0044] It should be understood that, through this positive consistency constraint, the three-dimensional reconstruction results output by the network can meet the actual light propagation law, reducing the possible false structures, incorrect completions or non-physical reconstruction results in the pure data-driven method. At the same time, this loss is used to constrain the three-dimensional voxels output by the network to regenerate a ToF signal consistent with the actual measurement data after forward light transmission, thereby ensuring that the network reconstruction results conform to the real light propagation law.

[0045] Step S30: Input the target transient data into a preset reverse optical transmission model to reverse map the transient data to the hidden space to obtain a physical reconstruction voxel, and determine the reverse loss based on the physical reconstruction voxel and the three-dimensional prediction voxel; In some embodiments, it should be noted that relying solely on the constraints of the forward light transport projection of the forward light transport model can easily lead to an extremely large solution space and strong ambiguity in the deep neural network for 3D voxel reconstruction. Therefore, in this embodiment, a reverse light transport model is constructed to generate physically reconstructed voxels based on the input ToF transient measurement data. Its expression is as follows: ,in, This is a reverse light transmission model. V phys This represents the three-dimensional voxel prior obtained from the physical model.

[0046] The reverse optical transmission model employs a Wiener filtering deconvolution algorithm and a wave-based solution algorithm based on Rayleigh-Sommerfeld diffraction to reverse-transmit the transient data of the target, thereby calculating the three-dimensional structure of the object in the hidden space and obtaining the physically reconstructed voxels. The expression for calculating the position of the hidden object is as follows:

[0047] To hide the location in the space x The physical reconstruction voxel at that location represents the albedo or scattering intensity of the hidden target at that location; x The three-dimensional spatial coordinates of the voxel to be reconstructed in the hidden space; D This is a distance attenuation term or a geometric weighting term, used to characterize the energy attenuation caused by the distance the light travels; For scanning points or illumination points The transient light intensity signal collected at the location, i.e., time t The corresponding flight time measurement; The location of the scanning point or illumination point on the relay surface; t For photon flight time or time sampling variables; The Dirac function is used to constrain the relationship between the light propagation path length and the flight time. r The total optical path length from the relay surface to the hidden target and back to the detector; cThe speed of light; This indicates that the integration is performed over the position of the scanned points on the relay surface. This indicates integration over the time dimension.

[0048] For confocal mode:

[0049] For non-confocal modes:

[0050] In the formula, This indicates the location of the detection point.

[0051] In one example, raw time-of-flight transient measurement data is input into a physics-based prior wave field propagation module, employing conventional physical wave propagation algorithms. These algorithms can stably back-map transient data to the hidden space, generating initial inverse physical voxels. However, limited by discretization errors and extremely low signal-to-noise ratios, the directly generated inverse physical voxels inevitably retain some artifacts. If directly used for supervised learning in deep networks, the network is prone to overfitting to these redundant noises. To address this, an artifact recognition strategy based on adaptive threshold segmentation is implemented for the voxels. The initial voxels are projected along specific depth or spatial dimensions with maximum intensity, and binarized using relative intensity thresholds to extract a two-dimensional mask containing only the core target contour. This mask is then extended back to the three-dimensional dimension and subjected to spatial Hadamard product filtering with the original voxels. After this operation, the artifacts in the background region are zeroed out, resulting in a relatively pure physics prior. V phys Then the data flows into the inverse loss calculation module. This is to quantify the prediction voxels deconstructed from the deep neural network backbone. Ⅴ rec ) and physical priors ( V phys To approximate the three-dimensional spatial distribution between the two models, this model uses Mean Squared Error (MSE) as the loss evaluation function, and its expression is:

[0052] Or it can be expressed as:

[0053] It should be understood that, through joint driving with the positive loss in the previous section, the bidirectional physical constraint enables the entire framework to achieve a closed loop between data feature driving and underlying physical laws. At the same time, this loss is used to ensure that the output of the neural network is consistent with the reconstruction result of the physical model, and to avoid the network generating false structures that do not conform to the actual optical propagation laws during the unlabeled training process.

[0054] It should also be understood that traditional physical model methods are easily affected by weak light signals, background noise, scanning parameters, and model approximations, often resulting in background noise and artifacts in the reconstruction results. In this embodiment, the physical reconstruction results generated by the reverse light transmission model are introduced as prior constraints during the training process. Combined with the feature extraction and nonlinear expression capabilities of the neural network, the network further suppresses noise and artifacts while adhering to the physical priors, thereby improving the stability of the reconstruction results.

[0055] Step S40: The three-dimensional Sobel convolution operator is used to perform convolution processing on the three-dimensional predicted voxels in three orthogonal directions to obtain multiple voxel gradients. Based on the multiple voxel gradients, the regularization loss of the voxel space is determined. In some embodiments, to further enhance reconstruction quality, a voxel regularization loss (Lreg) based on the 3D Tenengrad function is introduced into the three-dimensional physical feature space. It should be noted that the classic Tenengrad function is widely used for sharpness evaluation of two-dimensional images, and its core idea is to maximize the local gradient energy of the image. In this embodiment, it is extended to the three-dimensional voxel space, and a custom-designed 3×3×3 scale three-dimensional Sobel convolution operator is used. Assume that the three-dimensional predicted voxels output by the network are V. rec Its local gradient tensors in the three orthogonal directions (horizontal X-direction, vertical Y-direction, and depth Z-direction) can be calculated using specific 3D convolutional blocks, and their calculation expressions are as follows:

[0056] In the formula, These are the three-dimensional Sobel convolution kernels in the x, y, and z directions, respectively. For convolution operations, For three-dimensional prediction voxels; In this embodiment, considering the physical anisotropy of the system in terms of lateral spatial resolution and longitudinal depth resolution, the regularization loss in voxel space is calculated based on voxel gradients in three directions. The regularization loss calculation expression is as follows:

[0057] In the formula, This is the gradient energy aggregation function, used to sum or average the squared gradients at each voxel location in the 3D voxel space to obtain the voxel space regularization loss in scalar form. Let x be the voxel gradient in the x-direction. Let be the voxel gradient in the y-direction. Let be the voxel gradient in the z-direction.

[0058] It should be noted that this regularization term is used to preserve the geometric edges and structural details of the hidden target while suppressing isolated noise points and background artifacts. The 3D Tenengrad regularization calculates gradients in each direction using the 3D Sobel operator, which helps to preserve key geometric edges and suppress noise artifacts.

[0059] Step S50: Based on the forward loss, the inverse loss, and the regularization loss, determine the total loss function value, and perform backpropagation based on the total loss function value to update the network parameters of the initial 3D voxel reconstruction network until the loss converges, thereby obtaining the target 3D voxel reconstruction network. In some embodiments, the total loss function value is determined based on the forward loss, the inverse loss, and the regularization loss. The calculation of the total loss function value is expressed as follows:

[0060] In the formula, For positive consistency loss, For reverse consistency loss, For regularization loss, The weighting coefficients for the positive consistency loss are... The weighting coefficients for the reverse consistency loss are... The weighting coefficients for the regularization loss.

[0061] For example, the training process adopts batch training, dividing the unlabeled transient data into multiple training batches. After each batch is input into the initial 3D voxel reconstruction network, the forward loss, inverse loss, and voxel space regularization loss of the corresponding batch are calculated in sequence. After the total loss function value is obtained by weighting, the backpropagation operation is performed through the adaptive moment estimation optimizer to calculate the gradient of the parameters of each layer of the network. The convolutional kernel weights and bias parameters of the network are updated according to the gradient values ​​and the preset learning rate. The above process is repeated until the loss function converges or the preset number of training rounds is reached.

[0062] It should be noted that, since existing supervised learning models typically rely on the distribution of training data, changes in the target category, scanning range, scanning resolution, and confocal or non-confocal acquisition method can easily lead to a decrease in reconstruction quality or structural distortion. By using both forward and reverse light transport models to constrain neural network training, the network learns physical features related to non-viewpoint imaging mechanisms, rather than simply memorizing the distribution of training samples. Therefore, it can better adapt to unseen target categories, different scanning parameters, and real-world experimental scenarios.

[0063] It should also be noted that traditional physical model methods have clear optical meanings, but the reconstruction quality is easily affected by noise and parameters; deep learning methods have strong feature representation capabilities, but they tend to rely on labeled data and lack physical constraints. This embodiment unifies differentiable forward light transmission, inverse physical reconstruction, and neural network reconstruction into the same training framework, so that the reconstruction results are both constrained by the laws of optical propagation and can utilize neural networks to improve denoising and detail recovery capabilities.

[0064] Step S60: Input the transient data to be tested into the target three-dimensional voxel reconstruction network and output the target three-dimensional voxel.

[0065] In some embodiments, once the trained 3D voxel reconstruction network is completed, the Time-of-Flight (ToF) transient data to be measured is input into the trained 3D voxel reconstruction network to output the 3D voxel result of the hidden target. It should be understood that, depending on the actual application requirements, the 3D voxel result can be further converted into a 2D projection map, depth map, albedo map, point cloud, 3D mesh model, or target recognition result. For scenarios requiring real-time processing, fast inference can be performed using only the trained neural network; for scenarios with high accuracy requirements, a differentiable forward light transmission module can be further used to perform physical consistency verification or iterative correction on the output result.

[0066] In the above method steps, by inputting 3D predicted voxels into the forward light transmission model to generate simulated transient data and calculating the forward loss, the 3D reconstruction results output by the network conform to the actual light propagation law, reducing the occurrence of false structures and non-physical reconstruction results. The target transient data is input into the reverse light transmission model to generate physically reconstructed voxels and the reverse loss is calculated, introducing physical prior constraints into the network training process. Combined with the feature extraction and nonlinear expression capabilities of the neural network, background noise and artifacts in the reconstruction results are suppressed, improving the stability of the reconstruction results. Furthermore, by calculating the forward and reverse light transmission consistency losses, the network can utilize unlabeled transient measurement data for self-supervised training, thereby reducing dependence on manually labeled data or real 3D labeled data and lowering the cost of acquiring training data. Simultaneously, a 3D Sobel convolution operator is used to calculate the voxel gradients of the 3D predicted voxels in three orthogonal directions and construct a regularization loss, enhancing the contour structure and geometric edges of the hidden target, reducing edge blurring and detail loss, and improving the clarity of the 3D reconstruction results. By jointly driving the network parameter updates through forward loss, inverse loss, and regularization loss, the network learns physical features related to non-view imaging mechanisms, rather than simply memorizing the distribution of training samples. This improves the model's generalization ability under different target categories and acquisition conditions, making it suitable for various non-view imaging systems and different application scenarios.

[0067] Please see Figure 3 and Figure 4This embodiment provides an experiment based on the physically guided non-viewpoint 3D imaging method described in the above embodiments and existing technical methods, wherein... Figure 3 The corresponding experimental setup includes a single-photon avalanche diode, a lens, a laser, a non-polarizing beam splitter, and a scanning mirror. Figure 4 (a) Experimental data obtained on a 0.8m × 0.8m relay surface, and (b) Experimental data obtained on a 1m × 1m relay surface. Specifically, transient data captured on the 0.8m × 0.8m relay surface were used to reconstruct the objects of "leaf center" and "running person" (minimum linewidth 3cm); while Figure 4 (b) shows the results of reconstructing the letters “Z” and “L&T” (minimum linewidth 10cm) from data captured on a 1m×1m relay plane. Since these scanning parameters are inconsistent with the data used to train the U-net and Transformer models, these two methods are almost impossible to reconstruct and are therefore not suitable for comparison. In the traditional algorithm, the letter “Z” and both letters “L&T” can be reconstructed relatively accurately. However, the light cone transform algorithm introduces some artifacts, resulting in blurred edges. The frequency-wavenumber domain algorithm suffers from excessive background noise, while the phasor field algorithm yields relatively clear results, but still has residual background noise and imperfections (e.g., the letter “T”). For smaller, more complex objects, such as the “heart of a leaf” and the “running person,” the performance of existing methods degrades. The output of the frequency-wavenumber domain algorithm is overwhelmed by background noise, leading to lost contours or even unrecognizable structures (e.g., the “heart of a leaf”). Although the light cone transform algorithm performs relatively well, it still cannot resolve finer details, such as the forearm of the “running person.” In comparison, the phasor field algorithm exhibits superior performance, providing relatively clear contours and structural details, but there is still room for optimization in terms of background noise and edge reconstruction. The adaptive axial correction algorithm, based on Wiener filtering deconvolution and Rayleigh-Sommerfeld diffraction wave-based solution, achieves overall performance comparable to the optical conic transform and frequency-wavenumber domain algorithms. However, the physically guided non-viewpoint 3D imaging method described in the embodiments of this application not only effectively adapts to various scanning configurations but also achieves superior reconstruction quality. It accurately and clearly captures contour structures, significantly suppresses background noise, and resolves fine details, demonstrating a clear advantage over existing methods.

[0068] Please see Figure 5 , Figure 5 This embodiment illustrates a physically guided non-viewpoint 3D imaging system, which is used to perform the above-described... Figure 1 The steps in the corresponding embodiments. Please refer to the details. Figure 1 as well as Figure 1The relevant descriptions in the corresponding embodiments are shown below. For ease of explanation, only the parts relevant to this embodiment are shown. See also... Figure 5 The physically guided non-line-of-sight 3D imaging system 200 includes: The voxel reconstruction module 210 is configured to acquire unlabeled target transient data, input the target transient data into a preset initial three-dimensional voxel reconstruction network, and output three-dimensional predicted voxels. The positive constraint module 220 is configured to input the three-dimensional predicted voxel into a preset positive light transmission model to remap the light propagation path, output simulated transient data, and determine the positive loss based on the simulated transient data and the target transient data. The reverse processing module 230 is configured to input the target transient data into a preset reverse optical transmission model to reverse map the transient data to the hidden space, obtain a physical reconstruction voxel, and determine the reverse loss based on the physical reconstruction voxel and the three-dimensional prediction voxel. The voxel regularization module 240 is configured to perform convolution processing on the three-dimensional predicted voxels in three orthogonal directions using a three-dimensional Sobel convolution operator to obtain multiple voxel gradients, and determine the regularization loss of the voxel space based on the multiple voxel gradients. The network training module 250 is configured to determine the total loss function value based on the forward loss, the inverse loss, and the regularization loss, and to perform backpropagation based on the total loss function value to update the network parameters of the initial three-dimensional voxel reconstruction network until the loss converges, thereby obtaining the target three-dimensional voxel reconstruction network. The reconstruction output module 260 is configured to input the transient data to be tested into the target three-dimensional voxel reconstruction network and output the target three-dimensional voxel.

[0069] The physically guided non-viewpoint 3D imaging system provided in this embodiment generates simulated transient data by inputting 3D predicted voxels into a forward light transmission model and calculating the forward loss. This ensures that the 3D reconstruction results output by the network conform to the actual light propagation laws, reducing the occurrence of false structures and non-physical reconstruction results. The system also inputs target transient data into a reverse light transmission model to generate physically reconstructed voxels and calculates the reverse loss. This introduces physical prior constraints on the network training process. Combined with the feature extraction and nonlinear expression capabilities of neural networks, this suppresses background noise and artifacts in the reconstruction results, improving their stability. Furthermore, by calculating the forward and reverse light transmission consistency losses, the network can utilize unlabeled transient measurement data for self-supervised training, reducing reliance on manually labeled data or real 3D labeled data and lowering the cost of acquiring training data. Simultaneously, a 3D Sobel convolution operator is used to calculate the voxel gradients of the 3D predicted voxels in three orthogonal directions and construct a regularization loss, enhancing the contour structure and geometric edges of the hidden target, reducing edge blurring and detail loss, and improving the clarity of the 3D reconstruction results. By jointly driving the network parameter updates through forward loss, inverse loss, and regularization loss, the network learns physical features related to non-view imaging mechanisms, rather than simply memorizing the distribution of training samples. This improves the model's generalization ability under different target categories and acquisition conditions, making it suitable for various non-view imaging systems and different application scenarios.

[0070] It should be understood that the modules in the system provided in this embodiment are used to execute... Figure 1 The steps in the corresponding embodiments, and for Figure 1 The steps in the corresponding embodiments have been explained in detail in the above embodiments. Please refer to them for details. Figure 1 as well as Figure 1 The relevant descriptions in the corresponding embodiments will not be repeated here.

[0071] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 6 As shown, the electronic device provided in this embodiment includes at least one processor 901 and a memory 902. Optionally, the electronic device further includes a communication component 903. The processor 901, memory 902, and communication component 903 are connected via a bus 904. In the specific implementation process, at least one processor 901 executes computer execution instructions stored in memory 902, causing at least one processor 901 to execute the aforementioned physically guided non-viewpoint three-dimensional imaging method.

[0072] The specific implementation process of processor 901 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0073] In the above embodiments, it should be understood that the processor 901 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0074] The memory 902 can be an internal storage unit of an electronic device, such as a hard drive or memory in a server. The memory 902 can also be an external storage terminal device of an electronic device, such as a plug-in hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc.

[0075] Furthermore, the memory 902 may include both internal storage units of the electronic device and external storage terminal devices. The memory 902 is used to store computer programs and other programs and data required by the turntable terminal device. The memory 902 can also be used to temporarily store data that has been output or will be output.

[0076] Bus 904 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0077] Those skilled in the art will understand that Figure 6 This is merely an example of an electronic device and does not constitute a limitation on the electronic device. It may include more or fewer components than shown, or combine certain components, or different components. For example, a turntable terminal device may also include input / output terminal devices, network access terminal devices, etc.

[0078] In one embodiment, a computer-readable storage medium is also provided, the computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the physically guided non-viewpoint three-dimensional imaging method as described in the above embodiments.

[0079] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0080] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in application-specific integrated circuits (ASICs). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0081] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0082] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0083] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer-readable storage medium can be non-volatile or volatile. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable storage medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0084] The terms "first," "second," "third," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects and not to describe a particular order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, it may include a series of steps or units, or optionally, steps or units not listed, or other steps or units inherent to these processes, methods, products, or devices.

[0085] The accompanying drawings show only the portions relevant to this application, not all of them. Before discussing exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as sequential processes, many of these operations may be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations may be rearranged. The process may be terminated when its operation is completed, but may also have additional steps not included in the drawings. The process may correspond to a method, function, procedure, subroutine, subprogram, etc.

[0086] The terms “component,” “module,” “system,” “unit,” etc., used in this specification are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a unit can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, a thread of execution, a program, and / or distributed between two or more computers. Furthermore, these units can be executed from various computer-readable media on which various data structures are stored. Units can communicate, for example, via signals having one or more data packets (e.g., data from a second unit interacting with another unit between a local system, a distributed system, and / or a network; for example, the Internet interacting with other systems via signals).

[0087] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example.

[0088] Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. The reference to "embodiment" herein means that a specific feature, structure, or characteristic described in connection with an embodiment can be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily indicate the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

Claims

1. A physically guided non-line-of-sight three-dimensional imaging method, characterized in that, include: Acquire unlabeled target transient data, input the target transient data into a preset initial three-dimensional voxel reconstruction network, and output three-dimensional predicted voxels; The three-dimensional predicted voxels are input into a preset forward light transmission model to remap the light propagation path, output simulated transient data, and the forward loss is determined based on the simulated transient data and the target transient data. The target transient data is input into a preset reverse optical transmission model to reverse map the transient data to the hidden space, thereby obtaining a physically reconstructed voxel. Based on the physically reconstructed voxel and the three-dimensional predicted voxel, the reverse loss is determined. The three-dimensional Sobel convolution operator is used to convolve the three-dimensional predicted voxels in three orthogonal directions to obtain multiple voxel gradients. Based on the multiple voxel gradients, the regularization loss of the voxel space is determined. Based on the forward loss, the inverse loss, and the regularization loss, the total loss function value is determined, and backpropagation is performed based on the total loss function value to update the network parameters of the initial 3D voxel reconstruction network until the loss converges, thus obtaining the target 3D voxel reconstruction network. The transient data to be tested is input into the target three-dimensional voxel reconstruction network, and the target three-dimensional voxel is output.

2. The physically guided non-viewpoint three-dimensional imaging method according to claim 1, characterized in that, In the step of acquiring unlabeled target transient data, inputting the target transient data into a preset initial three-dimensional voxel reconstruction network, and outputting three-dimensional predicted voxels, the initial three-dimensional voxel reconstruction network adopts a hybrid U-Net architecture that integrates two-dimensional and three-dimensional convolution. The hybrid U-Net architecture includes an encoder composed of multiple downsampling modules, a decoder composed of multiple upsampling modules, and skip connections set between corresponding levels of the encoder and the decoder.

3. The physically guided non-viewpoint three-dimensional imaging method according to claim 2, characterized in that, In the step of acquiring unlabeled target transient data, inputting the target transient data into a preset initial 3D voxel reconstruction network, and outputting 3D predicted voxels, the data processing procedure of the initial 3D voxel reconstruction network includes: The target transient data is subjected to three-dimensional convolution and ReLU activation to multiply its feature channel number and obtain an initial feature map; The initial feature map is subjected to alternating downsampling processes of max pooling and three-dimensional convolutional blocks, which halves the spatiotemporal resolution of the feature map layer by layer, resulting in a deep semantic feature map with a preset number of channels. The deep semantic feature map is upsampled, and the features of each level during the upsampling process are spliced ​​and fused with the downsampled features of the corresponding level to restore spatial resolution and fuse multi-scale information to obtain three-dimensional voxel features. The three-dimensional voxel features are subjected to a three-dimensional to two-dimensional feature dimensionality reduction mapping to map the time-domain flight features into spatial-domain depth slice features, and the three-dimensional predicted voxel of the hidden target is output.

4. The physically guided non-viewpoint three-dimensional imaging method according to claim 1, characterized in that, In the step of inputting the three-dimensional predicted voxels into a preset forward light transmission model to remap the light propagation path and output simulated transient data, the data processing procedure of the forward light transmission model includes: Based on the laser scanning point position, detection point position, voxel spatial coordinates, light speed, distance attenuation coefficient, and voxel albedo in the three-dimensional predicted voxel, the contribution value of each voxel to the transient signal is determined. The contribution values ​​are time-binned using a differentiable kernel density estimation method. The contribution values ​​of all voxels after time-binning are summed to obtain the simulated transient data for predicting flight time.

5. The physically guided non-viewpoint three-dimensional imaging method according to claim 4, characterized in that, In the step of inputting the 3D predicted voxels into a preset forward light transmission model to remap the light propagation path and output simulated transient data, the calculation expression for the physical rendering process of the forward light transmission model is as follows: In the formula, The transient light intensity probability distribution density. To hide the location in the space x Physical reconstruction voxels at the site, D For distance attenuation, n Where h is the number of samples, and h is the kernel density estimation bandwidth. For kernel function, c At the speed of light, x The three-dimensional spatial coordinates of the voxel to be calculated in the hidden space. For the first i The spatial coordinates of a laser scanning point or sampling point.

6. The physically guided non-viewpoint three-dimensional imaging method according to claim 1, characterized in that, The step of inputting the target transient data into a preset reverse optical transmission model to reverse-map the transient data to the hidden space and obtain the physically reconstructed voxel includes: The reverse optical transmission model employs a deconvolution algorithm and a wave algorithm to reverse-transmit the transient data of the target, thereby calculating the three-dimensional structure of the object in the hidden space and obtaining physically reconstructed voxels. The calculation expression for the physically reconstructed voxels is as follows: In the formula, For scanning points or illumination points The transient light intensity signal collected at the location, The location of the scanning point or illumination point on the relay surface. t For photon flight time or time sampling variables, For the Dirac function, r The total optical path length from the relay surface to the hidden target and back to the detector. To integrate the positions of the scanning points on the relay surface, To integrate over the time dimension; For confocal mode: For non-confocal modes: In the formula, This indicates the location of the detection point.

7. The physically guided non-viewpoint three-dimensional imaging method according to claim 1, characterized in that, The step of performing convolution processing on the 3D predicted voxels using a 3D Sobel convolution operator in three orthogonal directions to obtain multiple voxel gradients, and determining the voxel space regularization loss based on these multiple voxel gradients, includes: The 3D prediction voxels are convolved using a 3×3×3 scale Sobel convolution operator in three orthogonal directions in space. The calculation expression is as follows: In the formula, These are the three-dimensional Sobel convolution kernels in the x, y, and z directions, respectively. For convolution operations, For three-dimensional prediction voxels; Based on the voxel gradients in three directions, the regularization loss in the voxel space is calculated, where the regularization loss calculation expression is: In the formula, The gradient energy aggregation function, for x Voxel gradient in the direction, for y Voxel gradient in the direction, for z Voxel gradient in the direction.

8. A physically guided non-line-of-sight three-dimensional imaging system, characterized in that, include: The voxel reconstruction module is configured to acquire unlabeled target transient data, input the target transient data into a preset initial three-dimensional voxel reconstruction network, and output three-dimensional predicted voxels. The positive constraint module is configured to input the three-dimensional predicted voxel into a preset positive light transmission model to remap the light propagation path, output simulated transient data, and determine the positive loss based on the simulated transient data and the target transient data. The reverse processing module is configured to input the target transient data into a preset reverse optical transmission model to reverse map the transient data to the hidden space, obtain a physical reconstruction voxel, and determine the reverse loss based on the physical reconstruction voxel and the three-dimensional prediction voxel. The voxel regularization module is configured to perform convolution processing on the three-dimensional predicted voxels in three orthogonal directions using a three-dimensional Sobel convolution operator to obtain multiple voxel gradients, and determine the regularization loss of the voxel space based on the multiple voxel gradients. The network training module is configured to determine the total loss function value based on the forward loss, the inverse loss, and the regularization loss, and to perform backpropagation based on the total loss function value to update the network parameters of the initial 3D voxel reconstruction network until the loss converges, thereby obtaining the target 3D voxel reconstruction network. The reconstruction output module is configured to input the transient data to be tested into the target three-dimensional voxel reconstruction network and output the target three-dimensional voxel.

9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the steps of a physically guided non-viewpoint three-dimensional imaging method according to any one of claims 1-7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the steps of a physically guided non-viewpoint three-dimensional imaging method according to any one of claims 1-7.