A single-photon three-dimensional imaging method and system based on physical priori enhanced deep learning
Patent Information
- Application Number
- CN202611010868.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-29
AI Technical Summary
然而,现有的数据驱动型深度学习方法仍普遍存在以下技术瓶颈:一方面,传统的深度神经网络通常将图像重建视为一个“黑箱”回归过程,网络仅能从训练数据集中隐式学习噪声的统计特征,但忽略了单光子成像过程中固有的物理规律
[0043]在本发明所提供的基于物理先验增强深度学习的单光子三维成像方法及系统中,利用ITV正则化对待处理数据进行优化求解,使获得的第二预处理后数据包含结构一致性和平滑性的物理先验信息;同时,通过时间窗口增强处理获得第一预处理后数据,以增强目标回波附近的有效光子信息。对第一预处理后数据和第二预处理后数据分别进行编码,获得具有互补性的第一特征体和第二特征体。利用跨注意力特征融合模块对第一特征体和第二特征体进行跨注意力交互和通道重校准,使物理先验信息与深度学习提取的时空特征实现充分融合,获得增强后的三维特征体。利用混合多尺度重建网络对增强后的三维特征体进行深度重建,实现了物理先验约束与深度学习网络相结合、物理模型与数据驱动重建的协同优化、局部结构信息与全局上下文信息的协同建模,从而有效兼顾噪声抑制与细节保持,提高低光子计数、高噪声条件下三维重建的精度、结构保持能力及鲁棒性。
Smart Images

Figure CN122836767A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of single-photon three-dimensional imaging technology, specifically to a single-photon three-dimensional imaging method and system based on physical prior enhancement deep learning. Background Technology
[0002] Single-photon 3D imaging technology is based on the time-of-flight (ToF) principle. It uses a SPAD sensor in conjunction with a time-correlated single-photon counting (TCSPC) module to capture the timestamps of photon arrivals, thereby constructing a time-of-flight histogram. This technology can achieve deep reconstruction of target scenes under extreme photon-scarce conditions and has broad application prospects in fields such as autonomous driving, remote sensing, and biomedical imaging.
[0003] While the theoretical performance of single-photon 3D imaging is limited by photon shot noise, its actual 3D reconstruction quality increasingly depends on computational algorithms. In recent years, deep learning technology, with its powerful nonlinear mapping capabilities, has gradually surpassed traditional statistical optimization methods, enabling the direct recovery of clear depth images from noisy original photon histograms. However, existing data-driven deep learning methods still generally suffer from the following technical bottlenecks: On the one hand, traditional deep neural networks typically treat image reconstruction as a "black box" regression process, where the network can only implicitly learn the statistical characteristics of noise from the training dataset, neglecting the inherent physical laws in single-photon imaging. Therefore, when the geometry, photon budget, or signal-to-background ratio of the test scene differs significantly from the distribution of the training data, the model's generalization performance often drops sharply. On the other hand, in terms of network architecture design, existing networks struggle to effectively balance the contradiction between noise suppression and preservation of structural details. For example, multi-scale convolutional neural networks (CNNs) are prone to losing high-frequency textures and edge details of the target when smoothing background noise; while models based on the Transformer architecture lack a mechanism to efficiently couple local geometric cues with long-distance global dependencies under conditions of extremely sparse photons.
[0004] Therefore, in the field of single-photon 3D imaging, there is an urgent need for a method that can deeply integrate the prior knowledge of imaging physics with the advantages of deep learning networks. Summary of the Invention
[0005] This invention provides a single-photon 3D imaging method and system based on physical prior enhancement deep learning, which effectively improves the quality and robustness of 3D depth reconstruction in low photon count and high noise environments.
[0006] In a first aspect, the present invention provides a single-photon three-dimensional imaging method based on physical prior enhancement deep learning, the method comprising:
[0007] Acquire the data to be processed from the three-dimensional single-photon avalanche diode in the target scene;
[0008] The data to be processed is augmented and subjected to anisotropic and isotropic total variational regularization respectively to obtain the first preprocessed data and the second preprocessed data respectively. The first preprocessed data and the second preprocessed data are then encoded using a convolutional attention module to obtain complementary first feature bodies and second feature bodies respectively.
[0009] The first and second feature volumes are fused using a cross-attention feature fusion module to obtain an enhanced three-dimensional feature volume.
[0010] A hybrid multi-scale reconstruction network is used to perform depth reconstruction on the enhanced 3D feature volume, obtain the 3D feature volume, and convert it into a depth map.
[0011] In some embodiments of the present invention, regularization is performed on the data to be processed to obtain second preprocessed data, including:
[0012] Based on the Poisson statistical properties of the data to be processed, a data fidelity loss function is constructed, and a weighted regularization term is introduced to obtain the objective function. The weighted regularization term is the difference between the anisotropic total variation and the isotropic total variation.
[0013] The objective function is solved iteratively using the alternating direction multiplier algorithm to obtain the second preprocessed data.
[0014] In some embodiments of the present invention, the objective function is iteratively solved using the alternating direction multiplier algorithm to obtain second preprocessed data, including:
[0015] By introducing auxiliary variables corresponding to the data fidelity term and the gradient regularization term, the unconstrained objective function is transformed into a constrained optimization problem.
[0016] In each iteration, the auxiliary variable and the dual variable are updated respectively;
[0017] Determine whether the convergence condition is met based on the change in the results of adjacent iterations or the preset maximum number of iterations.
[0018] If the convergence condition is met, the second preprocessed data is obtained from the output optimal solution.
[0019] In some embodiments of the present invention, the data to be processed is enhanced to obtain first preprocessed data, including:
[0020] The time window is determined based on the pulse response width of the single-photon three-dimensional imaging system;
[0021] Within the time window, the photon counts of adjacent time slots in the data to be processed are aggregated to obtain the first preprocessed data.
[0022] In some embodiments of the present invention, the convolutional attention module includes a multi-scale three-dimensional convolutional aggregation branch and a lightweight global attention reweighting branch;
[0023] The first preprocessed data and the second preprocessed data are encoded using a convolutional attention module to obtain complementary first and second feature bodies, including:
[0024] In the multi-scale 3D convolution aggregation branch, features are extracted from the first preprocessed data and the second preprocessed data using 3D convolution with different receptive fields, and the extracted multi-scale features are concatenated to obtain the first concatenated features and the second concatenated features.
[0025] In the lightweight global attention reweighting branch, scalar attention weights are generated using global relevance scores. Based on the scalar attention weights, the first concatenated features and the second concatenated features are weighted and modulated respectively, and the corresponding first feature volume and second feature volume are output through residual connection.
[0026] In some embodiments of the present invention, the cross-attention feature fusion module includes a window-type three-dimensional cross-attention submodule and a global channel attention submodule;
[0027] The first and second feature volumes are fused using a cross-attention feature fusion module to obtain an enhanced 3D feature volume, including:
[0028] The first feature body is used as the query and the second feature body is used as the key and value. The query is input into the window-based 3D cross-attention submodule, and the first feature body and the second feature body are aligned using the window-based 3D cross-attention submodule.
[0029] In the global channel attention submodule, channel attention weights are generated based on the second feature body; and the channel attention weights are used to perform channel-level recalibration on the first feature body.
[0030] The aligned first and second feature bodies are fused with the calibrated first feature body to obtain the enhanced three-dimensional feature body.
[0031] In some embodiments of the present invention, the hybrid multi-scale reconstruction network includes a first-level encoder-decoder and a second-level encoder-decoder;
[0032] A hybrid multi-scale reconstruction network is used to perform depth reconstruction on the enhanced 3D feature volume to obtain the 3D feature volume, including:
[0033] The enhanced 3D feature volume is processed using a first-level encoder-decoder to obtain an initial depth representation;
[0034] The initial depth representation is refined at multiple scales using a second-level encoder-decoder to obtain a fine feature volume, which is then used as a 3D feature volume.
[0035] In some embodiments of the present invention, the loss function of the hybrid multi-scale reconstruction network during training includes a first loss and a second loss. The first loss is the Kullback-Leibler divergence loss, which measures the distribution difference between the predicted normalized histogram and the true histogram output by the hybrid multi-scale reconstruction network. The second loss is the total variational regularization loss, which is applied to the depth map extracted from the predicted normalized histogram by the Soft-Argmax operation.
[0036] In some embodiments of the present invention, the three-dimensional feature volume is a predicted photon count histogram, and the depth map is obtained as follows:
[0037] The expected probability distribution of each pixel in the time chamber is calculated using the Soft-Argmax operation on the predicted photon count histogram, and the calculated expected probability distribution is converted into a depth map.
[0038] Secondly, the present invention also provides a single-photon three-dimensional imaging system based on physical prior enhancement deep learning, the system comprising:
[0039] The data acquisition unit is used to acquire the data to be processed from the three-dimensional single-photon avalanche diode in the target scene;
[0040] The preprocessing and encoding unit is used to perform enhancement processing and anisotropic and isotropic total variational regularization processing on the data to be processed, respectively, to obtain the first preprocessed data and the second preprocessed data, and to encode the first preprocessed data and the second preprocessed data using the convolutional attention module, respectively, to obtain the complementary first feature body and the second feature body.
[0041] The fusion unit is used to fuse the first feature volume and the second feature volume using the cross-attention feature fusion module to obtain the enhanced three-dimensional feature volume;
[0042] The depth reconstruction unit is used to perform depth reconstruction on the enhanced 3D feature volume using a hybrid multi-scale reconstruction network, obtain the 3D feature volume, and convert it into a depth map.
[0043] In the single-photon 3D imaging method and system based on physical prior enhancement deep learning provided by this invention, ITV regularization is used to optimize the solution of the data to be processed, so that the obtained second preprocessed data contains physical prior information on structural consistency and smoothness. At the same time, time window enhancement processing is used to obtain the first preprocessed data to enhance the effective photon information near the target echo. The first and second preprocessed data are encoded respectively to obtain a first feature body and a second feature body with complementarity. A cross-attention feature fusion module is used to perform cross-attention interaction and channel recalibration on the first and second feature bodies, so that the physical prior information and the spatiotemporal features extracted by deep learning are fully fused to obtain the enhanced 3D feature body. A hybrid multi-scale reconstruction network is used to perform deep reconstruction on the enhanced 3D feature body, realizing the combination of physical prior constraints and deep learning network, the synergistic optimization of physical model and data-driven reconstruction, and the synergistic modeling of local structural information and global context information. This effectively balances noise suppression and detail preservation, and improves the accuracy, structure preservation ability and robustness of 3D reconstruction under low photon count and high noise conditions. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0045] Figure 1 This is a flowchart illustrating the single-photon three-dimensional imaging method based on physical prior enhancement deep learning provided in an embodiment of the present invention.
[0046] Figure 2 This is a schematic diagram of the structure of the convolutional attention module CAM provided in an embodiment of the present invention;
[0047] Figure 3 This is a schematic diagram of the cross-attention feature fusion module CACEM provided in an embodiment of the present invention;
[0048] Figure 4 This is a schematic diagram of the structure of the hybrid multi-scale reconstruction network SWNet provided in an embodiment of the present invention;
[0049] Figure 5 This is a comparison chart of depth reconstruction results on a simulated dataset provided in an embodiment of the present invention;
[0050] Figure 6 These are comparison images of depth reconstruction effects in real-world scenarios provided by embodiments of the present invention;
[0051] Figure 7This is a schematic diagram of the hardware structure of the TCSPC single-photon lidar system provided in an embodiment of the present invention;
[0052] Figure 8 This is one of the comparison images of depth reconstruction effect under laboratory scene data provided in the embodiments of the present invention;
[0053] Figure 9 This is the second comparison image of depth reconstruction effect under laboratory scene data provided in the embodiments of the present invention. Detailed Implementation
[0054] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0055] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0056] "A and / or B" includes the following three combinations: A only, B only, and a combination of A and B.
[0057] The use of "applies to" or "configured to" in this invention implies an open and inclusive language, which does not exclude the applicability to or configuration to devices performing additional tasks or steps. Additionally, the use of "based on" implies openness and inclusivity, because processes, steps, calculations, or other actions "based on" one or more conditions or values may in practice be based on additional conditions or values beyond those conditions.
[0058] In this invention, the term "exemplary" is used to mean "serving as an example, illustration, or description." Any embodiment described as "exemplary" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.
[0059] The following description, in conjunction with the accompanying drawings, introduces the single-photon three-dimensional imaging method and system based on physical prior enhancement deep learning provided by embodiments of the present invention.
[0060] like Figure 1 As shown, this embodiment of the invention provides a single-photon three-dimensional imaging method based on physical prior enhancement deep learning, which includes the following steps:
[0061] S101: Obtain the data to be processed from the 3D single-photon avalanche diode SPAD in the target scene.
[0062] In some examples, SPAD sensors are used to acquire 3D single-photon avalanche diode (SPAD) data of the target scene. The data is in the form of a 3D histogram tensor with dimensions of [missing information]. , and These represent the number of rows and columns of the pixel grid, respectively. This refers to the number of time slots.
[0063] S102, the data to be processed is enhanced and anisotropic and isotropic total variational AITV regularization is performed respectively to obtain the first preprocessed data and the second preprocessed data respectively. The first preprocessed data and the second preprocessed data are encoded by the convolutional attention module CAM respectively to obtain the complementary first feature body and the second feature body respectively.
[0064] S103, the first feature volume and the second feature volume are fused using the Cross-Attention Feature Fusion Module (CACEM) to obtain the enhanced three-dimensional feature volume.
[0065] S104 utilizes the hybrid multi-scale reconstruction network SWNet to perform depth reconstruction on the enhanced 3D feature volume, obtain the 3D feature volume, and convert it into a depth map.
[0066] The single-photon 3D imaging method based on physical prior enhancement deep learning provided in this invention utilizes ITV regularization to optimize the solution of the data to be processed, so that the obtained second preprocessed data contains physical prior information on structural consistency and smoothness. Simultaneously, a first preprocessed data is obtained through temporal window enhancement processing to enhance the effective photon information near the target echo. The first and second preprocessed data are encoded respectively to obtain complementary first and second feature bodies. A cross-attention feature fusion module is used to perform cross-attention interaction and channel recalibration on the first and second feature bodies, enabling full fusion of physical prior information and spatiotemporal features extracted by deep learning, resulting in an enhanced 3D feature body. A hybrid multi-scale reconstruction network is used to perform deep reconstruction of the enhanced 3D feature body, achieving a combination of physical prior constraints and deep learning networks, collaborative optimization of physical models and data-driven reconstruction, and collaborative modeling of local structural information and global contextual information. This effectively balances noise suppression and detail preservation, improving the accuracy, structure preservation capability, and robustness of 3D reconstruction under low photon count and high noise conditions.
[0067] In some embodiments of the present invention, S102 includes the following sub-steps:
[0068] S201. Based on the Poisson statistical properties of the data to be processed, a data fidelity loss function is constructed, and a weighted regularization term is introduced to obtain the objective function. The weighted regularization term is the difference between the anisotropic total variation and the isotropic total variation.
[0069] Schematic, the objective function is: , For the data fidelity loss function, For weighted regularization terms, = , For anisotropic total variation, For isotropic total variation, , For smoothing coefficients, For the Poisson fidelity parameter, , For pixels In the time capsule The observed photon count at the location, For linear projection operators, It is the depth map to be reconstructed (i.e., the data to be processed). This is background noise.
[0070] In some examples, the Poisson fidelity parameter λ is set to λ=5, and the smoothing coefficient τ is set to τ=0.4 to preserve edges while maintaining the overall smoothness of the scene structure. During the depth map reconstruction stage, since the sharpness of object edge structures is more critical in depth map reconstruction, the sparsity parameter can be set to τ=0.6. In solving this objective function, the value of τ can be appropriately increased to further enhance the ability to preserve depth map edges during the ADMM solution step.
[0071] S202, the objective function is solved iteratively using the Alternating Direction Multiplier (ADMM) algorithm to obtain the second preprocessed data.
[0072] Specifically, during the iterative solution process, auxiliary variables v and w are introduced to transform the unconstrained objective function. Transform into a constrained optimization problem: The constraints are set as u=v and ∇u=w; and dual variables (Lagrange multipliers) y and z and a quadratic penalty parameter β are introduced to transform the constrained optimization problem into an unconstrained augmented Lagrangian function form: .
[0073] Based on the augmented Lagrangian function described above, updates are performed on u−, v−, w−, and the dual variable in each iteration. The update formulas are as follows:
[0074] u-Update:
[0075] ;
[0076] v-Update:
[0077] ;
[0078] w-Update:
[0079] ;
[0080] Dual variable update:
[0081] , ;
[0082] In the formula, u represents the depth map to be reconstructed; v is the auxiliary variable corresponding to the data fidelity term; w is the auxiliary variable corresponding to the gradient regularization term; and β is the penalty parameter of ADMM.
[0083] In some examples, the initial penalty parameter δ of ADMM is set to δ= The iterative formula is The maximum number of iterations is 300, until the algorithm converges. The output optimal solution is a two-dimensional coarse-grained depth map. Subsequently, this two-dimensional coarse-grained depth map is upsampled to three-dimensional space to obtain the second preprocessed data.
[0084] In some embodiments of the present invention, S102 includes the following sub-steps:
[0085] S301, determine the time window based on the impulse response width of the single-photon three-dimensional imaging system. Wherein, the time window T... wind Set to a full width at half maximum (FWHM) close to the system impulse response.
[0086] S302, within the time window, aggregate the photon counts of adjacent time slots in the data to be processed to obtain the first preprocessed data.
[0087] The single-photon three-dimensional imaging method based on physical prior enhancement deep learning provided in this invention improves the signal-to-background sum-time ratio (SBR) and suppresses noise by aggregating photon counts within a short time window through three-dimensional convolution.
[0088] In some embodiments of the present invention, such as Figure 2 As shown, the Convolutional Attention Module (CAM) includes a multi-scale 3D convolutional aggregation branch and a lightweight global attention reweighting branch. The multi-scale 3D convolutional aggregation branch comprises four parallel convolutional paths, while the lightweight global attention reweighting branch includes an attention module used to generate scalar attention weights and perform weighted modulation on the features. Features are extracted from the first and second preprocessed data through the four parallel convolutional paths. The outputs of these four paths are then concatenated and projected using a 1×1×1 convolution, splitting into two paths: one as the input to the residual connection, and the other entering the lightweight global attention reweighting branch. Finally, the features output from the lightweight global attention reweighting branch are added to the residual connection, resulting in the final feature. Figure 2 The final output.
[0089] Accordingly, S102 includes the following sub-steps:
[0090] S401, in the multi-scale three-dimensional convolution aggregation branch, features are extracted from the first preprocessed data and the second preprocessed data by using three-dimensional convolution with different receptive fields, and the extracted multi-scale features are spliced together to obtain the first spliced features and the second spliced features respectively.
[0091] Understandably, the preprocessed data is input into a multi-scale 3D convolutional aggregation branch. Features are extracted through 3D convolutions of different receptive fields in this branch, and the extracted features at different scales are then concatenated using channel stitching. Convolutional projection is used to obtain the first stitched features. Similarly, the second preprocessed data is input into a multi-scale 3D convolutional aggregation branch. Features are extracted through 3D convolution with different receptive fields in the multi-scale 3D convolutional aggregation branch, and the extracted features at different scales are then stitched together via channels. Convolutional projection is used to obtain the second concatenated features.
[0092] In some examples, feature extraction is performed using 3D convolutions with four different receptive fields, followed by channel stitching and... Convolutional projection, where the extracted features at different scales are The features after splicing are .
[0093] S402, in the lightweight global attention reweighting branch, scalar attention weights are generated using global relevance scores; based on the scalar attention weights, the first concatenated features and the second concatenated features are weighted and modulated respectively, and the corresponding first feature volume and second feature volume are output through residual connection.
[0094] In some examples, the lightweight global attention reweighted branch receives the concatenated feature F from the output of the aforementioned multi-scale convolution aggregation branch, which is processed through three independent... Convolution generates feature embeddings Global query embedding Global Context Embedding Calculate the global relevance score Attention modulation features were obtained. Finally, through grouping Convolution and grouping normalization, combined with residual connections, output the first and second feature bodies. It should be noted that since the first and second preprocessed data are input to the CAM module separately, the attention calculation process is performed on the first and second concatenated features respectively, ultimately outputting independent and complementary first and second feature bodies.
[0095] The single-photon 3D imaging method based on physical prior enhancement deep learning provided in this invention extracts features containing multi-scale structural information by using convolutions of different receptive fields in a multi-scale 3D convolution aggregation branch. It then uses a lightweight global attention reweighting branch to adaptively reweight the feature maps by channel reweighting, thereby highlighting effective signal-related features and suppressing background noise interference, providing a rich and clean low-level feature body for subsequent CACEM cross-attention fusion.
[0096] In some embodiments of the present invention, such as Figure 3As shown, the Cross-Attention Feature Fusion Module (CACEM) includes a windowed 3D cross-attention submodule and a global channel attention submodule. The windowed 3D cross-attention submodule includes a 1×1×1 convolution to generate a query from the first feature volume, a 1×1×1 convolution to generate keys and values from the second feature volume, and a branch that performs windowed cross-attention calculation and outputs projected features. The global channel attention submodule includes a globally average pooling GAP layer, an adaptive one-dimensional convolutional layer, and a sigmoid activation function connected in sequence. It is used to generate channel attention weights based on the second feature volume and to perform channel-level recalibration on the first feature volume.
[0097] Accordingly, S103 includes the following sub-steps:
[0098] S1031, the first feature body As a query, the second feature body As the key and value, they are input into the window-based 3D cross-attention submodule, which is used to align the first and second feature volumes.
[0099] Schematic, the volume is divided into non-overlapping windows, and attention is computed within each window. , It is a query matrix. Here, d is the key matrix, and d is the feature dimension of the query matrix and the key matrix. The output is... The final output is the aligned first feature body. In one specific implementation, the size of the partition window is set to 16 (depth) × 4 (height) × 4 (width), and the embedding dimension is set to 8.
[0100] S1032, in the global channel attention submodule, channel attention weights are generated based on the second feature body; and the channel attention weights are used to perform channel-level recalibration on the first feature body.
[0101] Schematic, for the second feature body Perform global average pooling to obtain Then, a one-dimensional convolution operation is performed on z, and the channel attention weights are obtained through an activation function. Then, the channel attention weight w is combined with the first feature body. Perform element-wise multiplication and then superimpose the original first feature volume. As a residual join, for Perform channel-level recalibration. The calibrated characteristics are represented as follows: ,in This indicates element-wise multiplication.
[0102] S1033, the aligned first feature body and second feature body fused with the calibrated first feature To obtain the enhanced three-dimensional feature volume Indicatively, .
[0103] In some embodiments of the present invention, the hybrid multi-scale reconstruction network SWNet adopts a W-type dual encoder-decoder topology, such as... Figure 4 As shown, it includes a first-level encoder-decoder and a second-level encoder-decoder. Both the first-level encoder-decoder and the second-level encoder-decoder include local convolutional branches and global Swin-Transformer branches. The first-level encoder-decoder includes encoder 1 and decoder 1, and the second-level encoder-decoder includes encoder 2 and decoder 2.
[0104] Accordingly, S104 includes the following sub-steps:
[0105] S1041, the enhanced 3D feature volume is processed using the first-level encoder-decoder to obtain the initial depth representation.
[0106] In one specific implementation, the first-level encoder-decoder and the second-level encoder-decoder use the same network configuration. The number of feature channels in the local convolutional branches is configured from shallow to deep as 8, 16, 32, 64, and 128, respectively; the feature dimensions of the global Swin-Transformer branches are configured as 4, 8, 16, 32, and 64, respectively, and the number of attention heads at each level is set to 4. The global branch uses a 3D Swin-Transformer block with a window size of 4×2×2 and an MLP expansion factor of 4. The local and global branches fuse features at each scale through feature concatenation and nonlinear transformation.
[0107] The enhanced 3D feature volume is input into both the local convolutional branch and the global Swin-Transformer branch. At each scale, the local convolutional branch extracts features through 3D convolution and residual blocks to capture local high-resolution geometric details; simultaneously, the global Swin-Transformer branch extracts global context and long-range structural dependencies through a shift window attention mechanism. Subsequently, the features extracted by the two branches are concatenated and nonlinearly transformed at this scale to achieve feature-level fusion. The fused multi-scale features are fed into encoder 1, where deep semantic information is extracted through stepwise downsampling (convolution with a stride of 2); then, stepwise upsampling (deconvolution) is performed through decoder 1, and skip connections are used to fuse the shallow features of the encoder with the corresponding scale features of the decoder to restore spatial resolution and obtain the initial depth representation.
[0108] S1042 uses a second-level encoder-decoder to refine the initial depth representation at multiple scales, obtaining a fine feature volume as a 3D feature volume.
[0109] Similarly, the initial depth representation is used as input and fed into the local convolutional branch and global Swin-Transformer branch of the second-level encoder-decoder for secondary feature extraction and fusion. The initial depth representation already possesses a preliminary depth structure. Based on this, encoder 2 and decoder 2 in the second-level network perform targeted multi-scale refinement on regions prone to errors in depth reconstruction (such as object edges and weakly reflective surfaces). Through resampling by encoder 2 and layer-by-layer upsampling by decoder 2, combined with multi-scale feature fusion, residual depth biases and noise from the first-level reconstruction are corrected to obtain a refined feature volume.
[0110] Understandably, in the second-level encoder-decoder, the initial depth representation is refined at multiple scales to correct erroneous regions that appeared in the preceding steps and further enhance global geometric consistency.
[0111] In some embodiments of the present invention, the loss function of the hybrid multi-scale reconstruction network during training includes a first loss and a second loss. The first loss is the Kullback-Leibler divergence loss, which measures the distribution difference between the predicted normalized histogram and the true histogram output by the hybrid multi-scale reconstruction network. The second loss is the total variational regularization loss, which is applied to the depth map extracted from the predicted normalized histogram by the Soft-Argmax operation.
[0112] Schematic representation of the loss function during training of a hybrid multi-scale reconstruction network. First loss Second loss , It is a hyperparameter that controls the weights of the total variation (TV) regularization. It is the true normalized histogram of pixel position (i,j) in the nth time bin; It is a depth map extracted from the predicted normalized histogram through the Soft-Argmax operation.
[0113] In some embodiments of the present invention, the three-dimensional feature volume is a predicted photon count histogram, and the depth map is obtained as follows:
[0114] The expected probability distribution of each pixel in the time chamber is calculated using the Soft-Argmax operation on the predicted photon count histogram, and the calculated expected probability distribution is converted into a depth map.
[0115] Schematic diagram: The predicted histogram is converted into depth values via a Soft-Argmax operation. :
[0116]
[0117] In the formula, Let N be the normalized photon count histogram predicted by the hybrid multi-scale reconstruction network at pixel location (i,j), where N is the total number of time bins in the photon count histogram. It is the normalized prediction probability value corresponding to the nth time bin in the prediction histogram.
[0118] In some embodiments of this invention, the dataset is trained using the NYU v2 dataset and tested on the Middlebury dataset and real-world scenarios. Photon arrival times are sampled from a non-homogeneous Poisson process (1024 time bins, bin width 80 ps, system impulse response full width at half maximum (FWHM) 400 ps). The training set contains 16,000 samples with a spatial resolution of [missing information]. During the training and validation phases, each sample is randomly pruned to 1000 samples. The magnitude is used as input. It is generated under different signal-to-background ratio combinations (10:2, 10:10, 10:50, 5:2, 5:10, 5:50, 2:2, 2:10, 2:50).
[0119] For training configuration, the Adam optimizer is used, and a cosine annealing strategy is employed for the learning rate. to Batch size is 4, training lasts 12 epochs, TV regularization weight λ= The experiments were conducted on a workstation equipped with dual RTX 5090 GPUs and an i9-14900KF CPU. To ensure fairness in the comparison, all reproducible deep learning baseline methods used the same hybrid SBR training protocol and were trained three times with different random seeds. The best validation set checkpoint was selected, and the mean and standard deviation of the three results were reported. In terms of evaluation metrics, this invention adds a structural similarity (SSIM) metric to comprehensively examine the structural preservation ability of the reconstruction results.
[0120] Table 1. Performance comparison between the method of the present invention and existing methods under different signal-to-noise ratios.
[0121]
[0122] As shown in Table 1, existing classic optimization algorithms (such as the Shin algorithm) exhibit significant performance degradation in complex noisy environments, with accuracy approaching zero in some test scenarios. Compared to traditional algorithms, deep learning methods have a significant advantage in overall root mean square error (RMSE), but their ability to recover structural details remains limited in scenarios with extremely low signal-to-noise ratios (e.g., 2:50).
[0123] In comparison, the embodiments of this invention achieved the lowest RMSE and the highest accuracy under most SBR conditions. Especially in extreme scenarios with extremely low SBR (2:50), this invention still maintains an accuracy of 96.48%, outperforming the best comparable benchmark method (PRS-Net, 94.65%); simultaneously, the SSIM metric remains consistently within the excellent range of 0.97 to 0.99. Experimental data fully validate that this invention, by fusing physical imaging priors with deep learning networks, possesses superior structural detail reconstruction accuracy and good environmental robustness.
[0124] Table 2 Ablation Experiment Results
[0125]
[0126] As shown in Table 2, removing CACEM leads to a significant performance degradation, especially under extremely sparse (2:2) signal conditions, indicating that this cross-attention mechanism is crucial for the accurate alignment and information fusion of dual-stream features in this invention. Simultaneously, removing the AITV physical prior module or CAM also results in a consistent decrease in various metrics.
[0127] It is worth noting that the variant using Mamba global branches, while maintaining reconstruction accuracy similar to the full model of this invention, can bring a significant improvement in computational efficiency (reducing the number of parameters by about 1.6%, floating-point operations by about 1.3%, and inference time by about 5%). This fully demonstrates that the physical prior fusion framework proposed in this invention has excellent modular flexibility and versatility, and can adapt to global feature modeling modules of different architectures, rather than being limited to the SwinTransformer structure.
[0128] To intuitively demonstrate the actual technical effects of the method of the present invention, Figure 5 Comparison images of depth reconstruction results on simulated datasets are presented, using three representative scenes: Art, Doll, and Moebius. In scenes with complex edges and severe occlusion, such as the Doll scene, the method of this invention can more clearly separate multiple target objects and effectively suppress background noise. Furthermore, Figure 6The reconstruction results are demonstrated in publicly available real-world scene tests, covering multiple scenes including Elephant, Stuff, Rolling Ball, Kitchen, and Hallway. Comparisons show that the depth map reconstructed by the method of this invention significantly outperforms existing comparative methods in terms of object edge preservation and depth hierarchy segmentation.
[0129] Furthermore, this embodiment of the invention further constructs a physical system of a dual-axis TCSPC single-photon lidar to verify the practical hardware deployment capability of the proposed method. The hardware structure of the physical system of the dual-axis TCSPC single-photon lidar is as follows: Figure 7 As shown, the Laser is a pulsed laser source used to emit 830 nm pulsed laser light; L_A and L_B are beam expanders used to expand and collimate the laser beam in the emission optical path; M_1 is a reflector used to reflect the laser light and guide it to the scanner; GVS012 is a dual-axis galvanometer scanner used to perform two-dimensional point-by-point scanning of the target scene; L_C is a receiving lens used to gather scattered photons returned by the target object; BPF is an 830 nm bandpass filter used to filter out ambient background light noise; SPAD is a single-photon avalanche diode detector used to convert the received weak single-photon signal into an electrical pulse and output it. System measurement results show that under low-noise conditions with an average of 3 photons per pixel and high background noise conditions with 11 photons per pixel (e.g., ...), the system can achieve the following results. Figure 8 , Figure 9 As shown in the figure, the method of the present invention can still clearly separate the foreground object from the background, and the surface reconstruction has good coherence, effectively verifying its robustness to real hardware background noise and environmental interference.
[0130] Regarding model inference complexity, for a high-resolution input image of 576×704, the model parameters in this embodiment are approximately 5.08M, slightly higher than lightweight networks. However, its computational cost per forward inference is only 100.6 GFLOPs, significantly lower than U-Net (1543.0 GFLOPs) and PRS-Net (132.8 GFLOPs), indicating that this invention has high parameter utilization efficiency. With a fixed batch size of 4, the peak inference memory usage is approximately 2.27GB, allowing for smooth operation on mainstream GPUs.
[0131] In summary, the embodiments of the present invention demonstrate superior performance in various objective indicators and subjective visual reconstruction effects by deeply fusing physical imaging priors and deep network architecture, and also possess high-precision structural reconstruction capabilities, good environmental robustness, and excellent hardware deployment applicability.
[0132] In addition, embodiments of the present invention also provide a single-photon three-dimensional imaging system based on physical prior enhancement deep learning, the system comprising a data acquisition unit, a preprocessing and encoding unit, a fusion unit and a depth reconstruction unit.
[0133] The data acquisition unit is used to acquire the data to be processed from the three-dimensional single-photon avalanche diode in the target scene.
[0134] The preprocessing and encoding unit is used to perform enhancement processing and anisotropic and isotropic total variational regularization processing on the data to be processed, respectively, to obtain the first preprocessed data and the second preprocessed data. The convolutional attention module is used to encode the first preprocessed data and the second preprocessed data respectively to obtain the complementary first feature body and the second feature body.
[0135] The fusion unit is used to fuse the first feature volume and the second feature volume using the cross-attention feature fusion module to obtain an enhanced three-dimensional feature volume.
[0136] The depth reconstruction unit is used to perform depth reconstruction on the enhanced 3D feature volume using a hybrid multi-scale reconstruction network, obtain the 3D feature volume, and convert it into a depth map.
[0137] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0138] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.
[0139] The foregoing has provided a detailed description of a single-photon three-dimensional imaging method and system based on physical prior enhancement deep learning provided by the embodiments of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A single-photon three-dimensional imaging method based on physical prior enhancement deep learning, characterized in that, The method includes: Acquire the data to be processed from the three-dimensional single-photon avalanche diode in the target scene; The data to be processed is subjected to enhancement processing and anisotropic and isotropic total variational regularization processing respectively to obtain the first preprocessed data and the second preprocessed data respectively. The first preprocessed data and the second preprocessed data are encoded by the convolutional attention module respectively to obtain the first feature body and the second feature body respectively. The first feature volume and the second feature volume are fused using a cross-attention feature fusion module to obtain an enhanced three-dimensional feature volume; The enhanced 3D feature volume is reconstructed using a hybrid multi-scale reconstruction network to obtain the 3D feature volume, which is then converted into a depth map.
2. The single-photon three-dimensional imaging method based on physical prior enhancement deep learning according to claim 1, characterized in that, The data to be processed is regularized to obtain the second preprocessed data, including: Based on the Poisson statistical characteristics of the data to be processed, a data fidelity loss function is constructed, and a weighted regularization term is introduced to obtain the objective function. The weighted regularization term is the difference between the anisotropic total variation and the isotropic total variation. The objective function is iteratively solved using the alternating direction multiplier algorithm to obtain the second preprocessed data.
3. The single-photon three-dimensional imaging method based on physical prior enhancement deep learning according to claim 2, characterized in that, The step of iteratively solving the objective function using the alternating direction multiplier algorithm to obtain the second preprocessed data includes: By introducing auxiliary variables corresponding to the data fidelity term and the gradient regularization term, the unconstrained objective function is transformed into a constrained optimization problem. In each iteration, the auxiliary variable and the dual variable are updated respectively; Determine whether the convergence condition is met based on the change in the results of adjacent iterations or the preset maximum number of iterations. If the convergence condition is met, the second preprocessed data is obtained from the output optimal solution.
4. The single-photon three-dimensional imaging method based on physical prior enhancement deep learning according to claim 1, characterized in that, The data to be processed is enhanced to obtain the first preprocessed data, including: The time window is determined based on the pulse response width of the single-photon three-dimensional imaging system; Within the time window, the photon counts of adjacent time slots in the data to be processed are aggregated to obtain the first preprocessed data.
5. The single-photon three-dimensional imaging method based on physical prior enhancement deep learning according to claim 1, characterized in that, The convolutional attention module includes a multi-scale 3D convolutional aggregation branch and a lightweight global attention reweighting branch; The step of encoding the first preprocessed data and the second preprocessed data using a convolutional attention module to obtain complementary first and second feature bodies includes: In the multi-scale three-dimensional convolution aggregation branch, features are extracted from the first preprocessed data and the second preprocessed data using three-dimensional convolution with different receptive fields, and the extracted multi-scale features are spliced together to obtain the first spliced feature and the second spliced feature. In the lightweight global attention reweighted branch, scalar attention weights are generated using global relevance scores; based on the scalar attention weights, the first concatenated feature and the second concatenated feature are weighted and modulated respectively, and the corresponding first feature volume and second feature volume are output through residual connection.
6. The single-photon three-dimensional imaging method based on physical prior enhancement deep learning according to claim 1, characterized in that, The cross-attention feature fusion module includes a window-based three-dimensional cross-attention submodule and a global channel attention submodule; The method of fusing the first feature volume and the second feature volume using a cross-attention feature fusion module to obtain an enhanced three-dimensional feature volume includes: The first feature body is used as a query, and the second feature body is used as a key and value. The query is input into the window-type three-dimensional cross-attention submodule, and the first feature body and the second feature body are aligned using the window-type three-dimensional cross-attention submodule. In the global channel attention submodule, channel attention weights are generated based on the second feature body; and the channel attention weights are used to perform channel-level recalibration on the first feature body. The aligned first and second feature bodies are fused with the calibrated first feature body to obtain the enhanced three-dimensional feature body.
7. The single-photon three-dimensional imaging method based on physical prior enhancement deep learning according to claim 1, characterized in that, The hybrid multi-scale reconstruction network includes a first-level encoder-decoder and a second-level encoder-decoder; The process of using a hybrid multi-scale reconstruction network to perform depth reconstruction on the enhanced 3D feature volume to obtain the 3D feature volume includes: The enhanced 3D feature volume is processed using the first-level encoder-decoder to obtain an initial depth representation; The initial depth representation is refined at multiple scales using the second-level encoder-decoder to obtain a fine feature volume, which is then used as the three-dimensional feature volume.
8. The single-photon three-dimensional imaging method based on physical prior enhancement deep learning according to claim 7, characterized in that, The loss function of the hybrid multi-scale reconstruction network during training includes a first loss and a second loss. The first loss is the Kullback-Leibler divergence loss, which measures the distribution difference between the predicted normalized histogram and the true histogram output by the hybrid multi-scale reconstruction network. The second loss is the total variational regularization loss, which is applied to the depth map extracted from the predicted normalized histogram by the Soft-Argmax operation.
9. The single-photon three-dimensional imaging method based on physical prior enhancement deep learning according to any one of claims 1 to 8, characterized in that, The three-dimensional feature volume is a predicted photon count histogram, and the depth map is obtained in the following manner: The expected probability distribution of each pixel in the time chamber is calculated using the Soft-Argmax operation on the predicted photon count histogram, and the calculated expected probability distribution is converted into the depth map.
10. A single-photon three-dimensional imaging system based on physical prior reinforcement deep learning, characterized in that, The system includes: The data acquisition unit is used to acquire the data to be processed from the three-dimensional single-photon avalanche diode in the target scene; The preprocessing and encoding unit is used to perform enhancement processing and anisotropic and isotropic total variational regularization processing on the data to be processed, respectively, to obtain the first preprocessed data and the second preprocessed data, and to encode the first preprocessed data and the second preprocessed data using the convolutional attention module, respectively, to obtain the complementary first feature body and the second feature body. The fusion unit is used to fuse the first feature volume and the second feature volume using the cross-attention feature fusion module to obtain an enhanced three-dimensional feature volume; The depth reconstruction unit is used to perform depth reconstruction on the enhanced 3D feature volume using a hybrid multi-scale reconstruction network, obtain the 3D feature volume, and convert it into a depth map.