A Method and System for Multi-Target Reconstruction of LiDAR Based on Temporal Aggregation and State-Space Networks
Patent Information
- Application Number
- CN202610830164.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-10
- Publication Date
- 2026-09-01
AI Technical Summary
[0005]本发明的目的在于提供一种基于时间聚合与状态空间网络的激光雷达多目标重建方法及系统,以解决现有技术在强背景光、低信噪比和多目标回波混叠条件下存在的目标数量估计不准确、深度定位误差较大以及三维重建稳定性不足的问题
[0016] Compared with existing technologies, the beneficial effects of this invention are as follows: Addressing the problems of low photon counting signal-to-noise ratio, indistinct target echo peaks, and overlapping echoes from multiple targets under background light conditions, this invention uses a temporal aggregation window to locally enhance the photon counting sequence, enabling target echoes to form a more stable continuous response within adjacent time bins and reducing the impact of random background light noise on individual time bins. Simultaneously, this invention extracts multi-scale spatial context features through an expanded spatial fusion module, improving the representation ability of complex target structures and spatial neighborhood relationships; and models long-distance dependencies in both spatial and temporal directions through state space scanning operations, allowing the network to simultaneously utilize spatial structure information and temporal echo distribution information, thereby improving the accuracy of multi-target quantity estimation and depth localization. Furthermore, this invention employs an encoder-decoder structure and a skip connection mechanism, preserving shallow detail information while extracting deep semantic features, which is beneficial for recovering the depth distribution of target edges, weak echo regions, and multi-layered structural regions. By outputting the target presence probability distribution of each pixel position in different time bins, this invention can simultaneously determine the number of targets and the depth position of each target, making it suitable for LiDAR 3D reconstruction in scenarios with strong background light, low signal-to-noise ratio, and complex multi-target environments.
Smart Images

Figure CN122672058A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of lidar three-dimensional imaging and computational imaging technology, specifically to a lidar multi-target reconstruction method and system based on temporal aggregation and state-space networks. Background Technology
[0002] LiDAR (Light Detection and Ranging) achieves 3D imaging by emitting pulsed laser light towards a target scene and receiving the reflected echoes. It calculates the target distance based on the echo flight time. Single-photon lidar, due to its extremely high detection sensitivity, can detect long-range, low-power, and weakly reflective targets with minimal photon usage. It has been widely applied in long-range 3D imaging, autonomous driving, remote sensing, complex environment perception, and all-weather target identification.
[0003] In a typical single-photon lidar imaging process, the detector statistically analyzes the arrival time of photons received at each pixel location, forming a time-correlated single-photon count histogram. For a single-surface target, the reflected photons typically form a significant peak near the corresponding depth location, and the target depth can be recovered through peak detection or statistical estimation. However, in real-world applications, target scenes are often not single-surface structures. When transparent, semi-transparent, occluded, reflective, or multi-layered structures coexist, multiple target surfaces may correspond to the same pixel direction. The echoes from different target surfaces will be distributed in different time bins, resulting in multi-target echoes. In this case, traditional single-surface depth reconstruction methods struggle to simultaneously recover the number and depth positions of multiple targets. Furthermore, when single-photon lidar operates outdoors or in environments with strong background light, factors such as sunlight, ambient scattered light, and detector dark counting significantly increase the number of background noise photons, raising the photon count in non-target time bins. This increased background noise reduces the signal-to-background ratio, making the target echo peak less prominent and even causing confusion between false peaks and target peaks. Under low signal-to-noise ratio conditions, methods based on maximum photon count, traditional filtering, or single statistical models are prone to problems such as depth misjudgment, incorrect target number estimation, and omission of multiple target echoes.
[0004] In recent years, deep learning methods have been applied to depth reconstruction tasks in single-photon lidar, improving reconstruction performance under complex noise conditions through a data-driven approach. However, existing methods still have the following shortcomings: First, some methods are mainly geared towards single-target or single-surface depth reconstruction, making it difficult to adapt to situations where multiple target echoes exist on the same pixel; second, while convolutional neural networks have strong local feature extraction capabilities, their ability to model long-distance spatial dependencies and long-term dimensional dependencies is limited; third, Transformer-based methods can model global relationships, but they involve high computational costs and suffer from efficiency issues when processing high-dimensional photon counting tensors; fourth, photon counting data under background light interference exhibits significant temporal sparsity and random noise characteristics, which can easily cause target echo features to be submerged by background noise if directly input into the network. Therefore, there is an urgent need for a lidar multi-target reconstruction method that can fully utilize the temporal distribution characteristics of photon counting data and simultaneously model long-distance dependencies in both spatial and temporal dimensions, in order to improve the accuracy and stability of multi-target quantity estimation and depth location reconstruction under background light conditions. Summary of the Invention
[0005] The purpose of this invention is to provide a multi-target reconstruction method and system for lidar based on temporal aggregation and state-space networks, so as to solve the problems of inaccurate target number estimation, large depth positioning error and insufficient stability of three-dimensional reconstruction in the existing technology under conditions of strong background light, low signal-to-noise ratio and multi-target echo aliasing.
[0006] To achieve the above objectives, the present invention provides the following technical solution: A multi-target reconstruction method for lidar based on temporal aggregation and state-space networks, comprising the following steps: Step 1: Obtain the echo photon count data collected by the lidar, and construct a photon count tensor based on the pixel spatial location and time bin. The photon count tensor is used to represent the number of photons received by different spatial pixel locations in different time bins. Step 2: Preprocess the photon counting tensor to obtain a normalized photon counting tensor, while preserving the correspondence between its spatial and temporal dimensions; Step 3: Construct a temporal aggregation window and locally aggregate the temporal photon count sequence corresponding to each pixel position to obtain the temporal aggregation feature; Step 4: Input the temporal aggregation features into a state space network, which includes an encoder, a decoder, and a skip connection structure connecting the encoder and the decoder. Step 5: In the encoder, multi-scale spatial context features are extracted through the dilated spatial fusion module, and downsampling and spatiotemporal dependency modeling are completed through state space scanning of the visual downsampling module to obtain multi-scale spatiotemporal features; Step 6: In the state space modeling process, the multi-scale spatial context features are unfolded into a sequence along the spatial and temporal directions. The long-distance dependency between different pixel positions and different time bins is modeled through state space scanning operation, and the scanned sequence features are restored to spatiotemporal features. Step 7: In the decoder, the multi-scale spatiotemporal features are upsampled and restored through the visual sampling module and the transposed convolution module, and the corresponding scale features in the encoder are fused to obtain the reconstructed features; Step 8: Output the probability distribution of the presence of target echo at each pixel location in each time bin based on the reconstructed features; Step 9: Determine the number of targets at each pixel location based on the probability distribution of the target echo, and calculate the depth position of the corresponding target based on the time bin of the target echo, thereby realizing multi-target reconstruction by lidar.
[0007] Further, in step 1, the photon counting tensor X is represented as: , in, Let C be the set of real numbers, T be the number of input channels, H be the number of time bins, and W be the height and width of the spatial image, respectively. When the input is single-channel photon counting data, C = 1.
[0008] Furthermore, in step 3, the time aggregation window takes the current time bin as its center and performs a weighted summation or average aggregation on several adjacent time bins before and after it, as shown below: , in, This represents the temporal aggregation feature of pixel position (i,j) at the k-th time bin, where r represents the radius of the temporal aggregation window, and ω... q X represents the weight of the q-th time offset position within the window. i,j,k+q This represents the photon count value for the corresponding time bin; The time aggregation window is used to form a local continuous response near the target echo, so that randomly distributed background photons cancel each other out or are weakened after aggregation, while the target reflected photons are cumulatively enhanced in adjacent time bins.
[0009] Furthermore, in step 5, the dilated spatial fusion module includes a regular 3D convolution branch and a dilated 3D convolution branch; wherein, the regular 3D convolution branch is used to extract local spatial-temporal features, and the dilated 3D convolution branch is used to expand the receptive field and extract a wider range of contextual features; the outputs of the regular 3D convolution branch and the dilated 3D convolution branch are combined, added or fused by convolution to obtain fused features; The output of the expansion space fusion module is expressed as follows: , Where X0 represents the input features, F1 represents the features extracted by the ordinary 3D convolution branch, F2 represents the features extracted by the dilated 3D convolution branch, [F1, F2] represents the concatenation operation along the channel dimension, and Conv 1×1×1 Y represents the 3D convolutional fusion operation, and Y represents the fused output feature.
[0010] Furthermore, the visual downsampling module in step 5 includes a normalization layer, a linear mapping layer, a state space scanning layer, a residual connection layer, and a downsampling layer; the visual downsampling module reduces the feature space resolution while retaining the target echo distribution information in the time dimension; The visual sampling module in step 7 includes an upsampling layer, a state space scanning layer, a feature fusion layer, and a feedforward network layer. The upsampling layer is used to restore feature resolution, the state space scanning layer is used to further model the spatial-temporal dependencies of the reconstruction stage, the feature fusion layer is used to fuse skip connection features of the same scale in the encoder, and the feedforward network layer is used to perform nonlinear mapping and channel feature enhancement on the fused features.
[0011] Furthermore, in step 6, the state space scan operation includes the following steps: (a) Expand the input spatial-temporal features into one-dimensional sequences according to the horizontal, vertical and temporal directions respectively; (b) Perform state-space recursive modeling on the one-dimensional sequence respectively to obtain scanning features in different directions; (c) Restore the scanning features in different directions to their original spatial-temporal arrangement; (d) The recovered multi-directional scanning features are fused to obtain spatiotemporal features that include global spatial correlation and temporal correlation.
[0012] Furthermore, in step 8, the output target echo probability distribution is expressed as: , Among them, P i,j p represents the probability distribution of the target at pixel position (i,j) over all time bins. i,j,kThis represents the probability that a target echo exists at pixel position (i,j) in the k-th time bin, where T is the number of time bins.
[0013] Further, step 9 includes: When p i,j,k When the probability exceeds a preset threshold and is a local peak, the k-th time bin is determined as the target echo location; the number of time bins that meet the conditions is counted to obtain the target number at pixel position (i,j); the target depth is calculated based on the time bin where the target echo is located, expressed as: , in, Let represent the depth of the m-th target at pixel position (i,j), c represent the speed of light, Δt represent the time interval between adjacent time bins, and k represent the depth of the m-th target at pixel position (i,j). m This represents the time bin number corresponding to the m-th target echo.
[0014] Furthermore, the training loss function of the state space network includes single-target reconstruction loss and multi-target reconstruction loss; wherein, the single-target reconstruction loss is used to constrain the probability distribution of a single target echo, and the multi-target reconstruction loss is used to constrain the existence, depth position and arrangement order of multiple target echoes; The multi-target reconstruction loss includes binary cross-entropy loss, depth position loss, and order constraint loss; wherein, the binary cross-entropy loss is used to determine whether there is a target echo in each time bin, the depth position loss is used to constrain the deviation between the predicted depth and the true depth, and the order constraint loss is used to constrain the sequential relationship between different target depths.
[0015] This invention also provides a lidar multi-target reconstruction system based on temporal aggregation and state-space networks, used to implement the lidar multi-target reconstruction method described above, including: The data acquisition module is used to acquire echo photon count data collected by the lidar and construct the photon count tensor. The preprocessing module is used to normalize the photon counting tensor; The temporal aggregation module is used to locally aggregate the photon counting tensor in the time dimension to obtain temporal aggregation features. The spatial fusion module is used to extract and fuse multi-scale spatial context features through ordinary 3D convolution and dilated 3D convolution. The state space modeling module is used to perform state space scanning of features along the spatial and temporal directions to obtain spatiotemporal features with long-distance dependencies. The decoding and reconstruction module is used to upsample and recover the spatiotemporal features, and output the probability distribution of the presence of target echoes at each pixel position in different time bins; and The target determination module is used to determine the number of targets and the corresponding target depth at each pixel location based on the target echo probability distribution.
[0016] Compared with existing technologies, the beneficial effects of this invention are as follows: Addressing the problems of low photon counting signal-to-noise ratio, indistinct target echo peaks, and overlapping echoes from multiple targets under background light conditions, this invention uses a temporal aggregation window to locally enhance the photon counting sequence, enabling target echoes to form a more stable continuous response within adjacent time bins and reducing the impact of random background light noise on individual time bins. Simultaneously, this invention extracts multi-scale spatial context features through an expanded spatial fusion module, improving the representation ability of complex target structures and spatial neighborhood relationships; and models long-distance dependencies in both spatial and temporal directions through state space scanning operations, allowing the network to simultaneously utilize spatial structure information and temporal echo distribution information, thereby improving the accuracy of multi-target quantity estimation and depth localization. Furthermore, this invention employs an encoder-decoder structure and a skip connection mechanism, preserving shallow detail information while extracting deep semantic features, which is beneficial for recovering the depth distribution of target edges, weak echo regions, and multi-layered structural regions. By outputting the target presence probability distribution of each pixel position in different time bins, this invention can simultaneously determine the number of targets and the depth position of each target, making it suitable for LiDAR 3D reconstruction in scenarios with strong background light, low signal-to-noise ratio, and complex multi-target environments. Attached Figure Description
[0017] Figure 1 This is an overall flowchart of a lidar multi-target reconstruction method based on temporal aggregation and state space network according to the present invention.
[0018] Figure 2 This is a schematic diagram of the overall structure of the state-space network in this invention;
[0019] Figure 3 This is a schematic diagram of the time aggregation window processing procedure in this invention;
[0020] Figure 4 This is a schematic diagram of the expansion space fusion module in this invention;
[0021] Figure 5 This is a schematic diagram of the structure of the visual downsampling module and the visual sampling module, and the three-dimensional state space scanning process in this invention;
[0022] Figure 6 This is a schematic diagram of the multi-target reconstruction effect of the method of the present invention under the background light intensity condition of SBR=1:100. Among them, (a) is the depth map recovered by the method of the present invention, and (b) is the error map, with white dots indicating the error.
[0023] Figure 7In the diagram, (a) shows the actual single-photon count distribution, and (b) shows the target echo probability distribution output by the method of this invention.
[0024] Figure 8 This is a schematic diagram illustrating the multi-target reconstruction effect of the method of the present invention under different background light intensity conditions. Detailed Implementation
[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] This invention provides a multi-target reconstruction method for lidar based on temporal aggregation and state-space networks. The overall process is as follows: Figure 1 As shown, this method first acquires photon counting data collected by a lidar system and constructs a photon counting tensor; then, it enhances the target echo response through a time aggregation window; further, it inputs the enhanced features into a state space network, and obtains reconstructed features through an encoder, an expanded space fusion module, a visual downsampling module, a state space scanning module, a visual sampling module, and a decoder; finally, it outputs the target echo probability distribution and determines the target quantity and depth location based on the probability distribution.
[0027] (1) Construct the photon counting tensor.
[0028] In this embodiment, the lidar system is preferably a single-photon lidar system. This system emits pulsed laser light towards the target scene and receives the reflected photons using a single-photon detector. For each pixel location, the system counts the photon arrival time and discretizes the photon arrival time into several time bins to obtain the corresponding photon count histogram.
[0029] For the entire imaging region, the photon count histograms of all pixel locations are arranged according to spatial coordinates to construct the photon count tensor X:
[0030]
[0031] in, Let T be the set of real numbers, where T represents the number of time bins, H and W are the height and width of the spatial image, respectively, H×W is the spatial resolution, and C is the number of input channels.
[0032] Under background light conditions, the photon counting tensor contains both target-reflected photons and background noise photons. Target-reflected photons are typically concentrated near the time bin corresponding to the target depth, while background noise photons are usually randomly distributed across time bins. Therefore, this invention further aggregates the time dimension to enhance target echoes and suppress the influence of random background noise.
[0033] (2) Perform preprocessing.
[0034] The input photon count tensor is normalized to ensure that the photon counts at different pixel locations and time bins are within a stable range. The normalized photon count tensor retains the original spatial location-time bin correspondence to ensure that subsequent networks can utilize spatial structure and temporal distribution information.
[0035] (3) Construct a time aggregation window.
[0036] For each pixel position (i,j), take its photon count sequence in the time dimension:
[0037]
[0038] X i,j,T The photon count value at pixel position (i,j) in the k-th time bin is represented. A time aggregation window is formed by selecting several adjacent time bins centered on the current time bin. The photon counts within the window are weighted and summed or averaged to obtain the time aggregation feature.
[0039]
[0040] in, This represents the temporal aggregation feature of pixel position (i,j) at the k-th time bin. The window radius r can be set according to the laser pulse width, detector temporal resolution, and target echo spread range. ω q X represents the weight of the q-th time offset position within the window. i,j,k+q This represents the photon count value for the corresponding time bin.
[0041] like Figure 3 As shown, the target reflected photons exhibit a locally continuous distribution near the time bin corresponding to the target depth, while background noise photons are typically randomly distributed along the time dimension. The temporal aggregation window, centered on the current time bin, performs local weighted aggregation on several adjacent time bins, cumulatively enhancing the effective photon response within the target echo neighborhood. In contrast, random background noise, lacking a stable locally continuous structure, has its relative influence smoothed or weakened after aggregation. Therefore, the features after temporal aggregation are more beneficial for subsequent state-space networks to identify the true target echo location.
[0042] (4) Input state space network.
[0043] The features aggregated over time are input into the state space network, such as... Figure 2 As shown, the state-space network employs an encoder-decoder structure. The encoder extracts multi-scale spatiotemporal features, while the decoder recovers spatial resolution and temporal distribution information. Skip connections are established between the encoder and decoder to fuse shallow detail features and deep semantic features.
[0044] In this embodiment, the state space network includes an initial feature fusion layer, a dilated space fusion module, a visual downsampling module, a visual sampling module, a transposed convolution module, and a final convolutional layer.
[0045] (5) Extract multi-scale features through the expansion space fusion module.
[0046] The dilated spatial fusion module is used to expand the receptive field without significantly increasing computational cost. This module includes both regular 3D convolution branches and dilated 3D convolution branches. Figure 4 As shown, the dilated spatial fusion module comprises two parallel branches: one branch extracts local spatial-temporal structural features through ordinary 3D convolution, and the other branch expands the receptive field and extracts a wider range of contextual features through dilated 3D convolution. The output features of the two branches are added, concatenated, or fused by convolution to obtain the fused features. Through this structure, the network can simultaneously utilize local echo details and large-scale spatial contextual information, thereby improving the representation ability of complex multi-target structures and weak echo regions.
[0047] Ordinary 3D convolutional branches extract local spatial-temporal structural features:
[0048]
[0049] Dilated 3D convolution branches extract broader contextual features:
[0050]
[0051] Among them, Conv 3×3×3 This represents a regular 3D convolution. Let d represent dilated 3D convolution, d represent the dilation rate, and σ represent the nonlinear activation function.
[0052] By merging the outputs of the two branches, we get:
[0053]
[0054] This structure enables the network to simultaneously obtain local target echo details and large-scale spatial context information, which is beneficial for the reconstruction of complex multi-target structures.
[0055] Where X0 represents the input features, F1 represents the features extracted by the ordinary 3D convolution branch, F2 represents the features extracted by the dilated 3D convolution branch, [F1, F2] represents the concatenation operation along the channel dimension, and Conv 1×1×1 Y represents the 3D convolutional fusion operation, and Y represents the fused output feature.
[0056] (6) Multi-scale coding is performed through the visual downsampling module.
[0057] The visual downsampling module is used to reduce the feature space resolution and extract higher-level spatiotemporal features. This module includes a normalization layer, a linear mapping layer, a state space scan layer, a residual connection layer, and a downsampling layer, such as... Figure 5 As shown.
[0058] During the encoding process, the input features are normalized and linearly mapped before entering the state space scanning layer. The state space scanning layer models the spatial and temporal dependencies of the features, thereby obtaining multi-scale features that incorporate the global context. Downsampling operations are used to reduce the spatial size and improve the abstract expressive power of deep features.
[0059] (7) Perform a three-dimensional state space scan.
[0060] For the input spatial-temporal feature F, it is expanded into a one-dimensional sequence along the horizontal, vertical, and temporal directions, respectively. A state-space recursive model is then performed for the sequence in each direction:
[0061]
[0062]
[0063] Where, x t h represents the t-th input feature in the sequence. t h t Let y represent the hidden states at time t and time t-1, respectively. t The output features are represented by A, B, C, and D, which are state-space parameters.
[0064] Through the aforementioned state-space recursion, the network can capture long-range dependencies with lower computational complexity. Compared to local modeling methods that rely solely on convolution, state-space scanning can better correlate distant pixel locations with distant temporal bins; compared to global attention mechanisms, this approach has better computational efficiency and is suitable for processing high-dimensional photon counting tensors. After completing scans in multiple directions, the scan results from each direction are restored to their original spatial-temporal arrangement and fused to obtain global spatiotemporal features, such as... Figure 5 As shown.
[0065] (8) Reconstruct features by using the visual sampling module and decoder.
[0066] During the decoding phase, the visual sampling module upsamples and recovers the encoded features, and further utilizes the state-space scanning layer to model the spatiotemporal dependencies of the reconstruction phase. The decoder simultaneously fuses skip connection features of the same scale from the encoder to recover detailed information about target edges, fine structures, and weak echo regions. The transposed convolution module progressively recovers the feature resolution, and the final convolutional layer outputs the target echo probability distribution.
[0067] The visual sampling module includes an upsampling layer, a state space scanning layer, a feature fusion layer, and a feedforward network layer. The upsampling layer is used to restore feature resolution, the state space scanning layer is used to further model the spatial-temporal dependencies in the reconstruction stage, the feature fusion layer is used to fuse skip connection features of the same scale in the encoder, and the feedforward network layer is used to perform nonlinear mapping and channel feature enhancement on the fused features to improve the reconstructed feature representation capability.
[0068] (9) Output target echo probability distribution.
[0069] For each pixel location (i,j), the network outputs the probability distribution of the target's presence across all time bins:
[0070]
[0071] Among them, P i,j p represents the probability distribution of the target at pixel position (i,j) over all time bins. i,j,k This represents the probability that a target echo exists at pixel position (i,j) at time bin k.
[0072] Unlike methods that directly output a single depth value, this invention outputs a time-bin-level probability distribution of target presence, which can simultaneously express the presence of multiple target echoes at the same pixel location. This is applicable to transparent, semi-transparent, occluded, and multi-layered structural scenes. Figure 7 As shown, in the case of a single-target echo, the actual single-photon count forms a peak near the time bin corresponding to the target depth, but there are also random counts caused by background light and dark counts. After processing by the state-space network of this invention, the output target echo probability distribution forms a significant probability peak at the time bin corresponding to the actual target, while the probability response at non-target time bins is suppressed. Therefore, based on this probability distribution, the time bin of the target echo can be determined more stably.
[0073] (10) Determine the number and depth of the target.
[0074] For the output probability distribution p i,jThe target echo location is determined based on a preset probability threshold and local peak values. When the probability value of a certain time bin is greater than the preset threshold and is a local peak value relative to its adjacent time bins, it is determined that a target echo exists in that time bin. The number of time bins that meet the above conditions is counted to obtain the number of targets at pixel position (i,j). For the m-th target, its depth position is calculated as follows:
[0075]
[0076] in, k represents the depth of the m-th target at pixel position (i,j). m Let be the time bin number corresponding to the m-th target, Δt be the time bin interval, and c be the speed of light. From this, the number of targets and the depth of each target at each pixel location can be obtained, ultimately forming the multi-target 3D reconstruction result.
[0077] (11) Network training process.
[0078] During training, training data including single-target and multi-target samples is constructed. For single-target samples, single-target reconstruction loss is used to constrain the consistency between the predicted probability distribution and the actual target echo location; for multi-target samples, multi-target reconstruction loss is used to constrain the target existence, depth location, and target order. The composite loss function L can be expressed as:
[0079]
[0080] Among them, L s L represents the single-objective reconstruction loss. m This indicates the loss during multi-target reconstruction.
[0081] The multi-objective reconstruction loss is further expressed as:
[0082]
[0083] Among them, L BCE L represents the binary cross-entropy loss, used to determine whether a target echo exists in each time bin; depth L represents the depth location loss, used to constrain the deviation between the predicted depth and the true depth. order λ1 and λ2 represent the order constraint loss, used to constrain the sequential order of multiple target depths; λ1 and λ2 are weighting coefficients.
[0084] Through the above training method, the network can learn the difference between target echoes and noise photons under background light conditions, and improve the performance of multi-target number estimation and depth localization.
[0085] (12) Implementation results.
[0086] When the background light is weak, this invention can accurately identify the target echo peak and recover the target depth. When the background light is enhanced and the signal-to-background ratio is reduced, this invention enhances the effective echo signal through a time aggregation window and captures global spatiotemporal correlation through state space scanning, thereby reducing misjudgments caused by false peaks and improving the recovery capability of weak echo targets. Figure 6 As shown, under strong background light conditions with SBR=1:100, the original photon counting data contains a large number of background noise photons, and the target echo peak is easily interfered with by noise. This invention enhances the effective response of the target echo neighborhood through a time aggregation window and models the long-distance dependencies in the spatial and temporal directions through a state-space network, thereby enabling the recovery of the target's spatial structure and depth distribution under strong background light conditions.
[0087] For multi-target scenarios, this invention does not output only a single depth, but rather the probability distribution of target presence at each time bin. Therefore, it can recover the depth positions of multiple target surfaces in the same pixel direction. This method is applicable to scenarios with complex background lighting, transparent targets, semi-transparent targets, occluded targets, multi-layered targets, and long-distance low signal-to-noise ratio detection, exhibiting good robustness and practical value. Figure 8 As shown, under different background light intensity conditions, the method of the present invention can recover the three-dimensional structure of multiple targets based on the target echo probability distribution.
[0088] This invention also provides a lidar multi-target reconstruction system based on temporal aggregation and state-space networks, used to implement the lidar multi-target reconstruction method described above. The system includes a data acquisition module, a preprocessing module, a temporal aggregation module, a spatial fusion module, a state-space modeling module, a decoding and reconstruction module, and a target determination module, all of which are computer programs.
[0089] The data acquisition module is used to acquire the echo photon count data collected by the lidar and construct the photon count tensor. For the specific implementation of the data acquisition module, please refer to step (1) of the lidar multi-target reconstruction method described above.
[0090] The preprocessing module is used to normalize the photon counting tensor. For the specific implementation of the preprocessing module, refer to step (2) of the above-described lidar multi-target reconstruction method.
[0091] The temporal aggregation module is used to locally aggregate the photon counting tensor in the time dimension to obtain temporal aggregation features. For the specific implementation of the temporal aggregation module, please refer to step (3) of the above-mentioned multi-target reconstruction method of lidar.
[0092] The spatial fusion module is used to extract and fuse multi-scale spatial context features through ordinary 3D convolution and dilated 3D convolution. For the specific implementation of the spatial fusion module, please refer to step (5) of the above-mentioned LiDAR multi-target reconstruction method.
[0093] The state-space modeling module is used to perform state-space scanning of features along the spatial and temporal directions to obtain spatiotemporal features with long-distance dependencies. For the specific implementation of the state-space modeling module, please refer to steps (6) and (7) of the above-mentioned LiDAR multi-target reconstruction method.
[0094] The decoding and reconstruction module is used to upsample and recover the spatiotemporal features and output the probability distribution of target echoes at different time bins for each pixel position. For the specific implementation of the decoding and reconstruction module, please refer to steps (8) and (9) of the above-mentioned multi-target reconstruction method of lidar.
[0095] The target determination module is used to determine the number of targets and the corresponding target depth at each pixel location based on the target echo probability distribution. For the specific implementation of the target determination module, refer to step (10) of the above-described multi-target reconstruction method for LiDAR.
[0096] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multi-target reconstruction method for lidar based on temporal aggregation and state-space networks, characterized in that, The method includes the following steps: Step 1: Obtain the echo photon count data collected by the lidar, and construct a photon count tensor based on the pixel spatial location and time bin. The photon count tensor is used to represent the number of photons received by different spatial pixel locations in different time bins. Step 2: Preprocess the photon counting tensor to obtain a normalized photon counting tensor, while preserving the correspondence between its spatial and temporal dimensions; Step 3: Construct a temporal aggregation window and locally aggregate the temporal photon count sequence corresponding to each pixel position to obtain the temporal aggregation feature; Step 4: Input the temporal aggregation features into a state space network, which includes an encoder, a decoder, and a skip connection structure connecting the encoder and the decoder. Step 5: In the encoder, multi-scale spatial context features are extracted through the dilated spatial fusion module, and downsampling and spatiotemporal dependency modeling are completed through state space scanning of the visual downsampling module to obtain multi-scale spatiotemporal features; Step 6: In the state space modeling process, the multi-scale spatial context features are unfolded into a sequence along the spatial and temporal directions. The long-distance dependency between different pixel positions and different time bins is modeled through state space scanning operation, and the scanned sequence features are restored to spatiotemporal features. Step 7: In the decoder, the multi-scale spatiotemporal features are upsampled and restored through the visual sampling module and the transposed convolution module, and the corresponding scale features in the encoder are fused to obtain the reconstructed features; Step 8: Output the probability distribution of the presence of target echo at each pixel location in each time bin based on the reconstructed features; Step 9: Determine the number of targets at each pixel location based on the probability distribution of the target echo, and calculate the depth position of the corresponding target based on the time bin of the target echo, thereby realizing multi-target reconstruction by lidar.
2. The lidar multi-target reconstruction method based on temporal aggregation and state-space networks according to claim 1, characterized in that, In step 1, the photon counting tensor X is represented as: , in, Let C be the set of real numbers, T be the number of input channels, H be the number of time bins, and W be the height and width of the spatial image, respectively. When the input is single-channel photon counting data, C = 1.
3. The lidar multi-target reconstruction method based on temporal aggregation and state-space networks according to claim 1, characterized in that, In step 3, the time aggregation window takes the current time bin as its center and performs a weighted summation or average aggregation on several adjacent time bins before and after it, as shown below: , in, This represents the temporal aggregation feature of pixel position (i,j) at the k-th time bin, where r represents the radius of the temporal aggregation window, and ω... q X represents the weight of the q-th time offset position within the window. i,j,k+q This represents the photon count value for the corresponding time bin; The time aggregation window is used to form a local continuous response near the target echo, so that randomly distributed background photons cancel each other out or are weakened after aggregation, while the target reflected photons are cumulatively enhanced in adjacent time bins.
4. The lidar multi-target reconstruction method based on temporal aggregation and state-space networks according to claim 1, characterized in that, In step 5, the dilated spatial fusion module includes a regular 3D convolution branch and a dilated 3D convolution branch; wherein, the regular 3D convolution branch is used to extract local spatial-temporal features, and the dilated 3D convolution branch is used to expand the receptive field and extract a wider range of contextual features. The outputs of the regular 3D convolution branch and the dilated 3D convolution branch are concatenated, added or fused by convolution to obtain fused features. The output of the expansion space fusion module is expressed as follows: , Where X0 represents the input features, F1 represents the features extracted by the ordinary 3D convolution branch, F2 represents the features extracted by the dilated 3D convolution branch, [F1, F2] represents the concatenation operation along the channel dimension, and Conv 1×1×1 Y represents the 3D convolutional fusion operation, and Y represents the fused output feature.
5. The multi-target reconstruction method for lidar based on temporal aggregation and state-space networks according to claim 1, characterized in that, The visual downsampling module in step 5 includes a normalization layer, a linear mapping layer, a state space scanning layer, a residual connection layer, and a downsampling layer; the visual downsampling module reduces the feature space resolution while retaining the target echo distribution information in the time dimension; The visual sampling module in step 7 includes an upsampling layer, a state space scanning layer, a feature fusion layer, and a feedforward network layer. The upsampling layer is used to restore feature resolution, the state space scanning layer is used to further model the spatial-temporal dependencies of the reconstruction stage, the feature fusion layer is used to fuse skip connection features of the same scale in the encoder, and the feedforward network layer is used to perform nonlinear mapping and channel feature enhancement on the fused features.
6. The multi-target reconstruction method for lidar based on temporal aggregation and state-space networks according to claim 1, characterized in that, Step 6, the state space scan operation includes the following steps: (a) Expand the input spatial-temporal features into one-dimensional sequences according to the horizontal, vertical and temporal directions respectively; (b) Perform state-space recursive modeling on the one-dimensional sequence respectively to obtain scanning features in different directions; (c) Restore the scanning features in different directions to their original spatial-temporal arrangement; (d) The recovered multi-directional scanning features are fused to obtain spatiotemporal features that include global spatial correlation and temporal correlation.
7. The lidar multi-target reconstruction method based on temporal aggregation and state-space networks according to claim 1, characterized in that, In step 8, the output target echo probability distribution is expressed as follows: , Among them, P i,j p represents the probability distribution of the target at pixel position (i,j) over all time bins. i,j,k This represents the probability that a target echo exists at pixel position (i,j) in the k-th time bin, where T is the number of time bins.
8. The multi-target reconstruction method for lidar based on temporal aggregation and state-space networks according to claim 7, characterized in that, Step 9 includes: When p i,j,k When the probability exceeds a preset threshold and is a local peak, the k-th time bin is determined as the target echo location; the number of time bins that meet the conditions is counted to obtain the target number at pixel position (i,j); the target depth is calculated based on the time bin where the target echo is located, expressed as: , in, Let represent the depth of the m-th target at pixel position (i,j), c represent the speed of light, Δt represent the time interval between adjacent time bins, and k represent the depth of the m-th target at pixel position (i,j). m This represents the time bin number corresponding to the m-th target echo.
9. The multi-target reconstruction method for lidar based on temporal aggregation and state-space networks according to claim 1, characterized in that, The training loss function of the state space network includes single-target reconstruction loss and multi-target reconstruction loss; wherein, single-target reconstruction loss is used to constrain the probability distribution of a single target echo, and multi-target reconstruction loss is used to constrain the existence, depth position and arrangement order of multiple target echoes. The multi-target reconstruction loss includes binary cross-entropy loss, depth position loss, and order constraint loss; wherein, the binary cross-entropy loss is used to determine whether there is a target echo in each time bin, the depth position loss is used to constrain the deviation between the predicted depth and the true depth, and the order constraint loss is used to constrain the sequential relationship between different target depths.
10. A lidar multi-target reconstruction system based on temporal aggregation and state-space networks, used to implement the lidar multi-target reconstruction method as described in any one of claims 1-9, characterized in that, include: The data acquisition module is used to acquire echo photon count data collected by the lidar and construct the photon count tensor. The preprocessing module is used to normalize the photon counting tensor; The temporal aggregation module is used to locally aggregate the photon counting tensor in the time dimension to obtain temporal aggregation features. The spatial fusion module is used to extract and fuse multi-scale spatial context features through ordinary 3D convolution and dilated 3D convolution. The state space modeling module is used to perform state space scanning of features along the spatial and temporal directions to obtain spatiotemporal features with long-distance dependencies. The decoding and reconstruction module is used to upsample and recover the spatiotemporal features and output the probability distribution of the presence of target echo at each pixel position in different time bins; as well as The target determination module is used to determine the number of targets and the corresponding target depth at each pixel location based on the target echo probability distribution.