SPAD depth optical flow estimation and reconstruction method based on counter overflow interval

By constructing a statistical sequence of counter overflow intervals and optical flow priors, and combining them with a pre-trained network for feature extraction and reconstruction, the challenges of image reconstruction and motion estimation under extremely low illumination and high-speed motion of passive SPADs are solved, achieving efficient image reconstruction and optical flow estimation.

CN122115604APending Publication Date: 2026-05-29XIDIAN UNIV HANGZHOU RES INST +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIDIAN UNIV HANGZHOU RES INST
Filing Date
2026-04-29
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Passive SPAD faces challenges in image reconstruction and motion estimation in extremely low-light and high-speed motion scenarios, including noise interference, blurring, fragmentation, and artifacts. Furthermore, directly using binary sequences as neural network inputs increases the computational burden and makes it difficult to stably perceive effective temporal structures.

Method used

By constructing a statistical sequence based on counter overflow intervals, optical flow priors are obtained using an optical flow algorithm. Feature extraction and reconstruction are performed using a pre-trained image reconstruction network, including a spiking neural network encoder, a convolutional neural network optical flow encoder, an image reconstruction decoder, and an optical flow decoder, to perform spatiotemporal feature extraction and alternating update of optical flow features.

Benefits of technology

It improves the temporal stability and reconstruction robustness under low photon count conditions, enhances image clarity and optical flow accuracy, adapts to the characteristics of passive SPAD sparse pulse data, and achieves high-quality reconstruction in high-speed dynamic scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115604A_ABST
    Figure CN122115604A_ABST
Patent Text Reader

Abstract

The application discloses a SPAD depth optical flow estimation and reconstruction method based on a counter overflow interval, belongs to the technical field of neuromorphic visual perception and image processing, and effectively improves the timing stability under a low photon counting condition through short-time statistical modeling and overlapping time window coding, so that the structural representation of a dynamic area is more reliable; a shared weight pulse neural network coding structure is adopted to reduce the parameter scale and enhance the consistency of cross-time window modeling, and the characteristics of passive SPAD sparse pulse data are adapted; an alternating update mechanism of image reconstruction and optical flow is introduced in the decoding stage, so that the optical flow can perform timing alignment and constraint on the reconstruction process, and the reconstruction feature can inversely refine the optical flow, thereby obtaining higher reconstruction robustness in a high-speed dynamic scene; through joint optimization of image reconstruction loss and optical flow loss, the structural recovery capability of an image at a key moment is improved, and the finally reconstructed image has higher clarity and stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of neuromorphic visual perception and image processing technology, specifically relating to a SPAD depth optical flow estimation and reconstruction method based on counter overflow interval. Background Technology

[0002] Passive single-photon avalanche diodes (SPADs) record photon-triggered events, capturing the arrival time of incident photons at the pixel level. The sensor outputs a binary data point at each sampling period, indicating whether a photon arrival was detected within that period. Successive binary frames constitute a high temporal resolution binary sequence, preserving the statistical characteristics of photon arrival in the scene even under extremely low illumination conditions. Passive SPADs rely solely on natural or ambient light for imaging, and the binary sequence directly reflects the true fluctuations in luminous flux.

[0003] For high temporal resolution sequences output by SPAD, image reconstruction mainly revolves around intensity reconstruction and motion estimation. One type of method utilizes the photon statistics contained in the binary sequence to recover the intensity image at the corresponding moment; another type of method infers the optical flow or motion field distribution of the scene from the temporal structure of the binary sequence.

[0004] Compared to traditional CMOS (Complementary Metal Oxide Semiconductor) image sensors, passive SPADs offer single-photon-level sensitivity, enabling operation in extremely low-light environments. Simultaneously, their output can achieve a time sampling rate far exceeding that of frame-based sensors, maintaining temporal continuity even in high-speed dynamic scenes. Furthermore, SPADs possess a wide dynamic range, allowing usable information to be obtained in both extremely low and high brightness regions through statistical analysis of photon arrival frequencies. These characteristics make passive SPADs an important carrier for low-light, high-speed imaging. However, passive SPAD data still faces significant challenges in reconstruction and motion analysis. First, within a single sampling period, the random arrival of photons exhibits high sparsity, making binary sequences susceptible to noise interference. In high-speed motion scenes, the instability of photon statistics further exacerbates blurring, fragmentation, or artifacts in the reconstructed image. Moreover, directly using binary sequences as neural network input increases computational burden and makes it difficult for the network to stably perceive effective temporal structure, particularly detrimental to time-sensitive tasks such as optical flow. Summary of the Invention

[0005] To address the aforementioned problems in the prior art, this invention provides a SPAD depth optical flow estimation and reconstruction method based on counter overflow intervals. The technical problem to be solved by this invention is achieved through the following technical solution: This invention provides a method for SPAD depth optical flow estimation and reconstruction based on counter overflow intervals, including: S1. Based on the binary sequence output by the single-photon camera, the number of photon arrivals is accumulated within a short time window using a pixel-level counter. By calculating the time interval between adjacent overflow events, a statistical sequence based on counter overflow is constructed. S2, based on statistical sequences, uses an optical flow algorithm to obtain optical flow priors; S3 utilizes a pre-trained image reconstruction network to extract features and reconstruct images from statistical sequences and optical flow priors, generating the final reconstructed image at preset key moments. The pre-trained image reconstruction network includes: a spiking neural network encoder, a convolutional neural network optical flow encoder, an image reconstruction decoder, and an optical flow decoder; The pulse neural network encoder is used to extract spatiotemporal features from statistical sequences to obtain spatiotemporal multi-scale features and spatiotemporal global features. The convolutional neural network optical flow encoder is used to extract multi-scale features from optical flow priors to obtain multi-scale optical flow features and global optical flow features. The image reconstruction decoder and optical flow decoder are used to perform alternating updates and mutual guidance between image reconstruction and optical flow from coarse to fine scale based on spatiotemporal multi-scale features, spatiotemporal global features, optical flow multi-scale features, and optical flow global features, generating the final reconstructed image at preset key moments.

[0006] The beneficial effects of this invention are: The solution provided by this invention effectively improves the temporal stability under low photon count conditions through short-time statistical modeling and overlapping time window coding, making the structural representation of dynamic regions more reliable. The shared-weight spiking neural network coding structure significantly reduces the parameter scale, enhances the consistency of cross-time window modeling, and adapts to the characteristics of passive SPAD sparse pulse data. An alternating update mechanism of image reconstruction and optical flow is introduced in the decoding stage, enabling optical flow to perform temporal alignment and constraint on the reconstruction process, while the reconstruction features can in turn refine the optical flow, thereby achieving higher reconstruction robustness in high-speed dynamic scenes. At the same time, the joint optimization of image reconstruction loss and optical flow loss improves the structural recovery capability of images at critical moments, resulting in higher clarity and stability of the final reconstructed image. Attached Figure Description

[0007] Figure 1 A schematic diagram illustrating the steps of a SPAD depth optical flow estimation and reconstruction method based on counter overflow interval provided in an embodiment of the present invention; Figure 2This is a schematic diagram of the image reconstruction network structure in a SPAD depth optical flow estimation and reconstruction method based on counter overflow interval provided in an embodiment of the present invention. Figure 3 This is a schematic diagram illustrating the training process of the image reconstruction network in a SPAD depth optical flow estimation and reconstruction method based on counter overflow interval provided in an embodiment of the present invention. Detailed Implementation

[0008] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0009] This invention provides a method for SPAD depth optical flow estimation and reconstruction based on counter overflow intervals, such as... Figure 1 As shown, it may include: S1. Based on the binary sequence output by the single-photon camera, the number of photon arrivals is accumulated within a short time window using a pixel-level counter. By calculating the time interval between adjacent overflow events, a statistical sequence based on counter overflow is constructed.

[0010] For step S1, it may include: S11, set a counter for each pixel of the binary sequence, accumulate the number of photon arrivals within a preset short time window, perform overflow judgment on the accumulated count value according to a preset threshold, and obtain the statistical binary sequence after the counter overflows. S12, based on the statistical binary sequence, calculate the time interval between two adjacent overflow events along the time axis to obtain the statistical sequence of each pixel, and use the statistical sequences of all pixels as the statistical sequence based on counter overflow.

[0011] Specifically, the binary sequence output by the single-photon camera can be represented as: ,in, Indicates the time step index. Indicates pixel position, This indicates that a photon arrival event occurs at this time step.

[0012] To obtain statistical characteristics of photon arrival, this embodiment of the invention sets an independent counter on each pixel and performs an accumulation overflow judgment on the number of photons arriving over time.

[0013] Initialize the counter at each pixel: ; When the time step is When the counter is updated, the update method is as follows: ; That is, it only accumulates once when a photon event is detected. A preset threshold is set. An overflow event is triggered when the accumulated count reaches a preset threshold; otherwise, the counter continues to increment. ; in, This represents the statistical binary sequence after the counter overflows.

[0014] If the time step is When the accumulated count reaches the threshold, an overflow event is triggered, and the counter is reset. ; Let the time step where the overflow event occurs be . Then the time interval Defined as: ; Thus, the resulting sequence .in, Indicates at time step An overflow event occurred at the counter. This represents the statistical sequence of the current pixel.

[0015] S2, based on statistical sequences, uses an optical flow algorithm to obtain optical flow priors, including: S21 transforms the statistical sequence into a continuous pseudo-grayscale image.

[0016] Specifically, statistical sequences It is continuous in time, and its value changes with variations in brightness or the incident photons caused by motion, thus it can be used as the input for calculating classical optical flow algorithms. : ; in, Indicates time interval index, This represents a grayscale mapping function that transforms a statistical sequence into a continuous pseudo-grayscale image.

[0017] S22, for pseudo grayscale images, use the classic optical flow algorithm to obtain optical flow priors.

[0018] for and Using classical optical flow algorithms Obtain optical flow prior : .

[0019] S3 utilizes a pre-trained image reconstruction network to extract features and reconstruct images from statistical sequences and optical flow priors, generating the final reconstructed image at preset key moments.

[0020] Among them, the pre-trained image reconstruction network, such as Figure 2 As shown, it may include: a spiking neural network encoder, a convolutional neural network optical flow encoder, an image reconstruction decoder, and an optical flow decoder; A spiking neural network encoder is used to extract spatiotemporal features from statistical sequences, obtaining spatiotemporal multi-scale features and spatiotemporal global features. A convolutional neural network optical flow encoder is used to extract multi-scale features from optical flow priors to obtain multi-scale optical flow features and global optical flow features. The image reconstruction decoder and optical flow decoder are used to perform alternating updates and mutual guidance between image reconstruction and optical flow from coarse to fine scale based on spatiotemporal multi-scale features, spatiotemporal global features, optical flow multi-scale features, and optical flow global features, generating the final reconstructed image at preset key moments.

[0021] The training process of a pre-trained image reconstruction network, such as... Figure 3 As shown, it may include: S01 organizes the statistical sequence according to the overlapping time window method and feeds it into the spiking neural network encoder with shared weights to extract spatiotemporal multi-scale features and spatiotemporal global features.

[0022] For step S01, it may include: S011, the statistical sequence is divided into time windows along the time axis, and multiple overlapping time window segments are constructed according to the preset time window length and time step; based on the characteristic that there is a non-zero overlap interval between adjacent overlapping time window segments, an input sequence composed of multiple overlapping time window segments is constructed.

[0023] Specifically, assuming the length of the preset time window segment is... L The time step size is b Then the first k The time range of a time window segment is defined as follows: ; The length of the overlapping region between adjacent windows is . In practical implementation, it is assumed that... L =11, b =8, then the overlap length of adjacent windows is 3.

[0024] The above method can be used to construct a system composed of... K The input sequence consists of overlapping time window segments , Indicates the first K A time window segment.

[0025] S012, the input sequence is fed into a shared-weight spiking neural network encoder, enabling each time window segment to extract spatiotemporal features under the same parameters, thereby obtaining spatiotemporal multi-scale features and spatiotemporal global features, which may include: Each time window segment is input as an independent branch to the same spiking neural network encoder; all branches use the same set of network parameters in each layer of the spiking neural network encoder to achieve feature extraction with shared weights across time windows; each layer of the spiking neural network encoder contains convolution operators and spiking neuron units. The convolution operator extracts the spatial features of each time window segment, uses the spatial features as input to the spiking neuron unit, and the spiking neuron unit performs membrane potential integration and threshold firing update, outputting the membrane potential state of the time step. The spatiotemporal features of the coding layer are obtained by summing and aggregating the membrane potential states of each time window segment within the same coding layer along the time dimension. By progressively reducing the spatial resolution of spatiotemporal features through strided convolution and downsampling structures, multiple spatial scale encoding feature layers are obtained. The spatiotemporal features of each spatial scale are output on the coding feature layer of each spatial scale to form spatiotemporal multi-scale features; The spatiotemporal features at the coarsest spatial scale are further fused across time windows to obtain spatiotemporal global features that characterize the overall temporal information.

[0026] Specifically, multiple overlapping time window segments are spliced ​​together to input the first coding layer of the spiking neural network (SNN) encoder: ; in, The time step membrane potential state is the output of the first layer of the SNN encoder. For the SNN encoder one layer, This represents the concatenated input of multiple time window segments along the feature dimension. For the ... i One coding layer: ; in, For the SNN encoder i The time-step membrane potential state output by the layer. For the SNN encoder i -1 layer output time step membrane potential state For the SNN encoder i Layer. This structure allows the temporal structure of the statistical sequence to accumulate continuously within the membrane potential, thereby enhancing its temporal expressiveness.

[0027] The window length of the time window segment can be 11, and the step size can be 8.

[0028] In this embodiment of the invention, each coding layer performs summation and aggregation on the membrane potential states of all time steps within a time window: ; in, Indicates the first i Each coding layer for time window segments The temporal aggregated features of all time window segments are densely connected along the feature dimension: ; This allows for the integration of short-term local structures at this spatial scale with long-term dependencies across windows.

[0029] The spiking neural network encoder incorporates strided convolution and downsampling operations to halve the spatial resolution layer by layer. ; in, and Indicates the first i Spatial resolution of each coding layer Indicates altitude, This represents the width. From this, we obtain the spatiotemporal multi-scale features: .in, B This represents the number of scales corresponding to spatiotemporal multiscale features.

[0030] The spatiotemporal multi-scale features at the coarsest spatial scale are further fused across time windows to obtain spatiotemporal global features that characterize the overall temporal information. .

[0031] S02, input the optical flow prior into the convolutional neural network optical flow encoder, perform multi-scale feature extraction on the optical flow prior, and obtain multi-scale optical flow features and global optical flow features, which may include: S021, the convolutional neural network optical flow encoder extracts optical flow priors layer by layer through a multi-level convolutional structure that includes downsampling operations, thereby obtaining optical flow context features at different spatial resolutions and forming multi-scale optical flow features from coarse to fine.

[0032] This yields multi-scale characteristics of optical flow at different scales: ; ; in, , This indicates that the optical flow encoder of the convolutional neural network is in the first... i Feature extraction modules at various spatial scales It is a priori for optical flow.

[0033] S022, at the coarsest spatial scale, feature fusion and residual update are performed on the multi-scale features of optical flow to obtain global optical flow features for global motion modeling. .

[0034] S03, the image reconstruction decoder and optical flow decoder, based on spatiotemporal multi-scale features, spatiotemporal global features, optical flow multi-scale features, optical flow global features, and optical flow priors, performs alternating updates and mutual guidance between image reconstruction and optical flow on a coarse-to-fine scale basis, outputting the corresponding reconstructed image and optical flow field, which may include: S031, at the coarsest spatial scale, the spatiotemporal multi-scale features and spatiotemporal global features of the coarsest spatial scale are input into the image reconstruction decoder to obtain the reconstruction features of the coarsest scale.

[0035] Specifically, the spatiotemporal multi-scale features at the coarsest spatial scale and the spatiotemporal global features are input into the image reconstruction decoder. The spatiotemporal multi-scale features at the current spatial scale are then aggregated along the temporal dimension to obtain aligned spatiotemporal multi-scale features. , To aggregate the spatiotemporal multi-scale features of the current spatial scale in the time dimension; in, This represents the spatiotemporal multiscale characteristics at the coarsest spatial scale.

[0036] At the coarsest spatial scale, the image reconstruction decoder receives aligned spatiotemporal multi-scale features and spatiotemporal global features. And it is fused along the channel dimension: .in, This represents the fusion characteristics at the coarsest spatial scale.

[0037] The feature update module obtains the reconstructed features at the coarsest spatial scale. : .

[0038] in, This represents the feature update module in the image reconstruction decoder.

[0039] S032, at the coarsest spatial scale, the reconstructed features of the coarsest spatial scale are input into the optical flow decoder, and the optical flow multi-scale features, global features and priors corresponding to the coarsest spatial scale are introduced. Correlation matching and feature aggregation are performed on the reconstructed features of the coarsest spatial scale, and the optical flow features and optical flow of the coarsest spatial scale are output.

[0040] At the coarsest spatial scale, the reconstructed features, multi-scale features, global features, and prior information of optical flow at this spatial scale are input into the optical flow decoder. Through multi-layer convolution and feature aggregation, the optical flow features and optical flow at the coarsest spatial scale are output.

[0041] In this embodiment of the invention, the reconstructed features are cascaded with multi-scale optical flow features and global optical flow features along the channel dimension. The input is given to an optical flow estimation module consisting of multi-layer convolution and feature aggregation units, and the output is the optical flow features and optical flow at the coarsest spatial scale. .in, Represents the optical flow characteristics at the coarsest spatial scale. Represents the optical flow at the coarsest spatial scale. This represents the optical flow estimation module of the optical flow decoder at the coarsest spatial scale.

[0042] S033, for each spatial scale other than the coarsest spatial scale, the alternating update and mutual guidance process between image reconstruction and optical flow includes: upsampling the optical flow features, optical flow, and reconstruction features of the previous spatial scale corresponding to the current spatial scale to the current spatial scale in order from coarse to fine; inputting the spatiotemporal multi-scale features of the current spatial scale into the image reconstruction decoder, using the upsampled optical flow to temporally align the spatiotemporal multi-scale features of the current spatial scale, and fusing them with the upsampled reconstruction features to obtain the reconstruction features of the current spatial scale; inputting the reconstruction features of the current spatial scale into the optical flow decoder, combining the optical flow multi-scale features of the current spatial scale, performing correlation matching and feature aggregation with the upsampled optical flow and optical flow features, and outputting the optical flow features and optical flow of the current spatial scale.

[0043] Specifically, at a finer spatial scale than the coarsest spatial scale. Above, the previous spatial scale Optical flow features, optical flow, and reconstructed features are upsampled to the current spatial scale: ; in, Indicates spatial scale The corresponding upsampled optical flow characteristics, Indicates upsampling, Indicates spatial scale The corresponding upsampled optical flow characteristics, Indicates spatial scale The corresponding upsampled optical flow, Indicates spatial scale The corresponding upsampled optical flow, Indicates upsampling to spatial scale The reconstruction features, Indicates upsampling to spatial scale The reconstruction features at the current spatial scale are input into the optical flow decoder, utilizing the spatial scale... Corresponding upsampled optical flow Align the reconstructed features: ; in, Indicates the current spatial scale The reconstructed features after alignment Indicates upsampling to the current spatial scale exist x Optical flow in direction Indicates upsampling to the current spatial scale exist y Optical flow in direction The data is concatenated along the channel dimension, input into a multi-layer convolutional network for cascaded feature aggregation, and outputs the current spatial scale. Optical flow characteristics and optical flow : ; in, This indicates the optical flow decoder at the current spatial scale. Optical flow estimation module on the device.

[0044] Utilizing optical flow at the current spatial scale For the current spatial scale Spatiotemporal multi-scale characteristics Perform temporal alignment to obtain aligned spatiotemporal multi-scale features. : ; in, Indicates the current spatial scale exist x Optical flow in direction Indicates the current spatial scale exist y Light flow in a specific direction.

[0045] The aligned spatiotemporal multi-scale features are then fused with the reconstructed features inherited from the previous spatial scale: ; Feature updates are performed using convolutional layers and residual structures to obtain the current spatial scale. Reconstruction characteristics: ; in, Indicates the current spatial scale The fusion characteristics Indicates upsampling to the current spatial scale The reconstruction features.

[0046] Following a coarse-to-fine spatial scale order, each spatial scale except the coarsest one is processed accordingly. This allows the image reconstruction decoder to align and aggregate spatiotemporal multi-scale features using the current optical flow, while the optical flow decoder refines the optical flow using the latest reconstructed features, thus creating an alternating update and mutual guidance between image reconstruction and optical flow. The optical flow decoder uses the latest reconstructed features to construct more reliable correlation features and refines the optical flow, gradually converging from a coarse global motion field to high-resolution detailed motion. The image reconstruction decoder uses the latest optical flow to temporally align and aggregate spatiotemporal multi-scale features at each spatial scale, gradually completing texture details and edge structures while reducing motion artifacts and ghosting.

[0047] S034, input the optical flow features corresponding to the finest spatial scale into the optical flow thinning module, use the convolutional layer to perform spatial context aggregation on the optical flow features corresponding to the finest spatial scale, and output the optical flow field between the thinned key frames; input the reconstructed features of the finest spatial scale into the image thinning module, further enhance the local texture and edge structure through residual convolution, and output the reconstructed image at the preset key moment.

[0048] Specifically, after completing multi-scale decoding, the finest spatial scale is obtained. Optical flow characteristics on With reconstruction features .

[0049] The optical flow features are input into an optical flow refinement module, which includes multiple convolutional layers with different dilation rates, and outputs a refined, high-resolution optical flow field. .

[0050] in, This is the optical flow refinement module.

[0051] Simultaneously, the reconstructed features at the finest spatial scale are input into the image thinning module. The image thinning module includes several residual convolutional structures, which are used to enhance local texture and edge information while preserving the global structure. This can be represented as: ; in, For image thinning module, To pre-set key moments, In order to be in critical moments The output reconstructed image.

[0052] S04, based on the optical flow field between the reconstructed image at a preset key moment and the key frame corresponding to the preset key moment, construct the total loss function of the image reconstruction network according to the time-weighted image reconstruction loss and optical flow loss; use the total loss function to jointly optimize the relevant parameters of the image reconstruction network to obtain the pre-trained image reconstruction network, which may include: S041, based on the pixel-level difference between the reconstructed image output by the image reconstruction decoder at a preset key moment and the corresponding real image, a time weight higher than that of non-preset key moments is assigned to the preset key moment to form a time-weighted image reconstruction loss; the pixel-level difference includes: mean square error loss and VGG loss.

[0053] For spatial scale ,time Define pixel-level mean square error loss for: ; in, Indicates time At the current spatial scale The total number of pixels involved in the calculation. Indicates time At the current spatial scale The reconstructed image output at pixel location ( x , y The pixel value at () This indicates the pixel position of the corresponding real image ( x , y The pixel value at ().

[0054] set up This indicates that the pre-trained VGG network in Each level of feature mapping, then the spatial scale VGG loss It can be defined as: ; in, Indicates the first The total number of pixels in the output feature map of each level. This indicates that the pre-trained VGG network is in the... Each level extracts feature maps from the reconstructed image. This indicates that the pre-trained VGG network is in the... Each level represents a feature map extracted from a real image. This represents the L2-norm squared operation in the feature space. The pre-trained VGG network is an auxiliary feature extraction network used to extract perceptual features of images. It is independent of the image reconstruction network and is only used to construct the VGG loss during the training phase.

[0055] Image reconstruction loss as follows: ; in, The weighting coefficients represent the pixel-level mean square error loss. This represents the weighting coefficient of the VGG loss term.

[0056] Preferably, It can be set to 1.0. It can be set to 0.05. It can be adjusted within the range of 0.5 to 2.0. It can be adjusted within the range of 0.01 to 0.1 to adapt to different datasets and noise conditions.

[0057] S042, construct optical flow smoothing constraints based on the optical flow field output by the optical flow decoder between key frames corresponding to preset key moments, calculate the total variation loss based on the first-order spatial gradient of the optical flow field in the horizontal and vertical directions, and align the reconstructed image through optical flow to obtain the reconstructed image aligned to the preset key moments corresponding to the key frames; compare the aligned reconstructed image with the corresponding real image at the pixel level to confirm the photometric consistency loss; obtain the optical flow loss based on the total variation loss and the photometric consistency loss.

[0058] An optical flow loss is constructed for the optical flow field output by the optical flow decoder between keyframes. This optical flow loss includes total variation loss and photometric consistency loss.

[0059] Specifically, the network is recorded from the input sequence The predicted optical flow is Define total variation loss. for: ; in, Indicates optical flow Gradient in spatial dimensions, This represents L1 norm operations.

[0060] For the photometric consistency loss, let the starting image of the keyframe obtained by the network reconstruction be denoted as . The final image is Using predicted optical flow ,Will Align to Let the alignment operator be . Then the loss of photometric uniformity Defined as: ; in, This represents the square of the Frobenius norm.

[0061] Optical flow loss as follows: .

[0062] in, The weighting coefficients represent the total variation loss. The weighting coefficient represents the loss of photometric uniformity.

[0063] Preferably, It can be set to 0.1. It can be set to 10.0. It can be adjusted within the range of 0.01 to 1.0. It can be adjusted within the range of 1 to 20 to adapt to different motion amplitudes and noise conditions in different scenarios.

[0064] S043, the image reconstruction loss and optical flow loss are weighted and combined according to preset weights to obtain the total loss function of the image reconstruction network; the total loss function is used to jointly optimize the relevant parameters of the spiking neural network encoder, the convolutional neural network optical flow encoder, the image reconstruction decoder and the optical flow decoder to obtain the pre-trained image reconstruction network.

[0065] Total loss function of image reconstruction network as follows: ; in, The weighting coefficients represent the relationship between image reconstruction loss and optical flow loss, used to adjust the relative contributions of the two types of losses to the total loss function. Preferably, It can be set to 1.0; it can also be adjusted according to actual needs under different datasets or noise conditions.

[0066] During the network training phase, backpropagation and gradient updates are performed on the total loss function to jointly optimize the learnable relevant parameters in the spiking neural network encoder, convolutional neural network optical flow encoder, image reconstruction decoder, and optical flow decoder.

[0067] After obtaining the pre-trained image reconstruction network, the network is used to extract features and reconstruct images from statistical sequences and optical flow priors, generating the final reconstructed image at preset key moments.

[0068] Specifically, based on the statistical sequence and optical flow prior obtained in steps S1 and S2, the corresponding input data is organized into the input of the pre-trained image reconstruction network in the same format as in the training phase, generating the final reconstructed image corresponding to the preset key moment.

[0069] The SPAD depth optical flow estimation and reconstruction method based on counter overflow intervals provided in this invention, compared to traditional methods that only perform image reconstruction or optical flow alone, comprehensively utilizes the temporal correlation inherent in the short-time statistical sequence based on counter overflow, the spatial prior constraints brought by scene structure and motion boundaries, and the synergistic effect of optical flow priors and multi-scale spatiotemporal decoders. This significantly improves image reconstruction quality and optical flow accuracy at critical moments in extremely challenging passive imaging scenarios such as extremely low photon counts and high-speed motion. Furthermore, this method can also be used as a basic framework for fusion with other depth reconstruction or motion estimation networks, thereby further improving the overall performance of passive SPAD imaging systems.

[0070] The experimental data for this SPAD depth optical flow estimation and reconstruction method were generated using a SPAD data simulation process. Images from the REDS dataset were used as the background and cropped to a resolution of 256×448. Three 3D objects from ShapeNet were randomly selected for the foreground, and the objects were rotated approximately... to Translation approx. Pixel to Pixels, the camera only performs approximately translation Pixel to Pixels; then, based on the NViSII rendering engine, the time series of each sample is rendered under the assumption of uniform motion, thus obtaining simulated data for training and evaluation. Training and testing are performed on the simulated data. Please refer to Table 1. The network input of the Baseline method is the SPAD sequence, and its network consists of a spiking neural network encoder and an image reconstruction decoder; Baseline+ISI (Interspike Interval) performs a counter overflow time interval operation on the input SPAD sequence, and inputs the resulting statistical sequence into the network, which consists of a spiking neural network encoder and an image reconstruction decoder; Baseline + ISI + Optical Flow is the method provided in the embodiments of the present invention, which performs a counter overflow time interval operation on the input SPAD sequence, obtains a statistical sequence, and then uses an optical flow algorithm to obtain an optical flow prior, inputting the statistical sequence and the optical flow prior into the network, which consists of a spiking neural network encoder, a convolutional neural network optical flow encoder, an image reconstruction decoder, and an optical flow decoder. Table 1 shows the effective improvements of the SPAD depth optical flow estimation and reconstruction method provided in the embodiments of the present invention.

[0071] Table 1. Experimental results of the method provided in the embodiments of the present invention.

[0072] This invention provides a SPAD depth optical flow estimation and reconstruction method based on counter overflow intervals. Short-time statistical modeling and overlapping time window coding effectively improve temporal stability under low photon count conditions, making the structural representation of dynamic regions more reliable. The shared-weight spiking neural network coding structure significantly reduces the parameter scale, enhances the consistency of cross-time window modeling, and adapts to the characteristics of sparse pulse data from passive SPADs. An alternating update mechanism for image reconstruction and optical flow is introduced in the decoding stage, enabling optical flow to temporally align and constrain the reconstruction process, while the reconstructed features can in turn refine the optical flow, thus achieving higher reconstruction robustness in high-speed dynamic scenes. Simultaneously, through joint optimization of image reconstruction loss and optical flow loss, the structural recovery capability of images at critical moments is improved, resulting in a final reconstructed image with higher clarity and stability.

[0073] The SPAD depth optical flow estimation and reconstruction method provided in this invention uses a pixel-level counter to accumulate the number of photon arrivals within a short time window based on the high frame rate binary sequence output by a single-photon camera. When the count reaches a preset threshold, an overflow event is generated, and the time interval between adjacent overflow events is calculated to construct a statistical sequence based on the counter overflow. Based on the short-time statistical sequence, an optical flow algorithm is used to obtain the optical flow prior. The statistical sequence is organized in an overlapping time window manner and fed into a spiking neural network encoder to extract spatiotemporal multi-scale features. The optical flow prior is input into a convolutional neural network optical flow encoder to extract multi-scale features of the optical flow prior, obtaining multi-scale features of optical flow. The spatiotemporal multi-scale features and the optical flow multi-scale features are input into an image reconstruction decoder and an optical flow decoder, respectively. The image reconstruction decoder outputs the reconstructed image at a preset key moment, and the optical flow decoder outputs the optical flow between key frames. A time-aware image reconstruction loss and an optical flow loss are constructed and jointly optimized for the image reconstruction network. The optimized image reconstruction network generates the final reconstructed image at the preset key moment. This invention utilizes short-time statistical feature modeling and keyframe-aware loss design to achieve joint optimization of intensity reconstruction and optical flow without requiring supervision of optical flow real values, thereby improving the performance of image reconstruction from high-frame passive SPAD data.

[0074] It should be noted that, in the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0075] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A method for SPAD depth optical flow estimation and reconstruction based on counter overflow interval, characterized in that, include: S1. Based on the binary sequence output by the single-photon camera, the number of photon arrivals is accumulated within a short time window using a pixel-level counter. By calculating the time interval between adjacent overflow events, a statistical sequence based on counter overflow is constructed. S2, based on statistical sequences, uses an optical flow algorithm to obtain optical flow priors; S3 utilizes a pre-trained image reconstruction network to extract features and reconstruct images from statistical sequences and optical flow priors, generating the final reconstructed image at preset key moments. The pre-trained image reconstruction network includes: a spiking neural network encoder, a convolutional neural network optical flow encoder, an image reconstruction decoder, and an optical flow decoder; The pulse neural network encoder is used to extract spatiotemporal features from statistical sequences to obtain spatiotemporal multi-scale features and spatiotemporal global features. The convolutional neural network optical flow encoder is used to extract multi-scale features from optical flow priors to obtain multi-scale optical flow features and global optical flow features. The image reconstruction decoder and optical flow decoder are used to perform alternating updates and mutual guidance between image reconstruction and optical flow from coarse to fine scale based on spatiotemporal multi-scale features, spatiotemporal global features, optical flow multi-scale features, and optical flow global features, generating the final reconstructed image at preset key moments.

2. The SPAD depth optical flow estimation and reconstruction method based on counter overflow interval according to claim 1, characterized in that, S1 includes: A counter is set for each pixel of the binary sequence, and the number of photon arrivals is accumulated within a short time window. An overflow judgment is performed on the accumulated count value according to a preset threshold to obtain a statistical binary sequence after the counter overflows. Based on the statistical binary sequence, the time interval between two adjacent overflow events is calculated along the time axis to obtain the statistical sequence of each pixel. The statistical sequences of all pixels are then used as the statistical sequence based on the counter overflow.

3. The SPAD depth optical flow estimation and reconstruction method based on counter overflow interval according to claim 1, characterized in that, S2 includes: Transform the statistical sequence into a continuous pseudo-grayscale image; For pseudo-grayscale images, the optical flow prior is obtained using the classic optical flow algorithm.

4. The SPAD depth optical flow estimation and reconstruction method based on counter overflow interval according to claim 1, characterized in that, The training process of the pre-trained image reconstruction network includes: S01, the statistical sequence is organized according to the overlapping time window method and fed into the spiking neural network encoder with shared weights to extract spatiotemporal multi-scale features and spatiotemporal global features; S02, input the optical flow prior into the convolutional neural network optical flow encoder, perform multi-scale feature extraction on the optical flow prior, and obtain the multi-scale features and global features of optical flow; S03, the image reconstruction decoder and optical flow decoder, based on spatiotemporal multi-scale features, spatiotemporal global features, optical flow multi-scale features, optical flow global features, and optical flow priors, performs alternating updates and mutual guidance between image reconstruction and optical flow from coarse to fine scale, and outputs the corresponding reconstructed image and optical flow field; S04. Based on the optical flow field between the reconstructed image at the preset key moment and the key frame corresponding to the preset key moment, construct the total loss function of the image reconstruction network according to the time-weighted image reconstruction loss and optical flow loss; use the total loss function to jointly optimize the relevant parameters of the image reconstruction network to obtain the pre-trained image reconstruction network.

5. The SPAD depth optical flow estimation and reconstruction method based on counter overflow interval according to claim 4, characterized in that, S01 includes: The statistical sequence is divided into time windows along the time axis, and multiple overlapping time window segments are constructed according to the preset time window length and time step. Based on the feature that there are non-zero overlap intervals between adjacent overlapping time window segments, an input sequence composed of multiple overlapping time window segments is constructed. The input sequence is fed into a spiking neural network encoder with shared weights, so that spatiotemporal features are extracted from each time window segment under the same parameters, thereby obtaining spatiotemporal multi-scale features and spatiotemporal global features.

6. The SPAD depth optical flow estimation and reconstruction method based on counter overflow interval according to claim 4, characterized in that, S02 includes: The convolutional neural network optical flow encoder extracts optical flow priors layer by layer through a multi-level convolutional structure that includes downsampling operations, thereby obtaining optical flow context features at different spatial resolutions and forming multi-scale optical flow features from coarse to fine. At the coarsest spatial scale, feature fusion and residual update are performed on the multi-scale features of optical flow to obtain global optical flow features for global motion modeling.

7. The SPAD depth optical flow estimation and reconstruction method based on counter overflow interval according to claim 4, characterized in that, S03 includes: S031, At the coarsest spatial scale, the spatiotemporal multi-scale features and spatiotemporal global features of the coarsest spatial scale are input into the image reconstruction decoder to obtain the coarsest scale reconstruction features. S032, at the coarsest spatial scale, the reconstructed features of the coarsest spatial scale are input into the optical flow decoder, and the optical flow multi-scale features, global features and priors corresponding to the coarsest spatial scale are introduced. Correlation matching and feature aggregation are performed on the reconstructed features of the coarsest spatial scale, and the optical flow features and optical flow of the coarsest spatial scale are output. S033, for each spatial scale except the coarsest one, the alternating update and mutual guidance process between image reconstruction and optical flow includes: upsampling the optical flow features, optical flow, and reconstruction features of the previous spatial scale corresponding to the current spatial scale to the current spatial scale in order from coarse to fine; inputting the spatiotemporal multi-scale features of the current spatial scale into the image reconstruction decoder, using the upsampled optical flow to temporally align the spatiotemporal multi-scale features of the current spatial scale, and fusing them with the upsampled reconstruction features to obtain the reconstruction features of the current spatial scale; inputting the reconstruction features of the current spatial scale into the optical flow decoder, combining the optical flow multi-scale features of the current spatial scale, performing correlation matching and feature aggregation with the upsampled optical flow and optical flow features, and outputting the optical flow features and optical flow of the current spatial scale. S034, input the optical flow features corresponding to the finest spatial scale into the optical flow thinning module, use the convolutional layer to perform spatial context aggregation on the optical flow features corresponding to the finest spatial scale, and output the optical flow field between the thinned key frames; input the reconstructed features of the finest spatial scale into the image thinning module, further enhance the local texture and edge structure through residual convolution, and output the reconstructed image at the preset key moment.

8. The SPAD depth optical flow estimation and reconstruction method based on counter overflow interval according to claim 4, characterized in that, S04 includes: Based on the pixel-level difference between the reconstructed image output by the image reconstruction decoder at a preset key moment and the corresponding real image, a time weight higher than that of non-preset key moments is assigned to the preset key moment to form a time-weighted image reconstruction loss; the pixel-level difference includes: mean square error loss and VGG loss. Optical flow smoothing constraints are constructed based on the optical flow field output by the optical flow decoder between key frames corresponding to preset key moments. The total variation loss is calculated based on the first-order spatial gradient of the optical flow field in the horizontal and vertical directions. Simultaneously, the reconstructed image is aligned using optical flow to obtain the reconstructed image aligned to the preset key moments corresponding to the key frames. The aligned reconstructed image and the corresponding real image are compared at the pixel level to confirm the photometric consistency loss. The optical flow loss is obtained based on the total variation loss and the photometric consistency loss. The image reconstruction loss and optical flow loss are weighted and combined according to preset weights to obtain the total loss function of the image reconstruction network. The total loss function is then used to jointly optimize the relevant parameters of the spiking neural network encoder, the convolutional neural network optical flow encoder, the image reconstruction decoder, and the optical flow decoder to obtain the pre-trained image reconstruction network.