Method for reconstructing video frame sequence based on event stream
By generating event representations of short-range and long-range physical memory patterns and multi-branch end-to-end neural networks, the problems of event camera information loss and recursive error accumulation in sparse areas are solved, and the robust reconstruction of event flow in the fields of robot vision and autonomous driving is achieved.
Patent Information
- Application Number
- CN202510410778.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-08-26
AI Technical Summary
In the prior art, the event stream generated by event cameras is difficult to directly apply to advanced visual tasks, especially in sparse area information is lost and recursive reconstruction frameworks are prone to accumulate errors, which limits its application potential in the fields of robot vision and autonomous driving.
The video frame sequence reconstruction method based on event stream is adopted, and short-range and long-range physical memory modes are generated through the event representation module of dynamic physical memory, and feature tensors are extracted in combination with multi-branch end-to-end neural networks to realize the reconstruction of video frames.
It effectively overcomes the problems of information loss and recursive error accumulation in event sparse areas, and achieves robust reconstruction performance under high time resolution and low latency.
Smart Images

Figure CN120547447A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and more particularly to a method for reconstructing a video frame sequence based on an event stream. Background Art
[0002] A dynamic vision sensor, also known as an event camera, is a biomimetic vision sensor. Each time a change in light intensity or a change in the logarithm of light intensity (i.e., a difference) exceeds a predefined threshold, an event (x, y, t, p) is triggered and output. This event is recorded with the pixel coordinates (x, y), a timestamp t, and a polarity p∈{-1,+1} indicating a decrease or increase in light intensity. The pixel of a dynamic vision sensor is the basic detection unit. These events are output when triggered, reflecting the temporal changes of the scene through a continuous data stream (called an event stream) rather than a series of static frames. The operating principle of an event camera enables extremely high temporal resolution (microseconds), a wide dynamic range, and low-latency dynamic visual perception. Therefore, event cameras have shown significant advantages in fields such as robotic vision, autonomous driving, and augmented reality. However, the event stream generated by an event camera is a discrete and sparse four-dimensional spatiotemporal signal, making it difficult to directly apply existing image processing techniques (which are typically based on continuous, dense pixel data) to this data type. This makes it difficult to directly apply event cameras to advanced vision tasks such as object detection and image segmentation, limiting their potential in these areas. Therefore, it is very important to reconstruct video frame images based on pure event streams, which can be immediately adapted to vision tasks.
[0003] In real-world scenarios, event cameras often produce event streams with highly uneven spatiotemporal distributions. Due to the lack of model memory, end-to-end direct mapping architectures suffer from severe information loss in regions of low event density. Currently, recursive reconstruction frameworks dominate the field of event stream-based video reconstruction, incorporating memory and recursively leveraging past states to compensate for the missing information. Model memory is able to recover lost details in regions of low event density. However, when learned in a data-driven manner, this memory can lead to the accumulation of errors in regions of persistently low event density. For example, in driving scenes, the sky and ground often result in persistently sparse event triggers. This leads to poor reconstruction due to the lack of distinct texture features and visual artifacts caused by memory effects that propagate across frames. Interestingly, end-to-end architectures without model memory are able to alleviate this problem. Notably, neither paradigm can address the persistently low event density distribution caused by narrow temporal windows. Unlocking the full potential of event streams—such as their high temporal resolution, wide dynamic range, and low latency—remains an open challenge.
[0004] In contrast, biological systems maintain perceptual stability through physical adaptive mechanisms: retinal neurons dynamically adjust their spike frequency—with transient bursts for rapid brightness changes and sustained low-frequency responses for gradual changes—to achieve robust environmental perception. This has inspired a paradigm shift: integrating physics-based memory into event-stream-based video reconstruction to circumvent data-driven uncertainty and recursive error propagation. Summary of the Invention
[0005] The present invention aims to overcome the shortcomings of the prior art, such as the end-to-end architecture suffering from information loss in event-sparse areas due to lack of memory, and the recursive architecture relying on a data-driven memory mechanism that may accumulate errors, making it difficult to overcome the persistent low-density distribution caused by a narrow time window, which seriously limits the potential of event data in high-speed and low-latency scenarios. A method for reconstructing video frame sequences based on event streams is provided.
[0006] In order to solve the above technical problems, the technical solutions of the present invention are as follows:
[0007] In a first aspect, the present invention provides a method for reconstructing a video frame sequence based on an event stream, comprising the following steps:
[0008] Step 1. The dynamic vision sensor receives light information from the outside world through the relay device.
[0009] The dynamic vision sensor is composed of an arrangement of basic detection units for receiving light information, and each basic detection unit independently senses changes in the received light intensity, and triggers an event and outputs an event stream when the perceived light intensity change exceeds a preset threshold.
[0010] Furthermore, the relay device modulates and guides external light information to be incident on the dynamic vision sensor;
[0011] Step 2. An event stream representation module with dynamic physical memory generates two types of pattern representations from the event stream output by the dynamic visual sensor: a type I pattern representing the spatiotemporal information of event data with short-range physical memory, and a type II pattern representing the image sequence with long-range physical memory;
[0012] Among them, short-range physical memory refers to any time point along the time dimension of the event flow, including the spatiotemporal information of event data within a time window at that time point. This time window is named the short-range physical memory time window corresponding to that time point;
[0013] Moreover, long-range physical memory refers to any time point along the time dimension of the event flow. Within a time window containing this time point, information is accumulated by each basic detection unit as the incremental information. The accumulated information is named the long-range physical memory time window corresponding to this time point.
[0014] The Class I mode has AI ≥ 1 modes, including: modes 510-A1, 510-A2, ..., 510-AI;
[0015] Moreover, the Class II mode has BI≥1 modes, including: modes 520-B1, 520-B2, ..., 520-BI;
[0016] Step 3. Select at least one of the AI modes of the Class I mode and at least one of the BI modes of the Class II mode to form MX ≥ 1 mixed type combinations;
[0017] Select SI≥1 combinations from AI patterns of Class I patterns to form SI single-type combinations of Class I;
[0018] or / and, select SII≥1 combinations from B1 patterns of Class II patterns to form SII single-type combinations of Class II;
[0019] Set each combination to correspond to an end-to-end neural network to independently extract its own feature tensor;
[0020] Step 4. Select SI'≥0 combinations from the AI patterns of the Class I mode, and select SII'≥0 combinations from the BI patterns of the Class II mode. Input them together with the feature tensor extracted in Step 3 into an end-to-end neural network to generate the final feature fusion reconstructed light intensity image at a time point, that is, reconstruct the video frame;
[0021] Step 5: Repeat steps 1 to 4 to obtain the reconstructed light intensity image corresponding to each time point until the task of reconstructing the video frame sequence based on the event stream is completed;
[0022] The end-to-end neural network has no time dimension memory for the input data;
[0023] In the method for reconstructing a video frame sequence based on an event stream, the end-to-end neural network used in steps 3 and 4 and its input-output association relationship constitute an end-to-end neural network reconstruction module.
[0024] As a preferred solution, the event stream output by the dynamic vision sensor is the original event stream data output by the dynamic vision sensor, or the event stream after the original event stream is denoised, or the event stream after the original event stream is edited.
[0025] As a preferred embodiment, the Class I pattern for representing the spatiotemporal information of event data with short-term memory includes but is not limited to event voxel grid representation, event graph representation, and time surface, which is used to encode the spatiotemporal information of each event; the Class II pattern for representing the image sequence with long-term physical memory, based on the attenuation response function, accumulates the event information to obtain the light intensity or logarithmic value of the light intensity at each time point along the time dimension for each basic detection unit; wherein the attenuation response function is an exponential attenuation function, or a non-exponential attenuation function, or an exponential-non-exponential mixed type; and the accumulation is integration, or summation, or a mixture of integration and summation.
[0026] As a preferred embodiment, the exponential decay function includes but is not limited to a single exponential decay function and a multi-exponential decay function, and the non-exponential decay function includes but is not limited to a power-law decay function, a Mittag-Leffler decay function, a fractional-order Mittag-Leffler decay function, a linear decay function, a fractal decay response function, or a combination thereof.
[0027] As a preferred solution, at any time point t along the time dimension of the event stream n , using formula (1) to accumulate the light information carried by the events in the corresponding time window of each basic detection unit as the incremental value, the expression is:
[0028] logA′(x,y,t k )=H(t k -t k-1 )·logA′(x,y,t k-1 )+p k ΔC (1)
[0029] The obtained time point t n The total amount of accumulated information is called the amount of information accumulated by each basic detection unit at time point t n The intensity image of is represented by A′(x,y);
[0030] At any time point t n The corresponding time window corresponds to a basic detection unit attenuation mask Named as dynamic global "remember-forget" mask;
[0031] Then, each basic detection unit at any time point t n The characterization of the type II pattern is done with Its mathematical expression is:
[0032]
[0033] Among them, any event e in the event stream k (x,y,t k ,pk ) is timestamp t k , polarity is p k (Increase in light intensity is +1, decrease is -1), the coordinates are (x, y); A′(x, y, t k ) represents each basic detection unit at timestamp t k Light intensity image at time H(t k -t k-1 ) is the attenuation response function, ΔC is the preset light intensity logarithmic change threshold for the dynamic vision sensor trigger event; t k-1 and t k Represents two adjacent timestamps in the event stream;
[0034] Among them, t n With t k One-to-one correspondence, partial correspondence, or no correspondence at all.
[0035] As a preferred solution, the attenuation mask Using formula (3):
[0036]
[0037] Among them, for any basic detection unit, when it is at any time point t n When no event occurs within the corresponding time window, the parameter α is 0; when an event occurs, the value of the parameter α is dynamically adjusted according to the relative motion state of the dynamic vision sensor: when the dynamic vision sensor is moving and / or there is relative motion with the scene, the value of the parameter α is greater than the first preset threshold, so that the function exhibits an accelerated attenuation characteristic to enhance the degree of "forgetting" of the light information obtained in the current time window; when the dynamic vision sensor is close to stillness and / or there is no relative motion with the scene, the value of the parameter α is less than the second preset threshold, so that the function exhibits a slowed attenuation characteristic to enhance the degree of "memory" of the light information obtained in the current time window; the specific values of the first preset threshold and the second preset threshold are dynamically adjusted according to the signal attenuation requirements in the actual application scenario.
[0038] As a preferred solution, a dynamic time window event counting analysis strategy is also included. This strategy analyzes the event distribution over the first j time windows. Its mathematical formula is:
[0039]
[0040] If the number of events in the current time window exceeds n times the average number of events in the previous j time windows, then α = γ a , to enhance the degree of “forgetting” of light information acquired in the current time window; if the number of events in the current time window does not exceed η times the average number of events in the previous j time windows, then α=γb , to enhance the "memory" of the light information obtained in the current time window; where j ≥ 1; 0 < η < 1, and γ a >γ b ; is the current time point t n The number of events in the corresponding time window.
[0041] As a preferred solution, the attenuation response function H(t k -t k-1 ) is an exponential decay function along the time dimension;
[0042] Formula (1-1) is used to accumulate the optical information carried by the event stream in each basic detection unit:
[0043] logA′(x,y,t k )=exp[-λ(t k -t k-1 )]·logA′(x,y,t k-1 )+p k ΔC (1-1)
[0044] It is called the event leakage-integral model; where λ is the attenuation factor;
[0045] Or, use formula (1-2) to accumulate the optical information carried by the event stream in each basic detection unit:
[0046]
[0047] It is called the event leakage-integral model with a harvesting parameter h; h ≥ 0 is the harvesting parameter.
[0048] As a preferred solution, the attenuation response function is a non-exponential attenuation function, or a combination of an exponential attenuation function and a non-exponential attenuation function;
[0049] Alternatively, a fractional-order Mittag-Leffler type function is used as the attenuation response function, and by introducing fractional-order parameters, the attenuation process from exponential attenuation (short time, <τ) to power-law attenuation (long time, >τ) is unified under a mathematical framework to embed memory; the attenuation response function contains or does not contain a harvesting restriction, contains or does not contain a total amount restriction; wherein τ is the characteristic time constant of the attenuation.
[0050] As a preferred solution, the architecture of each end-to-end neural network constituting the end-to-end neural network reconstruction module includes but is not limited to convolutional neural networks, Transformer, Mamba and other deep learning architectures with feature extraction functions, or a combination thereof; the architecture of each end-to-end neural network is entirely the same, partially the same, or different; its input layer matches the representation output by the event representation module with dynamic physical memory; the output layer includes a feature tensor output layer, a predicted image output layer, or a combination thereof.
[0051] As a preferred solution, the end-to-end neural network reconstruction module includes three end-to-end neural network branches: and branch road Take MX=1 mixed type combination as input and extract feature tensor ω1; branch Take SII = 1 type II single type combination or SI = 1 type I single type combination as input and extract the feature tensor ω2; branch Taking the feature tensors ω1 and ω2, and SII'=1 type II single type combination or SI'=1 type I single type combination as input, the final feature fusion reconstructed intensity image sequence, that is, the reconstructed video frame sequence, is generated.
[0052] As a preferred solution, the three end-to-end neural network branches and Share the same architecture, including the head convolution layer Conv1, recursive residual group RRG and image prediction layer Conv2; and The head convolution layer Conv1 and recursive residual group RRG in each architecture are used and express; and The output of is the extracted feature tensor; and The output of is a light intensity image; and and The corresponding components of the architecture may be all the same, partially the same, or completely different.
[0053] As a preferred solution, the end-to-end neural network reconstruction module adopts an end-to-end training strategy; preferably, the branches are parallel; each branch has an independent optimizer and loss function; during the backpropagation process, each optimizer updates the parameters of its respective branch.
[0054] As a preferred embodiment, the method further introduces a controllable blink module, wherein the dynamic vision sensor receives light information from the outside world via the controllable blink module and a relay device; wherein the controllable blink module includes a photomask, an occlusion rate control unit for adjusting the occlusion rate of the photomask, an occlusion area control unit for adjusting the occlusion area of the photomask, and a general control unit; in the controllable blink module, the general control unit issues a command, and the occlusion area control unit or the occlusion rate control unit drives the photomask to perform a "close-open" operation, that is, the degree of occlusion of the incident light of the dynamic vision sensor by the photomask changes from small to large to maintain to small to maintain; and the occlusion rate of the incident light of the photomask in the controllable blink module is equal to or less than 100%;
[0055] Then, the method further comprises the following steps:
[0056] Step z1. Determine the initial time window of the “close-open” operation [T s ,T e ]; where T s is the starting time point of the closing process in the “close-open” operation, T e The end time point of the open holding process in the "close-open" operation;
[0057] Step z2 sets the occlusion area and / or occlusion rate related parameters, and the general control unit issues a command to implement a corresponding "close - open" operation;
[0058] Step z3. Determine the opening time window of a "close-open" operation [T Os ,T Oe ], that is, the time interval corresponding to the opening process;
[0059] The time window [T Os ,T Oe The event triggered by the change of the shielding area or shielding rate of the photomask within ] is characterized as the time point T by the following process. Oe The corresponding static background light intensity image:
[0060] Formulas (5), (6) and (7) are used to perform linear or nearly linear event integration algorithm for each basic detection unit to obtain the value of each basic detection unit at time point T Oe The absolute light intensity value of all basic detection units at time point T Oe The absolute light intensity value at time point T Oe The corresponding static background light intensity image; its expression is:
[0061]
[0062] H(t i -t i-1 )=exp[-β·(t i -t i-1 )] (7)
[0063] Where I represents the time point T Oe The intensity image, i.e. the logarithmic value, ε{e i} means that when opening the time window [T Os ,T Oe ] is a set of events triggered by changes in the occlusion area or occlusion rate of the photomask; Formula (5) represents the event set for opening the time window [T Os ,T Oe ] is integrated by the basic detection unit event integration algorithm to obtain the time point T of each basic detection unit. Oe The logarithm of the absolute light intensity value;
[0064] Formula (6) is the update rule of the linear or nearly linear event integration algorithm for each basic detection unit; I(x, y, t i ) represents the coordinates (x, y) of the basic detection unit of the dynamic vision sensor at the time stamp t i The logarithm of the absolute light intensity value, t i represents ε{e i}The i-th event e i Timestamp of p i Represents the i-th event e i The event polarity; ΔC represents the preset constant threshold for event triggering;
[0065] Formula (7) is the expression of the response function H(t) of the basic detection unit event integration algorithm; β is the attenuation factor, and by adjusting the value of β, formula (5) becomes a linear or nearly linear event integration function;
[0066] Step z4. Set the time window of a “close-open” operation [T Oe ,T e ] The event stream output by the dynamic visual sensor is processed by steps 1 to 5 of the method for reconstructing a video frame sequence based on the event stream;
[0067] Step z5. Use T e +δt is the starting time point T of the closing process in the next "close-open" operation s , update the time window for the next "close-open" operation; repeat steps z1 to z4 until the task of reconstructing the video frame sequence based on the event stream is completed, and the light intensity image corresponding to each time point is obtained; where δt≥0 is the interruption duration.
[0068] As a preferred solution, in the time window of a "close-open" operation [T Oe ,T e ] In the event characterization module with dynamic physical memory, the characterization of the Class II pattern takes the light intensity image at the time point TOe in the step z3 as the initial value at the time point.
[0069] As a preferred embodiment, the relay device is a lens, or a lens group, or a diffraction optical device, or a diffraction optical device group, or a prism, or a reflector, or a polarizer, or a color filter, or a color filter array, or various combinations of the above devices; the photomask is imaged on the surface where the basic detection unit of the dynamic vision sensor is located through the relay device, or is imaged at a position far away from the dynamic vision sensor along the transmission direction of the incident light.
[0070] As a preferred solution, an aperture stop is introduced between the dynamic vision sensor and the outside world, and its size is adjusted according to the scene light intensity, so that the entire system can work in scenes with a dynamic contrast range greater than the intrinsic dynamic contrast range of the dynamic vision sensor.
[0071] As a preferred solution, an attenuation plate with an adjustable attenuation coefficient is introduced between the dynamic vision sensor and the outside world to adjust the incident light flux of the dynamic vision sensor.
[0072] As a preferred solution, in the event stream representation module with dynamic physical memory, the method of generating a Class II pattern representing an image sequence with long-range physical memory based on the event stream is independently used as the method of reconstructing a video frame sequence based on the event stream; and the image sequence with long-range physical memory is independently used as the reconstructed video frame sequence based on the event stream.
[0073] As a preferred solution, step 3 further includes: not selecting from the A1 modes of the Class I mode and not selecting from the B1 modes of the Class II mode, but only selecting at least one from the A1 modes of the Class I mode and at least one from the B1 modes of the Class II mode to form MX ≥ 1 mixed type combinations;
[0074] Each of the combinations is set to correspond to an end-to-end neural network, and each feature tensor is independently extracted.
[0075] In a second aspect, the present invention further proposes a database construction method, which is applied to the method for reconstructing a video frame sequence based on an event stream described in the present invention, and is used to train the end-to-end neural network reconstruction module, comprising the following steps:
[0076] An active pixel sensor is introduced; the active pixel sensor and the dynamic vision sensor share a basic detection unit array, which is used to output image data in a continuous frame mode within a preset time interval, and all basic detection units receive light signals with the same exposure time within the same time period;
[0077] The dynamic vision sensor and the active pixel sensor receive light information from the outside world via the relay device and / or the controllable blink module;
[0078] In a real space scene, the active pixel sensor outputs continuous frame images, and the dynamic vision sensor outputs an event stream;
[0079] The event stream and continuous frame images of each scene are acquired to form a database; wherein the continuous frame images serve as the label truth value GT of the reconstructed video frame image sequence of its corresponding event stream.
[0080] As a preferred solution, to improve dataset quality, during data collection, the scenarios described minimize overexposure, motion blur, and other label truth information loss in the active pixel sensor. To this end, the method further includes quality control of the collected data using overexposure detection and / or motion blur filtering; the scenarios include indoor and outdoor environments, varying camera motion states, independent object motion, and / or objects with varying textures; and the temporal and spatial density characteristics of the event distribution in the collected dataset overlap with the corresponding characteristics of the scene to be inferred.
[0081] As a preferred solution, when the dynamic contrast range and / or motion speed of the scene exceeds the dynamic contrast range and temporal resolution of the active pixel sensor, the frame image captured by a high-speed camera and / or a high dynamic contrast camera is corrected and used as the true label value GT.
[0082] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0083] The present invention proposes an end-to-end physically guided network architecture. Through multi-branch feature extraction and fusion, it overcomes the spatial and temporal density imbalance in the event stream, avoids the problem of information loss in event-sparse areas caused by the lack of memory in existing end-to-end architectures, and avoids the problem of recursive architectures relying on data-driven memory mechanisms that may accumulate errors, thereby achieving state-of-the-art reconstruction performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] Figure 1 An architectural diagram of the method for reconstructing video frame sequences based on event streams.
[0085] Figure 2 Schematic diagram of the data flow of the method for reconstructing a video frame sequence based on an event stream.
[0086] Figure 3 Schematic diagram of an example of an event flow representation module with dynamic physical memory.
[0087] Figure 4 Schematic diagram of an example of the overall training strategy for the end-to-end neural network reconstruction module.
[0088] Figure 5 A schematic diagram of an example of the overall inference architecture of the end-to-end neural network reconstruction module.
[0089] Figure 6 This is a framework diagram for the method of building a database.
[0090] Figure 7 Another architectural diagram of a method for reconstructing a video frame sequence based on an event stream.
[0091] Figure 8 Diagram of the basic optical structure of another architecture for the method of reconstructing video frame sequences based on event streams.
[0092] Figure 9 Schematic diagram of the time points of the video frames to be reconstructed and the event timestamps along the time dimension of the event stream. DETAILED DESCRIPTION
[0093] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.
[0094] The terms used in this invention are for the purpose of describing specific embodiments only and are not intended to limit the invention. The singular forms "a," "the," and "the" used in this invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0095] It should be understood that although the terms "first," "second," "third," etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, first information may also be referred to as second information, and similarly, second information may also be referred to as first information, without departing from the scope of the present invention. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to determining."
[0096] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0097] Example 1
[0098] See Figure 1 , which is an architectural diagram of a method for reconstructing video frame sequences based on event streams. Here, a dynamic vision sensor 10 receives light information from the outside world via a relay device 30 and outputs an event stream. This is then fed into an event representation module 50 with dynamic physical memory. The output of this event representation module 50 serves as input to an end-to-end neural network reconstruction module 60, which then reconstructs the video frame sequence.
[0099] For the convenience of description, the system composed of the dynamic vision sensor 10, the relay device 30 and the controllable blink module 20 that receives light information from the outside world and outputs an event stream is referred to as an event-based imaging system.
[0100] Each basic detection unit of the dynamic vision sensor 10 independently senses the change in received light intensity and triggers an event when the perceived light intensity change exceeds its preset threshold. The basic detection unit can be a sub-pixel, or a pixel formed by arranging sub-pixels on a surface, or a pixel formed by superimposing sub-pixels. In other words, each basic detection unit of the dynamic vision sensor 10 processes the change in incident light intensity independently, continuously, and asynchronously. When the absolute value of the change in light intensity received by the basic detection unit, or the change in the logarithm of light intensity, or the change in other functions of light intensity exceeds the preset constant threshold ΔC, an event e is triggered. k , optionally with (x,y,t k ,p k ) is represented by (x, y) where (x, y) is the unit coordinate of the basic detection; the timestamp t k is the time corresponding to the logarithmic change of light intensity exceeding ΔC; p k ∈{-1,+1} is polarity, p k =1 indicates an increase in light intensity, which is called a positive event; p k = -1 indicates a decrease in light intensity, known as a negative event. Events are output when they occur, and a continuous data stream (called an event stream) records the scene changes over time. Typically, a preset threshold ΔC for triggering time is set based on the absolute value of the logarithmic change in light intensity.
[0101] The event stream may be the original event stream data output by the dynamic vision sensor 10, or the event stream after the original event stream has been denoised, or the event stream after the original event stream has been edited. The event editing includes, but is not limited to, extracting a portion of one event stream sequence and inserting it into another event stream sequence, thereby reconstructing a video frame sequence based on the two event stream sequences.
[0102] The relay device 30 is used to modulate and guide external light to enter the dynamic vision sensor 10 .
[0103] The relay device 30 can be various optical devices, such as a diffractive optical device, a group of diffractive optical devices, a prism, a reflector, a polarizer, a color filter, or various combinations of the above devices, to modulate and guide light incident from the scene to the dynamic vision sensor 10. For example, the diffractive device or the group of diffractive devices can guide spectral information from the observed object to the dynamic vision sensor 10, or guide light of different frequencies to be incident on corresponding areas of the dynamic vision sensor 10; for example, the prism and the reflector can modulate the geometric structure of the imaging system by deflecting the optical path; for example, the polarizer can select corresponding deflection characteristic light from the observed object to be incident on the dynamic vision sensor 10; for example, the color filter can select corresponding color light from the observed object to be incident on the dynamic vision sensor 10. Generally, the relay device 30 can be various optical devices, such as a lens or lens group with imaging function, or a diffractive device or free-form surface group with imaging function, which are referred to as imaging relay devices, and images the observed object onto the basic detection unit surface of the dynamic vision sensor 10.
[0104] like Figure 1 As shown, the relay device 30 may further include an aperture stop 310, the size of which is adaptively adjusted based on the scene light intensity, enabling the entire system to operate in scenes with a dynamic contrast range greater than the intrinsic dynamic contrast range of the dynamic vision sensor 10. The aperture stop 310 may also be a separate device introduced as a component of the relay device 30, with the size of its clear aperture adjusted manually or electronically.
[0105] like Figure 1 As shown, the relay device 30 may further include an attenuation sheet 40 with an adjustable attenuation coefficient. The attenuation sheet 40 is placed on the incident light transmission path of the dynamic vision sensor 10, and the position between the attenuation sheet 40 and other relay devices 30 can be adjusted as needed, not just as shown in FIG. Figure 1 It is shown positioned between the relay device 30 and the dynamic vision sensor 10 .
[0106] The exemplary data flow diagram of the method for reconstructing a video frame sequence based on an event stream in this embodiment is as follows: Figure 2 As shown, it includes an event representation module 50 with dynamic physical memory and an end-to-end neural network reconstruction module 60.
[0107] In steps 1 and 2, the event stream output by the dynamic vision sensor 10 is input into the event representation module 50 with dynamic physical memory to generate representations of two types of patterns: Type I pattern representing the spatiotemporal information of event data with short-range physical memory, and Type II pattern representing the image sequence with long-range physical memory.
[0108] The representation of the Class I pattern should encode and store the spatiotemporal optical information carried by asynchronous, discrete events in the event stream as accurately and completely as possible; such representation determines the upper limit of the quality of the video frame sequence reconstructed based on the event stream.
[0109] The representation of the Class II pattern encodes the contextual association information of the spatiotemporal distribution between asynchronous and discrete events in the event stream, and compensates for the discontinuity and instability problems associated with the unique asynchrony and discreteness of event data through information accumulation; this type of representation ensures the stability of the video frame sequence reconstructed based on the event stream and determines the lower limit of the reconstruction quality.
[0110] Among them, the Class I mode has AI≥1 modes: mode 510-A1, 510-A2, ..., 510-AI; and the Class II mode has BI≥1 modes: mode 520-B1, 520-B2, ..., 520-BI.
[0111] Step 3: Select at least one of the AI modes of the Class I mode and at least one of the BI modes of the Class II mode to form MX ≥ 1 mixed type combinations;
[0112] Select SI≥1 combinations from AI patterns of Class I patterns to form SI single-type combinations of Class I;
[0113] Or / and, select SII≥1 combinations from B1 patterns of Class II patterns to form SII single-type combinations of Class II.
[0114] For example, see Figure 2 In the example of , one mode 510 - A2 is selected from the type I mode, and one mode 520 - B1 is selected from the type II mode, forming MX=1 mixed type combinations.
[0115] It can be understood that the selection and combination methods that meet the requirements are all within the scope of the requirements of this patent. For example, two modes are selected from Class I modes, such as 510-A1 and 510-A3, and three modes are selected from Class II modes, such as 520-B1, 520-B2 and 520-B3, to form a mixed type combination; one mode is selected from Class I modes, such as 510-A1, and two modes are selected from Class II modes, such as 520-B1 and 520-B2, to form a mixed type combination; then a total of MX=2 mixed combinations are formed.
[0116] For example, see Figure 2 In the example, no Class I single type combination is selected from the Class I pattern and formed, and one Class II single type combination (i.e., 520-B1) is selected from the Class II pattern and formed; then a total of SII=1 Class II single type combination is formed.
[0117] It is understood that any selection method that meets the requirements is within the scope of the present patent. For example, 510-A2 is selected from the Class I pattern to form a Class I single type combination, and 520-B1 and 520-B3 are selected from the Class II pattern to form a Class II single type combination; thus, SI = 1 Class I single type combination, and SII = 1 Class II single type combination are formed.
[0118] Furthermore, each combination is set to correspond to an end-to-end neural network, and each feature tensor is extracted independently.
[0119] For example, see Figure 2 In the example above, SII+MX=2 end-to-end neural networks are required to independently extract their own feature tensors. It can be understood that the combination selected in the above example requires SI+SII+MX=4 end-to-end neural networks to independently extract their own feature tensors;
[0120] Step 4: Select SI' ≥ 0 combinations from the AI modes of the Class I mode, and select SII' ≥ 0 combinations from the BI modes of the Class II mode, and input them together with the feature tensor extracted in Step 3 into an end-to-end neural network to generate the final feature fusion reconstructed light intensity image at a time point, that is, reconstruct the video frame.
[0121] For example, see Figure 2 For example, SI'=0 single type combinations of type I are selected from AI types of types of type I, and SII'=1 single type combination of type II is selected from BI types of types of type II. These are input together with the feature tensor extracted in step 3 into an end-to-end neural network to generate the final feature fusion reconstructed light intensity image at a time point, i.e., to reconstruct the video frame. It can be seen that Figure 2 In the example, the end-to-end neural network reconstruction module 60 includes 3 end-to-end neural networks.
[0122] It is understandable that any selection and combination method that meets the requirements is within the scope of the present patent. For example, in step 4, 510-A1 is selected from the AI patterns of the Class I mode to form SI'=1 Class I single type combination, and 520-B2 is selected from the BI patterns of the Class II mode to form SII'=1 Class II single type combination, which is input together with the feature tensor extracted in step 3 into an end-to-end neural network to generate the final feature fusion reconstructed light intensity image at a time point, that is, to reconstruct the video frame; it can be seen that in this case, combined with the above examples, the end-to-end neural network reconstruction module 60 includes 5 end-to-end neural networks.
[0123] Repeat steps 1 to 4 to obtain the reconstructed light intensity image corresponding to each time point until the task of reconstructing the video frame sequence based on the event stream is completed; each time the execution is repeated, the corresponding MX, SI, SII, SI', and SII' values can be the same, or they can be partially or completely different.
[0124] The end-to-end neural network in this embodiment has no time dimension memory for the input data.
[0125] In the method for reconstructing a video frame sequence based on an event stream, the end-to-end neural network and its input-output association relationship used in steps 3 and 4 constitute an end-to-end neural network reconstruction module 60.
[0126] As a special case, step 3 neither selects from the AI modes of Class I mode nor from the BI modes of Class II mode, but only selects at least one from the AI modes of Class I mode and at least one from the BI modes of Class II mode, forming MX≥1 mixed type combinations.
[0127] In an optional embodiment, the architecture of each end-to-end neural network constituting the end-to-end neural network reconstruction module includes but is not limited to a deep learning architecture with feature extraction function such as a convolutional neural network (CNN), a Transformer, a Mamba, or a combination thereof; the architecture of each end-to-end neural network is entirely the same, partially the same, or different; its input layer matches the representation output by the event representation module 50 with dynamic physical memory; the output layer includes a feature tensor output layer, a predicted image output layer, or a combination thereof.
[0128] It can be understood that the method of reconstructing a video sequence based on an event stream can first output an event stream for a period of time or an entire period of time through the dynamic visual sensor 10, without processing the original event data, or performing denoising, editing, etc., and then adopting the method of reconstructing a video sequence based on an event stream to obtain a reconstructed video frame sequence; or at the same time as the dynamic visual sensor 10 outputs the event stream, the relevant process of the method of reconstructing a video sequence based on an event stream is started to obtain a reconstructed video frame sequence; during online processing, although the characterization of the two types of modes of the event stream can be compatible with the incremental method of processing as it arrives, considering the diversity of computing power platform configurations, as well as the diversity of end-to-end neural network architectures, parameters, etc., there is a possibility that the inference speed of the reconstruction module is synchronized or delayed with the generation rate of the event stream.
[0129] Figure 3This is an example of an event stream representation module 50 with dynamic physical memory. The event stream representation module 50 with dynamic physical memory generates two types of pattern representations from the event stream output by the dynamic vision sensor 10: Type I patterns representing the spatiotemporal information of event data with short-term physical memory, and Type II patterns representing image sequences with long-term physical memory.
[0130] Among them, short-range physical memory refers to any time point along the time dimension of the event flow, including the spatiotemporal information of event data within a time window at that time point. This time window is named the short-range physical memory time window corresponding to that time point;
[0131] Long-range physical memory refers to any time point along the time dimension of the event flow. Within a time window containing this time point, information is accumulated by each basic detection unit as the new amount, and the accumulated information is named the long-range physical memory time window corresponding to this time point.
[0132] The time points set in Class I mode and the time points set in Class II mode can be completely the same, partially different, or completely different; the time windows corresponding to the time points set in Class I mode and the time windows corresponding to the time points set in Class II mode can be completely the same, partially different, or completely different.
[0133] Preferably, the time points set in the Class I mode correspond one-to-one with the time points set in the Class II mode, and maintain a one-to-one correspondence with the time points of each frame in the to-be-reconstructed video frame sequence; the time windows corresponding to each time point also maintain a one-to-one correspondence.
[0134] It is understood that for any event stream, the number of time points set in the Type I mode, the number of time points set in the Type II mode, and the number of time points corresponding to each frame in the to-be-reconstructed video frame sequence may be the same or different. For example, the number of time points corresponding to each frame in the to-be-reconstructed video frame sequence may be less than the number of time points set in the Type I mode and / or the number of time points set in the Type II mode.
[0135] The Class I patterns, including but not limited to Event Voxel Grid representation (EVG representation), Event Graph representation (EG representation), Time Surface, etc., encode the spatiotemporal information of each event and can provide high-frequency light information—texture details—for reconstructing video frame sequences based on event streams.
[0136] The Class II mode utilizes an attenuation response function to accumulate event information to obtain the light intensity or logarithmic value of the light intensity at each time point along the time dimension for each basic detection unit; the attenuation response function is an exponential attenuation function, or a non-exponential attenuation function, or an exponential-non-exponential mixed type; the exponential attenuation function includes a single exponential attenuation function, or a multi-exponential attenuation function; the non-exponential attenuation function includes a power-law attenuation function, or a Mittag-Leffler attenuation function, or a fractional-order Mittag-Leffler attenuation function, or a linear attenuation function, or a fractal attenuation response function, or a combination thereof; the attenuation response function may contain or include a harvesting restriction, a total amount restriction or not, etc.; the accumulation is integration, or summation, or a mixed type of integration-summation; the Class II mode can provide low-frequency light information-features for reconstructing a video frame sequence based on an event stream, and stabilize the reconstruction process.
[0137] For example, the implementation of a mode of the Class I mode is described by taking EVG characterization as an example. n , assuming that the length of the corresponding time window is fixed to ΔT, then the time window is expressed as [t n -ΔT,t n ]; the event stream and its index in the time window are recorded as N e is the number of events in the time window, and each event is represented by four coordinates (x i ,y i ,t i ,p i ) represents, where t i ∈[t n -ΔT,t n ].
[0138] It is understandable that at any time point t n The corresponding time window length ΔT is not necessarily a fixed value; corresponding to the time point t of the frame to be reconstructed n The current time window can also be [t n -ΔT,t n ], or [t n ,t n +ΔT], or a time window of length ΔT that includes this time point.
[0139] In order to adapt to the deep learning processing paradigm, such as convolutional neural networks (CNNs), the event stream is encoded into a spatiotemporal voxel grid tensor of size B×H×W, where B is the number of tensors (e.g. B=5), and H and W are the height and width of the image. Then the time point t n The EVG characterization encoding process is mathematically expressed as:
[0140]
[0141] Where δ(·) is the Kronecker delta function, t i * is the normalized timestamp.
[0142] It can be understood from this example that the characterization of type I patterns aims to effectively store the spatiotemporal distribution of optical information within the current time window (i.e., the short-range physical memory).
[0143] For example, for the Class II mode, an exponential decay function is used in its implementation to accumulate event light information, that is, the decay response function H(t k -t k-1 ) takes an exponential decay function.
[0144] like Figure 3 Figure 1 shows an example of an event stream representation module with dynamic physical memory. This example uses the Alpha image sequence representation as an example to illustrate the implementation of a Type II pattern. Specifically, this implementation involves aligning and fusing two branches along the temporal dimension to form a Type II pattern, called the Alpha image sequence representation.
[0145] In order to compensate for the information loss in the sparse regions of the event stream, the entire time period to be reconstructed is considered so that the historical information from the previous frames can supplement these areas where events are insufficient. The event stream in the entire time period to be reconstructed is represented as Among them, if k∈[N r -N e +1,N r ],but Where i represents the event index in the current time window, and k represents the event index in the entire time period; t n Refers to the time point at which the frame is reconstructed.
[0146] For example, a simple rule is used: the pixel intensity at position (x, y) is updated by processing the event e generated in the current time window k (x,y,t k ,p k ) to deduce any time point t n The light intensity at . The event stream can be viewed as a spike train, where the intensity information is encoded by the timestamp and polarity of each event. Inspired by the Leaky Integrate-and-Fire (LIF) neuron model in neural dynamics, we adopt the Event Leaky-Integration (ELI) model to process spike trains.
[0147] Specifically, branch ① uses the event leakage-integration model to accumulate the optical information carried by the event stream in each basic detection unit, and its mathematical expression is:
[0148] logA′(x,y,t k )=exp[-λ(t k -t k-1 )]·logA′(x,y,t k-1 )+p k ΔC(s3-1)
[0149] Wherein, λ is the attenuation factor, ΔC is the preset light intensity logarithmic change threshold for the dynamic vision sensor 10 to trigger the event, for example, λ=0.03, ΔC=0.002; t k-1 and t k It represents the two adjacent timestamps before and after the trigger event stream of the dynamic vision sensor 10, (x, y) is the coordinate of the basic detection unit, A′(x, y, t k ) is the basic detection unit at the time stamp t k Light intensity image at time .
[0150] Or, the mathematical expression (s3-2) is used to accumulate the optical information carried by the event stream in each basic detection unit:
[0151]
[0152] Formula (s3-2) is called the event leakage-integral model with harvesting parameter h; where λ is the attenuation factor and h≥0 is the harvesting parameter.
[0153] To minimize redundant computation, the ELI model uses the accumulated intensity image from the previous time point as an intermediate variable for recursive updates. While this physical memory mechanism compensates for event sparsity through historical integration, the persistence of this memory can lead to severe artifacts: during periods of inactivity, static regions devoid of new triggering events retain outdated motion features, resulting in persistent artifacts. These distortions disrupt spatiotemporal continuity and significantly increase the ambiguity of feature extraction.
[0154] Because the motion of the dynamic vision sensor 10 is dynamic and complex, a balance needs to be struck between "forgetting" and "memory." When the dynamic vision sensor 10 exhibits significant relative motion relative to the scene, the model should enhance its "forgetting" capability to mitigate motion artifacts. Conversely, when relative motion is negligible, the model should strengthen its "memory" to prevent information loss in sparse event regions.
[0155] To achieve this, this embodiment proposes a dynamic global attenuation algorithm, namely, the dynamic global "memory-forget" mask of branch ②. This algorithm is integrated into the ELI model of branch ①, and is aligned and fused along the time dimension to form an image sequence representation pattern with dynamic long-range physical memory. Its mathematical formula is:
[0156]
[0157] in, corresponds to time point t n The time window (e.g. [t n -ΔT,t n ]) of the basic detection unit attenuation mask; for any basic detection unit, when no event occurs in the time window Set to 1 or a value close to 1, when an event occurs Set to a value greater than 0 but less than 1; is the basic detection unit at time point t n The two branches are aligned and fused along the time dimension to form an image sequence representation pattern with dynamic long-range physical memory.
[0158] When the dynamic global "memory-forget" mask Using the mathematical expression (s5-1), the resulting representation of the Class II pattern is called the Alpha image sequence representation:
[0159]
[0160] For any basic detection unit, when it is at any time point t n When no event occurs in the corresponding time window, let the parameter α be equal to 0; when an event occurs, the value of the parameter α is dynamically adjusted according to the relative motion state of the dynamic vision sensor 10; when the dynamic vision sensor 10 is moving and / or there is relative motion with the scene, the value of the parameter α is greater than the first preset threshold, and the function exhibits an accelerated attenuation characteristic to enhance the degree of "forgetting" of the light information obtained in the current time window; when the dynamic vision sensor 10 is close to stillness and / or there is no relative motion with the scene, the value of the parameter α is less than the second preset threshold, and the function exhibits a slowed attenuation characteristic to enhance the degree of "memory" of the light information obtained in the current time window; the specific values of the first preset threshold and the second preset threshold are dynamically adjusted according to the signal attenuation requirements in the actual application scenario.
[0161] Preferably, this embodiment proposes a dynamic time window event counting analysis strategy, which analyzes the event distribution over the first j time windows to infer the current motion state of the dynamic vision sensor 10. The mathematical formula is:
[0162]
[0163] If the number of events in the current time window exceeds n times the average number of events in the previous j time windows, it is inferred that the dynamic vision sensor 10 is moving, then α = γ a , to enhance the degree of “forgetting” of light information acquired in the current time window; if the number of events in the current time window does not exceed η times the average number of events in the previous j time windows, it is inferred that the dynamic vision sensor 10 is close to being stationary, then α = γ b , to enhance the "memory" of the light information obtained in the current time window; where j ≥ 1; 0 < η < 1, and γ a >γ b ; is the current time point t n The number of events in the corresponding time window.
[0164] For example, γ a , γ b and η are set to 3×10 -2 and 3×10 -7 , η is set to 0.45.
[0165] like Figure 4 and Figure 5 As shown, an example of an end-to-end neural network reconstruction module 60 is shown, wherein the total number of the end-to-end neural networks is 3; Figure 4 Its overall training strategy is described exemplarily, Figure 5 Its overall reasoning architecture is described exemplarily.
[0166] like Figure 4 and Figure 5 In the example shown, the end-to-end neural network reconstruction module 60 is composed of three end-to-end neural networks, namely, and The three branches together constitute a multimodal feature fusion network. For example, EVG representation is selected as a mode from Class I mode, and Alpha image sequence representation is selected as a mode from Class II mode.
[0167] Let t n is any time point of the video frame to be reconstructed. and They represent a pattern of the type I pattern (also referred to as event frame E) and a pattern of the type II pattern (also referred to as image frame A) corresponding to the time point; t k is the timestamp of the event stream; t n With t k There is one-to-one correspondence, partial correspondence, or no correspondence at all.
[0168] For example, any time point t n The length of the corresponding time window is fixed to ΔT, and the corresponding time window is expressed as [t n -ΔT,t n ]; the event stream and its index in the time window are recorded as N e is the number of events in the time window, and each event is represented by four coordinates (x i ,y i ,t i ,p i ) represents, where t i ∈[t n -ΔT,t n ],(x i ,y i ) is the spatial coordinate of the basic detection unit, t i is the timestamp of the event indexed in the time window, p i ∈[-1,+1] represents the event polarity.
[0169] It is understandable that at any time point t n The corresponding time window length ΔT is not necessarily a fixed value; corresponding to the time point t of the frame to be reconstructed n The current time window can also be [t n -ΔT,t n ], or [t n ,t n +ΔT], or a time window of length ΔT that includes this time point.
[0170] The end-to-end neural network reconstruction module 60 adopts an end-to-end training strategy; each branch is trained in parallel, and each branch has an independent optimizer and loss function; during the backpropagation process, each optimizer updates the parameters of its respective branch.
[0171] As an example, Figure 4As shown, each branch has an independent optimizer and loss function. For example, the Adaptive Moment Estimation (ADAM) optimizer. The training losses of the three branches are set to Loss1, Loss2, and Loss3, respectively, that is, the deviation between the predicted light intensity image I output by each branch and the real image Y; the training loss function of the entire end-to-end neural network reconstruction module is obtained by combining the losses of the three branches, and the loss of each branch has its own weight; for example, if the weights of Loss1, Loss2, and Loss3 are designed to be the same, the arithmetic mean of the three is taken. Since Class II patterns, such as Alpha image sequence representations, can provide long-range physical memory, that is, context association chains along the time dimension, time consistency loss is not required in the loss function design, eliminating the accumulation of errors and training difficulties. Therefore, illustratively, image reconstruction losses are used for all three branches, including but not limited to absolute error loss, SSIM loss, and LPIPS loss. The loss function of each branch is mathematically expressed as:
[0172] Loss i =|I i -Y|1+SSIM(I i ,Y)+LPIPS(I i ,Y),i=1,2,3 (s8)
[0173] Among them, SSIM represents structural similarity, LPIPS represents the learned perceptual image patch similarity, and Y is the real image.
[0174] It is understandable that the composition of the loss function of each branch can be different, and the weight of the loss of each branch in the total loss can also be different.
[0175] Preferably, and The three branches share the same architecture.
[0176] For example, the end-to-end neural network of each branch includes a head convolution layer Conv1, a recursive residual group RRG, and an image prediction layer Conv2. Among them, RRG integrates multi-scale residual blocks to maintain spatially accurate high-resolution representation and provide rich contextual features. and Before the head convolution layer of each branch, there is a functional layer Concat, which is used to splice multiple inputs on the channel. and The head convolution layer Conv1 and the recursive residual group RRG in the two-branch end-to-end neural network are respectively and Indicates that its outputs are the extracted feature tensors ω1 and ω2 respectively; and Three branches, whose outputs are predicted image sequences; in general, and The two branches of the end-to-end neural network each have two output layers: the feature tensor output layer (corresponding to and Output of ), predicted image output layer (corresponding to and The output of each image prediction layer Conv2); The branch has an output layer, which is the predicted image output layer.
[0177] It is understandable that the end-to-end neural network architecture of each branch can also choose partially or completely different architectures, or the corresponding components of the architecture can be all the same, partially or completely different, as well as non-CNN type, non-RRG type architecture.
[0178] For example, one mode in the type I mode, EVG representation, and one mode in the type II mode (such as Alpha image sequence representation) are selected to form a mixed type combination as Branch input, extract the complementary feature tensor ω1;
[0179] Select one of the Class II patterns (e.g. Alpha image sequence representation) to form a Class II single type combination as The input of the branch extracts a stable feature tensor ω2; its stability can reduce the temporal consistency error in video generation.
[0180] Among them, the EVG characterization mode provides high-frequency details, while the Alpha image sequence characterization mode provides stable low-frequency information.
[0181] Select one of the II-type patterns (e.g., Alpha image sequence representation) to form a II-type single type combination, and use it together with the above-mentioned bimodal characteristic tensors ω1 and ω2 as The input of the branches is fused to generate the final feature fusion reconstructed light intensity image sequence, that is, the reconstructed video frame sequence.
[0182] It is understandable that A single type combination of type II in the input of the branch and A Class II single type combination of the inputs to the branches is the same or different.
[0183] As an example, Figure 5 As shown, when reasoning, the branch and The image prediction layer in is removed, that is, it becomes a branch and Its output is a feature tensor. During the inference process, the event stream is first converted into a type I pattern (such as EVG representation) and a type II pattern (such as Alpha image sequence representation); then, a hybrid combination of the two is input to the optimized branch. Output extracted feature tensor ω1; a type II single type combination formed by a type II pattern (such as Alpha image sequence representation), input optimized branch Output the extracted feature tensor ω2; ω1 and ω2 form a bimodal feature tensor, and a type II single type combination formed by a type II mode (such as Alpha image sequence representation), and jointly input the optimized The branch outputs the final feature fusion reconstructed light intensity image sequence, that is, the reconstructed video frame sequence.
[0184] It is understandable that A single type combination of type II in the input of the branch and A Class II single type combination of the inputs to the branches is the same or different.
[0185] In another optional embodiment, in the implementation of the Class II mode, the attenuation response function adopts a non-exponential attenuation function, or a combination of an exponential attenuation function and a non-exponential attenuation function.
[0186] For example, a fractional-order Mittag-Leffler function is used as the decay response function. By introducing fractional-order parameters, the decay process from exponential decay (short time, <τ) to power-law decay (long time, >τ) is unified within a mathematical framework, naturally embedding memory. The decay response function can be either with or without a harvesting constraint, or with or without a total quantity constraint. τ is the characteristic time constant of the decay. In practical applications, τ can be set to the frame interval of the video to be reconstructed, or approximately thereabouts.
[0187] Exemplarily, a solution of a fractional-order differential equation is used as the attenuation response function.
[0188] The process of reconstructing video frames from event streams can be modeled using formula (s7);
[0189]
[0190] The equation has both continuous variables and discrete random variables p k .
[0191] It is understandable that there are many specific expressions of fractional-order differential equations and forms of their solutions. This is only illustrated here as an example, rather than limiting its specific form.
[0192] Let's first consider the simplest fractional differential equation as an example:
[0193]
[0194] Where λ = 1 / τ α Represents the attenuation factor; α here represents the fractional differential, which is different from the exponent α in formula (s3).
[0195] Fractional differential equations can be solved using analytical and numerical methods. Common analytical methods include the Mittag-Leffler function method and the Laplace transform method. Here, the Mittag-Leffler function method is used as an example.
[0196] Formulas (s7) and (s8) are both fractional-order linear differential equations, whose general solutions are usually expressed as linear combinations of Mittag-Leffler functions:
[0197] H(t)=E α,1 (-λt α )+Bt α-1 E α,α (-λt α ) (s9)
[0198] Among them, E α,β (·) is the two-parameter Mittag-Leffler function; B is a parameter obtained by the boundary conditions.
[0199] The general solution in formula (s9) can be simplified to the following form,
[0200]
[0201] When t<<τ, let α=1, then formula (s10-1) is exponential decay.
[0202] Formula (s10-1) can also be expressed in non-exponential form:
[0203]
[0204] The decay response characteristics shown in the above formula are: initial rapid decay, exponential decay when α = 1, slower than exponential decay when α < 1, and power-law decay over long periods of time. If there is a harvest during the decay process, refer to formula (s3-2).
[0205] On the other hand, this embodiment also proposes a database construction method for training the end-to-end neural network reconstruction module 60. Figure 6 The diagram shown is a framework diagram of the method for building a database.
[0206] An active pixel sensor 120 is introduced into the event-based imaging system; the active pixel sensor 120 is characterized in that it shares a basic detection unit with the dynamic vision sensor 10, outputs image data in a continuous frame mode within a predetermined time interval, and all basic detection units receive light signals with the same exposure time within the same time period, so that the light intensity information of each frame of the image is nearly consistent.
[0207] In a real space scene, the event-based imaging system is used to acquire scene light information, the active pixel sensor 120 outputs continuous frame images, and the dynamic vision sensor 10 outputs an event stream.
[0208] The event stream and continuous frame images of each scene acquired by the event-based imaging system constitute a database, wherein the continuous frame images serve as the ground truth (GT) of the reconstructed image sequence of their corresponding event stream.
[0209] In order to improve the quality of the dataset, during the data collection process, the scene should try to avoid the loss of label true value information such as overexposure and motion blur of the active pixel sensor 120; overexposure detection and motion blur filtering are used to implement strict quality control on the collected data; the scene should be diverse, including indoor and outdoor environments, different camera motion states (such as acceleration, pause), independent object motion and different texture objects, etc.; in the collected dataset, the time and spatial density characteristics of the event distribution cover the corresponding characteristics of the scene to be inferred.
[0210] When the dynamic contrast range and / or motion speed of the scene exceeds the dynamic contrast range and temporal resolution of the active pixel sensor 120 , the frame image captured by the high-speed camera and the high dynamic contrast camera is corrected and used as the true value (GT).
[0211] Example 2
[0212] This embodiment makes improvements based on the method of reconstructing video frame sequence based on event stream proposed in embodiment 1. Figure 7 As shown in FIG, another architecture diagram of the method for reconstructing a video frame sequence based on an event stream is shown. Figure 1 Compared with the architecture diagram in FIG, this embodiment introduces a controllable blinking module 20.
[0213] Among them, through the controllable blinking module 20 and the relay device 30, the dynamic vision sensor 10 obtains scene light information and generates an event stream, and then reconstructs the video of the event stream through the event representation module 50 with the dynamic physical memory and the end-to-end neural network reconstruction module 60.
[0214] like Figure 8As shown, the controllable blink module 20 includes a photomask 210, an occlusion rate control unit 220 for adjusting the occlusion rate of the photomask 210, and / or an occlusion area control unit 230 for adjusting the occlusion area of the photomask 210, and a general control unit 240. The photomask 210 is located in the propagation path of the incident light from the dynamic vision sensor 10. The general control unit 240 issues a command, and the occlusion area control unit 230 or the occlusion rate control unit 220 drives the photomask 210 to perform a "close-open" operation. During this "close-open" operation, the degree of occlusion of the incident light from the dynamic vision sensor 10 by the photomask 210 changes from small to large to maintained to small to maintained.
[0215] In this embodiment and subsequent embodiments, the shielding rate of the photomask 210 to the incident light itself is always less than or equal to 100%, and the following description will not be repeated.
[0216] For example, for a photomask 210 with a 60% shading ratio, a light beam entering the photomask 210 will only have (1-60%) = 40% of the incident light intensity remaining when it exits the photomask 210. The shading ratio of the photomask 210 can be changed by the shading ratio control unit 220, or the shading area of the photomask 210 can be changed by the shading area control unit 230. The overall control unit 240 is signal-connected to the shading ratio control unit 220, the shading area control unit 230, and the dynamic vision sensor 10, and controls them.
[0217] like Figure 8 As shown, the occlusion rate control unit 220, the occlusion area control unit 230, and the general control unit 240 are shown as three separate devices; obviously, they can also be combined into a control module to change the occlusion rate of the photomask 210, change the occlusion area of the photomask 210, or drive the dynamic vision sensor 10 to work, or electrically change the size of the light-transmitting aperture of the aperture stop 310.
[0218] The relay device 30 is used to modulate and guide external light to enter the dynamic vision sensor 10 .
[0219] The relay device 30 can be various optical devices, such as a diffractive optical device, a group of diffractive optical devices, a prism, a reflector, a polarizer, a color filter, or various combinations of the above devices, to modulate and guide light incident from the scene to the dynamic vision sensor 10. For example, the diffractive device or the group of diffractive devices can guide spectral information from the observed object to the dynamic vision sensor 10, or guide light of different frequencies to be incident on corresponding areas of the dynamic vision sensor 10; for example, the prism and the reflector can modulate the geometric structure of the imaging system by deflecting the optical path; for example, the polarizer can select corresponding deflection characteristic light from the observed object to be incident on the dynamic vision sensor 10; for example, the color filter can select corresponding color light from the observed object to be incident on the dynamic vision sensor 10. Generally, the relay device 30 can be various optical devices, such as a lens or lens group with imaging function, or a diffractive device or free-form surface group with imaging function, which are referred to as imaging relay devices, and images the observed object onto the basic detection unit surface of the dynamic vision sensor 10. This imaging relay device can also image the photomask 210, either onto the basic detection unit surface of the dynamic vision sensor 10 or along the direction of incident light transmission to a location away from the basic detection unit surface of the dynamic vision sensor 10, such as in front of or behind the basic detection unit surface. For example, a combination of multiple lenses can simultaneously image the observed object and the photomask 210.
[0220] Figure 7 and 8 In the example, the relay device 30 is positioned between the photomask 210 and the dynamic vision sensor 10. In practice, both the photomask 210 and the relay device 30 are located in the incident path of the scene light on the dynamic vision sensor 10, and their positional relationship can be adjusted as needed. Even if the relay device 30 includes multiple optical components, the photomask 210 can be positioned between the optical components of the relay device 30. Alternatively, the photomask 210 can be attached to the basic detection unit of the dynamic vision sensor 10.
[0221] The relay device 30 may also be removed, for example, so that the scene light is directly incident on the dynamic vision sensor 10 .
[0222] The relay device 30 may also include an aperture stop 310, whose size is adaptively adjusted based on the scene light intensity, enabling the entire system to operate in scenes with a dynamic contrast range greater than the intrinsic dynamic contrast range of the dynamic vision sensor 10. The aperture stop 310 may also be a separate component introduced as a component of the relay device 30, with its clear aperture size adjusted manually or by the main control unit 240.
[0223] The system may also include an attenuation sheet 40 with an adjustable attenuation coefficient. The attenuation sheet 40 is placed on the incident light transmission path of the dynamic vision sensor 10, and the position between the attenuation sheet 40 and the photomask 210 and the relay device 30 can be adjusted as needed, not just as Figure 8 It is shown positioned between the relay device 30 and the dynamic vision sensor 10 .
[0224] In the controllable blink module 20, the main control unit 240 issues a command, causing the occlusion area control unit 230 or the occlusion rate control unit 220 to drive the photomask 210 to perform a "close-open" operation. During this "close-open" operation, the degree of occlusion of the incident light from the dynamic vision sensor 10 by the photomask 210 changes from small to large to maintained to small to maintained, i.e., it includes a closing process, a closing-maintaining process, an opening process, and an opening-maintaining process. The closing process and the closing-maintaining process are referred to as the closing operation, while the opening process and the opening-maintaining process are referred to as the opening operation.
[0225] When the dynamic vision sensor 10 receives a change in light intensity exceeding a preset threshold, an event is triggered. This triggered event refers to the coordinate information of the triggered basic detection unit, and / or time information, and / or event polarity information. The preset threshold is a preset absolute value of the light intensity change, a preset absolute value of the logarithmic change of light intensity, or the absolute value of a function of light intensity that enhances the sensitivity and accuracy of detecting light intensity changes and improves the dynamic contrast range of the perceived scene.
[0226] Common closing processes in the "close-open" operation include a change in the occlusion area of the photomask 210 from small to large, or a change in the occlusion rate from small to large. Common opening processes include a change in the occlusion area of the photomask 210 from large to small, or a change in the occlusion rate from large to small. At this time, events triggered by the dynamic vision sensor 10 include: events triggered by changes in the occlusion area of the photomask 210 during the closing and opening processes of the "close-open" operation; events triggered by changes in the occlusion rate of the photomask 210 during the closing and opening processes of the "close-open" operation; events triggered by the movement of objects in the scene; events triggered by the movement of the dynamic vision sensor 10 itself; events triggered by changes in scene illumination; events triggered by changes in the scene, etc. Ego-motion of the dynamic vision sensor 10 includes local movement of the dynamic vision sensor 10 to produce relative motion with the scene, thereby triggering an event, known as "micro-eye movement"; and overall movement of the dynamic vision sensor 10, such as when it is placed on a moving vehicle (e.g., a car, drone, robot, etc.).
[0227] according to Figure 7 and Figure 8The flowchart of the method for reconstructing a video frame sequence based on an event stream is shown. For example, the dynamic vision sensor 10 is fixed to a vehicle (such as a car, a drone, a robot, etc.) for explanation: through the main control unit 240, the controllable blinking module 20 is commanded to perform a "close-open" operation, and the light mask 210 is driven to implement a cyclic change of the degree of blocking of the incident light of the dynamic vision sensor 10 from small to large to maintain to small to maintain.
[0228] At the initial moment, before the vehicle starts to move, perform ≥1 "close-open" operations to obtain the light information of the scene, including static background information and motion light field information in the scene; after the vehicle starts to move, the photomask 210 is preferably in the open and maintained process. Based on the event triggering principle of the dynamic vision sensor 10, the scene light information is obtained due to the relative motion between the scene and the dynamic vision sensor 10. However, considering that there may be moving objects in the scene that are relatively still or nearly relatively still with the dynamic vision sensor 10, they will not trigger events or only trigger very few events. Therefore, ≥1 "close-open" operations should be appropriately performed to obtain the dynamic and static light information of the scene; when the vehicle temporarily stops moving (such as when a car encounters a red light), the photomask 210 is preferably in the open and maintained process. Based on the event triggering principle of the dynamic vision sensor 10, the scene light information is obtained due to the relative motion between the scene and the dynamic vision sensor 10. However, considering that there may be moving objects in the scene that are relatively still or nearly still with the dynamic vision sensor 10, they will not trigger events or only trigger very few events. Therefore, ≥1 "close-open" operations should be appropriately performed to obtain the dynamic and static light information of the scene. The vehicle starts moving again (such as when the vehicle is parked with a light on, a drone is hovering, or a robot is stationary), and performs ≥1 "close-open" operations to obtain the dynamic and static light information of the scene, including static background information and motion light field information in the scene; after the vehicle starts moving again, the photomask 210 is preferably in the open and maintained process. Based on the event triggering principle of the dynamic vision sensor 10, the scene light information is obtained due to the relative motion between the scene and the dynamic vision sensor 10. Similarly, considering that there may be moving objects in the scene that are relatively still or nearly relatively still with the dynamic vision sensor 10, they will not trigger events or only trigger very few events. Therefore, ≥1 "close-open" operations should be appropriately performed to obtain the dynamic and static light information of the scene.
[0229] Furthermore, the dynamic vision sensor 10, mounted on the carrier, coordinates with the relay device 30, the photomask 210, and the attenuation sheet 40 to perform localized movements, known as "micro-eye movements," to expand the field of view. These micro-eye movements and the "close-open" operation can be performed sequentially or simultaneously.
[0230] Taking the change in the occlusion region of the photomask 210 as an example, a specific "close-open" operation is described, and this operation is named an occlusion region-type "close-open" operation. The closing operation in an occlusion region-type "close-open" operation is called an occlusion region-type closing operation, the closing process is called an occlusion region-type closing process, the opening operation is called an occlusion region-type opening operation, and the opening process is called an occlusion region-type opening process. Specifically, under the control of the occlusion region control unit 230, any occlusion region-type "close-open" operation includes a closing process in which the occlusion region of the photomask 210 changes from the minimum occlusion region value in the current closing operation to the maximum occlusion region value in the current closing operation, followed by a closing and holding process in which the occlusion region remains at the maximum occlusion region value in the current closing operation; and an opening process in which the occlusion region of the photomask 210 changes from the maximum occlusion region value in the current opening operation to the minimum occlusion region value in the current closing operation, followed by an opening and holding process in which the occlusion region remains at the minimum occlusion region value in the current closing operation after the opening process is completed.
[0231] The maximum occlusion area of the photomask 210 includes the maximum occlusion area value when the area is partially occluded or the occlusion area value when the area is fully occluded. The minimum occlusion area of the photomask 210 includes the minimum occlusion area value when the area is partially occluded or the occlusion area value when the area is completely unoccluded, i.e., the occlusion area value is zero. For example, the maximum occlusion area of the photomask 210 is set to a full occlusion state, i.e., the coverage of the photomask 210 over the entire aperture area is 100%, while the minimum occlusion area is set to a no occlusion state, i.e., the coverage of the photomask 210 over the entire aperture area is zero. In practice, the maximum occlusion area of the photomask 210 can be occlusion with a coverage rate other than 100%, and the minimum occlusion area can be partial occlusion with a coverage rate other than zero. Furthermore, the occlusion state at the start of the closing process of a "close-open" operation can be different from the occlusion state at the end of the opening process. Of course, the occlusion state at the start of the closing process of a "close-open" operation can also be the same as the occlusion state at the end of the opening process.
[0232] For example, the blocking area of the photomask 210 changes along one direction, which is a common rolling shutter photomask 210. This can be implemented by physical devices such as a mechanically movable baffle, an electrically controlled liquid crystal light valve, and a flexible, retractable membrane. The rolling shutter-like closing and opening along one direction can make the total number of events triggered by the blocking area change during the closing and opening processes increase nearly linearly. The blocking area of the photomask 210 can also be changed by closing and opening in a radially enclosing manner, changing in a circular area, closing and opening in a point-by-point or line-by-line scanning manner, or synchronously closing and opening the entire maximum blocking area, or various combinations of the above. In fact, the blocking area of the photomask 210 can be changed in various possible ways, even using a combination of different changing methods. This application does not limit the way in which the blocking area of the photomask 210 changes.
[0233] In fact, it is preferred that the occlusion rate control unit 220 sets the occlusion rate of the photomask 210 so that the number of events triggered by the dynamic vision sensor 10 during the corresponding closing process and opening process changes linearly with time.
[0234] The photomask 210 may include only one mask layer or may be composed of more than one mask layer. For example, each of the more than one mask layers may have different occlusion characteristics, or different color filtering characteristics, or different area characteristics, or different polarization characteristics, or various combinations of these characteristics. The overlapping of these more than one mask layers can give the photomask 210 more flexible occlusion characteristics. For example, a photomask 210 composed of two mask layers with different occlusion ratios and different areas may have its occlusion area divided into two sub-areas with different occlusion ratios. A photomask 210 composed of two mask layers with different color filtering characteristics may have its occlusion area reconstruct image information of different color components of the target scene separately. A photomask 210 composed of two mask layers with different polarization characteristics, spliced together, can reconstruct images of the target scene with different polarizations.
[0235] During a block-area "close-open" operation, the block area of the photomask 210 changes between the closing and opening stages. Furthermore, the blockage ratio of the photomask 210 can be set to different values during the different stages of the "close-open" operation. These different blockage ratios can be implemented by the blockage ratio control unit 220, which is driven by the overall control unit 240.
[0236] In the above description, the starting time of a "close-open" operation is determined by the starting time of its closing process. In reality, the "close-open" operation is repeated over and over again. Although various parameters, including the total time duration of the "close-open" operation, the closing process duration, the closing and holding process duration, the opening process duration, the opening and holding process duration, and the closing method of the photomask 210, may vary, different "close-open" operations can be considered to occur cyclically. In this case, the starting time of any "close-open" operation can be any other time point during the closing process, the closing and holding process, the opening process, or the opening and holding process, rather than necessarily the starting time of the closing process. Furthermore, an additional system startup process may be required when the system is turned on.
[0237] The "close-open" operation can also be an operation in which the occlusion rate of the photomask 210 changes cyclically, which is named an occlusion rate type "close-open" operation, and its closing operation is called an occlusion rate type closing operation, and its closing process is called an occlusion rate type closing process, and its opening operation is called an occlusion rate type opening operation, and its opening process is called an occlusion rate type opening process.
[0238] In practice, an occlusion rate-based "close-open" operation can be inserted into an occlusion area-based "close-open" operation. For example, during the occlusion area-based closing operation, the occlusion rate of the photomask 210 is P1. After the closing operation is completed, during the closed-holding process, the main control unit 240 issues a command, and the occlusion rate control unit 220 drives the photomask 210 to perform ≥1 occlusion rate-based "close-open" operations. The photomask 210 implements a cyclic change process of decreasing → increasing → holding → decreasing → holding to the degree of occlusion of the incident light from the dynamic vision sensor 10. The occlusion rate P2 at the end of the operation serves as the occlusion rate of the photomask 210 during the occlusion area-based opening operation.
[0239] The basic detection units of the dynamic vision sensor 10 can also be divided into G>1 basic detection unit blocks arranged on a flat or curved surface. No basic detection units are shared among the G basic detection unit blocks at any given time. The basic detection unit resolution, basic detection unit density, array arrangement pattern, physical properties, and synchronization of different basic detection unit blocks can be the same, partially identical, or completely different.
[0240] The physical properties refer to color properties of different wavelengths or polarization properties of different polarization states. Each basic detection unit block with different physical properties only receives light with the corresponding physical properties.
[0241] For the convenience of description, the system composed of the dynamic vision sensor 10, the relay device 30 and the controllable blink module 20 that receives light information from the outside and outputs an event stream is called an event-based imaging system.
[0242] In specific implementation, embodiment 1 and this embodiment may also be composed of two event-based imaging systems arranged in a plane or a curve to form a binocular system.
[0243] In a specific implementation, more than two event-based imaging systems such as those proposed in Example 1 and this embodiment are arranged in a planar or curved manner to form a multi-eye system.
[0244] It can be understood that the system of this embodiment corresponds to the event-based imaging system of the above-mentioned embodiment 1, and the optional items in the above-mentioned embodiment 1 are also applicable to this embodiment, so they will not be described again here.
[0245] like Figure 7 and Figure 8 As shown, the method for reconstructing a video frame sequence based on an event stream of this embodiment introduces a controllable blink module (20), and the dynamic vision sensor (10) receives light information from the outside world via the controllable blink module (20) and the relay device (30);
[0246] The controllable blink module (20) comprises a photomask (210), a blocking rate control unit (220) for controlling the blocking rate of the photomask, a blocking area control unit (230) for controlling the blocking area of the photomask, and a general control unit (240); in the controllable blink module (20), the general control unit (240) issues a command once, and the blocking area control unit (230) or the blocking rate control unit (220) drives the photomask (210) to perform a "close-open" operation, that is, the photomask (210) implements a change process of small→large→maintain→small→maintain in the blocking degree of the incident light of the dynamic vision sensor (10);
[0247] Furthermore, the shielding rate of the light mask (210) in the controllable blinking module (20) to its own incident light is equal to or less than 100%.
[0248] Then, the method further comprises the following steps:
[0249] Step z1. Determine the initial time window of the “close-open” operation [T s ,T e ]; where T s is the starting time point of the closing process in the “close-open” operation, T e The end time point of the opening and holding process in the "close-open" operation;
[0250] Step z2 sets the occlusion area and / or occlusion rate related parameters, and the general control unit 240 issues a command to implement a corresponding "close - open" operation;
[0251] Step z3. Determine the opening time window of a "close-open" operation [T Os ,T Oe ], that is, the time interval corresponding to the opening process;
[0252] The time window [T Os ,T Oe The event triggered by the change of the shielding area or shielding rate of the photomask 210 in the ] is characterized as the time point T by the following process. Oe The corresponding static background light intensity image:
[0253] Formulas (5), (6) and (7) are used to perform linear or nearly linear event integration algorithm for each basic detection unit to obtain the value of each basic detection unit at time point T Oe The absolute light intensity value of all basic detection units at time point T Oe The absolute light intensity value at time point T Oe The corresponding static background light intensity image; its expression is:
[0254]
[0255] H(t i -t i-1 )=exp[-β·(t i -t i-1 )] (7)
[0256] Where I represents the time point T Oe The intensity image, i.e. the logarithmic value, ε{e i} means that when opening the time window [T Os ,T Oe ] is a set of events triggered by changes in the occlusion area or occlusion rate of the photomask; Formula (5) represents the event set for opening the time window [T Os ,T Oe ] is integrated by the basic detection unit event integration algorithm to obtain the time point T of each basic detection unit. Oe The logarithm of the absolute light intensity value.
[0257] Formula (6) is the update rule of the linear or nearly linear event integration algorithm for each basic detection unit; I(x, y, t i ) represents the coordinates (x, y) of the basic detection unit of the dynamic vision sensor at the time stamp t i The logarithm of the absolute light intensity value, t i represents ε{ei}The i-th event e i Timestamp of p i Represents the i-th event e i The event polarity; ΔC represents the preset constant threshold for event triggering.
[0258] Formula (7) is the expression of the response function H(t) of the basic detection unit event integration algorithm; β is the attenuation factor. By adjusting the value of β, Formula (5) becomes a linear or nearly linear event integration function.
[0259] Step z4. Set the time window of a “close-open” operation [T Oe ,T e The event stream output by the dynamic vision sensor 10 in ] is processed by steps 1 to 5 of the method for reconstructing a video frame sequence based on an event stream;
[0260] Step z5. Use T e +δt is the starting time point T of the closing process in the next "close-open" operation s , update the time window for the next "close-open" operation; repeat steps z1 to z4 until the task of reconstructing the video frame sequence based on the event stream is completed, and the light intensity image corresponding to each time point is obtained; where δt≥0 is the interruption duration.
[0261] Specifically, in the time window of a “close-open” operation [T Oe ,T e ], when the method for reconstructing a video frame sequence based on an event stream is executed and the representation of the Class II pattern is obtained by the event representation module 50 with dynamic physical memory, step z3 is performed at time point T Oe The intensity image (logarithmic value) I is used as the initial value of the integration from this time point. It can be understood that the set time point t n With the timestamp t of the event k Not necessarily corresponding, the integration from this time point refers to the timestamp near this time point.
[0262] It can be understood that the options in the above embodiment 1 are also applicable to this embodiment, so they will not be described again here.
[0263] Example 3
[0264] In the method for reconstructing a video frame sequence based on an event stream described in Examples 1 and 2, the length of the time window ΔT n It can be set according to needs and can remain fixed during the entire reconstruction process, or it can be non-fixed or adaptively change the length.
[0265] Exemplarily, the end-to-end neural network reconstruction module 60 is trained on the database described in Example 1 with a fixed time window of 33 ms, and its generalization ability is evaluated at extreme time scales (e.g., 0.2-60 ms) using a high-speed driving dataset.
[0266] It is understandable that training can also be performed in non-fixed time windows.
[0267] When there is no ground-truth (GT) label, evaluation metrics without reference images can be used to evaluate the intensity image or intensity image sequence reconstructed from the event stream.
[0268] It is understandable that the evaluation index of the no-reference image can also be used as a component of the loss function for training the end-to-end neural network reconstruction module 60.
[0269] Example 4
[0270] This embodiment applies similar Figure 3 In the process, if the dynamic global "memory-forget" mask The corresponding branch ② adopts the mathematical expression (s11), and the final generated type II pattern is called the Gamma image sequence representation pattern.
[0271]
[0272] Where γ is the mask attenuation factor; if the basic pixel unit (x, y) is at time point t n If no event occurs in the corresponding time window, Δt is the time point t n The timestamp of the last event before time t n The right endpoint of the corresponding time window; if the time point t n There is an event at the basic detection unit in the corresponding time window and at time point t n-1 There is no event in the corresponding time window, then Δt is the time point t n The timestamp of the last event before time t n The right endpoint of the corresponding time window; if the time point t n Corresponding time window and time point t n-1 If there are events in the corresponding time window, then let or a value close to 1.
[0273] For example, Figure 9 As shown, when calculating the time point t n The corresponding light intensity image is the corresponding time window When calculating time point t n+1 The corresponding light intensity image, the corresponding time window In the expression of Δt=tn+1 -t k+2 =t n -t k+2 +ΔT n+1 .
[0274] It is understandable that Δt can also take an approximate value, for example, In the expression, Δt≈ΔT n+1 .
[0275] Furthermore, the Gamma image sequence representation mode in the type II mode can be represented by introducing a virtual p=0 polarity, corresponding to the event stream timestamp t k Overlapping set time point t n ; Then, formulas (1), (2), and (3) can be combined into one mathematical expression:
[0276]
[0277] in, Refers to any set time point t n No event is triggered at ; in formula (s12), if t k-1 and t k The two timestamps do not span at least two set time points t n , then follow the upper part of formula (s12); if t k-1 and t k The two timestamps span at least two set time points t n , then follow the lower half of formula (s12).
[0278] Furthermore, if the decay function H(t k -t k-1 ) takes an exponential decay function, then formula (s12) can be written as:
[0279]
[0280] Formula (s13) is equivalent to masking the dynamic global "memory-forget" Unified into the event leakage-integration model framework; where γ = λ or γ ≠ λ.
[0281] It is understandable that a similar method can also be used to mask the dynamic global "memory-forget" Unified to any decay function H(t k -t k-1 ) framework.
[0282] The above are only preferred embodiments of the present invention, but the design concept of the present invention is not limited to this. Any non-substantial modifications made to the present invention using this concept also fall within the scope of protection of the present invention. For example, the dynamic global "memory-forgetting" mask corresponding to each basic detection unit can take various attenuation coefficients or attenuation function expressions with values not greater than 1. For another example, the dynamic visual sensor described in this document takes the light intensity changes perceived by its basic detection units as an example; the basic detection units can also independently perceive and receive other light information changes, and when the perceived changes in other light information exceed their preset thresholds, trigger events and output event streams, such as simultaneously perceiving light intensity changes of light information of different wavelengths or frequencies. For another example, a Class II pattern representation of an image sequence with long-term physical memory, which is a Class II pattern representation expressed jointly by Formula (1) and Formula (2), can also be constructed by M tn (x, y) and / or H(t) are unified in a mathematical framework; or a mathematical model for accumulating long-range optical information is constructed to unify them in a mathematical framework. For another example, the dynamic vision sensor (DVS) described in this application can be a temporal contrast automatic pixel detection sensor that independently detects brightness changes at each pixel in the time domain and outputs event stream data containing pixel coordinates, timestamps, and the polarity of the brightness change (positive for brightness increases and negative for brightness changes). It can also be an asynchronous time-based image sensor (ATIS), which has the same underlying physical mechanism. The ATIS combines a DVS change detector and a conditional exposure measurement circuit. After detecting a certain brightness change in the field of view pixel by pixel, the change detector independently and asynchronously initiates a new exposure / grayscale value measurement and outputs a four-dimensional data stream of pixel coordinates, timestamps, and relative brightness values. Alternatively, it can be a multi-mode event camera, such as a dynamic and active-pixel vision sensor (DAVIS), which uses a hybrid architecture to implement the fusion output data of the event stream and the global shutter frame intensity image or optical flow vector on a single chip. Accordingly, all relevant embodiments fall within the scope of protection of the present invention. It will be apparent to those skilled in the art that other variations or modifications may be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.
Claims
1. A method for reconstructing a video frame sequence based on an event stream, characterized in that: include: Step 1. The dynamic vision sensor (10) receives light information from the outside world through the relay device (30). The dynamic vision sensor (10) is composed of an arrangement of basic detection units for receiving light information, and each basic detection unit independently senses changes in received light intensity, and triggers an event and outputs an event stream when the perceived light intensity change exceeds a preset threshold. Furthermore, the relay device (30) modulates and guides external light information to be incident on the dynamic vision sensor (10); Step 2. An event stream representation module (50) with dynamic physical memory generates two types of pattern representations of the event stream output by the dynamic vision sensor (10): a type I pattern representing the spatiotemporal information of event data with short-range physical memory, and a type II pattern representing the image sequence with long-range physical memory; Among them, short-range physical memory refers to any time point along the time dimension of the event flow, including the spatiotemporal information of event data within a time window at that time point. This time window is named the short-range physical memory time window corresponding to that time point; Moreover, long-range physical memory refers to any time point along the time dimension of the event flow. Within a time window containing this time point, information is accumulated by each basic detection unit as the incremental information. The accumulated information is named the long-range physical memory time window corresponding to this time point. The Class I mode has AI ≥ 1 modes, including: modes 510-A1, 510-A2, ..., 510-AI; Moreover, the Class II mode has BI≥1 modes, including: modes 520-B1, 520-B2, ..., 520-BI; Step 3. Select at least one of the AI modes of the Class I mode and at least one of the BI modes of the Class II mode to form MX ≥ 1 mixed type combinations; Select SI≥1 combinations from AI patterns of Class I patterns to form SI single-type combinations of Class I; or / and, select SII≥1 combinations from B1 patterns of Class II patterns to form SII single-type combinations of Class II; Set each combination to correspond to an end-to-end neural network to independently extract its own feature tensor; Step 4. Select SI'≥0 combinations from the AI patterns of the Class I mode, and select SII'≥0 combinations from the BI patterns of the Class II mode. Input them together with the feature tensor extracted in Step 3 into an end-to-end neural network to generate the final feature fusion reconstructed light intensity image at a time point, that is, reconstruct the video frame; Step 5: Repeat steps 1 to 4 to obtain the reconstructed light intensity image corresponding to each time point until the task of reconstructing the video frame sequence based on the event stream is completed; The end-to-end neural network has no time dimension memory for the input data; In the method for reconstructing a video frame sequence based on an event stream, the end-to-end neural network and its input-output association relationship used in steps 3 and 4 constitute an end-to-end neural network reconstruction module (60).
2. The method according to claim 1, characterized in that The event stream output by the dynamic vision sensor (10) is the original event stream data output by the dynamic vision sensor (10), or the event stream after the original event stream is denoised, or the event stream after the original event stream is edited.
3. The method according to claim 1, characterized in that The type I pattern of the spatiotemporal information representation of event data with short-term memory, including but not limited to event voxel grid representation, event graph representation, and time surface, is used to encode the spatiotemporal information of each event; The type II mode of image sequence representation with long-range physical memory accumulates the light intensity or logarithm of light intensity at each time point along the time dimension of each basic detection unit for event information based on the attenuation response function; Wherein, the attenuation response function is an exponential attenuation function, a non-exponential attenuation function, or an exponential-non-exponential mixed function; Furthermore, the accumulation is integration, or summation, or a mixture of integration and summation.
4. The method according to claim 3, characterized in that The exponential decay function includes but is not limited to a single exponential decay function and a multi-exponential decay function, and the non-exponential decay function includes but is not limited to a power-law decay function, a Mittag-Leffler decay function, a fractional-order Mittag-Leffler decay function, a linear decay function, a fractal decay response function, or a combination thereof.
5. The method according to claim 1, wherein At any time point t along the time dimension of the event stream n , using formula (1) to accumulate the light information carried by the events in the corresponding time window of each basic detection unit as the incremental value, the expression is: logA′(x,y,t k )=H(t k -t k-1 )·logA′(x,y,t k-1 )+p k ΔC (1) The obtained time point t n The total amount of accumulated information is called the amount of information accumulated by each basic detection unit at time point t n The intensity image of is represented by a′(x,y); At any time point t n The corresponding time window corresponds to a basic detection unit attenuation mask Named as dynamic global "remember-forget" mask; Then, each basic detection unit at any time point t n The characterization of the type II pattern is done with Its mathematical expression is: Among them, any event e in the event stream k (x,y,t k ,p k ) is timestamp t k , polarity is p k , the coordinates are (x,y); A′(x,y,t k ) represents each basic detection unit at timestamp t k Light intensity image at time H(t k -t k-1 ) is the attenuation response function, ΔC is the preset light intensity logarithmic change threshold for the dynamic vision sensor (10) triggering event; t k-1 and t k Represents two adjacent timestamps in the event stream; Among them, t n With t k One-to-one correspondence, partial correspondence, or no correspondence at all.
6. The method according to claim 5, characterized in that The attenuation mask Using formula (3): Among them, for any basic detection unit, when it is at any time point t n When no event occurs in the corresponding time window, the parameter α is 0; when an event occurs, the value of the parameter α is dynamically adjusted according to the relative motion state of the dynamic vision sensor (10): when the dynamic vision sensor (10) is moving and / or there is relative motion with the scene, the value of the parameter α is greater than a first preset threshold value, so that the function exhibits an accelerated decay characteristic, thereby enhancing the degree of "forgetting" of the light information obtained in the current time window; when the dynamic vision sensor (10) is close to being still and / or there is no relative motion with the scene, the value of the parameter α is less than a second preset threshold value, so that the function exhibits a slow decay characteristic, thereby enhancing the degree of "memory" of the light information obtained in the current time window; The specific values of the first preset threshold and the second preset threshold are dynamically adjusted according to the signal attenuation requirements in actual application scenarios.
7. The method according to claim 6, characterized in that It also includes a dynamic time window event counting analysis strategy, which analyzes the event distribution over the first j time windows. Its mathematical formula is: If the number of events in the current time window exceeds n times the average number of events in the previous j time windows, then α = γ a , to enhance the degree of "forgetting" of the light information obtained in the current time window; if the number of events in the current time window does not exceed η times the average number of events in the previous j time windows, then α=γ b , in order to enhance the "memory" of light information obtained in the current time window; where j≥1; 0<η<1, and γ a >γ b ;s tn is the current time point t n The number of events in the corresponding time window.
8. The method according to claim 5, characterized in that The attenuation response function H(t k -t k-1 ) is an exponential decay function along the time dimension; Formula (1-1) is used to accumulate the optical information carried by the event stream in each basic detection unit: logA′(x,y,t k )=exp[-λ(t k -t k-1 )]·logA′(x,y,t k-1 )+p k ΔC (1-1) It is called the event leakage-integral model; where λ is the attenuation factor; Or, use formula (1-2) to accumulate the optical information carried by the event stream in each basic detection unit: It is called the event leakage-integral model with a harvesting parameter h; h ≥ 0 is the harvesting parameter.
9. The method according to claim 5, characterized in that The attenuation response function is a non-exponential attenuation function, or a combination of an exponential attenuation function and a non-exponential attenuation function; Alternatively, a fractional-order Mittag-Leffler type function is used as the attenuation response function, and by introducing fractional-order parameters, the attenuation process from exponential attenuation, time interval <τ, to power-law attenuation, time interval >τ, is unified under a mathematical framework to embed memory; the attenuation response function contains or does not contain a harvesting restriction, contains or does not contain a total amount restriction; wherein τ is the characteristic time constant of the attenuation.
10. The method according to claim 1, characterized in that The architecture of each end-to-end neural network constituting the end-to-end neural network reconstruction module includes but is not limited to a convolutional neural network, a Transformer, a Mamba, or a combination thereof; the architecture of each end-to-end neural network is entirely the same, partially the same, or different from one another; its input layer matches the representation output by the event representation module (50) with dynamic physical memory; and the output layer includes a feature tensor output layer, a predicted image output layer, or a combination thereof.
11. The method according to claim 1, wherein The end-to-end neural network reconstruction module (60) includes three end-to-end neural network branches and branch road Take MX=1 mixed type combination as input and extract feature tensor ω1; branch Take SII = 1 type II single type combination or SI = 1 type I single type combination as input and extract the feature tensor ω2; branch Taking the feature tensors ω1 and ω2, and SII'=1 type II single type combination or SI'=1 type I single type combination as input, the final feature fusion reconstructed intensity image sequence, that is, the reconstructed video frame sequence, is generated.
12. The method according to claim 11, characterized in that The three end-to-end neural network branches and Share the same architecture, including the head convolution layer Conv1, recursive residual group RRG and image prediction layer Conv2; and The head convolution layer Conv1 and recursive residual group RRG in each architecture are used and express; and The output of is the extracted feature tensor; and The output of is a light intensity image; and and The corresponding components of the architecture may be all the same, partially the same, or completely different.
13. The method according to claim 12, characterized in that The end-to-end neural network reconstruction module (60) adopts an end-to-end training strategy; preferably, each branch is parallel; each branch has an independent optimizer and loss function; during the back-propagation process, each optimizer updates the parameters of its respective branch.
14. The method according to claim 1, wherein A controllable blinking module (20) is introduced, and the dynamic vision sensor (10) receives light information from the outside world via the controllable blinking module (20) and the relay device (30); Wherein, the controllable blink module (20) comprises a photomask (210), a blocking rate control unit (220) for controlling the blocking rate of the photomask, a blocking area control unit (230) for controlling the blocking area of the photomask, and a general control unit (240); in the controllable blink module (20), the general control unit (240) issues a command once, and the blocking area control unit (230) or the blocking rate control unit (220) drives the photomask (210) to perform a "close-open" operation, that is, the photomask (210) implements a change process of small→large→maintain→small→maintain on the blocking degree of the incident light of the dynamic vision sensor (10); and the blocking rate of the photomask (210) on its own incident light in the controllable blink module (20) is equal to or less than 100%; Then, the method further comprises the following steps: Step z1. Determine the initial time window of the "close-open" operation [T s ,T e ]; where T s is the starting time point of the closing process in the "close-open" operation, T e The end time point of the open holding process in the "close-open" operation; Step z2. Setting the occlusion area and / or occlusion rate related parameters, and the general control unit (240) issues a command to implement a corresponding "close-open" operation; Step z3. Determine the opening time window of a "close-open" operation [T Os ,T Oe ], that is, the time interval corresponding to the opening process; The time window [T Os ,T Oe The event triggered by the change of the shielding area or shielding rate of the photomask (210) in ] is characterized as the time point T by the following process. Oe The corresponding static background light intensity image: Formulas (5), (6) and (7) are used to perform linear or nearly linear event integration algorithm for each basic detection unit to obtain the value of each basic detection unit at time point T Oe The absolute light intensity value of all basic detection units at time point T Oe The absolute light intensity value at time point T Oe The corresponding static background light intensity image; its expression is: H(t i -t i-1 )=exp[-β·(t i -t i-1 )] (7) Where I represents the time point T Oe The intensity image, i.e. the logarithmic value, ε{e i } means that when opening the time window [T Os ,T Oe ] event set triggered by the change of the shading area or shading rate of the inner photomask (210); Formula (5) represents the ... Os ,T Oe ] is integrated by the basic detection unit event integration algorithm to obtain the time point T of each basic detection unit. Oe The logarithm of the absolute light intensity value; Formula (6) is the update rule of the linear or nearly linear event integration algorithm for each basic detection unit; I(x, y, t i ) represents the coordinates (x, y) of the basic detection unit of the dynamic vision sensor (10) at the time stamp t i The logarithm of the absolute light intensity value, t i represents ε{e i }The i-th event e i Timestamp of p i Represents the i-th event e i The event polarity; ΔC represents the preset constant threshold for event triggering; Formula (7) is the expression of the response function H(t) of the basic detection unit event integration algorithm; β is the attenuation factor, and by adjusting the value of β, Formula (5) becomes a linear or nearly linear event integration function; Step z4. Set the time window of a "close-open" operation [T Oe ,T e ] the event stream output by the dynamic visual sensor (10), and processing steps 1 to 5 of the method for reconstructing a video frame sequence based on the event stream; Step z5. Use T e +δt is the starting time point T of the closing process in the next "close-open" operation s , update the time window for the next "close-open" operation; repeat steps z1 to z4 until the task of reconstructing the video frame sequence based on the event stream is completed, and the light intensity image corresponding to each time point is obtained; where δt ≥ 0 is the interruption duration.
15. The method according to claim 14, characterized in that The time window of a "close-open" operation [T Oe ,T e ], the representation of the Class II pattern in the event representation module (50) with dynamic physical memory takes the light intensity image at time point TOe in step z3 as the initial value at that time point.
16. The method according to claim 14, characterized in that The relay device (30) is a lens, or a lens group, or a diffractive optical device, or a diffractive optical device group, or a prism, or a reflector, or a polarizer, or a color filter, or a color filter array, or various combinations of the above devices; The photomask (210) is imaged on the surface where the basic detection unit of the dynamic vision sensor (10) is located via the relay device, or is imaged at a position away from the dynamic vision sensor (10) along the transmission direction of the incident light.
17. The method according to claim 14, characterized in that An aperture stop (310) is introduced between the dynamic vision sensor (10) and the outside world, and its size is adjusted according to the scene light intensity, so that the entire system can operate in a scene with a dynamic contrast range greater than the intrinsic dynamic contrast range of the dynamic vision sensor (10).
18. The method according to claim 14, characterized in that An attenuation plate (40) with an adjustable attenuation coefficient is introduced between the dynamic vision sensor (10) and the outside world to adjust the incident light flux of the dynamic vision sensor (10).
19. The method according to any one of claims 1 to 18, characterized in that In the event stream representation module (50) with dynamic physical memory, the method of generating a Class II pattern of image sequence representation with long-range physical memory based on the event stream is independently used as the method of reconstructing a video frame sequence based on the event stream; and the image sequence with long-range physical memory is independently used as the reconstructed video frame sequence based on the event stream.
20. The method according to any one of claims 1 to 18, characterized in that: The step 3 further includes: not selecting from the A1 modes of the class I mode and not selecting from the B1 modes of the class II mode, but only selecting at least one from the A1 modes of the class I mode and at least one from the B1 modes of the class II mode, to form MX ≥ 1 mixed type combinations; Each of the combinations is set to correspond to an end-to-end neural network, and each feature tensor is independently extracted.
21. A database construction method, applied to the method for reconstructing a video frame sequence based on an event stream according to any one of claims 1 to 20, for training the end-to-end neural network reconstruction module (60), characterized in that: include: Introducing an active pixel sensor (120); The active pixel sensor (120) and the dynamic vision sensor (10) share a basic detection unit array, and are used to output image data in a continuous frame mode within a preset time interval, and all basic detection units receive light signals with the same exposure time within the same time period; The dynamic vision sensor (10) and the active pixel sensor (120) receive light information from the outside world via the relay device (30) and / or the controllable blink module (20); In a real space scene, the active pixel sensor (120) outputs continuous frame images, and the dynamic vision sensor (10) outputs an event stream; The event stream and continuous frame images of each scene are acquired to form a database; wherein the continuous frame images serve as the label truth value GT of the reconstructed video frame image sequence of its corresponding event stream.
22. The method according to claim 21, characterized in that Also includes: Overexposure detection and / or motion blur filtering are used to perform quality control on the collected data; the scenes include indoor and outdoor environments, different camera motion states, independent object motion and / or objects with different textures; in the collected data set, the temporal and spatial density characteristics of the event distribution cover the corresponding characteristics of the scene to be inferred.
23. The method according to claim 21 or 22, characterized in that When the dynamic contrast range and / or motion speed of the scene exceeds the dynamic contrast range and time resolution of the active pixel sensor (120), the frame image captured by the high-speed camera and / or the high dynamic contrast camera is corrected and used as the true value GT of the label.
Citation Information
Cited By
Polarization-event-radar multi-mode cooperative sensing method based on Mama
CN121883822A