Video reconstruction method and device based on event stream, electronic equipment and storage medium
By using a convolutional neural network to process event stream data, combined with a bidirectional convolutional long short-term memory module, a multi-scale spatial enhancement module and a space-time fusion attention module, the image blur and foggy artifact problems in complex dynamic scenarios are solved, and high-quality video reconstruction is achieved.
Patent Information
- Application Number
- CN202510369166.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-03-26
AI Technical Summary
When reconstructing video from event streams, existing deep learning methods cannot effectively solve the image blur and foggy artifacts caused by nonlinear temporal information and non-uniform spatial distribution in complex dynamic scenarios.
Using the video reconstruction method based on event stream, a preset convolutional neural network model is used, including a bidirectional convolutional long short-term memory module, a multi-scale spatial enhancement module, a space-time fusion attention module and a multi-scale feature aggregation module to process event frames in the form of multi-channel tensors and generate high-quality predicted image frames.
High-quality video reconstruction in complex dynamic scenarios is realized, image blur and foggy artifact problems are solved, and the robustness and image quality of video reconstruction are improved.
Smart Images

Figure CN120495125A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, device, electronic device and storage medium for video reconstruction based on event stream. Background Art
[0002] Traditional cameras have limitations in complex dynamic scenes. For example, in high dynamic range or fast-moving scenes, traditional cameras may suffer from overexposure, underexposure, or motion blur. However, unlike traditional frame-based cameras, event cameras asynchronously record the coordinate position, time, and polarity of the change at extremely high temporal resolution when the logarithmic brightness change of the scene exceeds a predetermined threshold. Event cameras have the characteristics of high temporal resolution, high dynamic range, and low latency, and can provide richer dynamic information. These characteristics give event cameras advantages in various visual tasks, including object detection, high-speed motion estimation, semantic segmentation, and video frame interpolation. However, the event data they output is in a different format than traditional video frames, so specialized algorithms need to be developed to process this data.
[0003] In the related art, existing deep learning methods have made remarkable progress in reconstructing videos from event streams. These deep learning methods typically employ recurrent fully convolutional networks to generate intensity frames from events. To enhance generalization to real event data, researchers have used statistical analysis of existing datasets to generate synthetic training datasets. Another solution involves self-supervised learning, which trains neural networks using estimated optical flow and event-based photometric consistency assumptions without the need for real reference images.
[0004] However, due to the characteristics of event streams and their inherent assumption of photometric consistency, these methods still face limitations, such as image blur caused by nonlinear temporal information in complex dynamic scenes, and fog artifacts caused by non-uniform spatial distribution in complex dynamic scenes, which need to be urgently addressed. Summary of the Invention
[0005] The present application provides a video reconstruction method, device, electronic device and storage medium based on event stream to solve the problems of image blur and fog artifacts caused by nonlinear temporal information and non-uniform spatial distribution in complex dynamic scenes, thereby meeting the demand for video reconstruction to develop higher image quality and greater robustness.
[0006] To achieve the above objectives, the first embodiment of the present application proposes a video reconstruction method based on event stream, comprising the following steps:
[0007] Acquire event stream data of a target dynamic scene, and convert the event stream data into event frames in the form of multi-channel tensors;
[0008] Inputting an event frame in the form of a multi-channel tensor into a preset convolutional neural network model, and outputting a cross-scale predicted image frame, wherein the preset convolutional neural network model is trained by historical event frames, and the preset convolutional neural network model includes a preset bidirectional convolutional long short-term memory module, a preset multi-scale spatial enhancement module, a preset spatiotemporal fusion attention module, and a preset multi-scale feature aggregation module;
[0009] Video data of the target dynamic scene is reconstructed based on the cross-scale predicted image frames.
[0010] According to one embodiment of the present application, converting the event stream data into an event frame in the form of a multi-channel tensor includes:
[0011] Create an all-zero matrix of preset dimensions;
[0012] Evenly dividing the time range of the event stream data into a plurality of sub-time intervals based on the preset dimension;
[0013] Determining a normalization interval based on the number of the multiple sub-time intervals, and normalizing the timestamp of each event in the event stream data to the normalization interval to obtain each normalized new event timestamp;
[0014] According to each normalized new event timestamp, a bilinear interpolation strategy is used with event polarity as a weight to interpolate each normalized event into the all-zero matrix to obtain the event frame in the form of the multi-channel tensor.
[0015] According to one embodiment of the present application, inputting the event frame in the form of a multi-channel tensor into a preset convolutional neural network and outputting a cross-scale predicted image frame includes:
[0016] Preprocessing the event frame in the form of a multi-channel tensor, and inputting the preprocessed event frame into the preset bidirectional convolutional long short-term memory module to obtain dynamic features across time;
[0017] Inputting the event frame in the form of a multi-channel tensor and the preset image frame into the preset multi-scale spatial enhancement module to obtain spatial complementary features of different scales;
[0018] Inputting the cross-temporal dynamic features and the spatial complementary features of different scales into the preset spatiotemporal fusion attention module to obtain fusion features of different scales;
[0019] The fused features of different scales are input into the preset multi-scale feature aggregation module to obtain multi-scale enhanced features of different scales, and the cross-scale predicted image frame is obtained based on the multi-scale enhanced features of different scales.
[0020] According to one embodiment of the present application, preprocessing the event frame in the form of a multi-channel tensor includes:
[0021] The event frame in the form of a multi-channel tensor is sequentially input into a 3×3 convolutional layer, an activation layer, and a residual block to obtain the preprocessed event frame.
[0022] According to one embodiment of the present application, inputting the preprocessed event frame into the preset bidirectional convolutional long short-term memory module to obtain dynamic features across time includes:
[0023] Inputting the preprocessed event frame into the forward convolution unit in the preset bidirectional convolutional long short-term memory module, processing the features of each time step based on a preset cycle step size, and generating a forward feature sequence;
[0024] Inputting the preprocessed event frame into the backward convolution unit in the preset bidirectional convolutional long short-term memory module, processing the features of each time step based on the preset cycle step size, and generating a backward feature sequence;
[0025] The forward feature sequence and the backward feature sequence are fused to obtain the cross-time dynamic feature.
[0026] According to one embodiment of the present application, inputting the event frame in the form of a multi-channel tensor and the preset image frame into the preset multi-scale spatial enhancement module to obtain spatial complementary features of different scales includes:
[0027] Extracting dynamic features from the event frame in the form of a multi-channel tensor, and extracting static features from the preset image frame;
[0028] The dynamic features and the static features are input into the preset multi-scale spatial enhancement module for connection operation to obtain the spatial complementary features of different scales.
[0029] According to one embodiment of the present application, the step of inputting the cross-temporal dynamic features and the spatial complementary features of different scales into the preset spatiotemporal fusion attention module to obtain fusion features of different scales includes:
[0030] Using the preset spatiotemporal fusion attention module, layer regularization and 1×1 convolution operations are performed on the cross-temporal dynamic features and the spatial complementary features of different scales to obtain a query matrix, a key matrix, and a value matrix;
[0031] Calculate an attention matrix based on the query matrix, the key matrix and the value matrix;
[0032] Multiplying the attention matrix and the value matrix, and adding the obtained product result to the spatial complementary features of different scales to obtain a sum result;
[0033] The sum results are sequentially input into a normalization layer and a multi-layer perceptron to obtain the fusion features of different scales.
[0034] According to one embodiment of the present application, inputting the fusion features of different scales into the preset multi-scale feature aggregation module to obtain multi-scale enhanced features of different scales includes:
[0035] Performing upsampling or downsampling operations on the fused features of different scales to obtain fused features of a preset scale;
[0036] The fusion features of the preset scale are input into the preset multi-scale feature aggregation module to obtain the multi-scale enhanced features of different scales.
[0037] According to one embodiment of the present application, before inputting the event frame in the form of a multi-channel tensor into a preset convolutional neural network model, the method further includes:
[0038] Acquire event stream data of multiple historical dynamic scenes and corresponding clear image sequences, and convert the event stream data of the multiple historical dynamic scenes into historical event frames in the form of multi-channel tensors;
[0039] Inputting the historical event frames as a training set into a preset convolutional neural network, and training the preset convolutional neural network to obtain an initial convolutional neural network model;
[0040] The initial convolutional neural network model is verified using the corresponding clear image sequence as a verification set until the initial convolutional neural network model meets the preset criteria, and the iterative training of the preset convolutional neural network is terminated to obtain the preset convolutional neural network model. Otherwise, based on the preset optimization algorithm, the model parameters are adjusted by back propagation and then the iterative training is continued.
[0041] According to the event stream-based video reconstruction method proposed in the embodiment of the present application, the features of the event stream data of the target dynamic scene are extracted through a multi-scale multi-output multi-input neural network, a bidirectional convolutional long short-term memory module is proposed, which can simultaneously integrate past and future event information, a multi-scale spatial enhancement module is proposed to extract multi-scale complementary spatial features from events and from the reconstructed image in the previous stage, and a spatiotemporal fusion attention module is proposed to integrate spatial features at different time scales. Thus, by recording the continuous event stream in the dynamic scene, high-quality video reconstruction is achieved, which solves the problem of image blur and fog artifacts caused by nonlinear time information and non-uniform spatial distribution in complex dynamic scenes, thereby meeting the demand for video reconstruction to develop higher image quality and more robustness.
[0042] To achieve the above objectives, a second embodiment of the present application proposes a video reconstruction device based on an event stream, comprising:
[0043] A conversion module, configured to obtain event stream data of a target dynamic scene and convert the event stream data into event frames in the form of multi-channel tensors;
[0044] An acquisition module is configured to input the event frame in the form of a multi-channel tensor into a preset convolutional neural network model, and output a cross-scale predicted image frame, wherein the preset convolutional neural network model is trained by historical event frames and includes a preset bidirectional convolutional long short-term memory module, a preset multi-scale spatial enhancement module, a preset spatiotemporal fusion attention module, and a preset multi-scale feature aggregation module;
[0045] A reconstruction module is used to reconstruct the video data of the target dynamic scene based on the cross-scale predicted image frame.
[0046] According to one embodiment of the present application, the conversion module is specifically configured to:
[0047] Create an all-zero matrix of preset dimensions;
[0048] Evenly dividing the time range of the event stream data into a plurality of sub-time intervals based on the preset dimension;
[0049] Determining a normalization interval based on the number of the multiple sub-time intervals, and normalizing the timestamp of each event in the event stream data to the normalization interval to obtain each normalized new event timestamp;
[0050] According to each normalized new event timestamp, a bilinear interpolation strategy is used with event polarity as a weight to interpolate each normalized event into the all-zero matrix to obtain the event frame in the form of the multi-channel tensor.
[0051] According to one embodiment of the present application, the obtaining module includes:
[0052] a first obtaining unit, configured to preprocess the event frame in the form of a multi-channel tensor, and input the preprocessed event frame into the preset bidirectional convolutional long short-term memory module to obtain dynamic features across time;
[0053] A second obtaining unit is configured to input the event frame in the form of a multi-channel tensor and the preset image frame into the preset multi-scale spatial enhancement module to obtain spatial complementary features of different scales;
[0054] a third obtaining unit, configured to input the cross-temporal dynamic features and the spatial complementary features of different scales into the preset spatiotemporal fusion attention module to obtain fused features of different scales;
[0055] The fourth obtaining unit is used to input the fusion features of different scales into the preset multi-scale feature aggregation module to obtain multi-scale enhanced features of different scales, and obtain the cross-scale predicted image frame based on the multi-scale enhanced features of different scales.
[0056] According to one embodiment of the present application, the obtaining module is specifically configured to:
[0057] The event frame in the form of a multi-channel tensor is sequentially input into a 3×3 convolutional layer, an activation layer, and a residual block to obtain the preprocessed event frame.
[0058] According to one embodiment of the present application, the first obtaining unit is specifically configured to:
[0059] Inputting the preprocessed event frame into the forward convolution unit in the preset bidirectional convolutional long short-term memory module, processing the features of each time step based on the preset cycle step size, and generating a forward feature sequence;
[0060] Inputting the preprocessed event frame into the backward convolution unit in the preset bidirectional convolutional long short-term memory module, processing the features of each time step based on the preset cycle step size, and generating a backward feature sequence;
[0061] The forward feature sequence and the backward feature sequence are fused to obtain the cross-time dynamic feature.
[0062] According to one embodiment of the present application, the second obtaining unit is specifically configured to:
[0063] Extracting dynamic features from the event frame in the form of a multi-channel tensor, and extracting static features from the preset image frame;
[0064] The dynamic features and the static features are input into the preset multi-scale spatial enhancement module for connection operation to obtain the spatial complementary features of different scales.
[0065] According to one embodiment of the present application, the third obtaining unit is specifically configured to:
[0066] Using the preset spatiotemporal fusion attention module, layer regularization and 1×1 convolution operations are performed on the cross-temporal dynamic features and the spatial complementary features of different scales to obtain a query matrix, a key matrix, and a value matrix;
[0067] Calculate an attention matrix based on the query matrix, the key matrix and the value matrix;
[0068] Multiplying the attention matrix and the value matrix, and adding the obtained product result to the spatial complementary features of different scales to obtain a sum result;
[0069] The sum results are sequentially input into a normalization layer and a multi-layer perceptron to obtain the fusion features of different scales.
[0070] According to one embodiment of the present application, the fourth obtaining unit is specifically configured to:
[0071] Performing upsampling or downsampling operations on the fused features of different scales to obtain fused features of a preset scale;
[0072] The fusion features of the preset scale are input into the preset multi-scale feature aggregation module to obtain the multi-scale enhanced features of different scales.
[0073] According to one embodiment of the present application, before inputting the event frame in the form of a multi-channel tensor into a preset convolutional neural network model, the obtaining module is further configured to:
[0074] Inputting the historical event frames as a training set into a preset convolutional neural network, and training the preset convolutional neural network to obtain an initial convolutional neural network model;
[0075] The initial convolutional neural network model is verified using the corresponding clear image sequence as a verification set until the initial convolutional neural network model meets the preset criteria, and the iterative training of the preset convolutional neural network is terminated to obtain the preset convolutional neural network model. Otherwise, based on the preset optimization algorithm, the model parameters are adjusted by back propagation and then the iterative training is continued.
[0076] According to the event stream-based video reconstruction device proposed in the embodiment of the present application, the features of the event stream data of the target dynamic scene are extracted through a multi-scale multi-output multi-input neural network, a bidirectional convolutional long short-term memory module is proposed, which can simultaneously integrate past and future event information, a multi-scale spatial enhancement module is proposed to extract multi-scale complementary spatial features from events and from the reconstructed image in the previous stage, and a spatiotemporal fusion attention module is proposed to integrate spatial features at different time scales. Thus, by recording the continuous event stream in the dynamic scene, high-quality video reconstruction is achieved, solving the image blur and fog artifact problems caused by nonlinear time information and non-uniform spatial distribution in complex dynamic scenes, thereby meeting the requirements of video reconstruction tasks for higher image quality and greater robustness.
[0077] To achieve the above-mentioned objectives, the third aspect of the present application proposes an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the event stream-based video reconstruction method as described in the above-mentioned embodiment.
[0078] To achieve the above-mentioned objectives, the fourth embodiment of the present application proposes a computer-readable storage medium on which a computer program is stored. The program is executed by a processor to implement the event stream-based video reconstruction method as described in the above-mentioned embodiment.
[0079] To achieve the above objectives, the fifth embodiment of the present application proposes a computer program product, which includes a computer program. When the computer program is executed by a processor, it is used to implement the event stream-based video reconstruction method as described in the above embodiment.
[0080] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0082] Figure 1 A flowchart of a method for video reconstruction based on event streams provided according to an embodiment of the present application;
[0083] Figure 2 is a flowchart of a method for video reconstruction using spatiotemporal dynamics of events according to one embodiment of the present application;
[0084] Figure 3 Schematic diagram of a preset bidirectional convolutional long short-term memory module according to one embodiment of the present application;
[0085] Figure 4 is a schematic diagram of a preset multi-scale spatial enhancement module according to one embodiment of the present application;
[0086] Figure 5 Schematic diagram of a preset spatiotemporal fusion attention module according to one embodiment of the present application;
[0087] Figure 6 Schematic diagram showing comparison results of the HQF (High-Quality Frames) dataset and the BS-ERGB (Beam Splitter Event and RGB Dataset) dataset according to one embodiment of the present application with the current algorithm for reconstructing video from event streams;
[0088] Figure 7 Schematic diagram of a block diagram of a video reconstruction device based on event stream according to an embodiment of the present application;
[0089] Figure 8 A schematic diagram of the structure of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0090] The following describes in detail embodiments of the present application. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0091] The following describes the event stream-based video reconstruction method, device, electronic device, and storage medium proposed in accordance with embodiments of the present application with reference to the accompanying drawings.
[0092] Figure 1 This is a flowchart of a method for video reconstruction based on event stream according to an embodiment of the present application.
[0093] Before introducing the event stream-based video reconstruction method proposed in the embodiment of the present application, the relevant technical background is first introduced.
[0094] Existing deep learning methods overcome the limitations of traditional optimization methods that rely on scene structure, motion constraints, and artificial integration or filters, thereby avoiding the serious loss of details in the reconstruction process. However, due to the characteristics of event streams and their inherent photometric consistency assumptions, these methods still face limitations that restrict their performance and application in real-world dynamic scenes. (1) Non-uniform spatial distribution in complex dynamic scenes. In the same scene, the degree of brightness change in different regions varies, resulting in different speeds and numbers of event generation. Fast-moving objects or high-contrast edges generate a large number of events, enriching the reconstruction details of these areas. In contrast, static or low-contrast areas generate few events, resulting in "blank" areas or loss of details in the reconstructed video. This spatial difference particularly affects the reconstruction of stationary or slow-moving objects, thereby reducing the overall quality and consistency of the reconstructed image. (2) Non-linear temporal information in complex dynamic scenes. In dynamic scenes, changes in background, illumination, and object motion affect the generation and distribution of continuous event data in time, such as irregular time intervals and changing event rates, resulting in frequent changes in the rate and polarity direction of event streams on pixels. Existing deep learning methods cannot cover a wide range of data distributions and only perform well on test datasets with similar distributions to the training set data. When the scenario becomes more complex, the performance degrades significantly.
[0095] In summary, existing methods rely heavily on pre-trained optical flow models or optical flow estimation within a contrast maximization framework. When the event data distribution of the training dataset differs significantly from that of real-world test scenes, inaccurate optical flow estimation can introduce artifacts in the reconstructed frames, leading to severe performance degradation.
[0096] Based on the above problems, the embodiment of the present application proposes a video reconstruction method based on event stream, combining Figure 2As shown in the figure, a bidirectional convolutional long short-term memory (BCLSTM) module is proposed to aggregate past and future information in the event stream, thereby modeling the dynamics of different sequence lengths and capturing temporal dependencies in the event stream. To accurately recover motionless or low-contrast regions, a multi-scale spatial enhancement module (MSSE) is proposed to extract dynamic features from events and static features from previously reconstructed images. These complementary spatial features enable the network to flexibly process non-uniformly distributed event data. Furthermore, a spatial-temporal fusion attention module (STFA) is proposed to integrate spatial information from different timestamps to perceive the overall motion and changes of the scene. Finally, to effectively combine structural information and details at different scales and maximize the utilization of multi-scale features, a multi-scale feature aggregation module (MSFA) is proposed. This module aggregates features from different scales at any given scale, thereby promoting cross-scale information flow within the network. Therefore, by recording the continuous event stream in dynamic scenes, high-quality video reconstruction is achieved, which solves the problems of image blur and fog artifacts caused by nonlinear temporal information and non-uniform spatial distribution in complex dynamic scenes, thereby meeting the needs of video reconstruction tasks towards higher image quality and greater robustness.
[0097] For example, Figure 1 As shown, the event stream-based video reconstruction method includes the following steps:
[0098] In step S101 , event stream data of a target dynamic scene is acquired, and the event stream data is converted into event frames in the form of multi-channel tensors.
[0099] It can be understood that a dynamic scene refers to a scene that changes over time in a video or continuous image sequence, including moving objects or a changing environment. In computer vision, event stream data typically refers to the change information between consecutive frames extracted from a video or image sequence. These changes can be pixel-level or higher-level feature changes, such as object movement or shape changes. In machine learning and deep learning, a tensor is a multidimensional array that can be thought of as a container for numbers. A multi-channel tensor means that the tensor has multiple dimensions. For example, in image processing, a color image can be represented as a three-dimensional tensor containing information for the three color channels of red, green, and blue. An event frame can be understood as a snapshot or frame extracted from event stream data at a specific point in time. These frames contain dynamic information about the scene at that point in time and are typically converted into a format suitable for processing by machine learning models.
[0100] Specifically, the embodiment of the present application can use the event camera to perform arbitrary nonlinear motion to capture event stream data in complex dynamic scenes, and convert these data into event frames V in the form of multi-channel tensors suitable for convolutional neural network structure processing. ε .
[0101] When the brightness change (logarithmic value) at a pixel location exceeds the event trigger threshold C, an event is generated. The polarity of the event indicates whether the brightness has brightened or dimmed, represented by +1 and -1 respectively. The event stream is a collection of these events, each containing a location, timestamp, polarity, and index.
[0102] The relationship between the event stream and scene brightness changes is shown in the following formula:
[0103] log(I(x i ,y i ,t i )-I(x i ,y i ,t i -Δt))=p i C;
[0104] Among them, p i is the polarity of the event, {x i ,y i ,p i ,t i} is the event set (i.e. event stream) triggered by scene brightness changes, (x i ,y i ) is the pixel position where the event is triggered, t i is the triggering time of the event, and i is the i-th triggered event.
[0105] The following details how to convert event stream data into event frames in the form of multi-channel tensors.
[0106] As a possible implementation method, in some embodiments, event stream data is converted into an event frame in the form of a multi-channel tensor, including: creating an all-zero matrix of a preset dimension; evenly dividing the time range of the event stream data into multiple sub-time intervals based on the preset dimension; determining a normalization interval based on the number of multiple sub-time intervals, and normalizing the timestamp of each event in the event stream data to the normalization interval to obtain each normalized new event timestamp; based on each normalized new event timestamp, using a bilinear interpolation strategy, with event polarity as a weight, interpolating each normalized event into the all-zero matrix to obtain an event frame in the form of a multi-channel tensor.
[0107] Event stream data is a series of data points recorded over time, each containing a timestamp and other relevant information about the event. Normalization is a data preprocessing method used to scale data to a standard range for easier processing and comparison.
[0108] Specifically, first, an all-zero matrix (i.e., all elements are zero) can be defined for the event stream data, and its dimension is Bin×Width×Height, where Width and Height are the width and height of the event camera resolution, respectively. Then, the start and end time range of the event stream data is divided into Bin sub-time intervals along the time dimension, and the timestamp of each event in the event stream data is normalized to the normalized interval [0, Bin-1]. Then, a bilinear interpolation strategy is used to assign each event to adjacent coordinates according to the pixel position, and the event polarity (+1 or -1) is superimposed on the corresponding position as a weight (for example, positive events are accumulated +1 in the corresponding position in the matrix, and negative events are accumulated -1), thereby forming an event frame in the form of a multi-channel tensor, which is presented in the form of a matrix. Preferably, the embodiment of the present application can set Bin to 5, that is, the start and end time range of the event stream data is divided into 5 sub-time intervals along the time dimension, and the normalized interval is [0,4].
[0109] In step S102, the event frame in the form of a multi-channel tensor is input into a preset convolutional neural network model, and a cross-scale predicted image frame is output, wherein the preset convolutional neural network model is trained by historical event frames, and the preset convolutional neural network model includes a preset bidirectional convolutional long short-term memory module, a preset multi-scale spatial enhancement module, a preset spatiotemporal fusion attention module and a preset multi-scale feature aggregation module.
[0110] Specifically, after obtaining the event frame in the form of a multi-channel tensor, the event frame in the form of a multi-channel tensor can be input into a preset convolutional neural network model, which can process and analyze these data, wherein the preset convolutional neural network model is trained by analyzing historical event frames.
[0111] In addition, the architecture of the preset convolutional neural network model includes several key components that work together to achieve the generation of cross-scale predicted image frames. First, the model includes a preset bidirectional convolutional long short-term memory module BCLSTM, such as Figure 3 As shown in , this module can process time series data, can model dynamics of different sequence lengths, capture long-term dependencies in event streams, and reconstruct the motion information and edge contours of the target using past and future event streams; secondly, the model also includes a preset multi-scale spatial enhancement module MSSE, as shown in Figure 4 As shown in , this module enhances the model's ability to recognize spatial features by analyzing images at different scales to solve the problem of rare or missing events in certain areas; in addition, in order to integrate the features of the same object that are close in time to strengthen the features in the target frame, the model integrates a preset spatiotemporal fusion attention module STFA at each scale, as shown in Figure 5 As shown in the figure, this module assigns different weights to different temporal and spatial features, allowing the model to focus more on important features, thereby improving prediction accuracy. Finally, the model also includes a pre-defined Multi-Scale Feature Aggregation (MSFA) module, which integrates feature information from different scales to generate a richer and more comprehensive predicted image frame. Through these carefully designed modules, the pre-defined convolutional neural network model can effectively process the input event frames and output high-quality cross-scale predicted image frames.
[0112] As a possible implementation method, in some embodiments, an event frame in the form of a multi-channel tensor is input into a preset convolutional neural network, and a cross-scale predicted image frame is output, including: preprocessing the event frame in the form of a multi-channel tensor, and inputting the preprocessed event frame into a preset bidirectional convolutional long short-term memory module to obtain cross-temporal dynamic features; inputting the event frame in the form of a multi-channel tensor and the preset image frame into a preset multi-scale spatial enhancement module to obtain spatial complementary features of different scales; inputting the cross-temporal dynamic features and the spatial complementary features of different scales into a preset spatiotemporal fusion attention module to obtain fusion features of different scales; inputting the fusion features of different scales into a preset multi-scale feature aggregation module to obtain multi-scale enhanced features of different scales, and obtaining a cross-scale predicted image frame based on the multi-scale enhanced features of different scales.
[0113] Specifically, in the process of inputting the event frame in the form of a multi-channel tensor into a preset convolutional neural network and outputting a cross-scale predicted image frame, the embodiment of the present application first performs a multi-channel tensor event frame V ε Perform preprocessing and convert the preprocessed event frame f V As the input of the preset bidirectional convolutional long short-term memory module, the module can model the dynamics of different sequence lengths and capture the long-term dependencies in the event stream to output the dynamic features F across time. T ; Secondly, the event frame V in the form of multi-channel tensor ε and the previously reconstructed image frame I L (i.e., the preset image frame) is used as the input of the preset multi-scale spatial enhancement module to obtain spatial complementary features F at different scales S ; Then, the dynamic features F across time are transformed along the time axis T and spatial complementary features F at different scales S As the input of the preset spatiotemporal fusion attention module, different temporal features are aggregated to obtain fusion features of different scales (i.e., enhanced spatiotemporal features) F; the fusion features F of different scales are then used as the input of the preset multi-scale feature aggregation module to aggregate features from different scales at any given scale to obtain multi-scale enhanced features F of different scales. MSFA , and finally based on the multi-scale enhanced features F of different scales MSFA The predicted image frame of the corresponding scale is output through the output layer.
[0114] Optionally, in some embodiments, the event frame in the form of a multi-channel tensor is preprocessed, including: inputting the event frame into a 3×3 convolutional layer, an activation layer, and a residual block in sequence to obtain a preprocessed event frame.
[0115] It is understandable that the convolution layer is a layer used to extract features in deep learning, and 3×3 represents the size of the convolution kernel, which is a commonly used convolution kernel size in image processing and can extract local features. The activation layer usually refers to the nonlinear activation function added after the convolution layer. Its purpose is to solve complex problems that linear models cannot solve. For example, the ReLU (Rectified Linear Unit) activation function can increase the nonlinear ability of the model. The residual block is a structure in deep neural networks that allows inputs to be passed directly to subsequent layers through jump connections. This structure helps alleviate the gradient vanishing problem in deep network training, allowing the network to train deeper models.
[0116] Specifically, in the process of preprocessing the event frame in the form of a multi-channel tensor, the present application can convert the event frame V in the form of a multi-channel tensor into εIt is sequentially input into a 3×3 convolutional layer, activation layer and residual block. After these steps, the preprocessed event frame f can be obtained. V , which can provide more suitable input data for subsequent deep learning models.
[0117] For ease of understanding, different stages of the process of obtaining a cross-scale predicted image frame are described in detail below.
[0118] Optionally, in some embodiments, the preprocessed event frame is input into a preset bidirectional convolutional long short-term memory module to obtain dynamic features across time, including: inputting the preprocessed event frame into the forward convolution unit in the preset bidirectional convolutional long short-term memory module, processing the features of each time step based on a preset loop step, and generating a forward feature sequence; inputting the preprocessed event frame into the backward convolution unit in the preset bidirectional convolutional long short-term memory module, processing the features of each time step based on a preset loop step, and generating a backward feature sequence; fusing the forward feature sequence and the backward feature sequence to obtain dynamic features across time.
[0119] Specifically, if Figure 3 As shown in the figure, the preset bidirectional convolutional long short-term memory module mainly includes two parts: the forward convolution unit LSTM and the backward convolution unit LSTM. The preprocessed event frame f V The forward convolution unit and the backward convolution unit of the preset bidirectional convolution long short-term memory module are input in sequence, and the features of each time step are processed based on the preset cycle step size. The forward feature sequence and the backward feature sequence can be obtained respectively. By fusing the forward feature sequence and the backward feature sequence, the enhanced features in the time domain can be obtained, that is, the dynamic features F across time. T :
[0120] F T =BCLSTM(f V ).
[0121] It should be noted that the embodiment of the present application preferably sets the preset loop step to 5, that is, the module can simultaneously aggregate event stream data of 5 past and future time steps each time in time series processing, thereby more comprehensively modeling the time dependency of dynamic scenes.
[0122] Optionally, in some embodiments, the event frame in the form of a multi-channel tensor and the preset image frame are input into a preset multi-scale spatial enhancement module to obtain spatial complementary features of different scales, including: extracting dynamic features from the event frame in the form of a multi-channel tensor, and extracting static features from the preset image frame; inputting the dynamic features and the static features into the preset multi-scale spatial enhancement module for connection operation to obtain spatial complementary features of different scales.
[0123] Specifically, since this application is dedicated to video reconstruction based on event streams, only 5 images are reconstructed each time, so the last reconstructed image frame I can be used. L (i.e., the preset image frame) extracts static features from it and extracts them from the event frame V in the form of a multi-channel tensor ε Extract dynamic features from the dynamic features, use the dynamic features and static features as the input of the preset multi-scale spatial enhancement module, and use the connection operation to fuse the dynamic features and static features to obtain spatial complementary features F at different scales. S ,For example, the preset multi-scale spatial enhancement module considers 3 scales, and ,then three spatial complementary features of corresponding scales can be ,obtained.
[0124] F S =MSSE(Cat(V ε ,I L ));
[0125] Among them, Cat is the connection operation.
[0126] For example, the preset multi-scale spatial enhancement module considers three scales {1, 1 / 2, 1 / 4} ("1" represents the original resolution scale, "1 / 2" represents reducing the spatial resolution to 1 / 2 of the original image, and "1 / 4" represents reducing the spatial resolution to 1 / 4 of the original image). When downsampling is required for the input, the nearest neighbor interpolation can be performed on the last connection result to obtain the downsampling result. For example, when going from 1 / 2 scale to 1 / 4 scale (downsampling), the feature map resolution is halved. Nearest neighbor interpolation is an image processing technique used to generate new pixel values. It adjusts the image size by finding the nearest pixel point and assigning its value to the new pixel point. After this process, dynamic and static spatial complementary features at multiple scales can be obtained, that is, spatial complementary features F at different scales. S .
[0127] Optionally, in some embodiments, the dynamic features across time and the spatial complementary features of different scales are input into a preset spatiotemporal fusion attention module to obtain fusion features of different scales, including: using the preset spatiotemporal fusion attention module to perform layer regularization and 1×1 convolution operations on the dynamic features across time and the spatial complementary features of different scales to obtain a query matrix, a key matrix and a value matrix; calculating the attention matrix based on the query matrix, the key matrix and the value matrix; multiplying the attention matrix and the value matrix, and adding the obtained product result to the spatial complementary features of different scales to obtain a sum result; inputting the sum result into the normalization layer and the multilayer perceptron in sequence to obtain fusion features of different scales.
[0128] It is understandable that layer regularization is a technique in deep learning that is used to prevent overfitting during model training by adding regularization terms to constrain the complexity of the model. 1×1 convolution is a special convolution operation that can perform linear transformations on features without changing the size of the feature map, and is often used to adjust the number of channels. The normalization layer (LayerNormalization, referred to as LN) is used to adjust the data distribution so that it has certain statistical properties, such as a mean of 0 and a standard deviation of 1, which contributes to the stability and convergence speed of the model. Multi-Layer Perceptro (MLP) is a simple feedforward neural network that usually contains multiple fully connected layers to process complex nonlinear relationships.
[0129] Specifically, if Figure 5 As shown in the figure, this stage can transform the dynamic features F across time into T and spatial complementary features of different scales D S As the input of the preset spatiotemporal fusion attention module, layer regularization and 1×1 convolution operations are performed to obtain the query matrix Q∈R hw×c (indicates the features that need attention), key matrix K∈R hw×c (representing features that can be paid attention to) and the value matrix V∈R hw×c (represents the value of the feature), where h is the height of the matrix, w is the width of the matrix, and c is the number of channels in the matrix. Then, based on these three matrices, the attention matrix Attention(Q,K,V) is calculated, which reflects the correlation between each feature:
[0130]
[0131] Among them, d k is the dimension size of the value matrix, which is set to c here.
[0132] Then, the attention matrix is multiplied by the value matrix to obtain a weighted value matrix, and then this weighted value matrix is added to the spatial complementary features of different scales. The result of the addition is then passed through the normalization layer and the multi-layer perceptron in sequence to obtain the fusion features F at different scales:
[0133] F=(LN(Attention(Q,K,V)V+F S )).
[0134] Optionally, in some embodiments, fused features of different scales are input into a preset multi-scale feature aggregation module to obtain multi-scale enhanced features of different scales, including: performing upsampling or downsampling operations on fused features of different scales to obtain fused features of a preset scale; inputting the fused features of a preset scale into a preset multi-scale feature aggregation module to obtain multi-scale enhanced features of different scales.
[0135] Specifically, according to the requirements of the current scale, the fusion features F of different scales can be S Upsample or downsample to the current scale to obtain the fusion features of the preset scale (the preset scale is the same as the current scale), and then use the fusion features of the preset scale as the input of the preset multi-scale feature aggregation module to obtain multi-scale enhanced features F of different scales. MSFA Among them, the preset multi-scale feature aggregation module is composed of a 1×1 convolution layer, an activation layer and a 3×3 convolution layer stacked in sequence.
[0136] For example, when the default scale is 1:
[0137]
[0138] When the preset scale is 1 / 2:
[0139]
[0140] Among them, ↑ is the upsampling operation, and ↓ is the downsampling operation.
[0141] It should be noted that when the preset scale is 1 / 4, the multi-scale enhanced feature F MSFA The reconstruction result at the scale of 1 / 4 can be obtained by first passing through a residual block and then through a 3×3 convolution layer. When the preset scale is 1 / 2 or 1, the output features of the residual block at the previous scale can be upsampled first, and then the upsampled results can be combined with the multi-scale enhanced features F at the current scale. MSFA The network is connected and passes through a 1×1 convolutional layer, a residual block, and a 3×3 convolutional layer in sequence to obtain the reconstruction result at the current scale.
[0142] Next, we will explain in detail how to obtain the preset convolutional neural network model.
[0143] As a possible implementation method, in some embodiments, before inputting the event frame in the form of a multi-channel tensor into the preset convolutional neural network model, it also includes: obtaining event stream data and corresponding clear image sequences of multiple historical dynamic scenes, and converting the event stream data of the multiple historical dynamic scenes into historical event frames in the form of a multi-channel tensor; inputting the historical event frames as a training set into the preset convolutional neural network, training the preset convolutional neural network to obtain an initial convolutional neural network model; using the corresponding clear image sequence as a verification set to verify the initial convolutional neural network model until the initial convolutional neural network model meets the preset standards, ending the iterative training of the preset convolutional neural network to obtain the preset convolutional neural network model; otherwise, based on the preset optimization algorithm, adjusting the model parameters through back propagation and continuing the iterative training.
[0144] Specifically, event cameras and traditional cameras can be used to perform arbitrary nonlinear motion to capture event stream data and clear images of multiple historical complex dynamic scenes, and a data set can be constructed based on the event stream data of multiple historical dynamic scenes and the corresponding clear image sequences for training a preset convolutional neural network model. Before using the event stream data of historical dynamic scenes for model training, the event stream data of multiple historical dynamic scenes can be converted into historical event frames in the form of multi-channel tensors. A series of historical event frames are input into the preset convolutional neural network as a training set, and the preset convolutional neural network is trained using a supervised framework to obtain an initial convolutional neural network model; then, a series of corresponding clear image sequences are used as a verification set to verify the initial convolutional neural network model. If the initial convolutional neural network model can meet the preset standards, then the iterative training of the preset convolutional neural network can be ended to obtain the preset convolutional neural network model. If the model fails to meet the preset criteria, the model parameters can be adjusted through back propagation based on a preset optimization algorithm (such as the AdamW optimization algorithm, which is a commonly used optimization algorithm in the fields of machine learning and deep learning for training neural networks. It combines the adaptive moment estimation (Adam) algorithm and weight decay (weight decay) technology to improve the efficiency and performance of model training). Then, iterative training is continued until a model that meets the preset criteria (that is, the preset convolutional neural network model) is obtained.
[0145] Among them, in the process of using the supervised framework to train the preset convolutional neural network, the full supervision loss is:
[0146]
[0147] Among them, I k is the lth predicted image frame output by the model; is the real l-th image; It is the L1 loss under multi-scale; is the similarity loss at multiple scales; is the perceptual loss under multi-scale; λ1 and λ2 are both weight factors, which are both set to 1 in the embodiment of the present application; L is the number of consecutive images of the video in the training phase, preferably, it is set to 10 in the embodiment of the present application.
[0148] Furthermore, when adjusting model parameters using a pre-defined optimization algorithm (such as the AdamW optimization algorithm), independent adaptive learning rates can be designed for different parameters by calculating the first-order and second-order moment estimates of the gradient. This approach allows the model to more accurately adjust the parameters of the neural network convolutional layer when iteratively updating based on training data, thereby improving model training efficiency and performance.
[0149] In step S103 , video data of the target dynamic scene is reconstructed based on the cross-scale predicted image frames.
[0150] Specifically, after obtaining cross-scale predicted image frames, video data of the target dynamic scene can be reconstructed based on cross-scale prediction techniques. The core of this method is to leverage information from different scales to improve prediction accuracy and reconstructed image quality. This approach can better understand and simulate dynamic changes in the target scene, providing richer and more accurate data support for video analysis, computer vision, and related fields.
[0151] To facilitate those skilled in the art to further understand the event stream-based video reconstruction method proposed in the embodiment of the present application, the following is a comparison of the reconstruction results of the present application method (ST-E2V (Spatio-Temporal Event-to-Video)) and other methods (such as Figure 6 Shown) demonstrates the validity of this application.
[0152] Among them, other methods are event stream video reconstruction methods based on deep learning and supervision, including E2VID+ (event camera to video method), SPADE-E2VID (Spatially-Adaptive Denormalization for Event-Based Video Reconstruction, a spatially varying reconstruction method), SSL-E2VID (Self-Supervised Learning for Event-to-Video Reconstruction, a self-supervised learning method for event stream video reconstruction), ET-Net (Event-Time Neural Network, event time neural network, a transformer reconstruction method) and HyperE2VID (Hypernetworks for Event-Based Video Reconstruction, a hypernetwork reconstruction method).
[0153] The video reconstruction results of the above six models on the HQF dataset and the BS-ERGB dataset are shown in Table 1:
[0154] Table 1
[0155]
[0156] By measuring quantitative indicators under the same dataset, including Mean Square Error (MSE), Structural Similarity Index (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS), the specific definitions are as follows:
[0157]
[0158]
[0159] Among them, μ k is the predicted image frame I k The mean value, μ k,GT For real images The mean of is the predicted image frame I k The variance of For real images The variance of image I k and The covariance of , c1, c2 are both constants with very small values (used to avoid the situation where the denominator is zero).
[0160] It can be understood that higher SSIM values and lower MSE and LPIPS values indicate better video reconstruction performance. As shown in the numerical results in Table 1, the embodiments of the present application demonstrate better performance in reconstructing videos from event streams, validating the effectiveness of the event stream-based video reconstruction method proposed in the embodiments of the present application.
[0161] According to the event stream-based video reconstruction method proposed in the embodiment of the present application, the features of the event stream data of the target dynamic scene are extracted through a multi-scale multi-output multi-input neural network, a bidirectional convolutional long short-term memory module is proposed, which can simultaneously integrate past and future event information, a multi-scale spatial enhancement module is proposed to extract multi-scale complementary spatial features from events and from the reconstructed image in the previous stage, and a spatiotemporal fusion attention module is proposed to integrate spatial features at different time scales. Thus, by recording the continuous event stream in the dynamic scene, high-quality video reconstruction is achieved, which solves the problem of image blur and fog artifacts caused by nonlinear time information and non-uniform spatial distribution in complex dynamic scenes, thereby meeting the demand for video reconstruction to develop higher image quality and more robustness.
[0162] Next, the event stream-based video reconstruction device proposed in accordance with an embodiment of the present application will be described with reference to the accompanying drawings.
[0163] Figure 7 It is a block diagram of a video reconstruction device based on event stream according to an embodiment of the present application.
[0164] like Figure 7 As shown, the event stream-based video reconstruction device 10 includes: a conversion module 100 , an acquisition module 200 and a reconstruction module 300 .
[0165] The conversion module 100 is used to obtain event stream data of the target dynamic scene and convert the event stream data into event frames in the form of multi-channel tensors;
[0166] An acquisition module 200 is configured to input an event frame in the form of a multi-channel tensor into a preset convolutional neural network model, and output a cross-scale predicted image frame, wherein the preset convolutional neural network model is trained by historical event frames and includes a preset bidirectional convolutional long short-term memory module, a preset multi-scale spatial enhancement module, a preset spatiotemporal fusion attention module, and a preset multi-scale feature aggregation module;
[0167] The reconstruction module 300 is configured to reconstruct video data of a target dynamic scene based on the cross-scale predicted image frames.
[0168] Furthermore, in some embodiments, the conversion module 100 is specifically configured to:
[0169] Create an all-zero matrix of preset dimensions;
[0170] Divide the time range of event stream data into multiple sub-time intervals evenly based on preset dimensions;
[0171] Determining a normalization interval based on the number of the multiple sub-time intervals, and normalizing the timestamp of each event in the event stream data to the normalization interval to obtain each normalized new event timestamp;
[0172] According to each normalized new event timestamp, a bilinear interpolation strategy is used with event polarity as weight to interpolate each normalized event into an all-zero matrix to obtain an event frame in the form of a multi-channel tensor.
[0173] Furthermore, in some embodiments, the obtaining module 200 includes:
[0174] A first acquisition unit is used to preprocess the event frame in the form of a multi-channel tensor and input the preprocessed event frame into a preset bidirectional convolutional long short-term memory module to obtain dynamic features across time;
[0175] The second obtaining unit is used to input the event frame in the form of a multi-channel tensor and the preset image frame into a preset multi-scale spatial enhancement module to obtain spatial complementary features of different scales;
[0176] The third acquisition unit is used to input the cross-temporal dynamic features and the spatial complementary features of different scales into the preset spatiotemporal fusion attention module to obtain fusion features of different scales;
[0177] The fourth obtaining unit is used to input the fusion features of different scales into a preset multi-scale feature aggregation module to obtain multi-scale enhanced features of different scales, and obtain a cross-scale predicted image frame based on the multi-scale enhanced features of different scales.
[0178] Furthermore, in some embodiments, the obtaining module 200 is specifically configured to:
[0179] The event frame in the form of a multi-channel tensor is sequentially input into the 3×3 convolutional layer, the activation layer, and the residual block to obtain the preprocessed event frame.
[0180] Furthermore, in some embodiments, the first obtaining unit is specifically configured to:
[0181] The preprocessed event frame is input into the forward convolution unit in the preset bidirectional convolutional long short-term memory module, and the features of each time step are processed based on the preset cycle step size to generate a forward feature sequence;
[0182] The preprocessed event frame is input into the backward convolution unit in the preset bidirectional convolutional long short-term memory module, and the features of each time step are processed based on the preset cycle step size to generate a backward feature sequence;
[0183] The forward feature sequence and the backward feature sequence are fused to obtain dynamic features across time.
[0184] Furthermore, in some embodiments, the second obtaining unit is specifically configured to:
[0185] Extract dynamic features from event frames in the form of multi-channel tensors and extract static features from pre-set image frames;
[0186] The dynamic features and static features are input into the preset multi-scale spatial enhancement module for connection operation to obtain spatial complementary features of different scales.
[0187] Furthermore, in some embodiments, the third obtaining unit is specifically configured to:
[0188] Use the preset spatiotemporal fusion attention module to perform layer regularization and 1×1 convolution operations on the dynamic features across time and the spatial complementary features of different scales to obtain the query matrix, key matrix and value matrix;
[0189] The attention matrix is calculated based on the query matrix, key matrix and value matrix;
[0190] Multiply the attention matrix and the value matrix, and add the resulting product to the spatial complementary features of different scales to obtain the sum result;
[0191] The sum results are input into the normalization layer and the multi-layer perceptron in turn to obtain fusion features of different scales.
[0192] Furthermore, in some embodiments, the fourth obtaining unit is specifically configured to:
[0193] Perform upsampling or downsampling operations on the fusion features of different scales to obtain the fusion features of the preset scale;
[0194] The fusion features of the preset scale are input into the preset multi-scale feature aggregation module to obtain multi-scale enhanced features of different scales.
[0195] Furthermore, in some embodiments, before inputting the event frame in the form of a multi-channel tensor into the preset convolutional neural network model, the obtaining module is further configured to:
[0196] Input the historical event frames as training sets into the preset convolutional neural network, and train the preset convolutional neural network to obtain an initial convolutional neural network model;
[0197] The initial convolutional neural network model is verified using the corresponding clear image sequence as a validation set until the initial convolutional neural network model meets the preset standards. The iterative training of the preset convolutional neural network is terminated to obtain the preset convolutional neural network model. Otherwise, based on the preset optimization algorithm, the model parameters are adjusted by back propagation and the iterative training is continued.
[0198] It should be noted that the aforementioned explanation of the embodiment of the event stream-based video reconstruction method is also applicable to the event stream-based video reconstruction device of this embodiment, and will not be repeated here.
[0199] According to the event stream-based video reconstruction device proposed in the embodiment of the present application, the features of the event stream data of the target dynamic scene are extracted through a multi-scale multi-output multi-input neural network, a bidirectional convolutional long short-term memory module is proposed, which can simultaneously integrate past and future event information, a multi-scale spatial enhancement module is proposed to extract multi-scale complementary spatial features from events and from the reconstructed image in the previous stage, and a spatiotemporal fusion attention module is proposed to integrate spatial features at different time scales. Thus, by recording the continuous event stream in the dynamic scene, high-quality video reconstruction is achieved, solving the image blur and fog artifact problems caused by nonlinear time information and non-uniform spatial distribution in complex dynamic scenes, thereby meeting the requirements of video reconstruction tasks for higher image quality and greater robustness.
[0200] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0201] A memory 801 , a processor 802 , and a computer program stored in the memory 801 and executable on the processor 802 .
[0202] When the processor 802 executes the program, the event stream-based video reconstruction method provided in the above embodiment is implemented.
[0203] Furthermore, the electronic device further includes:
[0204] The communication interface 803 is used for communication between the memory 801 and the processor 802 .
[0205] The memory 801 is used to store computer programs that can be run on the processor 802.
[0206] The memory 801 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.
[0207] If the memory 801, the processor 802, and the communication interface 803 are implemented independently, the communication interface 803, the memory 801, and the processor 802 can be connected to each other via a bus and communicate with each other. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0208] Optionally, in a specific implementation, if the memory 801, the processor 802 and the communication interface 803 are integrated on a chip, the memory 801, the processor 802 and the communication interface 803 can communicate with each other through an internal interface.
[0209] The processor 802 may be a CPU (Central Processing Unit), or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application.
[0210] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned event stream-based video reconstruction method.
[0211] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the above event stream-based video reconstruction method.
[0212] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0213] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0214] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A video reconstruction method based on event stream, characterized in that: The following steps are involved: Acquire event stream data of a target dynamic scene, and convert the event stream data into event frames in the form of multi-channel tensors; Inputting the event frame in the form of a multi-channel tensor into a preset convolutional neural network model, and outputting a cross-scale predicted image frame, wherein the preset convolutional neural network model is trained by historical event frames, and the preset convolutional neural network model includes a preset bidirectional convolutional long short-term memory module, a preset multi-scale spatial enhancement module, a preset spatiotemporal fusion attention module, and a preset multi-scale feature aggregation module; Video data of the target dynamic scene is reconstructed based on the cross-scale predicted image frames.
2. The method according to claim 1, characterized in that The converting the event stream data into an event frame in a multi-channel tensor form includes: Create an all-zero matrix of preset dimensions; Evenly dividing the time range of the event stream data into a plurality of sub-time intervals based on the preset dimension; Determining a normalization interval based on the number of the multiple sub-time intervals, and normalizing the timestamp of each event in the event stream data to the normalization interval to obtain each normalized new event timestamp; According to each normalized new event timestamp, a bilinear interpolation strategy is used with event polarity as a weight to interpolate each normalized event into the all-zero matrix to obtain the event frame in the form of the multi-channel tensor.
3. The method according to claim 1, characterized in that The step of inputting the event frame in the form of a multi-channel tensor into a preset convolutional neural network and outputting a cross-scale predicted image frame includes: Preprocessing the event frame in the form of a multi-channel tensor, and inputting the preprocessed event frame into the preset bidirectional convolutional long short-term memory module to obtain dynamic features across time; Inputting the event frame in the form of a multi-channel tensor and the preset image frame into the preset multi-scale spatial enhancement module to obtain spatial complementary features of different scales; Inputting the cross-temporal dynamic features and the spatial complementary features of different scales into the preset spatiotemporal fusion attention module to obtain fusion features of different scales; The fused features of different scales are input into the preset multi-scale feature aggregation module to obtain multi-scale enhanced features of different scales, and the cross-scale predicted image frame is obtained based on the multi-scale enhanced features of different scales.
4. The method according to claim 3, characterized in that Preprocessing the event frame in the form of a multi-channel tensor includes: The event frame in the form of a multi-channel tensor is sequentially input into a 3×3 convolutional layer, an activation layer, and a residual block to obtain the preprocessed event frame.
5. The method according to claim 3, characterized in that The pre-processed event frame is input into the preset bidirectional convolutional long short-term memory module to obtain dynamic features across time, including: Inputting the preprocessed event frame into the forward convolution unit in the preset bidirectional convolutional long short-term memory module, processing the features of each time step based on the preset cycle step size, and generating a forward feature sequence; Inputting the preprocessed event frame into the backward convolution unit in the preset bidirectional convolutional long short-term memory module, processing the features of each time step based on the preset cycle step size, and generating a backward feature sequence; The forward feature sequence and the backward feature sequence are fused to obtain the cross-time dynamic feature.
6. The method according to claim 3, characterized in that The step of inputting the event frame in the form of a multi-channel tensor and the preset image frame into the preset multi-scale spatial enhancement module to obtain spatial complementary features of different scales includes: Extracting dynamic features from the event frame in the form of a multi-channel tensor, and extracting static features from the preset image frame; The dynamic features and the static features are input into the preset multi-scale spatial enhancement module for connection operation to obtain the spatial complementary features of different scales.
7. The method according to claim 3, characterized in that The step of inputting the cross-temporal dynamic features and the spatial complementary features of different scales into the preset spatiotemporal fusion attention module to obtain fusion features of different scales includes: Using the preset spatiotemporal fusion attention module, layer regularization and 1×1 convolution operations are performed on the cross-temporal dynamic features and the spatial complementary features of different scales to obtain a query matrix, a key matrix, and a value matrix; Calculate an attention matrix based on the query matrix, the key matrix and the value matrix; Multiplying the attention matrix and the value matrix, and adding the obtained product result to the spatial complementary features of different scales to obtain a sum result; The sum results are sequentially input into a normalization layer and a multi-layer perceptron to obtain the fusion features of different scales.
8. The method according to claim 3, characterized in that Inputting the fusion features of different scales into the preset multi-scale feature aggregation module to obtain multi-scale enhanced features of different scales includes: Performing upsampling or downsampling operations on the fused features of different scales to obtain fused features of a preset scale; The fusion features of the preset scale are input into the preset multi-scale feature aggregation module to obtain the multi-scale enhanced features of different scales.
9. The method according to claim 1, characterized in that Before inputting the event frame in the form of a multi-channel tensor into a preset convolutional neural network model, the method further includes: Acquire event stream data of multiple historical dynamic scenes and corresponding clear image sequences, and convert the event stream data of the multiple historical dynamic scenes into historical event frames in the form of multi-channel tensors; Inputting the historical event frames as a training set into a preset convolutional neural network, and training the preset convolutional neural network to obtain an initial convolutional neural network model; The initial convolutional neural network model is verified using the corresponding clear image sequence as a verification set until the initial convolutional neural network model meets the preset criteria, and the iterative training of the preset convolutional neural network is terminated to obtain the preset convolutional neural network model. Otherwise, based on the preset optimization algorithm, the model parameters are adjusted by back propagation and then the iterative training is continued.
10. A video reconstruction device based on event stream, characterized in that: include: A conversion module, configured to obtain event stream data of a target dynamic scene and convert the event stream data into event frames in the form of multi-channel tensors; An acquisition module is configured to input the event frame in the form of a multi-channel tensor into a preset convolutional neural network model, and output a cross-scale predicted image frame, wherein the preset convolutional neural network model is trained by historical event frames and includes a preset bidirectional convolutional long short-term memory module, a preset multi-scale spatial enhancement module, a preset spatiotemporal fusion attention module, and a preset multi-scale feature aggregation module; A reconstruction module is used to reconstruct the video data of the target dynamic scene based on the cross-scale predicted image frame.
Citation Information
Patent Citations
High-quality and high-frame-rate image reconstruction method based on event camera
CN111667442A
Self-supervised high-frame-rate video reconstruction method and system based on event camera
CN118537258A
Aerial image detail enhancement method based on deep learning
CN119624810A
Method, medium and device for enhancing event camera image reconstruction by fusing visible images
US20240378699A1