A blurry video super-resolution method based on event data drive
Through the fuzzy video super-resolution method driven by event data, the coordinated defuzzing and optical flow alignment of intra-frame events and inter-frame events is used to solve the reconstruction distortion problems of motion blur and low-resolution video in the prior art, and high-quality video super-resolution reconstruction is achieved.
Patent Information
- Application Number
- CN202510020067.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-01-07
AI Technical Summary
Existing video super-resolution methods are difficult to effectively remove blur and restore detailed information when processing motion blur and low-resolution videos, resulting in distortion of reconstruction results.
The fuzzy video super-resolution method driven by event data is adopted to build a fuzzy video super-resolution neural network, and the intra-frame events and inter-frame events are used to coordinate the blurring, and the coordinated alignment of optical flow and events is combined to restore high-frequency information.
Effectively eliminate motion blur, restore high-frequency information, improve video quality, reduce artifacts and distortion, and achieve higher-precision video reconstruction.
Smart Images

Figure CN119850420B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, in particular to a fuzzy video super-resolution method driven by event data. Background Art
[0002] With the widespread adoption of high-definition and ultra-high-definition (4K, 8K) display technologies, the demand for higher-resolution video content continues to increase. This is especially true in applications such as large-screen TVs, virtual reality (VR), and augmented reality (AR), where video quality has become a critical factor in the user experience. However, many low-resolution video sources (such as content recorded by low-quality cameras and streamed video) cannot meet the requirements of these high-resolution display devices, resulting in blurry images and a loss of detail, which compromises the visual experience. In this context, video super-resolution technology, as an effective solution, can restore greater detail and clarity from low-resolution video, enhancing visual quality and meeting the demands of modern display technologies for high-resolution video content.
[0003] While existing video super-resolution methods have made some progress, they still face significant distortion and blurred textures when processing input videos with motion blur. Traditional video super-resolution methods rely primarily on combining multiple low-resolution frames to restore a high-resolution image. However, when motion blur is present in a video, these methods typically encounter two major problems: an inability to effectively remove the blur and an inability to recover the details caused by the motion blur.
[0004] To address this problem, some existing solutions attempt to first remove motion blur using a deblurring network, and then restore the image using a super-resolution network. However, this approach suffers from a key issue: if the deblurring network performs poorly, the incomplete blur removal will introduce additional errors, which are amplified in the super-resolution stage, ultimately leading to severe distortion in the reconstruction. Therefore, the coupling of deblurring and super-resolution makes this approach difficult to achieve stable restoration results in practical applications.
[0005] Another common approach is to design specialized blurry video super-resolution algorithms to simultaneously address motion blur and low resolution. While some studies have attempted this approach, most rely on traditional video frame signals, which are limited when processing inputs that are both motion blurry and low-resolution. Traditional video frame signals lack effective decoupling of motion information, making them unable to recover details lost due to motion blur and difficult to provide the high-resolution information required for super-resolution. Therefore, existing solutions, relying on traditional video frame signals, struggle to effectively address the super-resolution problem of blurry videos. Summary of the Invention
[0006] In order to address the shortcomings of the above-mentioned existing technologies, the present invention proposes a blurred video super-resolution method driven by event data, hoping that by introducing event data without motion blur, the loss of details caused by motion blur can be effectively removed, and at the same time, high-frequency information in low-resolution videos can be restored, thereby improving video quality and clarity, and reducing artifacts and distortion, ultimately achieving higher-precision video super-resolution reconstruction to meet the needs of high-quality visual experience.
[0007] In order to achieve the above-mentioned object, the present invention adopts the following technical solutions:
[0008] The present invention is characterized in that the method for fuzzy video super-resolution based on event data driving is performed according to the following steps:
[0009] Step 1: Obtain training video image dataset and its corresponding event sequence, and characterize the event sequence:
[0010] Step 1.1 Obtain training video image dataset :
[0011] Obtaining clear high-resolution video image datasets ,in, Indicates the A clear, high-resolution image. , is the total number of high-resolution images;
[0012] For clear high-resolution video images Perform blur and resolution reduction to obtain a blurred low-resolution video image set ,in, Indicates the A blurry low-resolution image;
[0013] make Represents the training video image dataset;
[0014] Step 1.2 Characterize the event sequence:
[0015] Acquire high-resolution video image sets Intra-frame event sequence and inter-frame event sequences ,in, and Respectively represent A clear, high-resolution image The corresponding intra-frame event sequence and inter-frame event sequence;
[0016] right and Perform resolution reduction processing respectively to obtain a low-resolution video image set Corresponding intra-frame event sequence and inter-frame event sequences ,in, and Respectively represent A blurry low-resolution image The corresponding intra-frame event sequence and inter-frame event sequence;
[0017] right Perform voxel representation to obtain a low-resolution video image set Corresponding intra-frame event voxel sequence ,in, express The event voxels within the frame, represents the number of event voxel channels, Indicates the image height, Indicates the image width;
[0018] For event sequence Perform voxel representation to obtain a low-resolution video image set Corresponding intra-frame event voxel sequence ,in, express inter-frame event voxels;
[0019] Step 2: Construct a blur video super-resolution neural network, including a feature extraction layer, a "frame-event" collaborative deblurring module, an "optical flow-event" collaborative time alignment module, and a frame reconstruction module.
[0020] Step 2.1, the feature extraction layer includes: image feature extraction layer, intra-frame event feature extraction layer and inter-frame event extraction layer, and respectively 、 and Perform feature extraction and obtain the corresponding Blurry frame features , No. Intra-frame event features Hedi Inter-frame event features ,in, Represents the extracted feature dimension;
[0021] Step 2.2, the "frame-event" collaborative deblurring module includes: a frame and event preprocessing layer, a frame-guided event deblurring layer, an event-guided frame deblurring layer and a fully connected fusion layer, and respectively and Processing, the corresponding Deblurred frame features Hedi Deblurred intra-frame event features ;
[0022] Step 2.3, the "optical flow-event" collaborative alignment module includes: an optical flow-guided alignment module, an event-guided alignment module and a deformable convolution layer, and low-resolution images Hedi low-resolution images Process it and get Feature maps ;
[0023] Step 2.4, the frame reconstruction module is composed of multiple deconvolution layers and upsampling layers connected in series, and low-resolution images Hedi Feature maps After processing, we get Super-resolution images ; Thus we get the super-resolution video set ;
[0024] Step 3: Use the loss function of formula (10) :
[0025] (10)
[0026] In formula (10), is a non-negative constant;
[0027] Step 4: Use the gradient descent method to train the blurred video super-resolution neural network and calculate the loss function To update the network parameters, when the number of training iterations reaches the set number or the loss function When convergence occurs, the training stops, and the optimal blurry video super-resolution model is obtained; it is used to process blurry low-resolution video images to obtain corresponding clear high-resolution video images.
[0028] The event data-driven fuzzy video super-resolution method of the present invention is also characterized in that step 2.2 is performed as follows:
[0029] Step 2.2.1, the frame and event preprocessing layer consists of a frame branch channel attention module and an event branch channel attention module, and respectively and Perform context semantic feature aggregation processing and obtain the first Fuzzy frame semantic features Hedi Intra-frame event semantic features ;
[0030] Step 2.2.2: The frame-guided event deblurring layer uses equations (1) to (4) to obtain the first Deblurred intra-frame event features :
[0031] (1)
[0032] (2)
[0033] (3)
[0034] (4)
[0035] In formula (1) to formula (4), for The query matrix, for The query weight matrix, for The key matrix, for The key weight matrix, for The value matrix of for The value weight matrix of represents the matrix transpose operation, and softmax represents the activation function;
[0036] Step 2.2.3: The event-guided frame deblurring layer uses equations (5) to (8) to obtain the first Deblurred frame features :
[0037] (5)
[0038] (6)
[0039] (7)
[0040] (8)
[0041] In formula (5) to formula (8), for The query matrix, for The query weight matrix, for The key matrix, for The key weight matrix, for The value matrix of for The value weight matrix of
[0042] Step 2.2.4, the fully connected fusion layer is composed of a frame branch fully connected module and an event branch fully connected module, and and Processing, the corresponding and .
[0043] Furthermore, the step 2.3 is performed as follows:
[0044] Step 2.3.1: The optical flow guided alignment module is composed of Layer downsampling convolution layer and The upsampling convolution layers are connected alternately, and the and Perform optical flow estimation and get the No. Optical flow ;
[0045] Using formula (9) Optical flow Hedi Feature maps conduct Transformation, thus obtaining the Features from optical flow alignment :
[0046] (9)
[0047] In formula (9), Represents image distortion transformation; when i=1, let ;
[0048] Step 2.3.2: The event-guided alignment module is composed of Convolutional layers and GeLU activation layers are used to Inter-frame event features Hedi Feature maps Process it and get Features from event alignment ;
[0049] Step 2.2.3, 、 、 、 、 and After splicing along the channel, it is input into the deformable convolution layer for processing to obtain the Feature maps .
[0050] The electronic device of the present invention includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the blurred video super-resolution method, and the processor is configured to execute the program stored in the memory.
[0051] The present invention provides a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program executes the steps of the blurred video super-resolution method when the computer program is executed by a processor.
[0052] Compared with the prior art, the present invention has the following beneficial effects:
[0053] 1. This paper designs a blurred video super-resolution network based on event data, integrating motion-blur-free event data into the blurred video super-resolution task. Compared with existing technologies, this invention effectively eliminates detail loss caused by motion blur, restores high-frequency information, and fully utilizes the high dynamic range of event data. This significantly improves video quality in low-light and complex dynamic scenes, reduces artifacts and distortion, and ultimately achieves higher-precision video reconstruction.
[0054] 2. This invention innovatively divides event data into "intra-frame events" and "inter-frame events." Intra-frame events are used for intra-frame deblurring, while inter-frame events are used for inter-frame alignment. This division improves deblurring accuracy and inter-frame consistency, enhancing video reconstruction.
[0055] 3. The “frame-event” collaborative deblurring module proposed in the present invention optimizes the deblurring process through the collaboration between intra-frame events and video frames, thereby improving the detail recovery effect.
[0056] 4. The "optical flow-event" collaborative alignment module proposed in the present invention combines the motion information of optical flow and inter-frame events, improves the timing alignment accuracy, and solves the inter-frame alignment problem in the prior art.
[0057] 5. The present invention adopts a supervised training method to deeply embed event information into the fuzzy video super-resolution network, thereby improving the quality of the output frame. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 The neural network diagram of the fuzzy video super-resolution method of the present invention;
[0059] Figure 2This is a diagram of the "frame-event" collaborative deblurring module of the present invention;
[0060] Figure 3 This is a diagram of the "optical flow-event" collaborative time alignment module of the present invention. DETAILED DESCRIPTION
[0061] In this embodiment, a blurred video super-resolution method based on event data drive is proposed. This method utilizes the high temporal resolution characteristics of event data without motion blur and constructs a blurred video super-resolution network to generate clear high-resolution video frames. The main features are that the event data is divided into intra-frame events and inter-frame events, and a "frame-event" collaborative deblurring design and an "optical flow-event" collaborative alignment design are adopted. Figure 1 As shown, the specific steps of this method are as follows:
[0062] Step 1: Obtain training video image dataset and its corresponding event sequence, and characterize the event sequence:
[0063] Step 1.1 Obtain training video image dataset :
[0064] Obtaining clear high-resolution video image datasets ,in, Indicates the A clear, high-resolution image. , is the total number of high-resolution images; in this example, the total number of images used for neural network training is .
[0065] For clear high-resolution video images Perform blur and resolution reduction to obtain a blurred low-resolution video image set ,in, Indicates the A blurred low-resolution image; in this example, the time averaging method is used to Blur processing and bilinear interpolation algorithm are used to Reduce resolution;
[0066] make Represents the training video image dataset.
[0067] Step 1.2 Characterize the event sequence:
[0068] Acquire high-resolution video image sets Intra-frame event sequence and inter-frame event sequences ,in, and Respectively represent A clear, high-resolution image The corresponding intra-frame event sequence and inter-frame event sequence; In this example, the event camera simulator Vid2E is used to directly convert the input video image set Simulate its event data.
[0069] right and Perform resolution reduction processing respectively to obtain a low-resolution video image set Corresponding intra-frame event sequence and inter-frame event sequences ,in, and Respectively represent A blurry low-resolution image The corresponding intra-frame event sequence and inter-frame event sequence; in this example, the The same bilinear interpolation algorithm for resolution reduction and Reduce resolution.
[0070] right Perform voxel representation to obtain a low-resolution video image set Corresponding intra-frame event voxel sequence ,in, express The event voxels within the frame, represents the number of event voxel channels, Indicates the image height, Indicates the image width; in this example, the total number of images during neural network training is , , ;
[0071] For event sequence Perform voxel representation to obtain a low-resolution video image set Corresponding intra-frame event voxel sequence ,in, express The inter-frame event voxels.
[0072] Step 2: Construct a fuzzy video super-resolution neural network, such as Figure 1 As shown, it includes: feature extraction layer, "frame-event" collaborative deblurring module, "optical flow-event" collaborative time alignment module, and frame reconstruction module;
[0073] Step 2.1, the feature extraction layer includes: image feature extraction layer, intra-frame event feature extraction layer and inter-frame event extraction layer, and respectively 、 and Perform feature extraction and obtain the corresponding Blurry frame features , No. Intra-frame event features Hedi Inter-frame event features ,in, Represents the extracted feature dimension; in this example, .
[0074] Step 2.2, such as Figure 2 As shown in Figure 2, the “frame-event” collaborative deblurring module consists of a frame and event preprocessing layer, a frame-guided event deblurring layer, an event-guided frame deblurring layer, and a fully connected fusion layer. and Processing, the corresponding Deblurred frame features Hedi Deblurred intra-frame event features ;
[0075] Step 2.2.1, the frame and event preprocessing layer consists of the frame branch channel attention module and the event branch channel attention module, and and Perform context semantic feature aggregation processing and obtain the first Fuzzy frame semantic features Hedi Intra-frame event semantic features .
[0076] Step 2.2.2: The frame-guided event deblurring layer uses equations (1) to (4) to obtain the first Deblurred intra-frame event features :
[0077] (1)
[0078] (2)
[0079] (3)
[0080] (4)
[0081] In formula (1) to formula (4), for The query matrix, for The query weight matrix, for The key matrix, for The key weight matrix, for The value matrix of for The value weight matrix of represents the matrix transpose operation, and softmax represents the activation function.
[0082] Step 2.2.3: The event-guided frame deblurring layer uses equations (5) to (8) to obtain the first Deblurred frame features :
[0083] (5)
[0084] (6)
[0085] (7)
[0086] (8)
[0087] In formula (5) to formula (8), for The query matrix, for The query weight matrix, for The key matrix, for The key weight matrix, for The value matrix of for The value weight matrix of .
[0088] Step 2.2.4, the fully connected fusion layer consists of a frame branch fully connected module and an event branch fully connected module, and and Processing, the corresponding and .
[0089] Step 2.3, such as Figure 3 As shown in Figure 2, the “optical flow-event” collaborative alignment module includes: an optical flow-guided alignment module, an event-guided alignment module, and a deformable convolutional layer. low-resolution images Hedi low-resolution images Process it and get Feature maps ;
[0090] Step 2.3.1, the optical flow guided alignment module is composed of Layer downsampling convolution layer and The upsampling convolution layers are connected alternately, and the and Perform optical flow estimation and get the No. Optical flow In this example, ;
[0091] Using formula (9) Optical flow Hedi Feature maps conduct Transformation, thus obtaining the Features from optical flow alignment :
[0092] (9)
[0093] In formula (9), Represents image distortion transformation; when i=1, let .
[0094] Step 2.3.2, the event-guided alignment module is composed of Convolutional layers and GeLU activation layers are used to Inter-frame event features Hedi Feature maps Process it and get Features from event alignment In this example, ;
[0095] Step 2.2.3, 、 、 、 、 and After splicing along the channel, it is input into the deformable convolution layer for processing to obtain the Feature maps .
[0096] Step 2.4: The frame reconstruction module is composed of multiple deconvolution layers and upsampling layers connected in series. low-resolution images Hedi Feature maps After processing, we get Super-resolution images ; Thus we get the super-resolution video set .
[0097] Step 3: Use the loss function of formula (10) :
[0098] (10)
[0099] In formula (10), is a non-negative constant; in this example, .
[0100] Step 4: Use the gradient descent method to train the blurred video super-resolution neural network and calculate the loss function To update the network parameters, when the number of training iterations reaches the set number or the loss function When convergence occurs, the training stops, and the optimal blurry video super-resolution model is obtained; it is used to process blurry low-resolution video images to obtain corresponding clear high-resolution video images.
[0101] In this embodiment, an electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0102] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are executed.
Claims
1. A fuzzy video super-resolution method driven by event data, characterized in that: The steps are as follows: Step 1: Obtain training video image dataset and its corresponding event sequence, and characterize the event sequence: Step 1.1 Obtain training video image dataset : Obtaining clear high-resolution video image datasets ,in, Indicates the A clear, high-resolution image. , is the total number of high-resolution images; For clear high-resolution video images Perform blur and resolution reduction to obtain a blurred low-resolution video image set ,in, Indicates the A blurry low-resolution image; make Represents the training video image dataset; Step 1.2 Characterize the event sequence: Acquire high-resolution video image sets Intra-frame event sequence and inter-frame event sequences ,in, and Respectively represent A clear, high-resolution image The corresponding intra-frame event sequence and inter-frame event sequence; right and Perform resolution reduction processing respectively to obtain a low-resolution video image set Corresponding intra-frame event sequence and inter-frame event sequences ,in, and Respectively represent A blurry low-resolution image The corresponding intra-frame event sequence and inter-frame event sequence; right Perform voxel representation to obtain a low-resolution video image set Corresponding intra-frame event voxel sequence ,in, express The event voxels within the frame, represents the number of event voxel channels, Indicates the image height, Indicates the image width; For event sequence Perform voxel representation to obtain a low-resolution video image set Corresponding intra-frame event voxel sequence ,in, express inter-frame event voxels; Step 2: Construct a blurred video super-resolution neural network, including a feature extraction layer, a "frame-event" collaborative deblurring module, an "optical flow-event" collaborative time alignment module, and a frame reconstruction module. Step 2.1, the feature extraction layer includes: image feature extraction layer, intra-frame event feature extraction layer and inter-frame event extraction layer, and respectively 、 and Perform feature extraction and obtain the corresponding Blurry frame features , No. Intra-frame event features Hedi Inter-frame event features ,in, Represents the extracted feature dimension; Step 2.2, the "frame-event" collaborative deblurring module includes: a frame and event preprocessing layer, a frame-guided event deblurring layer, an event-guided frame deblurring layer and a fully connected fusion layer, and respectively and Processing, the corresponding Deblurred frame features Hedi Deblurred intra-frame event features ; Step 2.3, the "optical flow-event" collaborative alignment module includes: an optical flow-guided alignment module, an event-guided alignment module and a deformable convolution layer, and low-resolution images Hedi low-resolution images Process it and get Feature maps ; Step 2.3.1: The optical flow guided alignment module is composed of Layer downsampling convolution layer and The upsampling convolution layers are connected alternately, and the and Perform optical flow estimation and get the No. Optical flow ; Using formula (9) Optical flow Hedi Feature maps conduct Transformation, thus obtaining the Features from optical flow alignment : (9) In formula (9), Represents image distortion transformation; when i=1, let ; Step 2.3.2: The event-guided alignment module is composed of Convolutional layers and GeLU activation layers are used to Inter-frame event features Hedi Feature maps Process it and get Features from event alignment ; Step 2.3.3, 、 、 、 、 and After splicing along the channel, it is input into the deformable convolution layer for processing to obtain the Feature maps ; Step 2.4, the frame reconstruction module is composed of multiple deconvolution layers and upsampling layers connected in series, and low-resolution images Hedi Feature maps After processing, we get Super-resolution images ; Thus we get the super-resolution video set ; Step 3: Use the loss function of formula (10) : (10) In formula (10), is a non-negative constant; Step 4: Use the gradient descent method to train the blurred video super-resolution neural network and calculate the loss function To update the network parameters, when the number of training iterations reaches the set number or the loss function When convergence occurs, the training stops, and the optimal blurry video super-resolution model is obtained; it is used to process blurry low-resolution video images to obtain corresponding clear high-resolution video images.
2. The method for fuzzy video super-resolution based on event data drive according to claim 1, characterized in that: The step 2.2 is carried out as follows: Step 2.2.1, the frame and event preprocessing layer consists of a frame branch channel attention module and an event branch channel attention module, and respectively and Perform context semantic feature aggregation processing and obtain the first Fuzzy frame semantic features Hedi Intra-frame event semantic features ; Step 2.2.2: The frame-guided event deblurring layer uses equations (1) to (4) to obtain the first Deblurred intra-frame event features : (1) (2) (3) (4) In formula (1) to formula (4), for The query matrix, for The query weight matrix, for The key matrix, for The key weight matrix, for The value matrix of for The value weight matrix of represents the matrix transpose operation, and softmax represents the activation function; Step 2.2.3: The event-guided frame deblurring layer uses equations (5) to (8) to obtain the first Deblurred frame features : (5) (6) (7) (8) In formula (5) to formula (8), for The query matrix, for The query weight matrix, for The key matrix, for The key weight matrix, for The value matrix of for The value weight matrix of Step 2.2.4, the fully connected fusion layer is composed of a frame branch fully connected module and an event branch fully connected module, and and Processing, the corresponding and .
3. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the blurred video super-resolution method according to any one of claims 1-2, and the processor is configured to execute the program stored in the memory.
4. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the blurred video super-resolution method according to any one of claims 1 to 2 are executed.
Citation Information
Patent Citations
A super-resolution implementation method and device of an image frame
CN113920010A
Video deblurring method based on event data driving
CN114463218A