Deep Neural Video Super-Resolution Without Optical Flow Artifacts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video super-resolution techniques often rely on optical flow models, which are prone to inaccuracies leading to artifacts and inconsistencies between neighboring frames, especially in real-time video streaming scenarios.
Innovation Solution
A neural network architecture utilizing Residual-in-Residual Dense Blocks (RRDBs) and Convolutional Gated Recurrent Units (GRUs) processes video frames to increase resolution without optical flow, ensuring temporal consistency and robustness across frames.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If optical flow models are used for video super-resolution, then resolution enhancement can be achieved, but temporal consistency and frame coherence deteriorate due to inaccuracies in optical flow estimation
Solution Approach 1:
The patent extracts and removes the optical flow model component from the video super-resolution system. By eliminating this problematic component, the patent avoids the temporal inconsistency issues it causes while maintaining resolution enhancement capabilities through alternative methods that do not rely on optical flow estimation.
Solution Approach 2:
The patent introduces a transformer-based encoder as an intermediary component that processes video frames without relying on optical flow. This intermediary mechanism enables resolution enhancement while maintaining temporal consistency by capturing temporal dependencies through self-attention mechanisms rather than optical flow estimation.
2Manufacturing precision
If existing video super-resolution techniques are used, then resolution can be increased, but computational complexity and processing time increase significantly
Solution Approach 1:
The patent segments the video processing task into frame-level processing with a transformer encoder that handles temporal dependencies efficiently. By dividing the computation into manageable segments and using parallel processing capabilities of transformers, the patent reduces overall computational complexity compared to traditional sequential methods.
Solution Approach 2:
The patent changes the computational parameters by using transformer architecture with self-attention mechanisms that can process multiple frames simultaneously. This parameter change enables efficient computation by leveraging parallel processing and reducing the sequential dependency chain, thereby lowering computational complexity while maintaining high resolution output.
3Loss of time
If real-time video processing is implemented, then latency is reduced, but processing accuracy and quality may deteriorate
Solution Approach 1:
The patent applies preliminary action by pre-processing video frames through the transformer encoder before final resolution enhancement. This preliminary processing captures temporal dependencies and prepares the data for efficient real-time processing, enabling both low latency and high accuracy by performing necessary computations in advance.
Solution Approach 2:
The patent implements dynamics by using a flexible transformer architecture that can adapt processing depth and complexity based on real-time requirements. The system dynamically adjusts the number of transformer layers and attention mechanisms activated, allowing it to maintain processing accuracy while reducing latency when needed through selective computation.
Data Source
AI summary
Methods and systems for obtaining an input video sequence comprising input video frames; determining i) an input resolution of input video frames and ii) a target output resolution of the plurality of input video frames, wherein the target output resolution is higher than the input resolution; and processing the input video sequence using a neural network to generate an output video sequence, comprising, for each of the plurality of input video frames: processing the input video frame to generate an output video frame having the target output resolution, comprising processing the input video frame using a subnetwork of the neural network corresponding to the input resolution of the plurality of input video frames, the neural network configured to process input video frames having one of a set of possible input resolutions and to generate output video frames having one of a set of possible output resolutions.


