Video Transcoding Spatial Importance Resolution Enhancement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video transcoding methods face challenges in efficiently enhancing image resolution due to the high computational complexity of deep learning-based techniques, which hinders their application in real-time video streaming and large-scale video storage, especially when balancing image quality and processing speed.
Innovation Solution
The approach involves determining spatial and temporal importance levels in video frames, applying deep neural network-based techniques for high-importance regions and less complex methods for low-importance regions, allowing for efficient resolution enhancement while maintaining high image quality and controlling processing speed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If deep learning-based resolution enhancement techniques are applied to the entire video, then image quality is improved, but computational complexity increases significantly
Solution Approach 1:
The patent applies different resolution enhancement techniques to different spatial regions based on their importance. High-importance regions (e.g., foreground objects, text) receive deep learning-based enhancement, while low-importance regions (e.g., uniform backgrounds) receive simpler enhancement methods. This resolves the contradiction by localizing the complex processing only where quality improvement is most valuable.
Solution Approach 2:
The patent segments the video content into regions of different importance levels using image segmentation techniques. This allows the system to identify and separately process high-importance regions with computationally intensive deep learning methods, while applying lighter processing to other regions, thereby balancing image quality improvement with computational complexity control.
2Manufacturing precision
If deep learning-based resolution enhancement techniques are applied to the entire video, then image quality is improved, but processing speed decreases
Solution Approach 1:
The patent applies different resolution enhancement techniques to different spatial regions based on their importance. High-importance regions (e.g., foreground objects, text) receive deep learning-based enhancement, while low-importance regions (e.g., uniform backgrounds) receive simpler enhancement methods. This resolves the contradiction by localizing the complex processing only where quality improvement is most valuable.
Solution Approach 2:
The patent applies deep learning-based enhancement selectively to only the necessary portions of the video (high-importance regions) rather than the entire video. This partial application of the complex technique maintains processing speed while still achieving quality improvement where it matters most, resolving the contradiction between image quality and processing speed.
3Manufacturing precision
If uniform resolution enhancement is applied to all video regions, then image quality is improved, but computational resources are wasted on low-importance regions
Solution Approach 1:
The patent applies different resolution enhancement techniques to different spatial regions based on their importance. High-importance regions (e.g., foreground objects, text) receive deep learning-based enhancement, while low-importance regions (e.g., uniform backgrounds) receive simpler enhancement methods. This resolves the contradiction by localizing the complex processing only where quality improvement is most valuable.
Solution Approach 2:
The patent segments the video content into regions of different importance levels using image segmentation techniques. This allows the system to identify and separately process high-importance regions with computationally intensive deep learning methods, while applying lighter processing to other regions, thereby balancing image quality improvement with computational complexity control.
Data Source
AI summary
Methods and apparatuses for video transcoding based on spatial or temporal importance include: in response to receiving an encoded video bitstream, decoding a picture from the encoded video bitstream; determining a first level of spatial importance for a first region of a background of the picture based on an image segmentation technique; applying to the first region a first resolution-enhancement technique associated with the first level of spatial importance for increasing resolution of the first region by a scaling factor, wherein the first resolution-enhancement technique is selected from a set of resolution-enhancement techniques having different computational complexity levels; and encoding the first region using a video coding standard.


