Scalable Video Encoding with Motion Vector Upscaling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Contemporary video encoding technologies face challenges in efficiently encoding both lower- and higher-resolution images within a single video stream, leading to higher bitrate consumption and synchronization issues between lower-resolution overviews and higher-resolution regions of interest (ROIs).
Innovation Solution
The method involves obtaining images from a camera, identifying regions of interest, and encoding a set of video frames that include a lower-resolution overview, a no-display frame with motion vectors for upscaling the ROIs, and higher-resolution ROI frames. This approach allows for efficient encoding and synchronization within a single video stream.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If higher-resolution JPEG crops of ROIs are produced and sent in parallel with lower-resolution overview, then higher-resolution ROI images are provided, but bitrate consumption increases and synchronization issues occur
Solution Approach 1:
The patent merges the lower-resolution overview and higher-resolution ROI images into a single unified video stream using scalable video coding. The base layer contains the lower-resolution overview, while enhancement layers contain the higher-resolution ROI data. This combining approach eliminates the need for separate parallel streams, reducing overall bitrate consumption while maintaining synchronization through unified encoding.
Solution Approach 2:
The patent implements a nested structure where enhancement layer data (higher-resolution ROI) is embedded within the overall video stream structure that already contains the base layer (lower-resolution overview). The enhancement layers are nested within the scalable video coding framework, allowing the ROI information to be contained within the same bitstream container as the overview, thereby reducing redundant data transmission and synchronization overhead.
2Measurement precision
If higher-resolution JPEG crops of ROIs are produced and sent in parallel with lower-resolution overview, then higher-resolution ROI images are provided, but synchronization issues occur between the two streams
Solution Approach 1:
By merging the overview and ROI images into a single unified video stream with integrated timing and synchronization mechanisms, the patent eliminates the synchronization issues that arise from parallel separate streams. The scalable video coding framework ensures that both base layer and enhancement layers are encoded and transmitted with consistent timing information, guaranteeing reliable synchronization between lower-resolution overview and higher-resolution ROI data.
3Adaptability or versatility
If contemporary video encoding technologies are used to encode both lower- and higher-resolution images, then both resolutions are provided, but device complexity increases
Solution Approach 1:
The patent changes the encoding parameter structure by using scalable video coding with base layers and enhancement layers instead of encoding separate resolution streams. This parameter change allows a single encoding device to produce multiple resolution outputs by varying the decoding layer selection, rather than requiring separate encoding paths for each resolution, thereby reducing device complexity while maintaining multi-resolution capability.
Data Source
AI summary
A method for encoding a video stream includes obtaining images of a scene captured by a camera at a first resolution; identifying regions of interest (ROIs) in an image; adding, as part of an encoded video stream, a first video frame encoding at least part of the image at a second resolution lower than the first resolution; adding a second video frame marked as a no-display frame, and being an inter-frame referencing the first video frame with motion vectors for upscaling of the ROIs; adding a third video frame encoding the ROIs at a third resolution higher than the second resolution, and being an inter-frame referencing the second video frame.


