Selective Object Zooming in Streaming Video via Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current streaming video technologies limit user flexibility in viewing options due to bandwidth constraints, resulting in excessive bandwidth consumption and unsatisfactory latency when switching between different video streams.
Innovation Solution
Implementing systems and methods that allow users to select and zoom into specific objects of interest within a video stream with minimal latency, using metadata to identify objects and retrieve associated zoomed streams, enabling local processing and simultaneous display of zoomed and un-zoomed content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple video streams are provided to allow user selection of different views, then user flexibility is improved, but bandwidth consumption increases excessively
Solution Approach 1:
The video content is segmented into a base layer and enhancement layers. The base layer contains the main video stream, while enhancement layers contain additional view information. This segmentation allows the system to provide multiple views without transmitting complete separate streams for each view, thereby reducing overall bandwidth consumption while maintaining user flexibility.
Solution Approach 2:
The patent introduces a new dimension of view selection by encoding multiple potential views within the same base video stream timeline. Instead of providing separate streams for different views (consuming more bandwidth), the system packs multiple view possibilities into one stream with associated metadata, allowing users to select different views without increasing bandwidth usage.
2Loss of energy
If a user requests different streams at different times, then bandwidth consumption is limited, but latency increases between stream request and display
Solution Approach 1:
The system performs preliminary action by pre-encoding multiple view possibilities into the base video stream and preparing enhancement layer data in advance. When a user requests a specific view, the client can quickly assemble the requested view by combining the already-available base stream with the corresponding enhancement layer, significantly reducing latency compared to requesting entirely new streams.
Solution Approach 2:
The patent introduces an intermediary mechanism in the form of enhancement layers that act as a bridge between the base video stream and the user's desired view. These enhancement layers contain differential information that can be rapidly applied to the base stream to generate different views, serving as a fast intermediary solution that reduces latency without requiring full stream retransmission.
3Measurement precision
If zoomed streams are provided for objects of interest, then viewing quality is improved, but bandwidth consumption increases
Solution Approach 1:
The patent applies local quality enhancement by providing high-quality zoomed views only for specific regions of interest within the video, rather than enhancing the entire video stream. The enhancement layers contain compressed differential data for specific object regions, allowing the client to reconstruct high-quality zoomed views locally without receiving and processing entire high-resolution streams, thereby reducing bandwidth consumption while maintaining viewing quality for areas of interest.
Data Source
AI summary
Systems and methods are described for enabling a consumer of streaming video to obtain different views of the video, such as zoomed views of one or more objects of interest. In an exemplary embodiment, a client device receives an original video stream along with data identifying objects of interest and their spatial locations within the original video. In one embodiment, in response to user selection of an object of interest, the client device switches to display of a cropped and scaled version of the original video to present a zoomed video of the object of interest. The zoomed video tracks the selected object even as the position of the selected object changes with respect to the original video. In some embodiments, the object of interest and the appropriate zoom factor are both selected with a single expanding-pinch gesture on a touch screen.


