Panoramic Video Streaming with Target Area Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional panoramic video streaming faces challenges with high bitrate requirements, limited bandwidth in networks, and the need for quick region of interest changes to prevent user experience delays and sickness, especially in video on demand scenarios where processing and bandwidth capabilities are overwhelmed.
Innovation Solution
The system generates and encodes multiple target areas of a panoramic video stream, allowing for a high-quality region of interest and a lower-quality full field of view to be streamed together, using different bitrates and compression parameters, and dynamically updates the stream to ensure efficient delivery over high latency and low bandwidth networks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If the entire panoramic field of view is streamed at high resolution, then the region of interest can be changed in response to user movement, but the bitrate requirement becomes excessively high and consumes large amounts of network bandwidth
Solution Approach 1:
The panoramic video stream is divided into multiple target areas covering different positions of the field of view. Each target area is encoded separately with appropriate bitrate allocation, allowing high quality for regions of interest while reducing overall bitrate by not uniformly encoding the entire 360-degree view at maximum resolution
Solution Approach 2:
Different encoding parameters and bitrates are applied to different target areas based on their importance. Regions that are more likely to be viewed (based on user behavior patterns) receive higher quality encoding, while less important regions use lower bitrate encoding, optimizing the balance between image quality and bandwidth consumption
2Adaptability or versatility
If conventional panoramic video streaming is used, then the entire field of view is available, but the processing and bandwidth capabilities of the central server are overwhelmed in video on demand scenarios
Solution Approach 1:
Multiple target areas are pre-encoded and prepared in advance during the video encoding process. When a user requests panoramic video, the system can quickly deliver the pre-prepared target areas without requiring real-time processing of the entire panoramic stream, significantly reducing server processing load during playback
Solution Approach 2:
Instead of providing the entire panoramic field of view at full quality to all users, the system encodes multiple overlapping target areas that cover the full 360 degrees. Each user receives only the specific target area(s) relevant to their viewing direction, delivering partial content that is sufficient for the user experience while reducing overall processing requirements
3Ease of operation
If the region of interest changes quickly in response to user movement, then user experience is improved, but the timing synchronization and processing requirements increase
Solution Approach 1:
Multiple target areas are pre-encoded and buffered in advance, covering different regions of the panoramic view. When the user moves and the region of interest changes, the system can immediately switch to delivering the appropriate pre-encoded target area without requiring time-consuming real-time encoding or processing, enabling quick responsiveness to user movement
Solution Approach 2:
Multiple target areas with different temporal information are combined into a single delivery system. The content delivery network can select and deliver the most appropriate target area based on current user viewing direction, while all target areas are synchronized to the same time reference, enabling seamless switching between regions without timing issues
Data Source
AI summary
An apparatus comprising an interface and a processor. The interface may (a) receive a panoramic video, (b) present a plurality of encoded target areas to a network and (c) present a downscaled panoramic video to the network. The processor may (a) generate target areas by cropping sections of the panoramic video, (b) encode each of the target areas using first parameters and (c) use second parameters to encode the downscaled panoramic video. Encoding using the first parameters generates a different bitrate than using the second parameters. The target areas cover an entire field of view of the panoramic video. Each of the target areas covers a different position of the panoramic video. The network presents one of the encoded target areas to a playback device corresponding to a region of interest of the playback device. The network presents the downscaled panoramic video to each playback device.


