Hierarchical Video Coding for Region-of-Interest Telepresence
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data streaming technologies struggle with efficiently navigating and rendering data sets, particularly in telepresence applications, due to difficulties in changing streams on-the-fly, skipping to different portions, and managing constrained datalinks, which often require sacrificing quality or introducing delay.
Innovation Solution
A hierarchical data coding scheme using SMPTE VC-6 and MPEG-5, Part-2 standards, which encodes data at multiple layers, allowing region-of-interest decoding and pan-or-zoom functionality by transmitting and decoding only the spatial resolutions and regions that a user is viewing, rather than the entire frame.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If Intra or Group of Pictures compression is used to stream video and images, then the entire frame or image can be transported and stored, but only a part of the image or cut-out cannot be efficiently accessed or streamed
Solution Approach 1:
The patent divides the video frame into multiple spatial regions or tiles, allowing selective transmission and decoding of only the regions of interest. This segmentation enables efficient access to specific portions of the image without requiring transmission of the entire frame, directly resolving the contradiction between ease of region access and data transmission volume.
Solution Approach 2:
The patent extracts and transmits only the necessary portions of the video stream corresponding to regions of interest, rather than transmitting the complete frame. This extraction approach reduces the quantity of data transmitted while maintaining the ability to access and display relevant information, addressing both aspects of the contradiction.
2Adaptability or versatility
If monolithic compression encodes the entire frame as a single block, then the encoding process is simple, but the data cannot be efficiently navigated or changed on-the-fly
Solution Approach 1:
The patent segments the video stream into multiple independent regions or tiles that can be selectively accessed and decoded. This segmentation enables on-the-fly changes in stream content by allowing the system to switch between different regions without requiring complete frame re-encoding, thereby achieving adaptability while managing complexity through modular design.
Solution Approach 2:
The patent implements a dynamic coding scheme where the selected regions or tiles can be changed in real-time based on user needs. This dynamic capability allows the system to adapt to varying requirements by switching between different spatial regions or quality levels without requiring complex re-encoding operations, balancing adaptability with manageable complexity.
3Reliability
If full resolution video is transmitted to ensure quality, then image quality is maintained, but bandwidth consumption increases and constrained datalinks cannot support the required data rate
Solution Approach 1:
The patent applies different quality levels to different spatial regions of the video frame. Regions of interest are transmitted at high resolution to maintain image quality, while other regions can be transmitted at lower resolution or omitted entirely. This local quality approach ensures that image quality is maintained where needed while significantly reducing overall bandwidth consumption, resolving the contradiction between reliability and quantity of data transmitted.
Solution Approach 2:
The patent transmits only the partial information necessary for regions of interest rather than the complete full-resolution frame. This partial action approach reduces bandwidth consumption by transmitting only essential data while maintaining adequate image quality in the areas that matter, effectively resolving the contradiction between quality maintenance and bandwidth efficiency.
4Productivity
If the entire frame must be decoded to access any part of the image, then decoding simplicity is maintained, but time is wasted decoding unnecessary regions
Solution Approach 1:
The patent segments the video frame into independent regions or tiles that can be decoded separately. This segmentation allows the decoder to process only the regions of interest, eliminating the time waste associated with decoding entire frames when only portions are needed. By maintaining decoding simplicity at the regional level while achieving frame-level efficiency, the system improves productivity without sacrificing decoding ease.
Data Source
AI summary
A medical telepresence system comprising: an interface to receive a plurality of data feeds from a live medical procedure, at least one data feed comprising a video signal capturing the live medical procedure; a hierarchical encoder to encode the plurality of data feeds using a first tier-based hierarchical data coding scheme, wherein encoded data from the hierarchical encoder is decodable by a first set of computing devices for viewing, the first set of computing devices being communicatively coupled to the hierarchical encoder using a first network connection; a transcoder to convert from the first tier-based hierarchical data coding scheme to a second tier-based hierarchical data coding scheme, wherein encoded data from the transcoder is receivable by a second set of computing devices for viewing, the second set of computing devices being communicatively coupled to the transcoder using a second network connection, the second network connection being of a lower quality than the first network connection; and a recorder to store the output of the hierarchical encoder as a set of tier-based files for later retrieval, wherein each of the set of tier-based files represent different levels of quality.


