Hierarchical Video Coding for Region-of-Interest Telepresence

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data streaming technologies struggle with efficiently navigating and rendering data sets, particularly in telepresence applications, due to difficulties in changing streams on-the-fly, skipping to different portions, and managing constrained datalinks, which often require sacrificing quality or introducing delay.

Innovation Solution

A hierarchical data coding scheme using SMPTE VC-6 and MPEG-5, Part-2 standards, which encodes data at multiple layers, allowing region-of-interest decoding and pan-or-zoom functionality by transmitting and decoding only the spatial resolutions and regions that a user is viewing, rather than the entire frame.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If Intra or Group of Pictures compression is used to stream video and images, then the entire frame or image can be transported and stored, but only a part of the image or cut-out cannot be efficiently accessed or streamed

Engineering Contradiction:
ImproveRegion of interest accessVSAvoidData transmission volume
Core Design Contradiction:
Ease of operationVSQuantity of substance

Solution Approach 1:

The patent divides the video frame into multiple spatial regions or tiles, allowing selective transmission and decoding of only the regions of interest. This segmentation enables efficient access to specific portions of the image without requiring transmission of the entire frame, directly resolving the contradiction between ease of region access and data transmission volume.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts and transmits only the necessary portions of the video stream corresponding to regions of interest, rather than transmitting the complete frame. This extraction approach reduces the quantity of data transmitted while maintaining the ability to access and display relevant information, addressing both aspects of the contradiction.

Inventive Principle:
Principle #2Taking out (Extraction)

2Adaptability or versatility

If monolithic compression encodes the entire frame as a single block, then the encoding process is simple, but the data cannot be efficiently navigated or changed on-the-fly

Engineering Contradiction:
ImproveOn-the-fly stream changingVSAvoidCoding scheme complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the video stream into multiple independent regions or tiles that can be selectively accessed and decoded. This segmentation enables on-the-fly changes in stream content by allowing the system to switch between different regions without requiring complete frame re-encoding, thereby achieving adaptability while managing complexity through modular design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a dynamic coding scheme where the selected regions or tiles can be changed in real-time based on user needs. This dynamic capability allows the system to adapt to varying requirements by switching between different spatial regions or quality levels without requiring complex re-encoding operations, balancing adaptability with manageable complexity.

Inventive Principle:
Principle #15Dynamics

3Reliability

If full resolution video is transmitted to ensure quality, then image quality is maintained, but bandwidth consumption increases and constrained datalinks cannot support the required data rate

Engineering Contradiction:
ImproveImage qualityVSAvoidBandwidth consumption
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent applies different quality levels to different spatial regions of the video frame. Regions of interest are transmitted at high resolution to maintain image quality, while other regions can be transmitted at lower resolution or omitted entirely. This local quality approach ensures that image quality is maintained where needed while significantly reducing overall bandwidth consumption, resolving the contradiction between reliability and quantity of data transmitted.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent transmits only the partial information necessary for regions of interest rather than the complete full-resolution frame. This partial action approach reduces bandwidth consumption by transmitting only essential data while maintaining adequate image quality in the areas that matter, effectively resolving the contradiction between quality maintenance and bandwidth efficiency.

Inventive Principle:
Principle #16Partial or excessive action

4Productivity

If the entire frame must be decoded to access any part of the image, then decoding simplicity is maintained, but time is wasted decoding unnecessary regions

Engineering Contradiction:
ImproveDecoding efficiencyVSAvoidTime for unnecessary decoding
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent segments the video frame into independent regions or tiles that can be decoded separately. This segmentation allows the decoder to process only the regions of interest, eliminating the time waste associated with decoding entire frames when only portions are needed. By maintaining decoding simplicity at the regional level while achieving frame-level efficiency, the system improves productivity without sacrificing decoding ease.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250287023A1Use of hierarchical video and image coding for telepresence
Publication Date: 2025.09.11 V NOVA INT LTD
  • US20250287023A1 patent drawing
  • US20250287023A1 patent drawing
  • US20250287023A1 patent drawing

AI summary

A medical telepresence system comprising: an interface to receive a plurality of data feeds from a live medical procedure, at least one data feed comprising a video signal capturing the live medical procedure; a hierarchical encoder to encode the plurality of data feeds using a first tier-based hierarchical data coding scheme, wherein encoded data from the hierarchical encoder is decodable by a first set of computing devices for viewing, the first set of computing devices being communicatively coupled to the hierarchical encoder using a first network connection; a transcoder to convert from the first tier-based hierarchical data coding scheme to a second tier-based hierarchical data coding scheme, wherein encoded data from the transcoder is receivable by a second set of computing devices for viewing, the second set of computing devices being communicatively coupled to the transcoder using a second network connection, the second network connection being of a lower quality than the first network connection; and a recorder to store the output of the hierarchical encoder as a set of tier-based files for later retrieval, wherein each of the set of tier-based files represent different levels of quality.