Region-of-Interest Encoding for Video Telephony

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video telephony technologies lack the ability to remotely control and enhance the encoding quality of specific regions of interest within video streams, limiting the clarity and interaction during video conferencing and streaming applications.

Innovation Solution

Implementing region-of-interest (ROI) processing techniques that allow devices to define, transmit, and apply preferential encoding to specific areas of a video scene, enabling remote control of ROI encoding and decoding to improve image quality and user interaction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If preferential encoding of ROI is implemented, then image quality of selected regions is improved, but transmission bandwidth and encoding complexity increase

Engineering Contradiction:
Improveencoding qualityVSAvoidencoding complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent implements different encoding qualities for different regions of the video frame by defining a region of interest (ROI) that receives preferential encoding treatment. The encoder applies higher quality encoding parameters specifically to the ROI area while using standard encoding for non-ROI areas, thereby improving local image quality where needed without uniformly increasing overall encoding complexity

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The video frame is segmented into at least two distinct regions: an ROI that requires higher quality encoding and a non-ROI area with standard encoding. This segmentation allows the encoding system to focus computational resources on the important region while maintaining efficiency in other areas, resolving the contradiction between quality and complexity

Inventive Principle:
Principle #1Segmentation

2Ease of operation

If ROI information is transmitted from recipient to sender, then remote control of encoding is enabled, but communication overhead increases

Engineering Contradiction:
Improveremote control capabilityVSAvoiddata transmission volume
Core Design Contradiction:
Ease of operationVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential ROI information (such as ROI boundaries, coordinates, or region masks) from the full video content and transmits this extracted information between devices. By taking out only the necessary control parameters rather than transmitting complete video data, the system enables remote encoding control while minimizing the quantity of transmitted data

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces ROI information as an intermediary data structure that mediates between the recipient's encoding preferences and the sender's encoding process. This intermediary representation allows indirect control of remote encoding without requiring direct manipulation of the full video stream, reducing communication overhead

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS8977063B2Region-of-interest extraction for video telephony
Publication Date: 2015.03.10 QUALCOMM INC
  • US8977063B2 patent drawing
  • US8977063B2 patent drawing
  • US8977063B2 patent drawing

AI summary

The disclosure is directed to techniques for region-of-interest (ROI) processing for video telephony (VT) applications. According to the disclosed techniques, a recipient device defines ROI information for video information transmitted by a sender device, i.e., far-end video information. The recipient device transmits the ROI information to the sender device. Using the ROI information transmitted by the recipient device, the sender device applies preferential encoding to an ROI within a video scene. ROI extraction may be applied to process a user description of a region of interest (ROI) to generate information specifying the ROI based on the description. The user description may be textual, graphical, or speech-based. An extraction module applies appropriate processing to generated the ROI information from the user description. The extraction module may locally reside with a video communication device, or reside in a distinct intermediate server configured for ROI extraction.