Panoramic Video Audio Zone Processing for Speaker Clarity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In videoconferencing, existing systems fail to effectively enhance audio clarity for specific speakers based on viewer interest, leading to suboptimal audio experience for remote participants.

Innovation Solution

The method involves associating audio zones with a camera's field of view and processing audio based on a region of interest (ROI) for each client device, enhancing audio for the relevant zone and reducing or muting others, using a videoconferencing server that processes audio streams from a microphone array and adjusts audio directionality according to viewer interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If audio from all speakers is transmitted equally, then all participants can hear all speakers, but audio clarity for specific speakers is poor when viewers are focused on different regions

Engineering Contradiction:
Improveaudio clarityVSAvoidaudio customization per viewer
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies local quality by differentiating audio transmission based on spatial location. Audio is processed and enhanced selectively for specific audio zones corresponding to regions of interest identified through viewer interaction data. This allows each viewer to receive optimized audio clarity for the speaker they are focusing on, rather than receiving uniform audio from all speakers.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The audio stream is segmented into multiple audio zones corresponding to different spatial regions. Each zone is processed independently based on the region of interest detected for each viewer. This segmentation enables selective enhancement of audio from specific zones while maintaining the ability to transmit audio from multiple zones to different viewers simultaneously.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If audio enhancement is applied to all zones, then audio clarity improves for all speakers, but background noise and irrelevant audio increases

Engineering Contradiction:
Improveaudio clarityVSAvoidbackground noise
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

Instead of uniformly enhancing all audio zones, the system applies local quality by enhancing only the audio from zones corresponding to the detected region of interest for each viewer. Audio from other zones is either not enhanced or is muted, thereby reducing background noise and irrelevant audio for viewers focused on specific speakers.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system extracts and enhances only the audio components from specific audio zones that correspond to the region of interest. Audio from zones outside the region of interest is excluded from the enhanced transmission, effectively removing background noise and irrelevant audio from the viewer's experience.

Inventive Principle:
Principle #2Taking out (Extraction)

3Ease of operation

If custom audio processing is performed for each client device, then audio experience is optimized for individual viewers, but processing complexity and computational resources increase

Engineering Contradiction:
Improveaudio experience qualityVSAvoidprocessing complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system performs preliminary action by pre-dividing the audio field into multiple zones and pre-establishing the mapping between viewer interaction data and audio zone enhancement. When a viewer interacts with the interface to indicate their region of interest, the system can quickly retrieve and apply the corresponding pre-defined audio enhancement parameters, reducing real-time processing complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary processing layer between the audio capture and delivery stages. This intermediary layer handles the complex tasks of region of interest detection, audio zone mapping, and selective audio enhancement. By centralizing this complex processing in a dedicated layer, the overall system architecture becomes more manageable and the complexity is isolated to a specific component rather than being distributed throughout the entire system.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10880519B2Panoramic streaming of video with user selected audio
Publication Date: 2020.12.29 GN HEARING AS
  • US10880519B2 patent drawing
  • US10880519B2 patent drawing
  • US10880519B2 patent drawing

AI summary

A panoramic video conferencing server is disclosed. The server is operational to capture an audio stream via an array of microphones. The captured audio stream is de-multiplexed to extract audio for a plurality of audio zones. A region of interest (ROI) is determined for a plurality of client devices to which a panoramic video is being streamed. The ROI is determined based on a zoom information which is inputted by any one of the client devices. A modified audio channel associated with the panoramic video being streamed to at least one of the client devices is generated for each of the client devices. The modified audio channel is adjusted to coincide with directionality of an audio corresponding to the associated region of interest (ROI) inputted by the at least one of said client device and the modified audio channel is adjusted to mute audio corresponding to the audio zones other than the audio zone associated with the ROI.