Longitudinal-Depth Audio Processing for Video Communication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current telepresence systems in video communication fail to realistically distinguish sounds from different rows, as they only map image and sound directions on the same plane, leading to an unrealistic on-site feeling effect.

Innovation Solution

A method and apparatus for processing audio in video communication that acquires audio data and source position information, performing longitudinal-depth processing to differentiate sounds from various positions, using algorithms or wave field synthesis to provide a realistic distance and depth perception.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If audio data from multiple rows is mapped to the same plane, then the system structure remains simple, but the ability to distinguish sound directions from different rows deteriorates

Engineering Contradiction:
Improvesystem structureVSAvoidsound direction information
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent introduces a longitudinal-depth dimension to the audio mapping system. Instead of mapping all audio sources to a single two-dimensional plane, the system creates multiple depth layers corresponding to different physical rows. Audio data from the front row is mapped to a first depth layer, while audio data from the back row is mapped to a second depth layer, enabling listeners to distinguish sound directions from different rows while maintaining manageable system complexity through structured layering

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If audio data is processed with longitudinal-depth information, then the realism of sound positioning is improved, but the processing complexity increases

Engineering Contradiction:
Improvesound positioning accuracyVSAvoidaudio processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the audio processing system into distinct modules: an audio data acquisition module that captures audio signals with associated position information, a longitudinal-depth processing module that applies depth-based transformations, and an output module that delivers processed audio. This segmentation allows the system to implement complex longitudinal-depth processing while keeping each individual module manageable and well-defined

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces longitudinal-depth information as an intermediary parameter that mediates between raw audio data and final audio output. This intermediary layer carries depth positioning information through the processing pipeline, enabling realistic sound positioning without requiring complex direct manipulation of audio signals, thus balancing processing complexity with positioning accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9113034B2Method and apparatus for processing audio in video communication
Publication Date: 2015.08.18 HONOR DEVICE CO LTD
  • US9113034B2 patent drawing
  • US9113034B2 patent drawing
  • US9113034B2 patent drawing

AI summary

Embodiments of the present invention provide a method and an apparatus for processing audio in video communication. The method includes: acquiring audio data and audio source position information corresponding to the audio data in the video communication, and performing longitudinal-depth processing on the audio data according to the audio source position information. According to the embodiments of the present invention, the audio data and the audio source position information corresponding to the audio data in the video communication are acquired first, and then the longitudinal-depth processing is performed on the audio data according to the acquired audio source position information to make it be audio data that provides a longitudinal-depth feeling that matches the audio source position information, so that sounds generated by objects at different front/back positions can be distinguished in the video communication.