Spatial Audio Signal Segmentation for Legacy Codec Compatibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Legacy telephones and communication devices are not configured to process audio signals with spatial characteristics, limiting the adoption of spatial audio communications due to incompatible multichannel audio codecs and requiring network upgrades, which can cause asymmetrical audio playback disorienting listeners.

Innovation Solution

A method and apparatus that capture and process audio signals with spatial characteristics, separating main mono signals and ambience signals, adjusting their virtual positions, and coding them to be compatible with existing codecs, allowing legacy devices to receive and process spatial audio signals, even if they are not originally configured for it.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If spatial audio is captured and transmitted using multichannel audio codecs, then voice quality and spatial characteristics are improved, but compatibility with legacy telephones and network infrastructure deteriorates

Engineering Contradiction:
Improvevoice qualityVSAvoiddevice compatibility
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The spatial audio signal is segmented into a main speech signal component and an ambience signal component. The main speech signal is processed using conventional mono narrowband or wideband codecs for broad compatibility, while the ambience signal is processed separately using multichannel audio codecs to preserve spatial characteristics. This segmentation allows legacy devices to receive the main speech signal while spatial audio-capable devices can additionally process the ambience signal for enhanced spatial audio experience.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing approach where the audio signal is decomposed into different components that can be transmitted through different coding paths. The main speech signal acts as an intermediary that bridges compatibility between legacy and modern systems, while the ambience signal carries the spatial information for advanced devices. This intermediary structure enables gradual adoption without requiring complete system replacement.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If asymmetric microphone arrangement is used to capture spatial audio, then spatial characteristics are improved, but audio asymmetry causes listener disorientation

Engineering Contradiction:
Improvespatial characteristicsVSAvoidlistener disorientation
Core Design Contradiction:
Manufacturing precisionVSObject-affected harmful factors

Solution Approach 1:

The ambience signal, which contains the spatial asymmetry information from the asymmetric microphone arrangement, is extracted and separated from the main speech signal. By taking out the ambience component, the problematic asymmetry is isolated and can be processed differently - either corrected for compatibility or preserved for spatial audio capability without causing disorientation in the main speech path.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Different quality treatments are applied to different components of the audio signal. The main speech signal receives conventional processing optimized for clarity and compatibility, while the ambience signal receives specialized processing to preserve or adjust spatial characteristics. This local quality approach allows each component to be optimized for its specific function without compromising the other.

Inventive Principle:
Principle #3Local quality

3Manufacturing precision

If network upgrades are implemented to support higher quality voice codecs, then spatial audio capability is improved, but deployment complexity and cost increase

Engineering Contradiction:
Improvespatial audio capabilityVSAvoidnetwork upgrade complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system dynamically adapts to different device capabilities and network conditions. The audio encoding and transmission process automatically adjusts between conventional and spatial audio modes based on the recipient device's capability. This dynamic approach allows the network to support spatial audio where available while maintaining backward compatibility, eliminating the need for comprehensive network upgrades.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent creates a universal audio transmission system that can handle both conventional and spatial audio through a single infrastructure. The dual-component audio stream (main speech + ambience) can be received and processed differently based on device capability, making the system universally compatible across legacy and modern devices without requiring separate network infrastructures.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9344826B2Method and apparatus for communicating with audio signals having corresponding spatial characteristics
Publication Date: 2016.05.17 NOKIA TECHNOLOGIES OY
  • US9344826B2 patent drawing
  • US9344826B2 patent drawing
  • US9344826B2 patent drawing

AI summary

A method, apparatus computer program product are provided to facilitate the utilization of the spatial position of audio signals in order to improve voice quality. In the context of a method, a main mono signal is determined from one or more audio signals that were received. The method also includes determining one or more ambience signals from the one or more audio signals that were received, such as following removal of the main mono signal therefrom. The method also adjusts at least one of a virtual position of the main mono signal for provision to a recipient device or the one or more ambience signals for provision to the recipient device.