Spatial Audio Signal Segmentation for Legacy Codec Compatibility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Legacy telephones and communication devices are not configured to process audio signals with spatial characteristics, limiting the adoption of spatial audio communications due to incompatible multichannel audio codecs and requiring network upgrades, which can cause asymmetrical audio playback disorienting listeners.
Innovation Solution
A method and apparatus that capture and process audio signals with spatial characteristics, separating main mono signals and ambience signals, adjusting their virtual positions, and coding them to be compatible with existing codecs, allowing legacy devices to receive and process spatial audio signals, even if they are not originally configured for it.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If spatial audio is captured and transmitted using multichannel audio codecs, then voice quality and spatial characteristics are improved, but compatibility with legacy telephones and network infrastructure deteriorates
Solution Approach 1:
The spatial audio signal is segmented into a main speech signal component and an ambience signal component. The main speech signal is processed using conventional mono narrowband or wideband codecs for broad compatibility, while the ambience signal is processed separately using multichannel audio codecs to preserve spatial characteristics. This segmentation allows legacy devices to receive the main speech signal while spatial audio-capable devices can additionally process the ambience signal for enhanced spatial audio experience.
Solution Approach 2:
The patent introduces an intermediary processing approach where the audio signal is decomposed into different components that can be transmitted through different coding paths. The main speech signal acts as an intermediary that bridges compatibility between legacy and modern systems, while the ambience signal carries the spatial information for advanced devices. This intermediary structure enables gradual adoption without requiring complete system replacement.
2Manufacturing precision
If asymmetric microphone arrangement is used to capture spatial audio, then spatial characteristics are improved, but audio asymmetry causes listener disorientation
Solution Approach 1:
The ambience signal, which contains the spatial asymmetry information from the asymmetric microphone arrangement, is extracted and separated from the main speech signal. By taking out the ambience component, the problematic asymmetry is isolated and can be processed differently - either corrected for compatibility or preserved for spatial audio capability without causing disorientation in the main speech path.
Solution Approach 2:
Different quality treatments are applied to different components of the audio signal. The main speech signal receives conventional processing optimized for clarity and compatibility, while the ambience signal receives specialized processing to preserve or adjust spatial characteristics. This local quality approach allows each component to be optimized for its specific function without compromising the other.
3Manufacturing precision
If network upgrades are implemented to support higher quality voice codecs, then spatial audio capability is improved, but deployment complexity and cost increase
Solution Approach 1:
The system dynamically adapts to different device capabilities and network conditions. The audio encoding and transmission process automatically adjusts between conventional and spatial audio modes based on the recipient device's capability. This dynamic approach allows the network to support spatial audio where available while maintaining backward compatibility, eliminating the need for comprehensive network upgrades.
Solution Approach 2:
The patent creates a universal audio transmission system that can handle both conventional and spatial audio through a single infrastructure. The dual-component audio stream (main speech + ambience) can be received and processed differently based on device capability, making the system universally compatible across legacy and modern devices without requiring separate network infrastructures.
Data Source
AI summary
A method, apparatus computer program product are provided to facilitate the utilization of the spatial position of audio signals in order to improve voice quality. In the context of a method, a main mono signal is determined from one or more audio signals that were received. The method also includes determining one or more ambience signals from the one or more audio signals that were received, such as following removal of the main mono signal therefrom. The method also adjusts at least one of a virtual position of the main mono signal for provision to a recipient device or the one or more ambience signals for provision to the recipient device.


