Near-Far Stereo Audio Encoding for Mobile Spatial Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current spatial audio capture technologies for mobile devices face challenges in providing immersive voice and audio services, particularly in maintaining clear separation and accurate spatial presentation of voice and ambience signals during device rotations and changes in capture modes, leading to user confusion and reduced immersion.
Innovation Solution
The proposed solution involves an apparatus and method that receive and process channel voice audio signals and ambience audio signals with associated metadata, enabling the generation of encoded multichannel audio signals that separate voice and ambience spatially, and adjust bit rates and coding levels based on transmission and rendering capabilities, while compensating for device rotations through metadata adjustments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If spatial audio capture is implemented on mobile devices, then immersive voice and audio services are enabled, but maintaining clear separation and accurate spatial presentation of voice and ambience signals during device rotations becomes difficult
Solution Approach 1:
The audio signal is segmented into distinct voice channels and ambience channels with separate spatial metadata. This segmentation allows independent processing of voice and ambience components, maintaining their spatial separation even during device rotations. The voice signal and ambience signal are encoded separately with their own spatial parameters, enabling the receiver to reconstruct accurate spatial presentation.
Solution Approach 2:
The system dynamically adjusts spatial parameters (such as pan law, distance, and spatial position metadata) based on device orientation and capture mode. When device rotation is detected, the spatial metadata is transformed accordingly to maintain accurate spatial presentation relative to the user rather than to the device, resolving the contradiction between adaptability and precision.
2Manufacturing precision
If multiple audio channels with spatial metadata are transmitted, then spatial presentation quality is improved, but transmission bit rate and system complexity increase
Solution Approach 1:
The system extracts only the essential spatial metadata parameters needed for spatial presentation rather than transmitting complete multichannel spatial audio data. By separating the audio signal into voice and ambience channels and transmitting only their spatial parameters (direction, distance, pan law) rather than full spatial audio streams, the system achieves spatial presentation quality with reduced data volume.
Solution Approach 2:
The spatial metadata structure is designed to be universal and adaptable to different output configurations (mono, stereo, spatial). The same encoded multichannel audio signal with spatial metadata can be rendered for different output types, eliminating the need for separate encoding streams for each output type and thereby reducing overall transmission data volume while maintaining spatial presentation quality.
3Manufacturing precision
If separate voice and ambience channels are encoded, then spatial independence is achieved, but encoding complexity and processing requirements increase
Solution Approach 1:
The system merges the encoding of voice and ambience channels into a unified encoded multichannel audio signal structure that includes both audio data and spatial metadata. Rather than maintaining completely separate encoding systems, the voice and ambience channels are combined into a single encoded stream with integrated spatial parameters, reducing encoding complexity while preserving spatial independence through the metadata structure.
4Adaptability or versatility
If spatial audio processing is enhanced for device rotations, then immersion is improved, but latency and processing time increase
Solution Approach 1:
The spatial metadata is pre-calculated and embedded in the encoded audio signal during encoding, rather than being calculated in real-time during playback. Device orientation transformations are pre-planned through metadata structure design that allows efficient runtime adjustment. This preliminary action reduces real-time processing requirements and latency while maintaining immersion during device rotations.
Data Source
AI summary
An apparatus including circuitry, including at least one processor and at least one memory, configured to: receive at least one channel voice audio signal and metadata, the at least one channel voice audio signal and metadata generated from at least one microphone audio signal; and receive at least one channel ambience audio signal and metadata, wherein the at least one channel ambience audio signal and metadata are generated based on a parametric analysis of at least one microphone audio signal, and the at least one channel ambience audio signal is associated with the at least one channel voice audio signal; generate an encoded multichannel audio signal based on the at least one channel voice audio signal and metadata and further the at least one channel ambience audio signal and metadata.


