Spatially Encoded Signal Format for Immersive Audio Capture
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies lack an end-to-end system for flexible capture, distribution, and reproduction of immersive audio recordings in a generic digital format compatible with both two-channel and multi-channel playback systems, particularly in consumer devices with non-standard microphone configurations.
Innovation Solution
The system processes microphone signals using an arbitrary microphone array configuration to encode audio into a Spatially Encoded Signal (SES) format, which is agnostic to the capture configuration, allowing for efficient storage and distribution, and can be decoded for various playback configurations without requiring specific information about the original capture setup.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If time-domain phase-amplitude matrix encoding is used to reduce data channels, then storage and distribution efficiency is improved, but spatial localization fidelity deteriorates
Solution Approach 1:
The patent introduces an intermediate frequency-domain representation as a mediator between the time-domain encoded signal and the final decoded output. By transforming the matrix-encoded signal into the frequency domain, applying spatial encoding coefficients in the frequency domain, and then transforming back, the system preserves spatial localization fidelity while maintaining storage efficiency. This intermediate frequency-domain step acts as a mediator that prevents the loss of spatial information inherent in direct time-domain decoding.
Solution Approach 2:
The patent changes the domain parameter from time-domain to frequency-domain processing. By performing spatial encoding and decoding operations in the frequency domain rather than the time domain, the system achieves both efficient data reduction and preserved spatial localization fidelity. The frequency-domain representation allows for more precise control of spatial parameters while maintaining compression efficiency.
2Measurement precision
If a specific microphone configuration is used for capture, then capture quality is improved, but adaptability to arbitrary playback configurations deteriorates
Solution Approach 1:
The patent creates a universal encoding format that can represent spatial audio information captured by any microphone configuration. The spatial encoding coefficients are designed to be agnostic to the specific capture configuration, allowing the same encoded data to be decoded for various playback configurations including headphones, stereo speakers, and surround sound systems. This universality enables a single capture to serve multiple playback purposes.
Solution Approach 2:
The patent introduces dynamic adaptability by allowing the decoding process to adjust to different playback configurations. The spatial encoding coefficients can be adapted based on the target playback configuration, enabling the system to dynamically optimize the reproduction for the specific output device being used, whether it's headphones, stereo speakers, or surround sound systems.
3Productivity
If channel count is reduced for distribution, then data transmission efficiency is improved, but spatial information completeness deteriorates
Solution Approach 1:
The patent extracts and separates the spatial encoding coefficients from the audio signal, storing them independently from the main audio data. This extraction allows the spatial information to be preserved separately while the audio signal itself is efficiently compressed. During decoding, the extracted spatial coefficients are reapplied to reconstruct the spatial audio field, ensuring spatial information completeness is maintained despite channel reduction.
Data Source
AI summary
A sound field coding system and method that provides flexible capture, distribution, and reproduction of immersive audio recordings encoded in a generic digital audio format compatible with standard two-channel or multi-channel reproduction systems. This end-to-end system and method mitigates any impractical need for standard multi-channel microphone array configurations in consumer mobile devices such as smart phones or cameras. The system and method capture and spatially encode two-channel or multi-channel immersive audio signals that are compatible with legacy playback systems from flexible multi-channel microphone array configurations.


