Geometry-Based Spatial Audio Coding for Flexible Sound Scene Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing spatial audio coding techniques lack flexibility in modifying the sound scene and are inefficient in transmission and storage, as they are limited by the recording position and require all microphone signals for synthesis, and cannot accurately describe complex sound scenes with multiple active sources.
Innovation Solution
An apparatus and method for generating audio output signals based on audio data streams that include sound pressure values, position values, and diffuseness values, allowing for the computation of a sound field representation independent of the recording position, enabling efficient transmission and storage, and allowing for modifications such as changing the listening position and loudspeaker setup.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional spatial audio coding techniques are used, then the sound scene can be recorded and reproduced, but the sound scene cannot be modified and the representation is limited to the recording position
Solution Approach 1:
The patent transforms the sound scene representation from fixed microphone signals to parametric form, encoding position values, diffuseness values, and sound pressure values that can be dynamically modified. This allows the sound scene to be adapted to different listening positions and loudspeaker setups without re-recording, resolving the contradiction between adaptability and complexity by using compact parameters instead of full signal representations.
Solution Approach 2:
The patent extracts essential spatial characteristics (position, diffuseness, sound pressure) from the full audio signals and transmits only these extracted parameters. This separation allows the core spatial information to be transmitted efficiently while enabling flexible modification at the reproduction stage, addressing the contradiction by removing unnecessary data while preserving adaptability.
2Reliability
If all microphone signals are transmitted for synthesis, then accurate sound reproduction is achieved, but transmission and storage efficiency is reduced
Solution Approach 1:
The patent extracts only the essential spatial parameters (position, diffuseness, sound pressure) from the full microphone signals and transmits these extracted features instead of the complete signals. This extraction maintains sufficient information for accurate spatial reproduction while dramatically reducing the data volume for transmission and storage, resolving the contradiction between reliability and energy efficiency.
Solution Approach 2:
The patent applies different levels of detail to different aspects of the sound scene representation. Full audio signals are processed to extract spatial parameters, but only the essential spatial characteristics are transmitted in detail while the actual audio content can be reconstructed or synthesized. This local differentiation of quality levels achieves accurate spatial reproduction with reduced overall data transmission.
3Adaptability or versatility
If parametric representations are used, then flexibility and compactness are improved, but the spatial image is fixed relative to the spatial microphone and cannot be varied
Solution Approach 1:
The patent introduces dynamic adaptability by encoding the spatial parameters (position, diffuseness) in a coordinate-independent manner that allows runtime transformation to different listening positions and acoustic viewpoints. The parametric representation is designed to be dynamically adjustable, enabling the acoustic viewpoint to be varied and listening position changed without re-recording, resolving the contradiction between flexibility and complexity through dynamic parameter transformation.
Data Source
AI summary
An apparatus for generating at least one audio output signal based on an audio data stream having audio data relating to one or more sound sources is provided. The apparatus has a receiver for receiving the audio data stream having the audio data. The audio data has one or more pressure values for each one of the sound sources. Furthermore, the audio data has one or more position values indicating a position of one of the sound sources for each one of the sound sources. Moreover, the apparatus has a synthesis module for generating the at least one audio output signal based on at least one of the one or more pressure values of the audio data of the audio data stream and based on at least one of the one or more position values of the audio data of the audio data stream.


