Object-Oriented Audio Scene Mapping for Spatial Sound Positioning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio systems, such as 5.1 and 7.1 systems, limit the flexibility in sound source positioning and reproduction, resulting in an unsatisfactory spatial sound experience, especially when trying to accurately place sound sources within a movie scene, due to their channel-oriented paradigm and fixed speaker positions.
Innovation Solution
An object-oriented approach for generating and editing audio scenes, where audio objects are defined by their starting and end time instants, allowing for flexible mapping to input channels, enabling efficient use of wave-field synthesis rendering units and providing a user-friendly interface for sound recordists to manage complex audio scenes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If channel-oriented audio systems (5.1, 7.1) are used with fixed speaker positions, then the system structure is simple and easy to implement, but the spatial sound reproduction quality and flexibility in sound source positioning deteriorate
Solution Approach 1:
The audio scene is segmented into multiple independent audio objects, each with its own spatial characteristics and time instants. This allows flexible positioning of sound sources without being constrained by fixed speaker positions, resolving the contradiction between positioning flexibility and system complexity
Solution Approach 2:
The patent introduces a temporal dimension by defining audio objects with starting and end time instants, transforming the traditional spatial-only audio representation into a spatio-temporal model. This enables dynamic sound source positioning over time while maintaining system manageability
2Measurement precision
If wave-field synthesis is used to achieve natural spatial sound impression across great area, then the spatial sound reproduction quality improves, but the computational requirements and processing power deteriorate
Solution Approach 1:
The patent extracts and processes only the essential spatial and temporal characteristics of audio objects, rather than computing the complete wave field. By focusing on key parameters like starting and end time instants and spatial positions, the system achieves accurate spatial reproduction with reduced computational load
3Reliability
If manual postproduction steps are used to match acoustic impression with visual scene, then the authenticity of the audio-visual experience improves, but the production time and cost deteriorate
Solution Approach 1:
The patent enables acoustic matching to be performed during the initial audio object definition phase, rather than requiring manual postproduction adjustments. By preliminarily assigning spatial and temporal characteristics to audio objects, the system achieves authentic audio-visual synchronization without time-consuming manual intervention
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach allows for a more transparent and efficient handling of audio objects, reducing the computational load on wave-field synthesis rendering units and enhancing creative potential by enabling flexible sound source positioning and management, improving the overall audio experience without requiring extensive re-mixing of entire scenes.
Implementation Method 1
The principles of this technology, the so-called wave-field synthesis (WFS), have been studied at the TU Delft and first presented in the late 80s
Implementation Method 2
The basic idea of WFS is based on the application of Huygens' principle of the wave theory: Each point caught by a wave is starting point of an elementary wave propagating in spherical or circular manner
Data Source
AI summary
An apparatus for generating, storing or editing an audio representation of an audio scene includes audio processing means for generating a plurality of speaker signals from a plurality of input channels as well as means for providing an object-oriented description of the audio scene, wherein the object-oriented description of the audio scene includes a plurality of audio objects, wherein an audio object is associated with an audio signal, a starting time instant and an end time instant. The apparatus for generating further distinguishes itself by mapping means for mapping the object-oriented description of the audio scene to the plurality of input channels, wherein an assignment of temporally overlapping audio objects to parallel input channels is performed by the mapping means, whereas temporally sequential audio objects are associated with the same channel. With this, an object-oriented representation is transferred into a channel-oriented representation, whereby on the object-oriented side the optimal representation of a scene may be used, whereas on channel-oriented side the channel-oriented concept users are used to may be maintained.


