Audio Object Level Control via Preset Metadata Regions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users find it inconvenient to directly control all object signals in a downmix signal, making it difficult to reproduce an optimal audio signal, especially for non-experts.
Innovation Solution
An apparatus and method that utilize preset information and metadata to control the level and position of objects in an audio signal, allowing users to select and apply preset metadata to all or specific data regions of a downmix signal based on sound source characteristics, enabling easy adjustment of object levels and positions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If users directly control all object signals in a downmix signal, then precise control of object levels and positions is achieved, but operation complexity and difficulty increase significantly for non-experts
Solution Approach 1:
The patent introduces preset information and preset metadata as intermediary elements between the user and the complex object control parameters. Instead of directly controlling individual object signals, users select from pre-configured metadata sets that contain optimized parameter combinations. This intermediary layer simplifies the user interface while maintaining access to precise control capabilities through the preset configurations.
Solution Approach 2:
The patent applies preliminary action by pre-configuring optimal parameter sets (preset metadata) for different audio scenarios and sound source characteristics. These presets are prepared in advance based on expert knowledge and specific audio conditions, allowing users to benefit from pre-optimized settings without performing complex manual adjustments. The system automatically selects and applies appropriate presets based on the audio content.
2Stability of the object's composition
If preset information is applied to all data regions of a downmix signal, then consistent audio quality is achieved across the entire signal, but adaptability to specific sound source characteristics is reduced
Solution Approach 1:
The patent implements local quality by allowing different preset information to be applied to different data regions of the downmix signal based on the characteristics of the sound source in each region. The system analyzes the audio content and selectively applies appropriate presets to specific segments, ensuring that each region receives optimized processing tailored to its particular characteristics while maintaining overall signal integrity.
Solution Approach 2:
The patent applies dynamics by making the preset application process adaptive and variable rather than static and uniform. The system dynamically selects and applies preset information based on real-time analysis of sound source characteristics in different data regions. This allows the audio processing to adapt to changing conditions throughout the signal, optimizing quality for each specific segment while maintaining overall consistency.
Data Source
AI summary
An apparatus for processing an audio signal and method thereof are disclosed. The preset invention includesreceiving a downmix signal including at least one object, preset information to render the downmix signal and preset attribute information indicating attribute of the preset information; rendering the downmix signal by applying the preset information to all data regions of the downmix signal, if the preset information is included in a configuration information region based on the preset attribute information; and rendering the downmix signal by applying the preset information to one corresponding data region of the downmix signal, if the preset information is included in a data region based on the preset attribute information, wherein the preset information is obtained based on preset number information indicating a number of the preset information and output channel information indicating a number of output channel of the rendered downmix signal.Accordingly, one of a plurality of preset information is selected using a plurality of preset metadata without user's setting on each object, whereby a level of an output channel of an object can be adjusted with ease.


