Spatial Cue Rendering for Flexible Multi-Object Audio Positioning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for rendering multi-object or multi-channel audio signals lack flexibility in controlling the spatial positioning of audio signals during decoding, relying on fixed positions and limited control over spatial cues.
Innovation Solution
An apparatus and method that includes a decoder and a spatial cue renderer, which extracts and controls Channel Level Difference (CLD) or other spatial cues to dynamically adjust the power gain of audio signals, allowing for flexible positioning of multi-object or multi-channel audio signals based on user input or external control information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional SAC decoding is used, then audio signals can be decoded from down-mixed signals, but the spatial positioning of audio signals is fixed and cannot be dynamically controlled
Solution Approach 1:
The decoding system is segmented into independent functional modules: a decoder that processes down-mixed signals and extracts spatial cues, and a spatial cue renderer that independently controls spatial positioning based on extracted cues. This segmentation allows flexible spatial control without requiring complete redesign of the decoding architecture.
Solution Approach 2:
The patent introduces dynamic control of spatial cues by allowing the spatial cue renderer to adjust rendering parameters in real-time based on extracted spatial cue information. This enables audio signals to be positioned dynamically in a virtual sound field rather than at fixed positions, resolving the contradiction between adaptability and device complexity.
2Measurement precision
If spatial cue information is extracted and controlled, then precise spatial positioning is achieved, but the processing complexity increases
Solution Approach 1:
The decoder performs preliminary extraction of spatial cue information from down-mixed signals before the rendering stage. By pre-processing and extracting spatial cues (such as inter-channel level differences, inter-channel time differences) in advance, the system reduces the computational burden on the spatial cue renderer and achieves precise spatial positioning without excessive processing complexity.
Solution Approach 2:
The patent introduces spatial cue information as an intermediary element that bridges the down-mixed signal and the final rendered output. This intermediary carries precise spatial positioning data extracted from the down-mixed signal, allowing the spatial cue renderer to achieve accurate spatial reconstruction without directly processing complex multi-channel signals, thus reducing overall system complexity.
Data Source
AI summary
The present research relates to controlling rendering of multi-object or multi-channel audio signals. The present research provides a method and apparatus for controlling rendering of multi-object or multi-channel audio signals based on spatial cues in a process of decoding the multi-object or multi-channel audio signals. To achieve the purpose, the method suggested in the research controls rendering in a spatial cue domain in the process of decoding the multi-object or multi-channel audio signals.


