Spatial Cue Rendering for Flexible Multi-Object Audio Positioning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for rendering multi-object or multi-channel audio signals lack flexibility in controlling the spatial positioning of audio signals during decoding, relying on fixed positions and limited control over spatial cues.

Innovation Solution

An apparatus and method that includes a decoder and a spatial cue renderer, which extracts and controls Channel Level Difference (CLD) or other spatial cues to dynamically adjust the power gain of audio signals, allowing for flexible positioning of multi-object or multi-channel audio signals based on user input or external control information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional SAC decoding is used, then audio signals can be decoded from down-mixed signals, but the spatial positioning of audio signals is fixed and cannot be dynamically controlled

Engineering Contradiction:
Improvespatial positioning control flexibilityVSAvoiddecoder structure complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The decoding system is segmented into independent functional modules: a decoder that processes down-mixed signals and extracts spatial cues, and a spatial cue renderer that independently controls spatial positioning based on extracted cues. This segmentation allows flexible spatial control without requiring complete redesign of the decoding architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces dynamic control of spatial cues by allowing the spatial cue renderer to adjust rendering parameters in real-time based on extracted spatial cue information. This enables audio signals to be positioned dynamically in a virtual sound field rather than at fixed positions, resolving the contradiction between adaptability and device complexity.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If spatial cue information is extracted and controlled, then precise spatial positioning is achieved, but the processing complexity increases

Engineering Contradiction:
Improvespatial cue extraction precisionVSAvoidspatial cue rendering complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The decoder performs preliminary extraction of spatial cue information from down-mixed signals before the rendering stage. By pre-processing and extracting spatial cues (such as inter-channel level differences, inter-channel time differences) in advance, the system reduces the computational burden on the spatial cue renderer and achieves precise spatial positioning without excessive processing complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces spatial cue information as an intermediary element that bridges the down-mixed signal and the final rendered output. This intermediary carries precise spatial positioning data extracted from the down-mixed signal, allowing the spatial cue renderer to achieve accurate spatial reconstruction without directly processing complex multi-channel signals, thus reducing overall system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11375331B2Method and apparatus for control of randering multiobject or multichannel audio signal using spatial cue
Publication Date: 2022.06.28 ELECTRONICS & TELECOMM RES INST
  • US11375331B2 patent drawing
  • US11375331B2 patent drawing
  • US11375331B2 patent drawing

AI summary

The present research relates to controlling rendering of multi-object or multi-channel audio signals. The present research provides a method and apparatus for controlling rendering of multi-object or multi-channel audio signals based on spatial cues in a process of decoding the multi-object or multi-channel audio signals. To achieve the purpose, the method suggested in the research controls rendering in a spatial cue domain in the process of decoding the multi-object or multi-channel audio signals.