Audio Object Remixing via Side Information Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing techniques lack the ability to individually modify audio objects within a stereo signal, such as instruments, without affecting the entire audio signal, and conventional spatial audio coding methods require unnecessary processing and limit flexibility in remixing.
Innovation Solution
A method that allows for the modification of attributes like pan and gain of audio objects by generating side information representing the relations between the original audio signal and source signals, enabling remixing of stereo or multi-channel audio signals using mix parameters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If conventional spatial audio coding techniques are used to represent stereo or multi-channel audio channels using inter-channel cues, then audio signal can be compressed and transmitted, but separate signals for each audio object must be transmitted even if not modified, resulting in unnecessary processing at encoder and decoder
Solution Approach 1:
The patent extracts only the necessary side information (gain factors and panning coefficients) from the audio signal representation, rather than transmitting complete separate signals for each audio object. This extraction approach eliminates unnecessary data transmission and processing while preserving the ability to perform remixing operations at the decoder.
Solution Approach 2:
The patent segments the audio processing into essential components by representing the audio signal in terms of individual audio objects with separate gain and panning parameters, allowing selective modification of specific objects without processing entire channels or requiring complex decomposition at the decoder.
2Measurement precision
If conventional spatial audio coding techniques require a separate signal for each audio object to be transmitted, then individual audio objects can be identified, but this results in unnecessary processing at the encoder and decoder
Solution Approach 1:
The patent extracts only the essential parameters (gain factors and panning coefficients) that define audio object characteristics, rather than transmitting complete separate signals. This maintains precise audio object identification capability while dramatically reducing the processing complexity at both encoder and decoder.
3Ease of manufacture
If conventional spatial audio coding techniques are used, then audio channels can be represented using inter-channel cues, but this limits encoder input to either stereo or multi-channel audio signal, resulting in reduced flexibility for remixing at the decoder
Solution Approach 1:
The patent creates a universal encoding approach where the encoder can accept any audio signal type (stereo, multi-channel, or even synthesized signals) and represent them in terms of audio objects with gain and panning parameters. This multi-functional representation enables flexible remixing at the decoder regardless of the original signal format, as the same parameter-based approach works for all input types.
4Reliability
If conventional spatial audio coding techniques require complex de-correlation processing at the decoder, then spatial audio effects can be achieved, but this makes such techniques unsuitable for some applications or devices
Solution Approach 1:
The patent performs all necessary spatial processing and de-correlation operations at the encoder side during the encoding phase. The decoder simply applies the pre-computed gain and panning parameters without requiring complex de-correlation processing, making the technique suitable for devices with limited processing capability while maintaining spatial audio effect quality.
Data Source
AI summary
One or more attributes (e.g., pan, gain, etc.) associated with one or more objects (e.g., an instrument) of a stereo or multi-channel audio signal can be modified to provide remix capability. In some implementations, a method can include obtaining a first plural-channel audio signal having one or more objects; obtaining side information, at least some of which represents a relation between the first plural-channel audio signal and the one or more objects; obtaining a set of mix parameters; and generating a second plural-channel audio signal using the side information and the set of mix parameters.


