Audio Signal Objectification Through Neural Source Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio processing techniques struggle with converting non-object-based audio content into object-based audio content, which is laborious and expensive, and often result in audio signals being decoded to a channel-based format, limiting the ability to manipulate individual audio sources.
Innovation Solution
Utilizing machine learning models, particularly deep neural networks, trained through supervised learning on short audio snippets, to extract and separately manipulate individual audio sources in real-time or near real-time, enabling dynamic control of audio sources such as spatial position and volume.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional audio processing techniques are used to convert non-object-based audio content into object-based audio content, then the conversion process becomes laborious and expensive, but the ability to manipulate individual audio sources is improved
Solution Approach 1:
The patent replaces manual, mechanical audio conversion processes with an automated machine learning system. The neural network automatically separates audio sources and enables individual manipulation without requiring expert intervention in complex conversion processes, thus resolving the contradiction between ease of operation and device complexity
Solution Approach 2:
The system performs self-service by automatically analyzing and separating audio sources using the trained neural network model. The audio processing system handles the complex conversion task autonomously without requiring external expert intervention, making object-based manipulation accessible while reducing operational complexity
2Adaptability or versatility
If non-object-based audio content is converted into object-based audio content using traditional methods, then individual audio sources can be manipulated independently, but the process is laborious and expensive requiring expert involvement
Solution Approach 1:
The patent substitutes manual expert processing with an automated neural network system that performs audio source separation. This automation enables independent manipulation of audio sources while eliminating the need for expert involvement, thus improving both adaptability and ease of manufacture
Solution Approach 2:
The system changes the fundamental parameter of audio representation from mixed channel-based signals to separated object-based sources through the neural network's automatic separation process. This parameter transformation enables versatile manipulation while making the process accessible without expert intervention
3Reliability
If audio signals are decoded to channel-based format, then compatibility with legacy systems is maintained, but the ability to separately manipulate individual audio sources is limited
Solution Approach 1:
The patent implements a dynamic system that can operate in multiple modes: maintaining channel-based compatibility for legacy systems while simultaneously providing object-based separation for enhanced manipulation. The neural network enables flexible switching between these modes, resolving the contradiction between reliability and ease of operation
Solution Approach 2:
The system achieves universality by being capable of both channel-based playback for compatibility and object-based separation for manipulation flexibility. The same neural network infrastructure supports both operational modes, allowing the system to adapt to different requirements without sacrificing either compatibility or manipulation capability
Data Source
AI summary
Techniques for dynamic audio objectification are described. Embodiments include providing a first audio snippet from an audio signal to a machine learning model trained based on audio snippets labeled with an audio source and receiving, from the machine learning model, a subset of the first audio snippet that is associated with the audio source. Embodiments include, after playing the reconstituted first audio snippet, receiving a changed configuration relating to the audio source. Embodiments include providing a second audio snippet from the audio signal to the machine learning model and receiving, from the machine learning model, a subset of the second audio snippet that is associated with the audio source. Embodiments include playing a reconstituted second audio snippet based on the subset of the second audio snippet and the changed configuration, wherein an audibly perceptible parameter of the audio source is changed in the reconstituted second audio snippet.


