Audio Separation With Integrated Background Sound Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio separation technologies fail to effectively separate candidate audios from sound sources while managing background sounds, often resulting in artifacts and unclear audio due to separate inference processes for feature extraction and background sound removal.
Innovation Solution
An audio separation system that extracts first and background sound features from a sound source, adjusts the background sound based on a control parameter, and generates separated audios by processing these features using convolutional blocks and up-convolutional blocks, allowing for targeted control of background sound levels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If a separate model is used to extract background sound characteristics after candidate audio separation, then background sound can be adjusted or removed, but artificial sounds (artifacts) are mixed into the final separated audio and audio quality deteriorates
Solution Approach 1:
The patent merges the candidate audio separation function and background sound extraction function into a single integrated audio separation model. This unified model performs both tasks simultaneously during one inference process, eliminating the need for separate post-processing steps that previously introduced artifacts and degraded audio quality.
Solution Approach 2:
The integrated model performs background sound extraction and candidate audio separation in a single preliminary processing step, rather than attempting to remove background sound after separation. This preliminary action prevents artifact introduction that occurs when trying to subtract background sound from already-separated audio.
2Object-affected harmful factors
If two inference processes are performed (one for candidate audio separation and another for background sound removal), then background sound can be controlled, but processing time increases and productivity decreases
Solution Approach 1:
The patent combines two separate inference processes into a single unified model that performs both candidate audio separation and background sound extraction simultaneously. This merging reduces processing time and improves productivity while maintaining the ability to control background sound levels through a single inference operation.
3Device complexity
If traditional audio separation focuses only on candidate audio feature extraction, then the separation process is simple, but background sound is included in the separated audio and separation precision is insufficient
Solution Approach 1:
The audio separation model is designed with multi-functionality, simultaneously performing candidate audio separation and background sound extraction. This universal model handles multiple tasks within a single architecture, improving separation precision without requiring overly complex multi-stage processing systems.
Data Source
AI summary
Provided is a method for separating one or more candidate audios in a sound source including the one or more candidate audios and a background sound, by using an audio separation system, the method including extracting a first audio feature from the sound source, extracting a background sound feature from the sound source, the background sound feature identifying a degree of association between the first audio feature and the background sound, generating a second audio feature based on the first audio feature, the background sound feature, and a background sound control parameter configured to control the background sound and generating one or more separated audios based on target information corresponding to the one or more candidate audios, the first audio feature, and the second audio feature in which the background sound is adjusted.


