Audio Background Sound Removal Using Multi-Model Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Video production faces challenges in separating background sound sources from audio data due to copyright issues during export, particularly when background sounds include both instrumentals and human voices, requiring an accurate method to distinguish between background and foreground voices.
Innovation Solution
A method and device using multiple separation models to isolate human voices, speech components, music components, and noise components from audio data, allowing for the removal of background sound sources by synthesizing speech and noise components.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a single separation model is used to remove background sound, then the process is simple, but the separation accuracy is insufficient especially when distinguishing between background and foreground voices
Solution Approach 1:
The patent divides the background sound removal task into multiple specialized separation models: a first separation model that separates human voice from other sounds, a second separation model that separates vocal components from speech components, and a third separation model that separates music from noise. This segmentation allows each model to specialize in specific sound types, significantly improving separation accuracy while managing complexity through functional division.
Solution Approach 2:
The patent introduces intermediate processing components including a voice detection model that identifies human voice in the mixed audio, and multiple specialized separation models that act as intermediaries between the raw mixed audio and the final output. These intermediary models progressively refine the separation, with each model handling a specific aspect of the separation task, thereby achieving high accuracy through staged processing.
2Measurement precision
If multiple separation models are used to achieve accurate separation, then the separation accuracy improves, but the processing time and computational resources increase
Solution Approach 1:
By segmenting the processing into specialized models that each handle specific sound types, the system can optimize each model for its specific task, potentially reducing the computational burden compared to a single comprehensive model. The segmented approach allows parallel processing of different sound components, which can reduce overall processing time despite using multiple models.
Solution Approach 2:
The voice detection model performs preliminary identification of human voice components before the separation models process the audio. This preliminary action allows the subsequent separation models to focus their computational resources on separating already-identified voice components from other sounds, reducing redundant processing and optimizing computational efficiency.
3Adaptability or versatility
If background sound is not removed, then the audio quality is maintained, but copyright issues prevent video export
Solution Approach 1:
The patent extracts the background sound components (music and noise) from the mixed audio using specialized separation models, while preserving the foreground speech components. This extraction allows the video to be exported without copyright-infringing background music, while maintaining the essential speech content that carries the video's informational value.
Solution Approach 2:
The patent applies different quality standards to different audio components: the speech component (which carries essential information) is preserved with high fidelity, while the background music component (which causes copyright issues) is removed or significantly reduced. This local quality approach maintains overall audio quality for essential content while eliminating problematic elements.
Data Source
AI summary
One aspect of the present disclosure provides a method for automatically removing a background sound source of audio data of a video, including separating the audio data including at least one sound source component into a first component related to a human voice and a second component related to sounds other than the human voice using a first separation model, separating the first component into a vocal component and a speech component using a second separation model, separating the second component into a music component and a noise component using a third separation model, and generating an audio data with the background sound source for the audio data of the video removed by synthesizing the speech component and the noise component.


