Audio Background Sound Removal Using Multi-Model Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Video production faces challenges in separating background sound sources from audio data due to copyright issues during export, particularly when background sounds include both instrumentals and human voices, requiring an accurate method to distinguish between background and foreground voices.

Innovation Solution

A method and device using multiple separation models to isolate human voices, speech components, music components, and noise components from audio data, allowing for the removal of background sound sources by synthesizing speech and noise components.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a single separation model is used to remove background sound, then the process is simple, but the separation accuracy is insufficient especially when distinguishing between background and foreground voices

Engineering Contradiction:
Improveseparation accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the background sound removal task into multiple specialized separation models: a first separation model that separates human voice from other sounds, a second separation model that separates vocal components from speech components, and a third separation model that separates music from noise. This segmentation allows each model to specialize in specific sound types, significantly improving separation accuracy while managing complexity through functional division.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate processing components including a voice detection model that identifies human voice in the mixed audio, and multiple specialized separation models that act as intermediaries between the raw mixed audio and the final output. These intermediary models progressively refine the separation, with each model handling a specific aspect of the separation task, thereby achieving high accuracy through staged processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If multiple separation models are used to achieve accurate separation, then the separation accuracy improves, but the processing time and computational resources increase

Engineering Contradiction:
Improveseparation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

By segmenting the processing into specialized models that each handle specific sound types, the system can optimize each model for its specific task, potentially reducing the computational burden compared to a single comprehensive model. The segmented approach allows parallel processing of different sound components, which can reduce overall processing time despite using multiple models.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The voice detection model performs preliminary identification of human voice components before the separation models process the audio. This preliminary action allows the subsequent separation models to focus their computational resources on separating already-identified voice components from other sounds, reducing redundant processing and optimizing computational efficiency.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If background sound is not removed, then the audio quality is maintained, but copyright issues prevent video export

Engineering Contradiction:
Improveexport capabilityVSAvoidaudio quality
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent extracts the background sound components (music and noise) from the mixed audio using specialized separation models, while preserving the foreground speech components. This extraction allows the video to be exported without copyright-infringing background music, while maintaining the essential speech content that carries the video's informational value.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different quality standards to different audio components: the speech component (which carries essential information) is preserved with high fidelity, while the background music component (which causes copyright issues) is removed or significantly reduced. This local quality approach maintains overall audio quality for essential content while eliminating problematic elements.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20240265932A1Device and method for automatically removing a background sound source of a video
Publication Date: 2024.08.08 SK TELECOM CO LTD
  • US20240265932A1 patent drawing
  • US20240265932A1 patent drawing
  • US20240265932A1 patent drawing

AI summary

One aspect of the present disclosure provides a method for automatically removing a background sound source of audio data of a video, including separating the audio data including at least one sound source component into a first component related to a human voice and a second component related to sounds other than the human voice using a first separation model, separating the first component into a vocal component and a speech component using a second separation model, separating the second component into a music component and a noise component using a third separation model, and generating an audio data with the background sound source for the audio data of the video removed by synthesizing the speech component and the noise component.