Acoustic Signal Separation Using Representative Ambient Sound Component

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional source separation technologies struggle to accurately estimate ambient sound spectra using a small number of samples, leading to reduced accuracy in separating voice signals from mixed acoustic signals, resulting in incomplete removal of ambient sound from extracted voice signals.

Innovation Solution

A signal processing apparatus that estimates non-stationary ambient sound components using a limited number of features, generates a filter based on both the estimated ambient sound and voice components, and separates the acoustic signal into voice and ambient sound signals, improving the separation performance by using a representative component of previously estimated ambient sound components.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a small number of samples (fewer than several seconds) are used to estimate ambient sound, then computational cost is reduced and processing delay is minimized, but the accuracy of ambient sound spectrum estimation deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidambient sound estimation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent pre-calculates and stores basis matrices for different ambient sound types (noise, music, speech) during system initialization or offline processing. When processing incoming acoustic signals, the system only needs to calculate coefficient matrices by comparing with pre-stored basis matrices, significantly reducing real-time computational complexity while maintaining accurate ambient sound estimation even with limited samples

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the approach from direct ambient sound estimation to estimating coefficients that represent the contribution of pre-defined basis matrices. This parameter transformation allows the system to work with fewer samples by leveraging the statistical properties captured in the pre-calculated basis matrices, thereby improving estimation accuracy without increasing processing demands

Inventive Principle:
Principle #35Parameter changes

2Loss of time

If conventional source separation methods are used with limited samples, then processing is faster, but separation performance deteriorates due to insufficient ambient sound estimation

Engineering Contradiction:
Improveprocessing delayVSAvoidseparation performance
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The system pre-calculates basis matrices representing different ambient sound characteristics before actual source separation. During real-time processing, it only needs to compute coefficient matrices by matching incoming signals against these pre-stored bases, enabling accurate separation with minimal delay even using few samples

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces basis matrices as intermediary representations that capture ambient sound characteristics. These basis matrices act as a bridge between the limited input samples and the final separation result, allowing the system to infer complete ambient sound profiles from incomplete data while maintaining high separation performance

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9478232B2Signal processing apparatus, signal processing method and computer program product for separating acoustic signals
Publication Date: 2016.10.25 TOSHIBA DIGITAL SOLUTIONS CORP
  • US9478232B2 patent drawing
  • US9478232B2 patent drawing
  • US9478232B2 patent drawing

AI summary

According to an embodiment, a signal processing apparatus includes an ambient sound estimating unit, a representative component estimating unit, a voice estimating unit, and a filter generating unit. The ambient sound estimating unit is configured to estimate, from the feature, an ambient sound component that is non-stationary among ambient sound components having a feature. The representative component estimating unit is configured to estimate a representative component representing ambient sound components estimated from one or more features for a time period, based on a largest value among the ambient sound components within the time period. The voice estimating unit is configured to estimate, from the feature, a voice component having the feature. The filter generating unit is configured to generate a filter for extracting a voice component and an ambient sound component from the feature, based on the voice component and the representative component.