Mixed Audio Separation Using Local Frequency Pattern Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio separation methods face challenges in setting both high temporal and frequency resolutions independently, leading to suboptimal performance in extracting specific audio from mixed audio due to their trade-off relationship.

Innovation Solution

A mixed audio separation apparatus that generates local frequency information using local reference waveforms with specific temporal and frequency resolutions, allowing for independent setting of temporal and frequency resolutions through pattern matching and signal generation techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the time width of reference waveform is increased to improve frequency resolution, then frequency resolution is improved, but temporal resolution deteriorates

Engineering Contradiction:
Improvefrequency resolutionVSAvoidtemporal resolution
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the analysis waveform into multiple local segments and performs frequency analysis on each segment independently using cross-correlation with reference waveforms. This segmentation allows the system to achieve high temporal resolution by analyzing short local segments while maintaining high frequency resolution through the cross-correlation method, effectively resolving the trade-off between temporal and frequency resolution that plagues conventional Fourier transform approaches.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If conventional Fourier transform is used for frequency analysis, then frequency analysis can be performed, but temporal and frequency resolutions are in trade-off relationship and cannot be set independently

Engineering Contradiction:
Improveindependent resolution settingVSAvoidresolution
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces cross-correlation as an intermediary method between the analysis waveform and reference waveforms to perform frequency analysis. Unlike Fourier transform which inherently couples temporal and frequency resolution, the cross-correlation approach allows independent optimization of both resolutions by selecting appropriate reference waveform parameters and segment sizes, thereby achieving adaptable resolution settings for different application requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS7974420B2Mixed audio separation apparatus
Publication Date: 2011.07.05 PANASONIC HOLDINGS CORP
  • US7974420B2 patent drawing
  • US7974420B2 patent drawing
  • US7974420B2 patent drawing

AI summary

A mixed audio separation system (100) which separates a specific audio from among a mixed audio (S100) includes a local frequency information generation unit (105) which obtains pieces of local frequency information (S103) corresponding to local reference waveforms (S102), based on the local reference waveforms (S102) and an analysis waveform which is the waveform of the mixed audio (S100). Each of the local reference waveforms (S102) (i) constitutes a part of a reference waveform for analyzing a predetermined frequency, (ii) has a predetermined temporal/spatial resolution and (iii) includes at least one of an amplification spectrum and a phase spectrum in the predetermined frequency. The system includes: a specific audio's frequency feature value extraction unit (106) which performs pattern matching between a first set which is the pieces of local frequency information and a second set of pieces of frequency information (S103) of a predetermined specific audio, and extracts the first set of the pieces of local frequency information (S103), based on a result of the pattern matching; and an audio signal generation unit which generates a signal of the specific audio, based on the first set of the pieces of local frequency information (S103) extracted by the specific audio's frequency feature value extraction unit.