Speech Dereverberation via Probabilistic Likelihood Maximization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for speech dereverberation, such as HERB and SBD, effectively utilize speech signal features but lack analytical frameworks for optimizing performance, and existing techniques struggle to improve automatic speech recognition (ASR) performance when reverberation time exceeds 0.5 seconds.

Innovation Solution

A speech dereverberation apparatus and method that employs a likelihood maximization unit to determine a source signal estimate by maximizing a likelihood function based on probabilistic models of source and room acoustics, using an iterative optimization algorithm like the Expectation-Maximization algorithm, integrating source signal features with room acoustics through EM iterations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional dereverberation methods (HERB, SBD) are used, then speech signal features are utilized, but ASR performance cannot be improved when reverberation time exceeds 0.5 seconds

Engineering Contradiction:
ImproveASR performanceVSAvoidreverberation time
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent changes the fundamental approach from feature-based processing to probabilistic modeling by introducing a likelihood function that models both source speech and room acoustics as probabilistic entities. This allows the system to handle reverberation times exceeding 0.5 seconds by optimizing the likelihood function rather than relying on fixed feature extraction methods.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary probabilistic framework that mediates between the observed reverberant signal and the source speech estimate. The likelihood function acts as an intermediary that integrates source features and room acoustics models to produce optimized source estimates, rather than directly processing the reverberant signal.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If iterative dereverberation methods are applied, then ASR performance improves for sufficiently long signals, but computational complexity increases

Engineering Contradiction:
ImproveASR performanceVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent employs iterative optimization of the likelihood function where each iteration refines the source estimate continuously. The process maintains useful action by repeatedly updating the source speech estimate using the current estimate and observed signal until convergence, ensuring continuous improvement of ASR performance while managing computational resources.

Inventive Principle:
Principle #20Continuity of useful action

3Measurement precision

If probabilistic models of source and room acoustics are integrated, then energy decay curves are enhanced and speech intelligibility improves, but device complexity increases

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges the source speech model and room acoustics model into a unified likelihood function. This combining of probabilistic models allows simultaneous optimization of both source estimation and reverberation compensation, enhancing energy decay curves and speech intelligibility while managing the complexity through integrated formulation.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP2013869B1Method and apparatus for speech dereverberation based on probabilistic models of source and room acoustics
Publication Date: 2017.12.13 GEORGIA TECH RES CORP
  • EP2013869B1 patent drawingFigure 1
  • EP2013869B1 patent drawingFigure 2
  • EP2013869B1 patent drawingFigure 3A~3B

AI summary

Speech dereverberation is achieved by accepting an observed signal for initialization (1000) and performing likelihood maximization (2000) which includes Fourier Trasforms (4000).