Complex Digital Filter Chain for Real-Time Speech Formant Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems struggle to accurately determine the instantaneous frequency and bandwidth of speech resonances in real-time, which is crucial for effective speech processing, as previous methods either focus solely on frequency estimation or fail to provide instantaneous bandwidth calculations, limiting their usefulness in real-time applications.

Innovation Solution

A system that digitally filters speech signals using a chain of overlapping complex digital filters to estimate both instantaneous frequency and bandwidth of speech resonances in real-time, employing a combination of real and imaginary components and single-lag delays to reconstruct and analyze the speech signal.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If frequency-oriented methods using instantaneous frequency are used for high time-resolution frequency estimates, then time resolution is improved, but bandwidth estimation capability deteriorates (cannot estimate bandwidth)

Engineering Contradiction:
Improvetime resolutionVSAvoidbandwidth estimation capability
Core Design Contradiction:
Loss of timeVSLoss of information

Solution Approach 1:

The patent segments the speech signal into multiple overlapping short-time windows and applies different analysis methods to each window. By dividing the signal processing into multiple segments with different window lengths, the system achieves both high time resolution (using short windows for instantaneous frequency) and bandwidth estimation capability (using longer windows for spectral analysis), thereby resolving the contradiction between time resolution and bandwidth estimation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from one-dimensional frequency analysis to two-dimensional time-frequency analysis by introducing the time dimension through short-time Fourier transform and wavelet transform. This dimensional expansion allows simultaneous estimation of both instantaneous frequency and bandwidth across different time points, overcoming the limitation of traditional frequency-oriented methods that could only provide frequency information.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If simultaneous determination of frequency and bandwidth of speech resonances is performed on the same time scale as speech production (milliseconds), then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improvesimultaneous frequency and bandwidth determination precisionVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary signal preprocessing including segmentation, windowing, and pre-computation of spectral characteristics before the main analysis. By preparing the signal in advance with appropriate window functions and organizing the data structure, the system reduces the computational complexity of the subsequent simultaneous frequency and bandwidth estimation, making real-time processing feasible.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces complex mechanical or analog signal processing systems with digital signal processing algorithms. By using digital computation for spectral analysis, the system achieves precise simultaneous measurement of frequency and bandwidth while reducing physical device complexity. The digital implementation allows flexible parameter adjustment and efficient computation through software-based signal processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If a chain of overlapping complex digital filters is used to filter speech signals, then measurement precision of resonance parameters is improved, but device complexity increases

Engineering Contradiction:
Improveresonance frequency and bandwidth estimation precisionVSAvoidfilter chain complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent designs the complex digital filters to serve multiple functions simultaneously: frequency selection, bandwidth control, and phase information extraction. Each filter in the chain is configured with specific center frequencies and bandwidths that correspond to formant regions, allowing the same filter structure to perform both signal separation and parameter estimation, thereby reducing overall system complexity while maintaining high measurement precision.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements a nested filter structure where multiple bands of filters are organized hierarchically. The filter chain processes signals through progressively narrower bandwidth stages, with each stage nested within the frequency range of the previous stage. This nested arrangement allows efficient decomposition of the speech spectrum into formant regions while minimizing the total number of filters required, thus reducing device complexity.

Inventive Principle:
Principle #7Nested doll (Nesting)

Data Source

PatentUS9311929B2Digital processor based complex acoustic resonance digital speech analysis system
Publication Date: 2016.04.12 ELIZA CORP
  • US9311929B2 patent drawing
  • US9311929B2 patent drawing
  • US9311929B2 patent drawing

AI summary

A speech analysis system uses one or more digital processors to reconstruct a speech signal by accurately extracting speech formants from a digitized version of the speech signal. The system extracts the formants by determining an estimated instantaneous frequency and an estimated instantaneous bandwidth of speech resonances of the digital version of the speech signal in real time. The system digitally filters the digital speech signal using a plurality of complex digital filters in parallel having overlapping bandwidths to ensure that substantially all of the bandwidth of the speech signal is covered. This virtual chain of overlapping complex digital filters produces a corresponding plurality of complex filtered signals. A first estimated frequency and a first estimated bandwidth is generated for each of the filtered signals, and speech resonances of the input speech signal are identified therefrom.