Voice Recognition Filter Bank Frequency Adjustment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice recognition systems use fixed filter banks that do not adequately account for variations in vocal tract length among speakers, leading to suboptimal speech recognition accuracy.

Innovation Solution

Performing multiple voice recognition analyses using filter banks with adjustable maximum and minimum frequencies, allowing for dynamic adjustment during speech analysis to produce multiple recognition probabilities, which are then combined to determine a final recognition probability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If fixed filter banks are used with fixed fmin and fmax values, then the voice recognition system is simple to implement, but speech recognition accuracy deteriorates due to inability to account for vocal tract length variations among speakers

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies dynamics by making the filter bank parameters (fmin and fmax) adjustable rather than fixed. The system dynamically adapts the frequency range of filter banks based on detected vocal tract length characteristics of the speaker, allowing the voice recognition system to optimize its filtering for each speaker's unique vocal anatomy, thereby improving recognition accuracy without requiring a completely complex redesign

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameters of the filter bank (specifically fmin and fmax values) to adapt to different speakers. By modifying these frequency parameters based on vocal tract length detection, the system accounts for individual speaker variations in formant frequencies, which directly improves speech recognition accuracy while maintaining a relatively simple system architecture

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If multiple voice recognition analyses are performed with different filter banks, then speech recognition accuracy is improved by accounting for speaker variations, but processing time and computational complexity increase

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by first detecting the vocal tract length characteristics of the speaker before performing the main voice recognition analysis. This preliminary detection allows the system to pre-adjust the filter bank parameters (fmin and fmax) to match the speaker's vocal characteristics, so that when the actual speech recognition is performed, the filtering is already optimized, reducing the need for multiple trial analyses and thus saving processing time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically adjusts filter bank parameters based on real-time detection of vocal tract length, allowing it to perform a single optimized analysis rather than multiple fixed analyses. This dynamic adaptation enables the system to achieve high accuracy with fewer processing passes, thereby reducing the time loss associated with multiple rigid analyses

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS8010358B2Voice recognition with parallel gender and age normalization
Publication Date: 2011.08.30 SONY INTERACTIVE ENTERTAINMENT LLC
  • US8010358B2 patent drawing
  • US8010358B2 patent drawing
  • US8010358B2 patent drawing

AI summary

Methods and apparatus for voice recognition are disclosed. A voice signal is obtained and two or more voice recognition analyses are performed on the voice signal. Each voice recognition analysis uses a filter bank defined by a different maximum frequency and a different minimum frequency and wherein each voice recognition analysis produces a recognition probability ri of recognition of one or more speech units, whereby there are two or more recognition probabilities ri. The maximum frequency and the minimum frequency may be adjusted every time speech is windowed and analyzed. A final recognition probability Pf is determined based on the two or more recognition probabilities ri.