Voice Recognition Filter Bank Frequency Adjustment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition systems use fixed filter banks that do not adequately account for variations in vocal tract length among speakers, leading to suboptimal speech recognition accuracy.
Innovation Solution
Performing multiple voice recognition analyses using filter banks with adjustable maximum and minimum frequencies, allowing for dynamic adjustment during speech analysis to produce multiple recognition probabilities, which are then combined to determine a final recognition probability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If fixed filter banks are used with fixed fmin and fmax values, then the voice recognition system is simple to implement, but speech recognition accuracy deteriorates due to inability to account for vocal tract length variations among speakers
Solution Approach 1:
The patent applies dynamics by making the filter bank parameters (fmin and fmax) adjustable rather than fixed. The system dynamically adapts the frequency range of filter banks based on detected vocal tract length characteristics of the speaker, allowing the voice recognition system to optimize its filtering for each speaker's unique vocal anatomy, thereby improving recognition accuracy without requiring a completely complex redesign
Solution Approach 2:
The patent changes the parameters of the filter bank (specifically fmin and fmax values) to adapt to different speakers. By modifying these frequency parameters based on vocal tract length detection, the system accounts for individual speaker variations in formant frequencies, which directly improves speech recognition accuracy while maintaining a relatively simple system architecture
2Measurement precision
If multiple voice recognition analyses are performed with different filter banks, then speech recognition accuracy is improved by accounting for speaker variations, but processing time and computational complexity increase
Solution Approach 1:
The patent applies preliminary action by first detecting the vocal tract length characteristics of the speaker before performing the main voice recognition analysis. This preliminary detection allows the system to pre-adjust the filter bank parameters (fmin and fmax) to match the speaker's vocal characteristics, so that when the actual speech recognition is performed, the filtering is already optimized, reducing the need for multiple trial analyses and thus saving processing time
Solution Approach 2:
The system dynamically adjusts filter bank parameters based on real-time detection of vocal tract length, allowing it to perform a single optimized analysis rather than multiple fixed analyses. This dynamic adaptation enables the system to achieve high accuracy with fewer processing passes, thereby reducing the time loss associated with multiple rigid analyses
Data Source
AI summary
Methods and apparatus for voice recognition are disclosed. A voice signal is obtained and two or more voice recognition analyses are performed on the voice signal. Each voice recognition analysis uses a filter bank defined by a different maximum frequency and a different minimum frequency and wherein each voice recognition analysis produces a recognition probability ri of recognition of one or more speech units, whereby there are two or more recognition probabilities ri. The maximum frequency and the minimum frequency may be adjusted every time speech is windowed and analyzed. A final recognition probability Pf is determined based on the two or more recognition probabilities ri.


