Voice Recognition Device Using Filter Bank Smoothing for Vocal Tract Variation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice recognition technologies face challenges in improving accuracy, particularly in handling variations in vocal tract lengths and acoustic features across different speakers, leading to suboptimal performance with limited audio data.
Innovation Solution
The method involves segmenting audio signals into frame units, applying a filter bank to determine energy components, smoothing these components, and extracting feature vectors to input into a voice recognition model, while also transforming frequency axes to represent vocal tract length variations and applying room impulse filters to simulate acoustic features, thereby enhancing recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional voice recognition methods are used, then the system is simple to implement, but the recognition accuracy is insufficient due to inability to handle vocal tract length variations and acoustic features
Solution Approach 1:
The audio signal is segmented into multiple frame units, and the frequency spectrum is divided into multiple frequency bands using filter banks. This segmentation allows independent processing of different temporal and spectral components, enabling accurate capture of vocal tract length variations and acoustic features while maintaining manageable computational complexity through localized processing.
Solution Approach 2:
The patent transforms the one-dimensional audio signal into a multi-dimensional feature space by applying filter banks across frequency bands and analyzing multiple frame units temporally. This dimensional expansion creates a comprehensive representation that captures both spectral characteristics (vocal tract length) and temporal dynamics (acoustic features), significantly improving recognition accuracy beyond simple signal processing.
2Measurement precision
If more audio data is collected to improve recognition accuracy, then the recognition rate increases, but the system requires more data storage and processing resources
Solution Approach 1:
The patent extracts essential acoustic features by applying filter banks to extract energy components in different frequency bands, then smoothing these components to remove noise. This extraction process isolates the most discriminative features related to vocal tract length and acoustic characteristics, achieving high recognition accuracy with minimal processed data rather than requiring large volumes of raw audio storage.
Solution Approach 2:
The patent applies smoothing parameters to the energy components extracted from filter banks, transforming raw energy values into smoothed features that retain essential information while reducing data volume. This parameter-based transformation maintains recognition accuracy by preserving critical acoustic patterns while significantly reducing the quantity of data that needs to be stored and processed.
3Measurement precision
If filter banks and smoothing operations are applied to each frame unit, then the feature extraction accuracy improves, but the processing time increases
Solution Approach 1:
The patent applies filter banks and smoothing operations selectively to extract only the most relevant energy components from each frame unit, rather than processing all possible features. This partial action approach focuses computational resources on the most discriminative frequency bands and temporal frames, achieving high feature extraction accuracy while minimizing unnecessary processing time for less relevant data.
Data Source
AI summary
The disclosure relates to an electronic apparatus for recognizing user voice and a method of recognizing, by the electronic apparatus, the user voice. According to an embodiment, the method of recognizing the user voice includes obtaining an audio signal segmented into a plurality of frame units, determining an energy component for each filter bank by applying a filter bank distributed according to a preset scale to a frequency spectrum of the audio signal segmented into the frame units, smoothing the determined energy component for each filter bank, extracting a feature vector of the audio signal based on the smoothed energy component for each filter bank, and recognizing the user voice in the audio signal by inputting the extracted feature vector to a voice recognition model.


