Voice Recognition Device Using Filter Bank Smoothing for Vocal Tract Variation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice recognition technologies face challenges in improving accuracy, particularly in handling variations in vocal tract lengths and acoustic features across different speakers, leading to suboptimal performance with limited audio data.

Innovation Solution

The method involves segmenting audio signals into frame units, applying a filter bank to determine energy components, smoothing these components, and extracting feature vectors to input into a voice recognition model, while also transforming frequency axes to represent vocal tract length variations and applying room impulse filters to simulate acoustic features, thereby enhancing recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional voice recognition methods are used, then the system is simple to implement, but the recognition accuracy is insufficient due to inability to handle vocal tract length variations and acoustic features

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidsignal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The audio signal is segmented into multiple frame units, and the frequency spectrum is divided into multiple frequency bands using filter banks. This segmentation allows independent processing of different temporal and spectral components, enabling accurate capture of vocal tract length variations and acoustic features while maintaining manageable computational complexity through localized processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the one-dimensional audio signal into a multi-dimensional feature space by applying filter banks across frequency bands and analyzing multiple frame units temporally. This dimensional expansion creates a comprehensive representation that captures both spectral characteristics (vocal tract length) and temporal dynamics (acoustic features), significantly improving recognition accuracy beyond simple signal processing.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If more audio data is collected to improve recognition accuracy, then the recognition rate increases, but the system requires more data storage and processing resources

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidaudio data volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts essential acoustic features by applying filter banks to extract energy components in different frequency bands, then smoothing these components to remove noise. This extraction process isolates the most discriminative features related to vocal tract length and acoustic characteristics, achieving high recognition accuracy with minimal processed data rather than requiring large volumes of raw audio storage.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies smoothing parameters to the energy components extracted from filter banks, transforming raw energy values into smoothed features that retain essential information while reducing data volume. This parameter-based transformation maintains recognition accuracy by preserving critical acoustic patterns while significantly reducing the quantity of data that needs to be stored and processed.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If filter banks and smoothing operations are applied to each frame unit, then the feature extraction accuracy improves, but the processing time increases

Engineering Contradiction:
Improvefeature extraction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies filter banks and smoothing operations selectively to extract only the most relevant energy components from each frame unit, rather than processing all possible features. This partial action approach focuses computational resources on the most discriminative frequency bands and temporal frames, achieving high feature extraction accuracy while minimizing unnecessary processing time for less relevant data.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11961522B2Voice recognition device and method
Publication Date: 2024.04.16 SAMSUNG ELECTRONICS CO LTD
  • US11961522B2 patent drawing
  • US11961522B2 patent drawing
  • US11961522B2 patent drawing

AI summary

The disclosure relates to an electronic apparatus for recognizing user voice and a method of recognizing, by the electronic apparatus, the user voice. According to an embodiment, the method of recognizing the user voice includes obtaining an audio signal segmented into a plurality of frame units, determining an energy component for each filter bank by applying a filter bank distributed according to a preset scale to a frequency spectrum of the audio signal segmented into the frame units, smoothing the determined energy component for each filter bank, extracting a feature vector of the audio signal based on the smoothed energy component for each filter bank, and recognizing the user voice in the audio signal by inputting the extracted feature vector to a voice recognition model.