Voice Keyword Table Adaptation for Speaker-Dependent Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice recognition systems, such as those used in Voice Keypads, face challenges in achieving high recognition performance, especially with speakers who have heavy accents or divergent speech patterns, and require burdensome user training for speaker-dependent systems, while lacking necessary resources for improved recognition.
Innovation Solution
A method and system that generate a Voice Keyword Table (VKT) with visual, spoken, and phonetic form data, and adapt voice models using TTS-guided pronunciation editing, confusion testing, and user-initiated or new-model-availability-initiated adaptations, to enhance voice recognition performance by updating and modifying voice models on the voice recognition device.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speaker independent voice recognition is used, then ease of operation is improved, but measurement precision deteriorates
Solution Approach 1:
The system performs preliminary speaker adaptation by collecting speech samples during an initialization phase and pre-computing adapted voice models before actual voice recognition tasks. This preliminary action allows the system to maintain speaker independence (no training required) while achieving speaker-dependent accuracy through pre-computed adaptation models.
Solution Approach 2:
The voice recognition system automatically performs speaker adaptation without user intervention by self-adjusting voice models based on collected speech data. The system serves itself by automatically updating acoustic models and pronunciation dictionaries, eliminating the need for manual training while improving recognition accuracy for each speaker.
2Measurement precision
If speaker dependent voice recognition is used, then measurement precision is improved, but device complexity worsens
Solution Approach 1:
The voice recognition system is segmented into modular components: acoustic model adaptation module, pronunciation dictionary module, and recognition engine. This segmentation allows the complex speaker-dependent functionality to be implemented as separate, manageable modules that can be independently optimized and maintained, reducing overall system complexity.
Solution Approach 2:
The system implements a universal adaptation framework that handles both speaker-independent and speaker-dependent recognition using the same core infrastructure. The voice model adaptation mechanism serves multiple purposes: initial speaker adaptation, continuous learning, and handling of accented speech, eliminating the need for separate systems and reducing complexity.
3Measurement precision
If speaker adapted voice recognition is used, then measurement precision is improved, but loss of time worsens
Solution Approach 1:
The system performs partial adaptation by focusing computational resources on adapting only the speaker-specific acoustic model parameters while keeping the general language model intact. This selective adaptation approach achieves speaker-specific accuracy without requiring complete retraining of the entire voice recognition system, significantly reducing adaptation time.
Solution Approach 2:
The system changes parameters by adjusting acoustic model parameters based on speaker characteristics rather than retraining the entire model. This parameter-level adaptation allows rapid speaker-specific customization by modifying only the necessary model parameters, achieving high recognition accuracy in minutes rather than hours.
4Measurement precision
If voice model adaptation is performed, then measurement precision is improved, but use of energy worsens
Solution Approach 1:
The system extracts only the speaker-specific features from speech signals for model adaptation, separating these from the general language understanding components. This extraction approach allows adaptation to focus computational energy only on speaker-specific acoustic characteristics rather than reprocessing the entire speech recognition pipeline, reducing energy consumption while maintaining accuracy.
Data Source
AI summary
An improved voice recognition system in which a Voice Keyword Table is generated and downloaded from a set-up device to a voice recognition device. The VKT includes visual form data, spoken form data, phonetic format data, and an entry corresponding to a keyword, and TTS-generated voice prompts and voice models corresponding to the phonetic format data. A voice recognition system on the voice recognition device is updated by the set-up device. Furthermore, voice models in the voice recognition device are modified by the set-up device.


