Voice Recognition Threshold Adaptation for Misrecognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice recognition systems using fixed threshold values for trigger recognition suffer from misrecognition due to variable user and environmental changes, leading to reduced recognition rates.
Innovation Solution
A voice recognition apparatus and method that adapts by dynamically changing the preset threshold value based on recognition results, storing successful and failed recognition data as speaker and background models, and re-adjusting the threshold value to improve accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a fixed threshold value is used for trigger recognition, then the system is simple and easy to operate, but the recognition accuracy deteriorates due to misrecognition in variable environments
Solution Approach 1:
The patent applies the dynamics principle by transitioning from a fixed threshold value to a dynamically adjusted threshold. The voice recognition processor automatically changes the threshold value based on recognition results and environmental factors, allowing the system to adapt to variable conditions while maintaining ease of operation through automated adjustment rather than manual configuration.
Solution Approach 2:
The patent implements feedback by using recognition results to adjust the threshold value. The voice recognition processor monitors recognition performance and modifies the threshold accordingly, creating a closed-loop system that continuously improves accuracy while requiring no additional user input or configuration.
2Device complexity
If a fixed threshold value is used for trigger recognition, then the device complexity is low, but the adaptability to environmental changes deteriorates
Solution Approach 1:
The patent applies self-service by enabling the system to automatically adjust its own threshold value based on recognition results without external intervention. The voice recognition processor monitors performance and modifies the threshold autonomously, providing adaptability while keeping the user interface simple and the overall system manageable.
Solution Approach 2:
The patent implements parameter changes by dynamically modifying the threshold value parameter based on environmental conditions and recognition performance. This allows the system to adapt to varying conditions while maintaining relatively simple device architecture, as the adjustment is performed through software-based parameter modification rather than hardware complexity.
3Reliability
If a high threshold value is used to prevent misrecognition, then the false alarm rate decreases, but the recognition rate deteriorates due to missed valid triggers
Solution Approach 1:
The patent resolves this contradiction by making the threshold dynamic rather than static. The voice recognition processor adjusts the threshold value based on current environmental conditions and recognition patterns, allowing the system to maintain high reliability when needed while improving recognition rate when conditions permit, thus balancing false alarm prevention with valid trigger detection.
Solution Approach 2:
The patent applies parameter changes by modifying the threshold parameter adaptively. When the system detects patterns indicating potential misrecognition or false alarms, it adjusts the threshold to reduce false positives. Conversely, when valid triggers are missed, the threshold is adjusted to improve detection sensitivity, thereby optimizing the balance between reliability and productivity.
Data Source
AI summary
A voice recognition apparatus, a voice recognition method, and a non-transitory computer readable recording medium are provided. The voice recognition apparatus includes a storage configured to store a preset threshold value for voice recognition; a voice receiver configured to receive a voice signal of an uttered voice; and a voice recognition processor configured to recognize a voice recognition starting word from the received voice signal, perform the voice recognition on the voice signal in response to a similarity score, which represents a recognition result of the recognized voice recognition starting word, being greater than or equal to the stored preset threshold value, and change the preset threshold value based on the recognition result of the voice recognition starting word.


