Dynamic Dictionary Model for Voice Recognition Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice recognition technology in multimedia devices has a low recognition rate for new words due to its reliance on static language models and dictionaries, and it requires time-consuming postprocessing to update dictionaries.
Innovation Solution
The method involves dynamically updating a dictionary model in real-time based on the current state of a multimedia device and applying a speech-to-text (STT) module, with the ability to adjust weights for data in the dictionary model based on data accuracy sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional statistical speech recognition systems based on existing language models and dictionaries are used, then the system structure is simple and easy to implement, but the recognition rate for new words is considerably low
Solution Approach 1:
The patent applies dynamics by making the dictionary model adaptable and changeable rather than static. The system dynamically updates the dictionary model with new words encountered during operation, allowing it to adapt to new vocabulary and contexts. This is achieved through mechanisms that automatically learn and incorporate new word patterns into the recognition system, resolving the contradiction between maintaining simple system structure and improving new word recognition capability.
Solution Approach 2:
The system implements self-service by automatically updating its own dictionary model without requiring external manual intervention. The speech recognition system learns new words autonomously from incoming audio data and updates its internal language model accordingly, enabling it to improve its own performance and expand its vocabulary independently over time.
2Reliability
If the dictionary model is updated each time to improve recognition accuracy, then the voice recognition rate improves, but it takes a lot of time to perform postprocessing operations
Solution Approach 1:
The patent applies preliminary action by performing dictionary model updates in advance or in the background rather than waiting for postprocessing. The system proactively learns and incorporates new words as they are encountered, updating the language model before recognition tasks are needed. This eliminates the need for time-consuming postprocessing operations by ensuring the dictionary is already updated and ready for use.
Solution Approach 2:
The system maintains continuous useful action by continuously updating the dictionary model in real-time as new data is received, rather than performing batch updates during postprocessing. This continuous learning process ensures the recognition system always operates with the most current and accurate language model, improving recognition rates without interrupting or delaying operations.
3Measurement precision
If different weights are set for each data in the dictionary model based on source accuracy, then the recognition accuracy improves, but the device complexity increases
Solution Approach 1:
The patent applies local quality by assigning different weights to different data sources based on their specific accuracy characteristics. Rather than treating all dictionary data uniformly, the system evaluates and weights individual data points or sources according to their reliability, allowing more accurate sources to have greater influence on the language model. This targeted approach improves recognition accuracy without requiring complete system restructuring.
Data Source
AI summary
A control method of a multimedia device, according to one embodiment of the present disclosure, comprises the steps of: displaying a video of content on a screen of the multimedia device; receiving audio data corresponding to a voice uttered by a user, recognizing the voice of the received audio data by referring to a memory; and executing a command according to a result of the voice recognition.


