Mobile Terminal Voice Recognition Model Update
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice recognition systems face inefficiencies due to high computational costs, time-consuming manual transcription processes, and limited performance improvement in actual environments, especially when using unselective data and relying on either learning or adaptive methods alone.
Innovation Solution
A terminal that classifies learnable data by reliability, generates learning and adaptive models using unsupervised learning, evaluates their performance, and determines whether to update existing acoustic models automatically, reducing human intervention and improving voice recognition performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual transcription process is used to transcribe sampled voice data, then voice recognition performance can be improved, but time consumption and cost increase significantly
Solution Approach 1:
The system performs self-learning by automatically transcribing voice data using the existing voice recognition model without requiring manual transcription. The model learns from its own recognition results, creating a self-service mechanism that eliminates the need for external human intervention in the transcription process.
Solution Approach 2:
The patent introduces an automatic transcription module as an intermediary between the voice recognition model and the learning process. This intermediary automatically transcribes voice data by utilizing the model's recognition capabilities, replacing the need for manual transcription while maintaining the learning quality.
2Quantity of substance
If unselective data is used for learning, then more data is available for training, but learning time rapidly increases
Solution Approach 1:
The system applies different quality standards to different data by classifying voice data into learnable and unlearnable categories based on recognition confidence levels. High-confidence data is selected for learning while low-confidence data is excluded, ensuring that only quality data is used for training without requiring manual review of all data.
Solution Approach 2:
Instead of processing all available voice data, the system selectively processes only the necessary portion - voice data that meets specific confidence thresholds. This partial action approach avoids the time cost of processing excessive data while still providing sufficient training material for model improvement.
3Device complexity
If only one learning method or adaptive method is used, then the system is simpler to implement, but improvement in voice recognition performance in actual environment is limited
Solution Approach 1:
The patent merges two previously separate processes - the voice recognition process and the learning process - into an integrated system. The voice recognition model simultaneously performs recognition and generates learning data, while the learning model continuously improves the recognition model, creating a synergistic combined system that enhances performance without requiring complex external infrastructure.
4Measurement precision
If a person directly configures test set to evaluate model performance, then evaluation accuracy is ensured, but time consumption and cost increase
Solution Approach 1:
The system performs self-evaluation by automatically generating test sets and evaluating model performance without human intervention. The voice recognition model autonomously creates evaluation data and assesses its own performance, eliminating the need for manual test set configuration while maintaining evaluation accuracy.
Data Source
AI summary
A terminal includes a memory configured to store voice data and a processor configured to measure reliability of learnable data stored in the memory, to classify the learnable data into learning data or adaptive data according to the measured reliability, to generate a learning model by performing unsupervised learning with respect to the learning data, to generate an adaptive model using the adaptive data, and to evaluate recognition performance of each of the learning model and the adaptive model.


