Voice Recognition Terminal Server System Personalized Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice recognition technologies face challenges in accommodating personalized characteristics without increasing server storage burden and computational load, particularly in implementing voice recognition systems that require two-stage processing between terminals and servers.
Innovation Solution
A voice recognition terminal and server system where the terminal extracts feature data, calculates acoustic model scores, and transmits them to the server for language network processing, reducing data transmission by selecting n-best candidates and performing acoustic model adaptation locally, thus minimizing server load and protecting user privacy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speaker-dependent recognition method is used to reflect personalized characteristics, then recognition accuracy for specific speakers is improved, but device complexity and training procedure requirements increase
Solution Approach 1:
The voice recognition system is segmented into two parts: a speaker-independent recognition device that handles general recognition without training, and a terminal that stores personalized vocabulary. This segmentation allows the system to maintain high accuracy for specific speakers through the terminal's personalized vocabulary while avoiding the complexity of full speaker-dependent training procedures.
Solution Approach 2:
Personalized vocabulary and acoustic models are pre-stored in the terminal before voice recognition is performed. This preliminary action eliminates the need for training procedures during actual use, as the terminal already contains speaker-specific information ready for immediate application in the two-stage recognition process.
2Adaptability or versatility
If online voice recognition method is used to accommodate large-vocabulary language model, then vocabulary coverage is improved, but server storage burden and computational load increase
Solution Approach 1:
The system segments the language model storage between terminal and server. The terminal stores personalized vocabulary and acoustic models locally, while the server stores only the general language model. This segmentation reduces server storage burden by eliminating the need to store personalized information, while still achieving large vocabulary coverage through the combination of both storage locations.
Solution Approach 2:
The terminal performs self-service by storing and processing personalized vocabulary locally without requiring server storage capacity. The terminal independently handles speaker-specific recognition tasks, freeing the server from storing personalized data and reducing its computational load for personalized processing.
3Measurement precision
If two-stage voice recognition is performed between terminal and server, then personalized characteristics are reflected, but processing time and system complexity increase
Solution Approach 1:
The terminal pre-calculates and stores acoustic models and personalized vocabulary information before recognition is needed. During actual voice recognition, the terminal can quickly retrieve and apply this pre-prepared information in the first stage, reducing the time required for the two-stage process while maintaining personalized recognition accuracy.
Data Source
AI summary
A voice recognition terminal, a voice recognition server, and a voice recognition method for performing personalized voice recognition. The voice recognition terminal includes a feature extraction unit for extracting feature data from an input voice signal, an acoustic score calculation unit for calculating acoustic model scores using the feature data, and a communication unit for transmitting the acoustic model scores and state information to a voice recognition server in units of one or more frames, and receiving transcription data from the voice recognition server, wherein the transcription data is recognized using a calculated path of a language network when the voice recognition server calculates the path of the language network using the acoustic model scores.


