Speech Recognition Training System Using Scheduled Audio Review
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition engines in dictation systems for call centers require cumbersome and time-consuming training, which can lead to reduced accuracy, and there is no effective method to predict whether training will increase or maintain accuracy.
Innovation Solution
A system that uses scheduled training with actual audio samples and tests the trained speech recognition engine to determine if accuracy increases or decreases, allowing for efficient training without disrupting customer service representatives, using a processor to process audio and text files and allowing for real-time or batch transcription.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech recognition engines are trained using traditional methods, then the system can adapt to user-specific speech patterns, but the training process becomes time-consuming and reduces accuracy
Solution Approach 1:
The system performs preliminary actions by collecting audio samples and generating training data in advance, before actual training is needed. Audio samples are captured during normal operations, transcribed using the speech recognition engine, and corrected during subsequent review periods, so that training data is ready when needed without disrupting service.
Solution Approach 2:
The training process is segmented into distinct phases: audio sample collection during normal operations, separate transcription generation, dedicated correction periods for accuracy improvement, and final training execution. This segmentation allows each phase to be optimized independently and prevents accuracy loss by ensuring corrections are made before training.
2Manufacturing precision
If training is performed continuously to maintain accuracy, then the speech recognition engine remains up-to-date, but customer service productivity is reduced
Solution Approach 1:
Training is performed periodically rather than continuously. The system operates in cycles: during normal service operations, audio samples are collected passively; then during scheduled intervals, transcriptions are reviewed and corrected; finally, training is executed during low-activity periods. This periodic approach maintains accuracy while minimizing disruption to customer service productivity.
Solution Approach 2:
The system performs self-service by automatically collecting audio samples during normal operations without requiring agent intervention. The speech recognition engine automatically transcribes audio, and the system manages the training data collection and processing autonomously, freeing customer service representatives to focus on their primary duties.
3Manufacturing precision
If training data is collected during normal operations, then real-world accuracy can be improved, but system complexity increases
Solution Approach 1:
The system achieves multi-functionality by using the same speech recognition engine for both real-time transcription during customer service and for generating training data. The audio collection mechanism serves dual purposes: capturing customer service interactions for business purposes and simultaneously gathering training data. This universal approach improves real-world accuracy without adding separate dedicated training hardware or systems.
Data Source
AI summary
A method and apparatus useful to train speech recognition engines is provided. Many of today's speech recognition engines require training to particular individuals to accurately convert speech to text. The training requires the use of significant resources for certain applications. To alleviate the resources, a trainer is provided with the text transcription and the audio file. The trainer updates the text based on the audio file. The changes are provided to the speech recognition to train the recognition engine and update the user profile. In certain aspects, the training is reversible as it is possible to over train the system such that the trained system is actually less proficient.


