Speech Recognition Dictionary Switching for Accurate Call Transcription
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Operators face challenges in selecting the appropriate speech recognition dictionary, leading to inaccurate outcomes when multiple dictionaries are available for use in contact centers.
Innovation Solution
An information processing system that includes a selection part to choose a speech recognition dictionary and a speech recognition part to generate text using the selected dictionary, with the ability to reprocess voices using the new dictionary after a switch, ensuring accurate recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple speech recognition dictionaries are provided for different purposes and languages, then speech recognition adaptability is improved, but operator selection difficulty increases
Solution Approach 1:
The speech recognition system automatically selects the appropriate dictionary based on the voice call context without requiring operator intervention. The system monitors the voice call content and autonomously switches between dictionaries, making the system self-serve the dictionary selection function and eliminating the operator's selection burden.
Solution Approach 2:
The system dynamically changes the dictionary parameter based on the detected speech content and context. By monitoring keywords, language patterns, and call type, the system adjusts the dictionary parameter in real-time to match the appropriate dictionary for accurate recognition.
2Device complexity
If a default general-purpose speech recognition dictionary is used, then system complexity is reduced, but speech recognition accuracy deteriorates
Solution Approach 1:
The system transitions from a static default dictionary approach to a dynamic dictionary selection mechanism. The dictionary changes dynamically based on real-time analysis of the voice call content, ensuring the most appropriate dictionary is used for each specific recognition task while maintaining system simplicity.
Solution Approach 2:
The system implements feedback by analyzing the voice call content and using that information to select the appropriate dictionary. The analysis results feed back into the dictionary selection process, creating a closed-loop system that continuously optimizes recognition accuracy based on actual call conditions.
3Speed
If speech recognition is performed in real-time during voice calls, then response speed is improved, but recognition accuracy with multiple dictionaries deteriorates
Solution Approach 1:
The system performs preliminary analysis of the voice call content to determine the appropriate dictionary before full speech recognition begins. This preliminary action allows the system to prepare the correct dictionary in advance, ensuring both real-time response speed and high recognition accuracy by having the appropriate recognition resources ready.
Data Source
AI summary
An information processing system includes:a selection part configured to select a speech recognition dictionary for use in speech recognition from among a plurality of speech recognition dictionaries; anda speech recognition part configured to generate speech recognition text by converting voices uttered during a voice call with a customer, into text, by speech recognition using the speech recognition dictionary selected by the selection part.In this information processing system, when a switchover to a different speech recognition dictionary selected is made by the selection part, the speech recognition part is configured to generate speech recognition text by converting voices uttered before the switchover is made, among the voices uttered during the voice call with the customer, into text, by speech recognition using the different speech recognition dictionary.


