Transcription Accuracy Verification via User Emotion Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Transcription systems in communication sessions often produce inaccuracies due to factors like quick speech, accents, and background noise, which can lead to misunderstandings, especially for hearing-impaired users relying on text captioned telephone systems.
Innovation Solution
A method and system that verify transcription accuracy by analyzing sound characteristics and emotions of users during communication sessions, using machine learning algorithms to identify inaccuracies based on tone, pitch, volume, and facial expressions, and providing feedback to improve transcription systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If transcription systems process speech quickly to maintain real-time communication, then communication speed is improved, but transcription accuracy deteriorates due to quick speech and background noise
Solution Approach 1:
The system performs preliminary transcription of the communication session, then subsequently analyzes sound characteristics and emotions to verify accuracy. This two-stage approach allows real-time communication to proceed while accuracy verification happens afterward, resolving the contradiction between speed and accuracy.
Solution Approach 2:
The system provides feedback about transcription accuracy by analyzing user emotions and speech patterns. This feedback mechanism allows the system to identify inaccurate transcriptions and potentially correct them, improving accuracy without compromising communication speed.
2Device complexity
If transcription systems use basic speech-to-text conversion, then device complexity is reduced, but transcription accuracy deteriorates due to accents and background noise
Solution Approach 1:
The system merges multiple analysis components including sound characteristic analysis, emotion recognition, and speech pattern analysis into a unified transcription verification system. This integration allows the system to improve accuracy through multiple factors while managing complexity through unified processing.
Solution Approach 2:
The system introduces an intermediary verification layer that analyzes user emotions and speech characteristics between the original speech and the final transcription. This intermediary analysis helps identify inaccuracies without requiring complete system redesign.
3Loss of time
If no verification mechanism is used, then processing time is reduced, but loss of information increases due to transcription errors
Solution Approach 1:
The system performs preliminary transcription without extensive verification to maintain fast processing, then applies targeted verification using emotion and speech pattern analysis only when needed. This approach minimizes processing time while preventing information loss through selective verification.
Data Source
AI summary
A method may include obtaining audio of a communication session between a first device of a first user and a second device of a second user. The method may further include obtaining a transcription of second speech of the second user. The method may also include identifying one or more first sound characteristics of first speech of the first user. The method may also include identifying one or more first words indicating a lack of understanding in the first speech. The method may further include determining an experienced emotion of the first user based on the one or more first sound characteristics. The method may also include determining an accuracy of the transcription of the second speech based on the experienced emotion and the one or more first words.


