Transcription Accuracy Verification via User Emotion Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Transcription systems in communication sessions often produce inaccuracies due to factors like quick speech, accents, and background noise, which can lead to misunderstandings, especially for hearing-impaired users relying on text captioned telephone systems.

Innovation Solution

A method and system that verify transcription accuracy by analyzing sound characteristics and emotions of users during communication sessions, using machine learning algorithms to identify inaccuracies based on tone, pitch, volume, and facial expressions, and providing feedback to improve transcription systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If transcription systems process speech quickly to maintain real-time communication, then communication speed is improved, but transcription accuracy deteriorates due to quick speech and background noise

Engineering Contradiction:
Improvecommunication speedVSAvoidtranscription accuracy
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The system performs preliminary transcription of the communication session, then subsequently analyzes sound characteristics and emotions to verify accuracy. This two-stage approach allows real-time communication to proceed while accuracy verification happens afterward, resolving the contradiction between speed and accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system provides feedback about transcription accuracy by analyzing user emotions and speech patterns. This feedback mechanism allows the system to identify inaccurate transcriptions and potentially correct them, improving accuracy without compromising communication speed.

Inventive Principle:
Principle #23Feedback

2Device complexity

If transcription systems use basic speech-to-text conversion, then device complexity is reduced, but transcription accuracy deteriorates due to accents and background noise

Engineering Contradiction:
Improvesystem complexityVSAvoidtranscription accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The system merges multiple analysis components including sound characteristic analysis, emotion recognition, and speech pattern analysis into a unified transcription verification system. This integration allows the system to improve accuracy through multiple factors while managing complexity through unified processing.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system introduces an intermediary verification layer that analyzes user emotions and speech characteristics between the original speech and the final transcription. This intermediary analysis helps identify inaccuracies without requiring complete system redesign.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of time

If no verification mechanism is used, then processing time is reduced, but loss of information increases due to transcription errors

Engineering Contradiction:
Improveprocessing timeVSAvoidinformation accuracy
Core Design Contradiction:
Loss of timeVSLoss of information

Solution Approach 1:

The system performs preliminary transcription without extensive verification to maintain fast processing, then applies targeted verification using emotion and speech pattern analysis only when needed. This approach minimizes processing time while preventing information loss through selective verification.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11699043B2Determination of transcription accuracy
Publication Date: 2023.07.11 SORENSON IP HOLDINGS LLC
  • US11699043B2 patent drawing
  • US11699043B2 patent drawing
  • US11699043B2 patent drawing

AI summary

A method may include obtaining audio of a communication session between a first device of a first user and a second device of a second user. The method may further include obtaining a transcription of second speech of the second user. The method may also include identifying one or more first sound characteristics of first speech of the first user. The method may also include identifying one or more first words indicating a lack of understanding in the first speech. The method may further include determining an experienced emotion of the first user based on the one or more first sound characteristics. The method may also include determining an accuracy of the transcription of the second speech based on the experienced emotion and the one or more first words.