Accent Detection for Automatic Closed Captioning in Video Conferencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition technologies fail to effectively address communication difficulties due to accents during cross-lingual and cross-cultural interactions in video conferencing, where participants with different accents may struggle to understand each other.
Innovation Solution
A method is implemented to automatically enable closed captioning in video conferencing when a heavy accent is detected, by receiving user preferences for language and ethnicity, determining the accent level of the speaker, and comparing it to the acceptable accent level, enabling closed captioning if the accent level does not match the user preference.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speech-to-accent conversion technology is used to help users communicate with accented speakers, then communication accessibility is improved, but the system complexity and resource consumption increase significantly
Solution Approach 1:
The patent extracts only the essential function needed for accessibility - detecting accent levels and enabling closed captioning - rather than implementing full speech-to-accent conversion. This selective extraction maintains communication accessibility while avoiding the complexity of comprehensive speech conversion systems.
Solution Approach 2:
The system automatically detects accent levels and enables closed captioning without requiring user intervention or complex configuration. The accent detection and captioning activation happen autonomously, reducing system complexity while maintaining ease of operation for users with accessibility needs.
2Reliability
If closed captioning is always enabled to ensure understanding of accented speech, then communication clarity is improved, but the user experience deteriorates due to unnecessary captions during clear speech
Solution Approach 1:
The system dynamically adjusts closed captioning based on real-time accent level detection. Instead of static always-on or always-off behavior, the captioning feature is dynamically enabled or disabled according to the detected accent level, optimizing both communication clarity and user experience for different speaking conditions.
Solution Approach 2:
The system uses feedback from accent level detection to control closed captioning activation. The accent detection continuously monitors speech characteristics and provides feedback that triggers appropriate captioning responses, ensuring captions appear only when communication clarity would benefit from them.
3Ease of operation
If accent detection and closed captioning activation are implemented, then communication accessibility for accented speakers is improved, but the processing time and computational resources increase
Solution Approach 1:
The system performs partial accent detection focused only on the essential acoustic features needed to determine accent level, rather than comprehensive speech analysis. This partial action approach provides sufficient information for closed captioning activation while minimizing processing time and computational overhead.
Solution Approach 2:
The system performs accent level detection during the initial phases of speech processing, before full speech-to-text conversion is needed. This preliminary action allows the system to determine whether closed captioning should be enabled early in the processing pipeline, avoiding redundant processing when captions are not needed.
Data Source
AI summary
A method for automatically enabling closed captioning in video conferencing when a heavy accent is detected from a current speaker is provided. Language background and/or ethnicity information is received as a user preference. An acceptable accent level is determined according to the user preference. An audio signal of a speaker speaking in a language is received. A pronunciation of the speaker in the audio signal is compared with standard pronunciation for the language. An accent level of the speaker is determined, and the accent level of the speaker is compared to the acceptable accent level. If the comparison determines that the accent level of the speaker does not comply with the acceptable accent level, closed captioning in enabled for the audio signal. If the comparison determines that the accent level of the speaker complies with the acceptable accent level, closed captioning is not enabled for the audio signal.


