Speech Impediment Detection Avatar for Real-Time Meeting Feedback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current virtual meeting platforms are inaccessible or difficult to use for users with speech disorders, exacerbating feelings of anxiety and impairing effective communication, and existing speech improvement features do not address speech disorders during live video meetings.
Innovation Solution
A system that utilizes real-time speech analysis and machine learning to detect speech impediments, generating an avatar that provides visual feedback and notifications to users, synchronizing with their speech in real-time to assist in communication.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If visual and auditory aspects of video conferencing are used, then communication effectiveness is improved, but anxiety and speech impediments are exacerbated for users with speech disorders
Solution Approach 1:
The patent introduces an avatar as an intermediary between the user and the communication environment. The avatar serves as a mediator that absorbs the visual-auditory feedback loop, allowing users with speech disorders to communicate without directly experiencing the full impact of visual and auditory cues that trigger anxiety. The avatar acts as a buffer that translates user input into visual representation, reducing the harmful feedback effect.
Solution Approach 2:
The system creates a visual copy (avatar) of the user's speech patterns and lip movements. This copy allows users to see their speech representation without the full emotional and psychological impact of real-time video feedback. The avatar serves as a simplified, controllable representation that maintains communication effectiveness while reducing anxiety-provoking elements.
2Measurement precision
If real-time speech analysis is performed to detect speech impediments, then detection accuracy is improved, but system complexity increases
Solution Approach 1:
The patent replaces complex mechanical signal processing systems with machine learning-based detection algorithms. Instead of using traditional signal processing methods that require multiple sensors and complex analysis, the system employs trained ML models that can detect speech impediments with high accuracy through automated feature extraction and pattern recognition, thereby reducing overall system complexity while maintaining or improving detection precision.
3Loss of information
If an avatar is generated and synchronized with user speech, then visual feedback is improved, but processing time and computational resources increase
Solution Approach 1:
The system performs preliminary actions by pre-processing and analyzing user speech patterns before full avatar generation is required. Speech features are extracted and analyzed in advance, allowing the avatar synchronization to occur more efficiently. The system prepares speech representations and lip movement data beforehand, reducing the real-time processing burden during actual communication.
Solution Approach 2:
The avatar synchronization system employs dynamic processing where the level of detail and processing intensity adjusts based on real-time requirements. During critical speech moments, the system focuses computational resources on key synchronization points rather than maintaining constant high-fidelity processing, thereby reducing overall processing time while maintaining visual feedback quality.
Data Source
AI summary
A system and method and for providing speech assistance during a virtual meeting includes receiving a request over a communication network to provide speech assistance during a virtual meeting between a plurality of participants and analyzing speech data of the virtual meeting, via a speech impediment detection engine, to detect a speech impediment for one of the plurality of participants. Upon detecting the speech impediment, an avatar is automatically generated for the participant experiencing speech impediment and the avatar is synchronized with the participant's speech in real-time during the communication session to provide real-time visual feedback to the participant. The avatar is then provided for display to the participant.


