Interaction Analytics System for Meeting Engagement Assessment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Video conferencing lacks real-time and post-interaction feedback on participant engagement and morale, making it difficult for meeting leaders to assess meeting effectiveness, especially in large meetings where manual analysis is limited and existing transcription services fail to recognize social cues.
Innovation Solution
A system and method that analyze audiovisual data from meetings to generate interaction analytics, including sentiment and engagement scores, which are displayed in real-time and post-meeting, using models that assess face sentiment, laughter, text sentiment, and other factors to provide actionable feedback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual analysis of participants is used, then measurement precision of participant engagement is improved, but device complexity and time consumption increase significantly
Solution Approach 1:
The patent replaces manual visual analysis with automated computer vision technology. The system uses deep learning models (ResNet, VGG, etc.) to automatically detect facial expressions, gestures, and body language from video feeds, substituting the mechanical/manual process of human observation with an automated digital system that processes visual data through neural networks.
Solution Approach 2:
The system enables self-service analysis by automatically processing meeting recordings and generating engagement metrics without requiring manual intervention. The automated pipeline captures video data, processes it through pre-trained models, and outputs actionable insights about participant engagement, allowing the system to serve itself rather than requiring human analysts to manually review each participant.
2Productivity
If real-time feedback is provided, then productivity of meeting decisions is improved, but loss of time for data processing increases
Solution Approach 1:
The system performs preliminary action by pre-processing video data in real-time during the meeting. The deep learning models continuously analyze facial expressions and gestures as they occur, preparing engagement metrics before the meeting concludes. This allows the system to have ready-made insights available immediately after the meeting, eliminating delays in post-meeting analysis.
Solution Approach 2:
The system maintains continuous useful action by continuously processing video streams throughout the meeting without interruption. The automated analysis runs concurrently with the meeting, constantly updating engagement metrics based on real-time visual data, ensuring that feedback is always current and ready for immediate use in meeting decisions.
3Loss of information
If existing transcription services are used, then loss of information from audio is reduced, but measurement precision of social cues and engagement remains insufficient
Solution Approach 1:
The patent merges multiple data sources including audio transcription, video facial expression analysis, gesture recognition, and body language detection into a unified engagement assessment system. By combining these complementary technologies, the system achieves both complete audio information capture and accurate social cue recognition, overcoming the limitations of transcription services alone.
Solution Approach 2:
The system uses composite materials in the sense of combining multiple analytical approaches: deep learning models for facial expression recognition, computer vision algorithms for gesture detection, and natural language processing for audio transcription. This composite approach integrates diverse technological components to achieve comprehensive and precise engagement measurement that no single technology could provide alone.
Data Source
AI summary
A method comprises receiving at least one of a transcript, a video recording, an audio recording, or an audiovisual recording of at least a portion of an interaction, and receiving an audiovisual score for a relevant portion of the interaction, the received audiovisual score being based on data received from at least a subset of participants in the interaction. A reaction metric is calculated based on the received audiovisual score. The at least one transcript, video recording, audio recording, or audiovisual recording is then displayed proximate to the reaction metric.


