Audio Video Analytics for Meeting Engagement Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Video conferencing lacks real-time and post-interaction feedback on participant engagement and morale, making it difficult for meeting leaders to assess meeting effectiveness, especially in large meetings where manual analysis is limited and existing transcription services fail to recognize social cues.
Innovation Solution
A system that combines audio and video analytics to calculate scores for participant engagement and sentiment, generating alerts and recommendations based on data from video and audio feeds, which can be displayed in real-time or post-interaction, using models that analyze factors like face sentiment, laughter, and talking status.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual analysis of participants is used, then analysis precision is improved, but device complexity and time consumption increase significantly
Solution Approach 1:
The patent replaces manual mechanical analysis with automated computer vision and audio processing systems. Video feeds are processed by machine learning models to detect facial expressions, body language, and engagement levels, while audio processing transcribes and analyzes speech patterns. This substitution eliminates the need for manual observation while providing precise, objective measurements of participant engagement.
Solution Approach 2:
The system performs self-service analysis by automatically processing video and audio data without human intervention. The automated models continuously monitor participant behavior, generate engagement scores, and provide real-time feedback, allowing the system to serve itself rather than requiring manual analysis for each interaction.
2Productivity
If real-time feedback is provided, then meeting effectiveness is improved, but data processing complexity and computational resources increase
Solution Approach 1:
The patent segments the data processing into distinct modules: video processing for visual engagement analysis, audio processing for speech and tone analysis, and integration modules that combine these streams. Each module processes specific aspects independently before consolidating results, which manages computational complexity while enabling real-time feedback.
Solution Approach 2:
The system applies partial action by focusing computational resources on the most impactful metrics rather than analyzing every possible parameter. It prioritizes key engagement indicators such as facial expressions, speech clarity, and interaction frequency, providing real-time feedback on these critical aspects without exhaustively processing all possible data dimensions.
3Productivity
If existing transcription services are used, then audio data processing is improved, but social cue recognition and engagement detection remain insufficient
Solution Approach 1:
The patent merges audio processing capabilities with video processing capabilities into a unified analysis system. While audio transcription provides the foundation for understanding speech content, the video stream is simultaneously analyzed for facial expressions, body language, and social cues. This combination allows the system to accurately detect engagement levels and social dynamics that audio alone cannot capture.
Solution Approach 2:
The system creates a composite analysis approach by integrating multiple data streams (video, audio, and processed metadata) into a comprehensive engagement assessment. The audio transcript data is combined with visual engagement metrics and social cue detections to produce a holistic view of participant interaction quality, achieving precision that neither component could achieve alone.
4Adaptability or versatility
If video conferencing is used for remote work, then work flexibility is improved, but participant engagement and morale deteriorate
Solution Approach 1:
The patent implements continuous feedback mechanisms that monitor participant engagement in real-time and provide actionable insights. The system generates engagement scores, identifies disengaged participants, and suggests interventions such as breaking up monotone sections or highlighting active contributors. This feedback loop helps meeting facilitators maintain participant morale and engagement while preserving the flexibility of remote work.
Data Source
AI summary
A method comprises receiving a set of data from at least one interaction, calculating, an audio score for any audio data received in the set of data, and calculating a video score for any video data received in the set of data. The audio score is joined with the video score to create an audiovisual score, and some subset of measurement, alert or recommendation is sent to a user, the set being based on the audiovisual score and configured to be displayed on a user device.


