Audio Video Analytics for Meeting Engagement Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Video conferencing lacks real-time and post-interaction feedback on participant engagement and morale, making it difficult for meeting leaders to assess meeting effectiveness, especially in large meetings where manual analysis is limited and existing transcription services fail to recognize social cues.

Innovation Solution

A system that combines audio and video analytics to calculate scores for participant engagement and sentiment, generating alerts and recommendations based on data from video and audio feeds, which can be displayed in real-time or post-interaction, using models that analyze factors like face sentiment, laughter, and talking status.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual analysis of participants is used, then analysis precision is improved, but device complexity and time consumption increase significantly

Engineering Contradiction:
Improveparticipant engagement analysis precisionVSAvoidanalysis system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces manual mechanical analysis with automated computer vision and audio processing systems. Video feeds are processed by machine learning models to detect facial expressions, body language, and engagement levels, while audio processing transcribes and analyzes speech patterns. This substitution eliminates the need for manual observation while providing precise, objective measurements of participant engagement.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-service analysis by automatically processing video and audio data without human intervention. The automated models continuously monitor participant behavior, generate engagement scores, and provide real-time feedback, allowing the system to serve itself rather than requiring manual analysis for each interaction.

Inventive Principle:
Principle #25Self-service

2Productivity

If real-time feedback is provided, then meeting effectiveness is improved, but data processing complexity and computational resources increase

Engineering Contradiction:
Improvemeeting effectivenessVSAvoiddata processing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the data processing into distinct modules: video processing for visual engagement analysis, audio processing for speech and tone analysis, and integration modules that combine these streams. Each module processes specific aspects independently before consolidating results, which manages computational complexity while enabling real-time feedback.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies partial action by focusing computational resources on the most impactful metrics rather than analyzing every possible parameter. It prioritizes key engagement indicators such as facial expressions, speech clarity, and interaction frequency, providing real-time feedback on these critical aspects without exhaustively processing all possible data dimensions.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If existing transcription services are used, then audio data processing is improved, but social cue recognition and engagement detection remain insufficient

Engineering Contradiction:
Improveaudio data processing efficiencyVSAvoidsocial cue recognition precision
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent merges audio processing capabilities with video processing capabilities into a unified analysis system. While audio transcription provides the foundation for understanding speech content, the video stream is simultaneously analyzed for facial expressions, body language, and social cues. This combination allows the system to accurately detect engagement levels and social dynamics that audio alone cannot capture.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system creates a composite analysis approach by integrating multiple data streams (video, audio, and processed metadata) into a comprehensive engagement assessment. The audio transcript data is combined with visual engagement metrics and social cue detections to produce a holistic view of participant interaction quality, achieving precision that neither component could achieve alone.

Inventive Principle:
Principle #40Composite materials

4Adaptability or versatility

If video conferencing is used for remote work, then work flexibility is improved, but participant engagement and morale deteriorate

Engineering Contradiction:
Improvework flexibilityVSAvoidparticipant engagement
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent implements continuous feedback mechanisms that monitor participant engagement in real-time and provide actionable insights. The system generates engagement scores, identifies disengaged participants, and suggests interventions such as breaking up monotone sections or highlighting active contributors. This feedback loop helps meeting facilitators maintain participant morale and engagement while preserving the flexibility of remote work.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11799679B2Systems and methods for creation and application of interaction analytics
Publication Date: 2023.10.24 READ AI INC
  • US11799679B2 patent drawing
  • US11799679B2 patent drawing
  • US11799679B2 patent drawing

AI summary

A method comprises receiving a set of data from at least one interaction, calculating, an audio score for any audio data received in the set of data, and calculating a video score for any video data received in the set of data. The audio score is joined with the video score to create an audiovisual score, and some subset of measurement, alert or recommendation is sent to a user, the set being based on the audiovisual score and configured to be displayed on a user device.