Unified Meeting Dynamics Detection via ML Audio-Visual Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing meeting analysis approaches fail to combine audio-visual streams from multiple locations to create a unified representation of meeting dynamics, including emotional states, agreement/disagreement on topics, and participant interactions, which hinders effective remote participation and coordination in online meetings.
Innovation Solution
The system uses machine learning algorithms to capture and analyze video and audio streams from multiple locations, tracking emotional, attentional, and dispositional states of participants, and takes actions based on this analysis, such as inviting missing participants or scheduling side meetings, using avatars to represent participant states and facilitate improved interaction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If existing meeting analysis approaches are used, then basic audio-visual recording is achieved, but unified representation of meeting dynamics including emotional states and participant interactions is not created
Solution Approach 1:
The patent merges audio streams, video streams, and sensor data from multiple locations into a unified meeting representation. The system combines audio-visual streams from different physical spaces and synthesizes them into a single comprehensive model that captures emotional states, agreement/disagreement dynamics, and participant interactions across all locations simultaneously.
Solution Approach 2:
The patent introduces machine learning models as intermediaries that process raw audio-visual sensor data and transform it into meaningful meeting dynamics representations. These ML models act as mediators between the complex multi-location data streams and the unified meeting representation, extracting emotional states, attentional focus, and interaction patterns without requiring direct complex processing of all raw sensors.
2Measurement precision
If machine learning algorithms are used to track emotional and attentional states, then participant engagement analysis is improved, but processing time and computational resources increase
Solution Approach 1:
The patent performs preliminary processing of audio and video streams by extracting relevant features (such as audio volume, speech patterns, facial expressions) before feeding them into the machine learning models. This preliminary feature extraction prepares the data in advance, enabling faster and more efficient emotional state detection during the actual meeting without requiring complex real-time processing of raw sensor data.
Solution Approach 2:
The patent segments the meeting analysis into distinct components: audio stream processing, video stream processing, sensor data processing, and integration into unified representation. Each segment is processed independently by specialized machine learning models, allowing parallel computation and reducing overall processing time while maintaining high measurement precision for each emotional and attentional state.
3Loss of information
If avatars are used to represent participant states, then interaction transparency is enhanced, but system complexity increases
Solution Approach 1:
The patent creates simplified digital avatar copies that represent each participant's emotional and attentional states. These avatars are graphical representations that mirror the actual participant's engagement level, emotional state, and interaction patterns. The avatars provide a simplified visual interface that conveys complex meeting dynamics information without requiring the full complexity of the underlying sensor and machine learning processing systems.
Data Source
AI summary
In one embodiment, in accordance with the present invention, a method, computer program product, and system for performing actions based on captured interpersonal interactions during a meeting is provided. One or more computer processors capture the interpersonal interactions between people in a physical space during a period of time, using machine learning algorithms to detect the interpersonal interactions and a state of each person based on vision and audio sensors in the physical space. The one or more computer processors analyze and categorize the interactions and state of each person, and tag representations of each person with the respectively analyzed and categorized interactions and states of the respective person over the period of time. The one or more computer processors then take an action based on the analysis.


