Multimodal Conversation Analysis for Coaching Insights
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems fail to effectively analyze and leverage multiparty conversation data from workplaces, particularly in remote collaborations, to develop employee skills through coaching, due to the difficulty in sorting through rich acoustic and video data generated during videoconference meetings.
Innovation Solution
A machine learning system that analyzes acoustic, video, and text data from multiparty conversations to generate conversation analysis indicators, using multimodal and sequential machine learning to synthesize data across modalities and time segments, providing insights for coaching improvement and skill development.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If acoustic and video data from multiparty conversations are collected and stored, then rich information about user interactions is available, but it becomes very difficult to sort through and leverage the data for coaching and skill development
Solution Approach 1:
The patent segments the complex task of conversation analysis into multiple specialized machine learning models, each handling specific aspects: acoustic data processing, video data processing, and text data processing. This segmentation allows the system to manage large volumes of multiparty conversation data by dividing it into manageable analytical components, addressing the contradiction between data quantity and analysis complexity
Solution Approach 2:
The patent introduces an intermediary layer of machine learning models that act as mediators between the raw acoustic/video/text data and the coaching insights. These models transform unstructured conversation data into structured analysis indicators, making the data leveragable for coaching purposes without requiring direct human sorting through the raw data
2Measurement precision
If machine learning models process acoustic, video, and text data through multiple layers, then comprehensive conversation analysis is achieved, but processing time and computational resources increase
Solution Approach 1:
The patent applies preliminary action by preprocessing acoustic, video, and text data into standardized formats before the main analysis phase. The system prepares the data by extracting relevant features and organizing them in advance, which reduces the computational burden during the actual analysis and speeds up the overall processing while maintaining comprehensive analysis capabilities
Data Source
AI summary
Technology is provided for conversation analysis. The technology includes, receiving multiple utterance representations, where each utterance representation represents a portion of a conversation performed by at least two users, and each utterance representation is associated with video data, acoustic data, and text data. The technology further includes generating a first utterance output by applying video data, acoustic data, and text data of the first utterance representation to a respective video processing part of the machine learning system to generate video, text, and acoustic-based outputs. A second utterance output is further generated for a second user. Conversation analysis indicators are generated by applying, to a sequential machine learning system the combined speaker features and a previous state of the sequential machine learning system.


