Multimodal Emotion Analysis via Synchronized Audio-Video Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in accurately analyzing human emotions and mental states using multimodal data, such as facial expressions and audio cues, particularly in diverse cultural and demographic contexts, due to limited annotated datasets and the complexity of synchronizing visual and audio emotional expressions.

Innovation Solution

A machine-trained analysis system that captures contemporaneous audio and video information using a multilayered convolutional network, learning trained weights simultaneously from both modalities to provide emotion metrics, and utilizing techniques like early fusion and semi-supervised learning to enhance emotional analysis across various demographics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional unimodal analysis methods are used, then system complexity is reduced, but measurement precision of emotional states deteriorates

Engineering Contradiction:
Improveemotion analysis accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple information channels (audio, video, physiological data) into a unified multimodal analysis system. The convolutional neural network integrates these different modalities simultaneously, allowing the system to leverage complementary information from each channel to improve emotion analysis accuracy while managing complexity through unified processing architecture.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system is designed to process multiple types of data (audio, video, physiological signals) through a single multimodal neural network architecture. This universal approach allows the same system to handle diverse emotional expression modalities, improving measurement precision across different types of emotional cues without requiring separate specialized systems for each modality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If multimodal data synchronization is implemented, then measurement precision improves, but difficulty of detecting and measuring increases

Engineering Contradiction:
Improveemotional expression detection accuracyVSAvoidsynchronization complexity
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The system performs preliminary synchronization of audio and video data before feeding them into the neural network. By pre-aligning the temporal relationships between different modalities and preprocessing the data to establish proper synchronization, the system reduces the computational burden during actual emotion detection while maintaining high measurement precision.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces intermediate processing layers that act as mediators between raw multimodal inputs and the final emotion classification. These intermediate layers handle the synchronization and alignment of different data streams, transforming complex synchronized multimodal data into structured representations that are easier to process while preserving the precise temporal relationships needed for accurate emotion detection.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If diverse annotated datasets are collected, then adaptability across demographics improves, but loss of time in data collection increases

Engineering Contradiction:
Improvecross-cultural generalization capabilityVSAvoiddata collection duration
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary training on diverse, annotated multimodal datasets to learn population-specific emotional expression patterns before deployment. By pre-collecting and processing diverse data from different cultures, genders, and demographics during the training phase, the system builds generalized knowledge that enables accurate emotion recognition across various populations without requiring extensive data collection during actual use.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The neural network automatically learns and adapts to diverse demographic patterns through self-supervised learning mechanisms. The system can identify and learn population-specific characteristics from the training data without manual intervention for each demographic group, reducing the time required for targeted data collection while maintaining high adaptability across different cultures and demographics.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10628741B2Multimodal machine learning for emotion metrics
Publication Date: 2020.04.21 AFFECTIVA
  • US10628741B2 patent drawing
  • US10628741B2 patent drawing
  • US10628741B2 patent drawing

AI summary

Techniques are described for machine-trained analysis for multimodal machine learning. A computing device captures a plurality of information channels, wherein the plurality of information channels includes contemporaneous audio information and video information from an individual. A multilayered convolutional computing system learns trained weights using the audio information and the video information from the plurality of information channels, wherein the trained weights cover both the audio information and the video information and are trained simultaneously, and wherein the learning facilitates emotional analysis of the audio information and the video information. A second computing device captures further information and analyzes the further information using trained weights to provide an emotion metric based on the further information. Additional information is collected with the plurality of information channels from a second individual and learning the trained weights factors in the additional information. The further information can include only video data or audio data.