Multimodal Emotion Recognition Fusion for Video Utterance Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current emotion recognition systems in multimedia content rely predominantly on textual data, neglecting the rich information present in visual and acoustic modalities, leading to inaccuracies in predicting emotional states.

Innovation Solution

A multi-modal fusion-based deep neural network that incorporates acoustic, textual, and visual features using a triplet network with adaptive margin, covariance, and variance loss functions to enhance emotion recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional text-based approaches are used for emotion recognition, then the system complexity is low, but the measurement precision of emotional state prediction deteriorates

Engineering Contradiction:
Improveemotion recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple modalities (text, audio, video) into a unified emotion recognition system. The neural network architecture integrates features from different sources: text embeddings from NLP models, audio features from acoustic analysis, and visual features from video processing, merging them to improve prediction accuracy while managing system complexity through structured integration.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system is designed to process multiple types of input data (text, audio, video) simultaneously using a multi-modal neural network architecture. The model can handle different modalities through dedicated processing branches that converge in the final prediction layer, enabling universal emotion recognition across diverse input types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If multi-modal fusion is implemented, then the measurement precision of emotion recognition improves, but the device complexity increases

Engineering Contradiction:
Improveemotion recognition accuracyVSAvoidnetwork architecture complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The neural network is segmented into distinct processing modules for each modality: text processing branch, audio processing branch, and video processing branch. Each branch independently processes its specific input type and extracts relevant features, which are then fused in subsequent layers. This segmentation manages complexity by organizing the multi-modal processing into manageable, specialized components.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If text transcript accuracy is improved, then the emotion recognition accuracy improves, but the loss of information in visual and acoustic signals increases

Engineering Contradiction:
Improveemotion recognition accuracyVSAvoidvisual and acoustic information loss
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system merges text, audio, and video modalities into a unified emotion recognition framework. By combining multiple information sources, the model compensates for limitations in any single modality. The fusion architecture ensures that visual and acoustic information is integrated alongside text, preventing information loss and providing redundant verification channels for emotion detection.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP4413565B1Emotion recognition in multimedia videos using multi-modal fusion-based deep neural network
Publication Date: 2026.03.18 SONY GROUP CORP
  • EP4413565B1 patent drawingFigure 1
  • EP4413565B1 patent drawingFigure 2
  • EP4413565B1 patent drawingFigure 3

AI summary

A system and method of landmark detection using emotion recognition in multimedia videos using multi-modal fusion based deep neural network is provided. The system includes circuitry and a memory configured to store a multimodal fusion network which includes one or more feature extractors, a network of transformer encoders, a fusion attention network, and an output network coupled to the fusion attention network. The system inputs a multimodal input to the one or more feature extractors. The multimodal input is associated with an utterance depicted in one or more videos. The system generates input embeddings as an output of the one or more feature extractors for the input and further generates a set of emotion-relevant features based on the input embeddings. The system further generates a fused-feature representation of the set of emotion-relevant features and predicts an emotion label for the utterance based on fused-feature representation.