Multimodal Emotion Classification via Guidance Vector Assimilation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing multimodal fusion methods for emotion analysis neglect the interactive relationship between multiple modalities, leading to unbalanced contributions and redundancy, which hampers the accuracy of emotion classification.

Innovation Solution

A method utilizing a TokenLearner module to establish a guidance vector through multi-head attention scores, ensuring orthogonality and complementary information across modalities, combined with supervised contrastive learning to refine the model's representation and balance modality contributions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing fusion models use attention mechanism to establish compact multimodal representation, then the model can perform emotion analysis based on extracted information from each modality, but the interactive relationship between information of multiple modalities is neglected and there is redundancy within each modality

Engineering Contradiction:
Improveemotion analysis accuracyVSAvoidinteractive relationship between modalities
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent introduces a guidance vector as an intermediary element that mediates the interaction between multiple modalities. This guidance vector is constructed by selecting important information from each modality and fusing them, then used to guide the attention mechanism to capture the interactive relationships between modalities while filtering out redundant information within each modality.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the multimodal fusion process into distinct stages: first extracting features from each modality separately, then constructing guidance vectors for each modality independently, and finally fusing these guided representations. This segmentation allows the model to handle each modality's redundancy separately while capturing their interactions systematically.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If transformation network is used to obtain cross-modal common subspace by transforming source modality distribution into target modality distribution, then the model can solve non-alignment between data of multiple modalities, but the solution space becomes overly dependent on target modality contribution

Engineering Contradiction:
Improvecross-modal alignment capabilityVSAvoidbalance of modality contributions
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

Instead of transforming source modalities into a target modality's distribution (one-way transformation), the patent inverts the approach by transforming each modality into a common subspace independently, then fusing them. This ensures each modality contributes equally to the solution space without being overly dependent on any single target modality.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent creates a universal common subspace that can accommodate multiple modalities (text, audio, video) simultaneously. Each modality is transformed into this universal space independently, allowing the model to handle any combination of modalities and ensuring balanced contributions from all sources.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If existing methods focus on transformation from text to audio and text to video, then the model can capture some cross-modal relationships, but the possibility of transformation of other modalities is neglected

Engineering Contradiction:
Improvecross-modal relationship captureVSAvoidmodality transformation coverage
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal guidance vector construction approach that works for any modality combination. Instead of limiting transformations to text-to-audio and text-to-video, the model can construct guidance vectors for any modality (audio-to-video, video-to-audio, etc.) by transforming each modality into the common subspace independently, providing full modality transformation coverage.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240119716A1Method for multimodal emotion classification based on modal space assimilation and contrastive learning
Publication Date: 2024.04.11 HANGZHOU DIANZI UNIV
  • US20240119716A1 patent drawing
  • US20240119716A1 patent drawing
  • US20240119716A1 patent drawing

AI summary

The present disclosure provides a method for multimodal emotion classification based on modal space assimilation and contrastive learning. The present disclosure introduces the concept of assimilation. A guidance vector composed of complementary information between modalities is utilized to guide each modality to simultaneously approach a solution space. This operation not only further improves the efficiency of searching for the solution space but also renders heterogeneous spaces of three modalities isomorphic. In a process of making spaces isomorphic, contributions of a plurality of modalities to a final solution space can be effectively balanced to a certain extent. When guiding each modality, this strategy enables a model to be more concerned about emotion features, thereby reducing intra-modal redundancy. Thus, the difficulty of establishing a multimodal representation is reduced.