Multimodal Sentiment Detection With Temporal Attention Weighting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems face challenges in accurately detecting user sentiment and emotions during interactions, which limits their ability to provide personalized and responsive human-computer interfaces.

Innovation Solution

The development of a cross-modal sentiment detection system that utilizes multimodal temporal attention models to process audio and image data, along with language inputs, to estimate sentiment scores and categories, enabling the system to adjust its operations based on the user's emotional state and improve interaction accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition systems use only audio data for sentiment detection, then the system complexity is low, but the sentiment detection accuracy is insufficient

Engineering Contradiction:
Improvesentiment detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple data modalities (audio, image, and language inputs) into a unified sentiment detection framework. The multimodal temporal attention model integrates these diverse inputs to achieve more accurate sentiment detection than single-modality systems, directly resolving the contradiction between accuracy and complexity by showing that the performance gain justifies the increased system complexity

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system is designed to process multiple types of inputs (audio signals, image data, and language inputs) through a single multimodal temporal attention model. This multi-functional approach allows the system to leverage diverse data sources for sentiment detection, improving accuracy while maintaining a cohesive system architecture that manages complexity effectively

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If the system processes multiple data modalities (audio, image, language), then the sentiment detection accuracy improves, but the computational resources and processing time increase

Engineering Contradiction:
Improvesentiment detection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary processing of audio, image, and language inputs before they are fed into the multimodal temporal attention model. By pre-processing and preparing the data in advance, the system reduces the computational burden during actual sentiment detection, thereby decreasing processing time while maintaining the benefits of multimodal analysis

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The temporal attention mechanism dynamically adjusts the weighting and processing of different modalities based on their relevance and quality. This dynamic approach allows the system to focus computational resources on the most informative inputs at each time step, improving efficiency and reducing overall processing time while maintaining high detection accuracy

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If the system uses multimodal temporal attention models to process multiple inputs, then the sentiment detection accuracy improves, but the device complexity increases

Engineering Contradiction:
Improvesentiment detection accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The multimodal temporal attention model is segmented into distinct processing components for audio, image, and language inputs, with separate attention mechanisms for each modality. This segmentation allows for more manageable and interpretable system architecture, reducing the perceived complexity while maintaining the accuracy benefits of integrated multimodal processing

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11501794B1Multimodal sentiment detection
Publication Date: 2022.11.15 AMAZON TECH INC
  • US11501794B1 patent drawing
  • US11501794B1 patent drawing
  • US11501794B1 patent drawing

AI summary

Described herein is a system for improving sentiment detection and/or recognition using multiple inputs. For example, an autonomously motile device is configured to generate audio data and/or image data and perform sentiment detection processing. The device may process the audio data and the image data using a multimodal temporal attention model to generate sentiment data that estimates a sentiment score and/or a sentiment category. In some examples, the device may also process language data (e.g., lexical information) using the multimodal temporal attention model. The device can adjust its operations based on the sentiment data. For example, the device may improve an interaction with the user by estimating the user's current emotional state, or can change a position of the device and/or sensor(s) of the device relative to the user to improve an accuracy of the sentiment data.