Multimodal Intent Detection via Voice and Image Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Electronic devices struggle to provide user-relevant information as they often fail to accurately determine user intent, emotion, or situation based solely on voice input, leading to irrelevant information being provided.

Innovation Solution

An electronic device equipped with a microphone, camera, and processor that analyzes both voice input and image data to select appropriate user reactions, user states, and conversational categories, ensuring information is provided that aligns with the user's intent, emotion, or situation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the electronic device provides information by analyzing user speech only, then the device can provide feedback based on user input, but the information provided becomes irrelevant to the user's intent, emotion, or situation

Engineering Contradiction:
Improveaccuracy of user intent detectionVSAvoidrelevance of provided information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent combines multiple input modalities (voice input through microphone, image input through camera) to comprehensively analyze user state. The processor integrates analysis results from both voice and image data to determine user intent, emotion, and situation, thereby improving the accuracy and relevance of information provision while compensating for the limitations of single-modality analysis

Inventive Principle:
Principle #5Merging (Combining)

2Device complexity

If the electronic device uses only voice input to determine user state, then the system complexity remains low, but the ability to accurately determine user intent, emotion, or situation deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidaccuracy of user state determination
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent divides the user state determination process into separate analysis modules: one for analyzing voice input characteristics and another for analyzing image input characteristics. Each module independently processes its respective data type and extracts relevant features, which are then integrated by the processor to comprehensively determine user intent, emotion, and situation. This segmentation allows the system to handle multiple data types without overwhelming complexity

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11532181B2Provision of targeted advertisements based on user intent, emotion and context
Publication Date: 2022.12.20 SAMSUNG ELECTRONICS CO LTD
  • US11532181B2 patent drawing
  • US11532181B2 patent drawing
  • US11532181B2 patent drawing

AI summary

An electronic device and method are disclosed herein. The electronic device includes a microphone, a camera, an output device, a memory, and a processor. The processor implements the method, including receiving a voice input and/or capturing an image, and analyze the first voice input or the image to determine at least one of a user's intent, emotion, and situation based on predefined keywords and expressions, identifying a category based on the input, selecting first information based on the category, selecting and outputting a first query prompting confirmation of output of the first information, detect a first responsive input to the first query, and when a condition to output the first information is satisfied, output a second query, detecting a second input responsive to the second query, and selectively outputting the first information based on the second input.