Multimodal Intent Detection via Voice and Image Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Electronic devices struggle to provide user-relevant information as they often fail to accurately determine user intent, emotion, or situation based solely on voice input, leading to irrelevant information being provided.
Innovation Solution
An electronic device equipped with a microphone, camera, and processor that analyzes both voice input and image data to select appropriate user reactions, user states, and conversational categories, ensuring information is provided that aligns with the user's intent, emotion, or situation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the electronic device provides information by analyzing user speech only, then the device can provide feedback based on user input, but the information provided becomes irrelevant to the user's intent, emotion, or situation
Solution Approach 1:
The patent combines multiple input modalities (voice input through microphone, image input through camera) to comprehensively analyze user state. The processor integrates analysis results from both voice and image data to determine user intent, emotion, and situation, thereby improving the accuracy and relevance of information provision while compensating for the limitations of single-modality analysis
2Device complexity
If the electronic device uses only voice input to determine user state, then the system complexity remains low, but the ability to accurately determine user intent, emotion, or situation deteriorates
Solution Approach 1:
The patent divides the user state determination process into separate analysis modules: one for analyzing voice input characteristics and another for analyzing image input characteristics. Each module independently processes its respective data type and extracts relevant features, which are then integrated by the processor to comprehensively determine user intent, emotion, and situation. This segmentation allows the system to handle multiple data types without overwhelming complexity
Data Source
AI summary
An electronic device and method are disclosed herein. The electronic device includes a microphone, a camera, an output device, a memory, and a processor. The processor implements the method, including receiving a voice input and/or capturing an image, and analyze the first voice input or the image to determine at least one of a user's intent, emotion, and situation based on predefined keywords and expressions, identifying a category based on the input, selecting first information based on the category, selecting and outputting a first query prompting confirmation of output of the first information, detect a first responsive input to the first query, and when a condition to output the first information is satisfied, output a second query, detecting a second input responsive to the second query, and selectively outputting the first information based on the second input.


