Multimodal Q&A Matching for Image and Video Answer Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice assistants can only provide answers in text form when processing multi-modal data, such as images and videos, limiting the output modality of responses.

Innovation Solution

An electronic device and method that acquires, indexes, and matches various content items within input data to generate queries and answers, enabling responses in multiple modalities like images and videos.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If the voice assistant processes multi-modal data including images and videos, then the information extraction capability is improved, but the output modality remains limited to text form

Engineering Contradiction:
Improveinformation extraction capabilityVSAvoidoutput modality versatility
Core Design Contradiction:
Loss of informationVSAdaptability or versatility

Solution Approach 1:

The voice assistant is enhanced to perform multiple output functions by integrating both text generation and image generation capabilities. The system can now process multi-modal input data and produce answers in diverse formats including text and images, making it universally adaptable to different types of user queries and information needs.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system transitions from a single-dimensional text-only output to multi-dimensional outputs by incorporating image generation capability. This allows the voice assistant to provide answers not only in textual form but also in visual form, adding a new dimension to the interaction modality.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Device complexity

If the voice assistant extracts only text from multi-modal input data, then the processing simplicity is maintained, but the answer completeness and relevance are reduced

Engineering Contradiction:
Improveprocessing simplicityVSAvoidanswer completeness
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The processing pipeline is segmented into distinct modules: a text processing component that handles textual information extraction and generation, and an image processing component that handles visual information extraction and generation. This segmentation allows the system to process different types of content through specialized pathways, maintaining processing simplicity while improving answer completeness.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

An intermediary processing layer is introduced that receives multi-modal input data, separates it into text and image components, and routes them to appropriate processing pathways. This intermediary structure enables the system to handle complex multi-modal data without significantly increasing overall processing complexity, while ensuring both text and image information are utilized for comprehensive answer generation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20260017256A1Electronic device, and question and answer provision method of electronic device
Publication Date: 2026.01.15 SAMSUNG ELECTRONICS CO LTD
  • US20260017256A1 patent drawing
  • US20260017256A1 patent drawing
  • US20260017256A1 patent drawing

AI summary

An electronic device is provided. The electronic device includes memory storing instructions, and at least one processor communicatively coupled to the memory. The instructions, when executed by the at least one processor individually or collectively, cause the electronic device to acquire input data including a plurality of content items, determine a type of each of the plurality of content items included in the acquired input data, index the plurality of content items of each type, generate a candidate query corresponding to the plurality of content items, select, from among the plurality of content items, at least one content item corresponding to the candidate query, match the candidate query and the candidate answer with each other, and store the matched candidate query and candidate answer.