Multi-modal Knowledge Conversation Data Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing AI conversation models are limited in generating conversation data that incorporates both text and image knowledge, and they struggle with coreference resolution labeling across multiple images and texts.

Innovation Solution

A conversation data collection system and method that searches and collects both text and image knowledge related to user utterances, enabling coreference resolution labeling between all images and texts in a conversation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If only text-based knowledge is used for conversation data collection, then the system is simple to operate, but it cannot collect image knowledge-based conversation data

Engineering Contradiction:
Improvemulti-modal knowledge supportVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The conversation data collection system is designed to handle multiple types of knowledge (text and images) through a unified interface and processing pipeline. The system can collect both text-based and image-based conversation data using the same basic framework, making it versatile without requiring separate systems for each modality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system transitions from single-modal (text-only) to multi-modal (text and image) by adding the image dimension. This is achieved by integrating image search functionality, image display capabilities, and image-based knowledge extraction while maintaining the existing text processing infrastructure.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If coreference resolution labeling is limited to texts only, then the labeling process is simple, but it cannot resolve coreference between images and texts

Engineering Contradiction:
Improvecoreference resolution accuracyVSAvoidlabeling complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The coreference resolution process is segmented into distinct stages: text mention identification, image object detection, cross-modal matching, and entity grouping. This segmentation allows the system to handle text-text, image-image, and text-image coreference relationships systematically without overwhelming complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary layer that bridges text and image modalities. This intermediary performs cross-modal coreference resolution by matching text mentions with image objects through shared semantic representations, enabling accurate coreference resolution across different modalities.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If knowledge selection is limited to displayed sentences, then the operation is straightforward, but the ability to select knowledge is limited

Engineering Contradiction:
Improveknowledge selection flexibilityVSAvoidoperation simplicity
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The knowledge selection interface is made dynamic by allowing users to navigate through search results, filter by relevance, and select multiple knowledge sources. The system adapts to user preferences and conversation context to present relevant knowledge options, making the selection process both flexible and user-friendly.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250139153A1Multi-modal knowledge-based conversation data generation and additional information labeling system
Publication Date: 2025.05.01 KOREA ELECTRONICS TECH INST
  • US20250139153A1 patent drawing
  • US20250139153A1 patent drawing
  • US20250139153A1 patent drawing

AI summary

There is provided a multi-modal knowledge-based conversation data generation and additional information labeling system. A conversation data generation method according to an embodiment includes: receiving input of a user utterance; searching pieces of text knowledge related to the user utterance; searching images related to the user utterance; receiving input of an answer to the user utterance referring to the searched text knowledge and images; and collecting the user utterances, the pieces of text knowledge, the images, and the answer to the user utterances as conversation data. Accordingly, open domain knowledge-based conversation data may be established by using both images and texts, and an image may be used as utterance and information acquired from the image may be used as utterance, so that AI conversation data similar to actual conversations can be established.