Composite Embedding Retrieval for Complex Multimodal Object Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing AI models struggle to accurately extract image objects related to complex natural language text due to variations in data complexity, leading to inconsistent performance in multimodal object extraction.

Innovation Solution

A device and method utilizing composite embedding that generates natural language and image embeddings, including overall sentence and key word/object embeddings, to measure similarity and extract relevant images from a large dataset, with error measurement and model parameter updates based on ground truth information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If AI models are trained using simple natural language text and simple image objects, then the model can accurately extract image objects for simple queries, but the model fails to accurately extract image objects for complex natural language text and complex images

Engineering Contradiction:
Improveaccuracy of multimodal object extractionVSAvoidability to handle complex natural language and image data
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the embedding process into multiple components: overall sentence embedding and key word embedding for natural language, and overall image embedding and key object embedding for images. This segmentation allows the model to handle both simple and complex inputs by focusing on relevant portions, thereby improving accuracy for complex queries without sacrificing performance on simple ones.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by generating key word embeddings and key object embeddings that capture the most important semantic information in the input. This allows the model to pay attention to specific critical elements in complex sentences and images, improving extraction accuracy for complex multimodal data while maintaining efficiency.

Inventive Principle:
Principle #3Local quality

2Productivity

If the AI model extracts only overall sentence embedding and overall image embedding, then the processing is simple and fast, but the accuracy decreases for complex natural language and complex images

Engineering Contradiction:
Improveprocessing speedVSAvoidaccuracy of multimodal object extraction
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent merges multiple embedding types (overall sentence embedding, key word embedding, overall image embedding, key object embedding) into a composite embedding structure. This combination allows the model to process complex information efficiently by integrating multiple levels of representation, thereby maintaining both speed and accuracy for complex multimodal extraction tasks.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates composite embeddings that combine different types of embeddings (overall and key-level) to form a richer representation. This composite approach enables the model to capture both global context and local details, improving accuracy for complex inputs while maintaining processing efficiency through structured integration.

Inventive Principle:
Principle #40Composite materials

3Adaptability or versatility

If the AI model is trained on diverse and complex data, then the model can handle complex natural language and images, but the training process becomes more difficult and time-consuming

Engineering Contradiction:
Improveability to handle complex natural language and image dataVSAvoidtraining time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-extracting and storing key word embeddings and key object embeddings during the training phase. This preliminary processing of training data allows the model to learn effective representations in advance, reducing the time needed for inference on complex queries while maintaining high adaptability to diverse data.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250384669A1Device and method for retrieving multimodal object based on composite embedding
Publication Date: 2025.12.18 ELECTRONICS & TELECOMM RES INST
  • US20250384669A1 patent drawing
  • US20250384669A1 patent drawing
  • US20250384669A1 patent drawing

AI summary

Provided are a device and method for extracting a multimodal object on the basis of composite embedding. The device extracts training natural language text and training images from a training data storage, generates image composite embeddings including embeddings of the training images and key objects included in the training images, generates natural language composite embeddings on the basis of the training natural language text, and measure multimodal similarities between the image composite embeddings and the natural language composite embeddings.