Context-Aware Image Retrieval from Text Using Dual AI Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems struggle to provide images that accurately match the context of user-input text without requiring separate keyword inputs or image tags, leading to inefficiencies in image selection.

Innovation Solution

An electronic device employs a first and second AI data recognition model to determine the relatedness between text and images, using deep neural networks to identify and display images that align with the context of the input text, either locally or through a server, thereby facilitating intelligent image selection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional keyword-based image search is used, then users can input specific search terms, but the system cannot accurately understand the context and intended meaning of the text

Engineering Contradiction:
Improveimage-text matching accuracyVSAvoiduser input complexity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent replaces traditional keyword-based mechanical search systems with AI-based semantic understanding systems. Deep learning models analyze the contextual meaning, sentiment, and semantic relationships in text to automatically identify and retrieve relevant images, eliminating the need for manual keyword input while significantly improving matching accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by automatically performing image retrieval based on user input text without requiring users to manually enter keywords or tags. The AI models autonomously interpret the text meaning and select appropriate images, making the process user-friendly while maintaining high precision.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If multiple AI models are used to determine image-relatedness, then image selection accuracy improves, but system complexity increases

Engineering Contradiction:
Improveimage selection accuracyVSAvoidsystem architecture complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple AI models (language understanding models, image recognition models, and relatedness determination models) into an integrated system. These models work together through a unified architecture that processes text and images simultaneously, determining their relatedness through coordinated analysis rather than separate sequential operations.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The AI system is designed with multi-functional models that can perform multiple tasks: text semantic analysis, image feature extraction, sentiment detection, and relatedness scoring. This universal approach allows a single integrated system to handle various image retrieval scenarios without requiring separate specialized systems for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12430561B2Method and electronic device for providing text-related image
Publication Date: 2025.09.30 SAMSUNG ELECTRONICS CO LTD
  • US12430561B2 patent drawing
  • US12430561B2 patent drawing
  • US12430561B2 patent drawing

AI summary

An artificial intelligence (AI) system for simulating functions such as recognition, determination, and so forth of human brains by using a mechanical learning algorithm like deep learning, or the like, and an application thereof is provided. A method of providing a text-related image is provided. The method includes obtaining a text, determining at least one image related to the obtained text based on a degree of relatedness between a result of applying a first AI data recognition model to the obtained text and a result of applying a second AI data recognition model to a user-accessible image, and displaying the determined at least one image to a user.