Visually Guided Embedding Space for Multi-Modal Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional language models lack visual intuition, leading to inaccuracies in tasks like text classification, natural language understanding, and digital content searches due to their inability to recognize visually similar digital images associated with text, resulting in inefficient navigation and resource usage.

Innovation Solution

A visually guided machine-learning model is trained to support a unified embedding space that clusters text describing similar visual concepts together, using a combination of text and digital image embeddings with relational operators for real-time search and focus adaptation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional language models are used for digital content search, then text classification and natural language understanding can be performed, but visual intuition is lost causing inaccurate search results

Engineering Contradiction:
Improvesearch accuracyVSAvoidmodel architecture
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges text encoding and image encoding into a unified embedding space architecture. The text encoder and image encoder both produce embeddings that can be directly compared in the same vector space, enabling the model to understand visual concepts through text and vice versa. This integration resolves the contradiction by combining multiple data modalities into a single coherent representation system.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The embedding space is designed to be universal, accommodating both text and image data in the same vector representation. This multi-functional embedding space can handle diverse input types (text queries, image queries, text-image pairs) and produce consistent results across different search scenarios, improving search accuracy without requiring separate processing pipelines.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If conventional search techniques use second order ranking, then search results can be refined, but real time output is hindered

Engineering Contradiction:
Improvesearch response timeVSAvoidsearch result accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary encoding of both text and image data into embeddings during the indexing phase, creating a pre-processed representation that is ready for direct comparison. This preliminary action eliminates the need for sequential re-ranking operations during query processing, enabling real-time search results while maintaining accuracy through the pre-established embedding relationships.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the mechanical sequential processing of second-order ranking with a direct vector-based comparison mechanism. Instead of iteratively refining results through multiple ranking passes, the system uses direct cosine similarity or dot product calculations between embeddings to immediately produce sorted results, achieving both speed and accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If conventional search techniques are used, then basic search functionality is provided, but flexibility to weight particular items is limited

Engineering Contradiction:
Improvefocus adaptationVSAvoidcomputational resource usage
Core Design Contradiction:
Adaptability or versatilityVSLoss of energy

Solution Approach 1:

The search system incorporates dynamic weighting capabilities where the importance of different embedding components can be adjusted in real-time. Users can assign different weights to text embeddings versus image embeddings, or emphasize specific visual features, allowing the search to adapt to different query intentions. This dynamic adjustment is achieved through simple scalar multiplication of embedding vectors, maintaining computational efficiency while enhancing versatility.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11604822B2Multi-modal differential search with real-time focus adaptation
Publication Date: 2023.03.14 ADOBE INC
  • US11604822B2 patent drawing
  • US11604822B2 patent drawing
  • US11604822B2 patent drawing

AI summary

Multi-modal differential search with real-time focus adaptation techniques are described that overcome the challenges of conventional techniques in a variety of ways. In one example, a model is trained to support a visually guided machine-learning embedding space that supports visual intuition as to “what” is represented by text. The visually guided language embedding space supported by the model, once trained, may then be used to support visual intuition as part of a variety of functionality. In one such example, the visually guided language embedding space as implemented by the model may be leveraged as part of a multi-modal differential search to support search of digital images and other digital content with real-time focus adaptation which overcomes the challenges of conventional techniques.