Visually Guided Language Model Embedding Space

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional language models lack visual intuition, leading to inaccuracies in tasks like text classification, natural language understanding, and digital content searches due to their inability to recognize visually similar digital images associated with text, resulting in inefficient navigation and resource usage.

Innovation Solution

A visually guided machine-learning language model is trained to support a unified embedding space that clusters text describing similar visual concepts together, using a fixed image embedding space and a text encoder trained with digital images, enabling direct comparison and similarity determination between text and image embeddings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional language models are used for text classification and digital content search, then text processing can be performed, but accuracy deteriorates due to lack of visual intuition

Engineering Contradiction:
ImproveaccuracyVSAvoidmodel complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges text encoding and image encoding into a unified embedding space. The text encoder and image encoder are trained to produce embeddings that can be directly compared, combining textual and visual information processing capabilities into a single integrated model architecture that improves accuracy for multimodal tasks.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The unified embedding space serves multiple functions: text classification, image search, and cross-modal retrieval all use the same embedding framework. This multi-functional approach allows the model to handle diverse tasks without requiring separate specialized models, improving reliability while managing complexity through shared resources.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If conventional search techniques use second order ranking, then search results can be refined, but productivity deteriorates due to repeated processing

Engineering Contradiction:
Improvesearch result precisionVSAvoidsearch processing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The model performs preliminary encoding of both text and images into a unified embedding space during training, so that during actual search queries, only simple distance calculations are needed. This pre-computation of embeddings eliminates the need for repeated second-order ranking processing, significantly improving search productivity while maintaining precision through the enriched embedding representations.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If conventional language models process text alone, then computational resources are reduced, but measurement precision deteriorates for visual concept recognition

Engineering Contradiction:
Improvevisual concept recognition accuracyVSAvoidcomputational energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies local quality by processing only the portions of data that are relevant to the specific task. For image search, the image encoder processes visual features locally, while for text classification, only text encoding is applied. This selective processing approach improves measurement precision for visual concepts when needed without unnecessarily increasing computational energy consumption for all tasks.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12079269B2Visually guided machine-learning language model
Publication Date: 2024.09.03 ADOBE INC
  • US12079269B2 patent drawing
  • US12079269B2 patent drawing
  • US12079269B2 patent drawing

AI summary

Visually guided machine-learning language model and embedding techniques are described that overcome the challenges of conventional techniques in a variety of ways. In one example, a model is trained to support a visually guided machine-learning embedding space that supports visual intuition as to “what” is represented by text. The visually guided language embedding space supported by the model, once trained, may then be used to support visual intuition as part of a variety of functionality. In one such example, the visually guided language embedding space as implemented by the model may be leveraged as part of a multi-modal differential search to support search of digital images and other digital content with real-time focus adaptation which overcomes the challenges of conventional techniques.