Magnitude-Invariant Image-Text Tokens for Low-Label UI Automation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing deep learning models require large amounts of labeled data for training, which is time-consuming and costly, and they struggle to generalize to new tasks without sufficient data updates, limiting their performance and efficiency.

Innovation Solution

A system for automating user interface workflows using a multimodal agent that integrates human-in-the-loop learning, leveraging synthetic data and core set construction methods to reduce data annotation needs and enhance model performance, while utilizing a Transformer-based architecture for parallel processing and general intelligence.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If large amounts of labeled data are used for training deep learning models, then model performance is improved, but data annotation time and cost increase significantly

Engineering Contradiction:
Improvemodel performanceVSAvoiddata annotation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent uses synthetic data generation to create virtual copies of training data that replicate real-world patterns without requiring actual annotated examples. The system generates synthetic images and text pairs that mimic real data distributions, allowing models to learn from these copies rather than requiring time-consuming manual annotation of real data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system employs self-supervised learning mechanisms where the model generates its own training data through internal processes. The multimodal agent creates synthetic data representations from raw inputs without external human annotation, enabling the system to serve its own training needs autonomously and eliminating the time loss associated with manual data preparation.

Inventive Principle:
Principle #25Self-service

2Reliability

If more training data is collected to improve model performance, then accuracy increases, but the rate of data growth lags behind model parameter growth

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata growth rate
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent generates synthetic data copies that can be produced at any rate needed, independent of real-world data collection constraints. The system creates virtual training samples through algorithmic generation, allowing unlimited scaling of training data without the productivity limitations of collecting real data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system pre-generates synthetic training data in advance before actual data collection becomes necessary. By preparing synthetic data representations beforehand, the system eliminates the bottleneck where data collection rate limits model development, allowing model parameters to scale independently of real data acquisition speed.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If existing deep learning models are used, then they can handle specific tasks, but they struggle to generalize to new tasks without sufficient data updates

Engineering Contradiction:
Improvetask generalization capabilityVSAvoidtraining data quantity
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent implements a universal multimodal agent architecture that can handle diverse tasks across different domains using a single unified model. The system integrates multiple capabilities (image processing, text generation, reasoning) into one versatile agent that adapts to new tasks without requiring task-specific training data, achieving generalization through architectural design rather than data quantity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system changes the fundamental parameters of how models are trained by using synthetic data representations and multimodal integration rather than relying on large quantities of real data. This parameter change in the training approach enables the model to generalize to new tasks by learning underlying patterns and relationships rather than memorizing specific examples.

Inventive Principle:
Principle #35Parameter changes

4Reliability

If human knowledge is integrated into the learning framework, then performance on sparse data improves, but the complexity of the learning system increases

Engineering Contradiction:
Improveperformance on sparse dataVSAvoidlearning framework complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces synthetic data as an intermediary layer between real-world tasks and model training. This intermediary synthetic representation layer encodes human knowledge and patterns in a compressed form, allowing the model to access sophisticated reasoning capabilities without directly integrating complex human expertise into the training framework, thus managing complexity while improving performance on sparse data.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20260105248A1Magnitude Invariant Multimodal Agent for Efficient Image-Text Interface Automation
Publication Date: 2026.04.16 ANTHROPIC PBC
  • US20260105248A1 patent drawing
  • US20260105248A1 patent drawing
  • US20260105248A1 patent drawing

AI summary

A system for magnitude-invariant image-text agentic interface automation is disclosed. A bit vectorization logic is configured to convert image patches in a plurality of image patches into magnitude-invariant bit vectors, and generate a plurality of lines of magnitude-invariant bit vectors. A tokenization logic is configured to translate the input text sequence into a sequence of input text tokens, and to translate the successive lines of magnitude-invariant bit vectors interleaved with a newline character into a sequence of input magnitude-invariant bit vector tokens. A linear projection logic is configured to linearly project a single token stream of the sequence of input text tokens and the sequence of input magnitude-invariant bit vector tokens into a decoder-only Transformer logic, wherein the linear projection of the single token stream bypasses any embedding lookup.