Spatial Interface For Multi-Modal AI Input Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing artificial intelligence models are limited to processing one-dimensional inputs, such as text or single images, which restricts their ability to ingest multiple prompts efficiently, requiring significant time and processing power.

Innovation Solution

A spatial interface that allows for multi-modal input, enabling users to select and manipulate objects within a resizable and reshaped window, combined with text or voice commands, which are then processed by an AI model to generate dynamic responses.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If AI models process multiple prompts individually in sequence, then the model can handle complex queries, but it requires significant time and processing power

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidtime for sequential processing
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent transforms one-dimensional sequential processing into two-dimensional parallel processing by introducing a spatial interface that displays multiple objects simultaneously. Users can select multiple objects in parallel through spatial manipulation, and the system processes these multi-modal inputs (visual selections + text prompts) concurrently through the AI model, thereby reducing processing time while maintaining complex query capability

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If AI models accept only one-dimensional input, then the model structure remains simple, but the ability to ingest multiple prompts is limited

Engineering Contradiction:
Improvemulti-modal input capabilityVSAvoidinput processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal input interface that accepts multiple modalities (visual object selections through spatial window manipulation and text prompts) and consolidates them into a unified multi-dimensional input format for the AI model. The spatial interface serves multiple functions: displaying objects, enabling selection, capturing spatial relationships, and integrating with text input, thereby enhancing versatility without proportionally increasing model complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system adds a spatial dimension to input processing by allowing users to interact with objects in a two-dimensional display space. This spatial interface enables simultaneous selection of multiple objects and capture of their spatial relationships, transforming single-dimensional text input into multi-dimensional input that includes visual, spatial, and textual components

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20250182423A1Spatial Interface For Multi-Modal Artificial Intelligence Model
Publication Date: 2025.06.05 GOOGLE LLC
  • US20250182423A1 patent drawing
  • US20250182423A1 patent drawing
  • US20250182423A1 patent drawing

AI summary

The technology described herein is directed to spatial interface for multi-modal input to artificial intelligence (AI) powered tools. The interface allows for a first mode of input, such as selection of one or more objects using a movable window that can be resized and reshaped by a user. In addition, the interface allows for a second mode of input, such as text or voice commands. The AI powered tools accept the inputs from the first and second modes and dynamically generates a response.