Multi-Modal Object Placement in Virtual Worlds

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current software applications lack the ability to understand and execute multi-modal inputs, such as natural language and gestures, for complex object placement in virtual worlds, making it difficult for users to interactively manipulate graphical objects without precise instructions.

Innovation Solution

A computer system and method that utilizes multi-modal inputs like natural language, gesture, text, and sketch to manipulate graphical objects in virtual worlds, employing a Communicative Agent for Spatio-Temporal Reasoning (CoASTeR) with components like sensors, actuators, and cognitive elements to interpret and execute user commands, allowing for object placement in three-dimensional virtual environments with constraints and rules.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If traditional mouse-based manipulation is used for object placement, then precise control is achieved, but the operation becomes arduous and complex

Engineering Contradiction:
Improveease of object placementVSAvoidcomplexity of manipulation
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical mouse-based manipulation system with a multi-modal input system that includes natural language processing, gesture recognition, and sketch processing. This substitution eliminates the need for precise mouse coordination while achieving object placement through more intuitive verbal and graphical commands.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an intermediary processing layer that includes a constraint satisfaction algorithm and a virtual agent that mediates between the user's high-level commands (natural language, gestures) and the low-level object placement operations. This intermediary translates ambiguous user intentions into precise placement actions automatically.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If multi-modal inputs are accepted, then user accessibility is improved, but the system complexity increases

Engineering Contradiction:
Improveacceptance of multiple input typesVSAvoidcomplexity of input processing
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal input processing framework that handles multiple input modalities (natural language, gestures, sketches, text) through a single integrated constraint satisfaction system. The same algorithmic framework processes all input types, translating them into unified placement constraints, thereby achieving versatility without proportionally increasing complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of operation

If natural language commands are used, then ease of operation improves, but precision of object placement may be reduced

Engineering Contradiction:
Improveease of command inputVSAvoidprecision of object placement
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent incorporates feedback mechanisms where the system provides suggestions and refinements based on the natural language command. The constraint satisfaction algorithm analyzes the command, generates potential placements, and can seek clarification or confirm the intended location, thereby maintaining precision while preserving the ease of natural language interaction.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8893048B2System and method for virtual object placement
Publication Date: 2014.11.18 KNEXUS RESEARCH CORP
  • US8893048B2 patent drawing
  • US8893048B2 patent drawing
  • US8893048B2 patent drawing

AI summary

A computer system and method according to the present invention can receive multi-modal inputs such as natural language, gesture, text, sketch and other inputs in order to manipulate graphical objects in a virtual world. The components of an agent as provided in accordance with the present invention can include one or more sensors, actuators, and cognition elements, such as interpreters, executive function elements, working memory, long term memory and reasoners for object placement approach. In one embodiment, the present invention can transform a user input into an object placement output. Further, the present invention provides, in part, an object placement algorithm, along with the command structure, vocabulary, and the dialog that an agent is designed to support in accordance with various embodiments of the present invention.