Multimodal Object Identification for Out-of-View Robot Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Robotics systems face challenges in disambiguating between multiple objects in their environment when a user provides a command referencing a specific object, especially when the object is out of view, leading to inaccurate identification and action execution.

Innovation Solution

The system combines speech and image/video inputs to analyze user commands and gestures, using an inventory of the environment to narrow down the search for the intended object based on spatial regions indicated by the gesture, thereby improving object-locating accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the robotics system searches the entire inventory for the referenced object, then it may find the object, but it may also incorrectly identify a different instance of the object located in a different region

Engineering Contradiction:
Improveobject identification accuracyVSAvoidsearch process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the search space by dividing the environment into multiple spatial regions (e.g., rooms, zones). Instead of searching the entire inventory, the system segments the search to only the relevant spatial region indicated by the user's gesture, thereby improving identification accuracy while reducing search complexity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by focusing the search effort on a specific local region rather than uniformly searching the entire environment. The system identifies the spatial region where the object is likely located based on gesture analysis, and then performs the search only within that localized area, improving precision without proportionally increasing overall system complexity

Inventive Principle:
Principle #3Local quality

2Measurement precision

If the robotics system uses only speech input to identify objects, then the system is simpler to operate, but the system cannot accurately identify objects when multiple instances exist in different locations

Engineering Contradiction:
Improveobject location precisionVSAvoiduser interaction complexity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent merges multiple input modalities (speech recognition and gesture recognition) into a unified object identification system. The speech input provides the object type while the gesture input provides spatial information, and combining these modalities enables accurate identification of specific object instances without significantly complicating user interaction

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements multi-functionality by designing a hybrid input system that can process both speech commands and gestural inputs. This universal interface allows the system to handle various types of object reference (by name, by gesture, by combination) while maintaining ease of operation through natural human communication modes

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If the robotics system searches the entire inventory for the object, then it ensures finding the object, but the search time increases when the object is located in a specific region

Engineering Contradiction:
Improveobject retrieval speedVSAvoidsearch time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-processing and organizing the environment into spatial regions and maintaining an inventory structured by location. When a user provides a gesture, the system has already prepared the spatial segmentation, allowing it to quickly filter and search only the relevant region rather than scanning the entire inventory, thus improving retrieval speed without sacrificing completeness

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250010482A1Multimodal Object Identification
Publication Date: 2025.01.09 GDM HOLDING LLC
  • US20250010482A1 patent drawing
  • US20250010482A1 patent drawing
  • US20250010482A1 patent drawing

AI summary

Methods, systems, and apparatus for receiving a command for controlling a robot, the command referencing an object, receiving sensor data for a portion of an environment of the robot, identifying, from the sensor data, a gesture of a human that indicates a spatial region located outside of the portion of the environment described by the sensor data, searching map data for the object, determining, based at least on searching the map data for the object referenced in the command, that the object referenced in the command is present in the spatial region, and in response to determining that the object referenced in the command is present in the spatial region, controlling the robot to perform an action with respect to the object referenced in the command.