Focus-of-Attention Encoding for Robot Learning in Clutter

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional Learning-from-Observation (LFO) systems fail to effectively teach robots in noisy, chaotic, or cluttered environments due to improper correlations between observed human movements and task sequences caused by other objects and unrelated movements in the vicinity.

Innovation Solution

Implementing an input-based focus-of-attention (FOA) model that filters human demonstrations using verbal or textual cues to identify target objects and movements, allowing robots to focus on the relevant actions and ignore distractions in cluttered environments through spatio-temporal filtering of time-series images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If traditional LFO systems observe all human movements in cluttered environments, then more movement data is captured, but improper correlations occur between observed movements and task sequences due to unrelated objects and movements

Engineering Contradiction:
Improvemovement dataVSAvoidcorrelation accuracy
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent extracts and isolates the target object from the cluttered environment using verbal cues and spatial filtering. By focusing only on the target object and its immediate vicinity, the system removes unrelated objects and movements that cause improper correlations, while retaining all relevant movement data associated with the target object.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies local quality by creating a focused attention region around the target object. Verbal cues define a specific spatial region of interest, and the system processes movement data with different quality levels - high detail for movements near the target object and filtered/removed data for movements in distant or unrelated regions.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If verbal cues are used to guide focus-of-attention filtering, then learning accuracy in cluttered environments improves, but system complexity increases due to input parsing and spatio-temporal filtering requirements

Engineering Contradiction:
Improvelearning accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces verbal cues as an intermediary that bridges the gap between the cluttered environment and the robot's learning process. The verbal cues serve as a mediator that defines the region of interest and guides the filtering process, simplifying the system's task of identifying relevant movements without requiring complex environmental analysis.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies preliminary action by processing and parsing verbal cues before the actual learning process begins. The system pre-defines the region of interest and filtering parameters based on verbal input, so that during movement observation, only pre-filtered data needs to be processed, reducing real-time computational complexity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11731271B2Verbal-based focus-of-attention task model encoder
Publication Date: 2023.08.22 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11731271B2 patent drawing
  • US11731271B2 patent drawing
  • US11731271B2 patent drawing

AI summary

Traditionally, robots may learn to perform tasks by observation in clean or sterile environments. However, robots are unable to accurately learn tasks by observation in real environments (e.g., cluttered, noisy, chaotic environments). Methods and systems are provided for teaching robots to learn tasks in real environments based on input (e.g., verbal or textual cues). In particular, a verbal-based Focus-of-Attention (FOA) model receives input, parses the input to recognize at least a task and a target object name. This information is used to spatio-temporally filter a demonstration of the task to allow the robot to focus on the target object and movements associated with the target object within a real environment. In this way, using the verbal-based FOA, a robot is able to recognize “where and when” to pay attention to the demonstration of the task, thereby enabling the robot to learn the task by observation in a real environment.