Stochastic Neural Network Policy Segmentation for Multi-Mode Robot Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional Learning from Demonstration (LfD) techniques fail to represent the multi-modal nature of data, leading to sub-optimal behavior when multiple tasks are demonstrated without careful task setup, labeling, and engineering, as they average different behavior modes instead of learning them separately.

Innovation Solution

The use of a stochastic deep neural network (SNN) with a stochastic attention module to learn multiple modes of behavior from visual data, allowing the system to represent the underlying intention as a stochastic activation and avoid averaging different behavior modes, enabling the system to learn and exhibit multiple modes of behavior without requiring careful task setup or labeling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional imitation learning techniques are used to learn from multiple task demonstrations without careful task setup and labeling, then the system can process diverse visual data, but it averages different behavior modes and produces sub-optimal behavior

Engineering Contradiction:
Improveability to learn from multiple tasksVSAvoidbehavior optimality
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments the learned policy into multiple distinct modes or behaviors, each corresponding to different task demonstrations. Instead of creating a single averaged policy, the system identifies and separates different behavior modes (e.g., different grasping techniques, different navigation strategies) that exist in the demonstration data, allowing each mode to be selected appropriately based on the current situation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a dynamic mode selection mechanism that allows the system to switch between different behavior modes based on the current state and context. This dynamic approach enables the robot to adaptively choose the most appropriate behavior mode for the current task situation, rather than relying on a static averaged policy.

Inventive Principle:
Principle #15Dynamics

2Device complexity

If traditional LfD techniques average multiple behavior modes, then the system simplifies the learning process, but it loses the ability to represent multi-modal intentions

Engineering Contradiction:
Improvelearning process simplicityVSAvoidmulti-modal intention representation
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent adds a new dimension to the policy representation by introducing a mode selection variable. Instead of representing the policy in the original state-action space only, the system extends it to include mode selection, creating a multi-dimensional policy structure that can represent both the continuous action parameters and the discrete behavior mode simultaneously.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Reliability

If careful task setup, labeling, and engineering are performed to demonstrate multiple tasks, then the system can learn distinct behavior modes, but it increases the time and effort required for demonstration

Engineering Contradiction:
Improvebehavior mode distinctionVSAvoiddemonstration preparation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent enables the system to automatically identify and separate different behavior modes from the demonstration data without requiring manual labeling or task categorization. The mode selection mechanism learns to automatically recognize which behavior mode is appropriate for each situation based on the demonstration data alone, eliminating the need for time-consuming manual task setup and labeling.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11507826B2Computerized imitation learning from visual data with multiple intentions
Publication Date: 2022.11.22 OSARO
  • US11507826B2 patent drawing
  • US11507826B2 patent drawing

AI summary

A computer system uses Learning from Demonstration (LfD) techniques in which a multitude of tasks are demonstrated without requiring careful task set up, labeling, and engineering, and learns multiple modes of behavior from visual data, rather than averaging the multiple modes. As a result, the computer system may be used to control a robot or other system to exhibit the multiple modes of behavior in appropriate circumstances.