Imitation Learning Data Augmentation via Behavioral Replication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Imitation learning faces challenges with insufficient demonstration data leading to policy failure and overfitting, as existing data augmentation techniques are not effectively applicable to sequential time-series data from human or robot actions.

Innovation Solution

A system and method for imitation learning that uses a data augmentation device with behavioral and inverse behavioral replication models, based on artificial neural networks, to generate augmented data sets from expert demonstration data, improving the learning process by inferring behavioral and state data through loss function optimization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If imitation learning uses existing demonstration data, then learning can be performed, but the policy fails or overfits when demonstration data is insufficient

Engineering Contradiction:
Improvequantity of training dataVSAvoidpolicy derivation reliability
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent creates synthetic copies of expert demonstration data by using a behavioral replication model to generate augmented data sets that replicate the structure and characteristics of real expert demonstrations. This allows the system to increase training data quantity while maintaining data quality and distribution characteristics.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms parameters of the demonstration data by using the behavioral replication model to generate variations of state-action pairs. The model changes parameters such as state representations and action sequences to create diverse augmented data that maintains the underlying behavioral patterns while increasing data variety.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If existing image data augmentation techniques are applied to sequential time-series data, then data quantity increases, but the techniques are not effectively applicable due to the sequential nature and wide action space of robotic actions

Engineering Contradiction:
Improvequantity of training dataVSAvoidadaptability of augmentation technique
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent replaces traditional image-based mechanical augmentation techniques with a neural network-based behavioral replication model. Instead of applying geometric transformations to images, the system uses a learned model to generate synthetic sequential data that respects the temporal and action-space constraints of robotic behaviors.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The behavioral replication model acts as an intermediary between the limited demonstration data and the augmented training data. It processes the original state-action sequences and transforms them into synthetic augmented data that maintains the sequential structure and action space characteristics while increasing data quantity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20230325712A1System and method for imitation learning
Publication Date: 2023.10.12 ELECTRONICS & TELECOMM RES INST
  • US20230325712A1 patent drawing
  • US20230325712A1 patent drawing
  • US20230325712A1 patent drawing

AI summary

The present disclosure relates to a system and method for imitation learning. The system for imitation learning may include a data augmentation device configured to acquire a plurality of augmented data sets from a plurality of demonstration data sets corresponding to an expert's demonstration behavior trajectory using a behavioral replication model that infers behavioral data from input state data and an inverse behavioral replication model that infers state data from input behavioral data, and an imitation learning device configured to perform imitation learning to derive a model that outputs behavioral data similar to an expert in a specific state using the plurality of demonstration data sets and the plurality of augmented data sets, in which the plurality of demonstration data sets and the plurality of augmented data sets each include a pair of corresponding state data and behavioral data.