Reinforcement Learning Data Augmentation for Asymmetric Packing Tasks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing reinforcement learning techniques for robotic motion, such as grasping and packing, are inefficient due to unsuitable data augmentation methods that do not account for the inherent asymmetry of starting and destination locations, limiting the effectiveness of learning.
Innovation Solution
An information processing device that augments data by swapping the starting and destination states in robotic motion tasks, generating multiple experience data sets to improve learning efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data augmentation by plane-symmetric coordinate conversion is applied, then the quantity of learning data is increased, but the quality and suitability of the data for reinforcement learning deteriorates
Solution Approach 1:
The patent applies asymmetry by recognizing that packing tasks have inherent directional asymmetry (source container to destination container) and designs data augmentation methods that preserve this asymmetry. Instead of using symmetric plane conversions that create unrealistic reversed scenarios, the system generates augmented data that maintains the correct directional relationship while varying other parameters like container positions and object characteristics.
Solution Approach 2:
The patent inverts the conventional approach by not simply reversing coordinates, but rather by inverting the augmentation logic itself - instead of asking 'how to reverse the data?', it asks 'how to generate varied but realistic data?'. This leads to augmentation techniques that swap non-critical elements (like container identities or background elements) while preserving the critical source-destination relationship.
2Adaptability or versatility
If conventional data augmentation methods are used, then the diversity of training data is improved, but the learning efficiency and occupancy rate deteriorate
Solution Approach 1:
The patent applies local quality by differentiating between elements that should be varied (local characteristics like container positions, object attributes, background elements) and elements that must remain consistent (global structure like source-destination relationships, task logic). This selective augmentation approach maintains data diversity while preserving learning effectiveness.
Solution Approach 2:
The patent utilizes parameter changes by systematically varying specific parameters in the training data (container positions, object properties, environmental conditions) while keeping the fundamental task parameters (source container, destination container, packing action) intact. This allows diverse training scenarios without compromising the core learning objective.
Data Source
AI summary
An information processing device includes processing circuitry. The processing circuitry is configured to acquire one or more pieces of first state information representing a state of each of one or more second subjects related to a first subject to be a subject of inference at first time, and one or more pieces of second state information representing a state of each of the one or more second subjects at second time; and generate learning data for use in reinforcement learning of a machine learning model for use in inference. The learning data includes the first state information at least part of which is replaced with any of the one or more pieces of the second state information, and the second state information at least part of which is replaced with any of the one or more pieces of the first state information.


