Multi-Peg-in-Hole Robot Assembly With Hierarchical Reinforcement Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Robots face inefficiencies in learning complex assembly tasks like multi-peg-in-hole assembly due to the complexity of objects and difficulty in shaping a reward function, leading to prolonged learning times.
Innovation Solution
A method and system utilizing hierarchical reinforcement learning and distributed learning, which constructs sub-process networks in different environments to update a master-control assembly-strategy model, comprising high-level and low-level strategy networks, to improve learning efficiency and reduce time for robotic multi-peg-in-hole assembly tasks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If ordinary reinforcement learning algorithm is used for robot assembly learning, then the robot can learn assembly tasks, but the learning time is excessively long and learning efficiency is low
Solution Approach 1:
The patent divides the complex multi-peg-in-hole assembly task into multiple sub-processes, each handled by a dedicated sub-process network. This segmentation allows the robot to learn and execute assembly operations in discrete, manageable stages rather than as a monolithic task, significantly improving learning efficiency and reducing overall learning time.
Solution Approach 2:
The patent introduces a hierarchical dimension to the reinforcement learning architecture by creating multiple layers of networks (master-control network, sub-process networks, and option-value networks). This dimensional transformation from flat to hierarchical structure enables more efficient learning by organizing assembly operations at different levels of abstraction.
2Adaptability or versatility
If complex assembly objects like multi-peg-in-hole are used, then the assembly skill becomes more comprehensive, but the learning difficulty increases and interactive data acquisition becomes difficult
Solution Approach 1:
The patent segments the complex assembly task into multiple sub-processes, each managed by a sub-process network. This breakdown reduces the complexity faced by individual learning modules while maintaining the ability to handle comprehensive assembly skills through coordinated sub-process execution.
Solution Approach 2:
The patent implements a nested hierarchical structure where option-value networks are contained within sub-process networks, which are in turn coordinated by the master-control network. This nesting allows complex assembly skills to be built from simpler, reusable components at each hierarchical level.
3Reliability
If reward function is shaped for complex assembly tasks, then learning guidance is improved, but the difficulty of shaping and implementing the reward function increases
Solution Approach 1:
The patent divides the reward function into multiple sub-reward functions, each corresponding to a specific sub-process. This segmentation allows for simpler, more targeted reward shaping at each sub-process level rather than attempting to design a single complex reward function for the entire assembly task.
Solution Approach 2:
The patent introduces a hierarchical dimension to reward function design by implementing reward mechanisms at multiple levels (master-control level and sub-process level). This dimensional transformation allows reward guidance to be distributed across hierarchical layers, reducing the complexity of individual reward functions while maintaining overall learning guidance effectiveness.
Data Source
AI summary
A method and system for robotic multi-peg-in-hole assembly based on hierarchical reinforcement and distributed learning, including: establishing a master-control assembly-strategy model based on deep reinforcement learning by using data of states and actions of a robot; constructing a plurality of sub-process networks based on different assembly interaction environments, updating and training the master-control assembly-strategy model by using interaction data of the robot obtained by the constructed plurality of sub-process networks, and obtaining a trained master-control assembly-strategy model; and controlling and instructing the robot to execute an assembly task of a robotic multi-peg-in-hole assembly by using the trained master-control assembly-strategy model. By utilizing a manner of constructing a sub-process network in a plurality of different environments to update an overall network, comparing with an ordinary reinforcement learning algorithm, the final effect of the robot learning can be improved, the efficiency of the robot learning is improved, and learning time is saved.


