Industrial Robot Control Strategy Using Simulation Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional robotic system development is hindered by long lead times for programming complex tasks, lack of flexibility to adapt to changing conditions, and inability to learn from mistakes without manual reprogramming, leading to production downtime and inefficiencies.

Innovation Solution

A self-learning industrial robotic system utilizing simulation-based reinforcement learning allows robots to learn tasks through trial and error in a virtual environment, eliminating the need for human intervention and enabling continuous performance improvement without manual updates.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual programming is used for complex robotic tasks, then the robot can perform the task, but the lead time becomes very long and skilled programmers are hard to find

Engineering Contradiction:
Improverobot task executionVSAvoidprogramming lead time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The robot performs self-programming through autonomous reinforcement learning in simulation environments, eliminating the need for human programmers to manually code tasks. The system learns optimal control policies independently through trial-and-error training, directly resolving the contradiction by making the robot self-sufficient in task acquisition.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual programming mechanics with automated machine learning systems. Instead of human programmers writing code, a reinforcement learning agent automatically generates and optimizes control policies through simulation-based training, substituting the mechanical process of manual coding with an automated computational approach.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If hard-coded robotic programs are used, then the robot can operate based on programmed conditions, but the system lacks flexibility to adapt to operating condition changes

Engineering Contradiction:
Improverobot operation stabilityVSAvoidadaptation to condition changes
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The control policy is transformed from a static hard-coded program to a dynamic learned policy that can adapt to changing conditions. The reinforcement learning approach enables the robot to learn flexible control strategies that respond to varying operating conditions, part variations, and environmental changes while maintaining stable performance through continuous learning.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs preliminary training in simulation environments that encompass diverse operating conditions and edge cases before deployment. This pre-exposure to various scenarios during training enables the robot to adapt to condition changes in real-world operation without requiring reprogramming, resolving the contradiction between stability and adaptability.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If manual programming is used, then the robot can be programmed for specific tasks, but the robot cannot learn from mistakes or improve performance automatically

Engineering Contradiction:
Improvetask completionVSAvoidperformance improvement rate
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The reinforcement learning system implements continuous feedback loops where the robot receives rewards or penalties based on task performance. This feedback mechanism enables automatic learning from mistakes, as the system adjusts its control policy based on performance outcomes, continuously improving task completion rates without human intervention.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The robot autonomously improves its own performance through self-directed learning in simulation environments. By independently analyzing its own mistakes and adjusting its control strategy based on reinforcement signals, the system achieves continuous performance improvement without requiring external reprogramming or human analysis of errors.

Inventive Principle:
Principle #25Self-service

4Productivity

If simulation-based reinforcement learning is used, then training can be done faster and cheaper without physical assets, but the system complexity increases

Engineering Contradiction:
Improvetraining speedVSAvoidsystem architecture
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent creates virtual copies of the physical robot and environment in simulation environments. These digital twins replicate the physics, sensors, and actuators of the real system, enabling fast and cheap training without risking physical assets. The copying approach resolves the contradiction by transferring the complex learning process to a virtual domain while keeping the actual hardware simple and safe.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11554482B2Self-learning industrial robotic system
Publication Date: 2023.01.17 HITACHI LTD
  • US11554482B2 patent drawing
  • US11554482B2 patent drawing
  • US11554482B2 patent drawing

AI summary

Example implementations described herein are directed to a simulation environment for a real world system involving one or more robots and one or more sensors. Scenarios are loaded into a simulation environment having one or more virtual robots corresponding to the one or more robots, and one or more virtual sensors corresponding to the one or more virtual system to train a control strategy model from reinforcement learning, which is subsequently deployed to the real world environment. In cases of failure of the real world environment, the failures are provided to the simulation environment to generate an updated control strategy model for the real world environment.