Automatic Machine Control Model Switching to Prevent Overlearning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge lies in reducing the error between actual machines and simulations in automatic operation control, particularly for large-size industrial machines like overhead cranes, due to overlearning issues caused by mathematically-described functions, which leads to fluctuations in control signal strings and reward values, making it difficult to achieve consistent automatic operation across varying environments.

Innovation Solution

An automatic operation control system is implemented, which sets a first model based on a mathematically-described function and switches to a second model after satisfying specific conditions to prevent overlearning, thereby optimizing the control signal generation and reducing errors between the actual machine and simulation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If adjustment by mathematically-described function is carried out to reduce error between actual machine and simulation, then manufacturing precision is improved, but overlearning occurs causing reliability to deteriorate

Engineering Contradiction:
Improvesimulation accuracyVSAvoidcontrol stability
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The system performs preliminary actions by collecting multiple pieces of actual machine data before generating the adjustment function. This includes acquiring control signal strings and operation result data from multiple actual operations, then using this accumulated data to create a more robust adjustment function that is less prone to overlearning. The preliminary data collection phase ensures the adjustment function is based on sufficient empirical evidence rather than fitting too closely to a single dataset.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback mechanisms by evaluating whether the adjustment function has overlearned using multiple evaluation indices. It compares simulation results with actual machine data across multiple dimensions and uses this feedback to determine if the adjustment function needs modification. The feedback loop includes checking if evaluation indices fall within predetermined ranges and using this information to adjust the adjustment function accordingly, preventing overlearning while maintaining accuracy.

Inventive Principle:
Principle #23Feedback

2Manufacturing precision

If simulation is made to closely match actual machine behavior, then manufacturing precision is improved, but adaptability deteriorates due to strong dependence on specific parameters

Engineering Contradiction:
Improvesimulation fidelityVSAvoidenvironmental adaptability
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The system applies parameter changes by modifying the adjustment function based on multiple pieces of actual machine data rather than fitting to a single parameter set. The adjustment function is generated to account for variations in control signal strings, conveyance distances, weights, and environmental conditions across multiple operations. This approach creates a more generalized adjustment function that maintains simulation fidelity while adapting to different operating parameters and environments.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If reinforcement learning is used to search for control signal string, then productivity is improved, but learning progress stagnates due to reward fluctuation

Engineering Contradiction:
Improvecontrol optimization speedVSAvoidlearning convergence
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system introduces an intermediary adjustment function that mediates between the reinforcement learning process and the simulation environment. This adjustment function acts as a bridge that reduces the discrepancy between simulation and actual machine behavior, thereby stabilizing the reward signals received by the reinforcement learning algorithm. By using this intermediary, the reinforcement learning can proceed more efficiently with less reward fluctuation, improving both productivity and learning convergence.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11619929B2Automatic operation control method and system
Publication Date: 2023.04.04 HITACHI LTD
  • US11619929B2 patent drawing
  • US11619929B2 patent drawing
  • US11619929B2 patent drawing

AI summary

An object of the present invention is to reduce an error between an actual machine and a simulation by removing the influence of overlearning of an adjustment by a mathematically-described function, and to optimize automatic operation control of the machine. An automatic operation control system for controlling an automatic operation of a machine sets a first model showing a relation between a control signal string input to the machine on the basis of a mathematically-described function and data output from the machine controlled in accordance with the control signal string. In a learning process including learning the automatic operation control of the machine, the system executes learning using the first model until a first condition is satisfied. After the first condition is satisfied, the learning is executed using a second model that is a model after the first model is changed one or more times until a second condition meaning overlearning is satisfied or the learning is finished without satisfying the second condition.