Reinforcement Learning Agents for Molding Condition Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Reinforcement learning methods for adjusting molding conditions in manufacturing devices often result in inappropriate settings, leading to abnormal operations and disadvantages for both the machine and operators, due to unrestricted search ranges.

Innovation Solution

A reinforcement learning method involving a first agent that adjusts manufacturing conditions based on observation data and a second agent with a functional model representing the relationship between observation data and conditions, which calculates reward data and limits the search range to ensure safe optimization of molding conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If reinforcement learning is used to adjust molding conditions without limiting search range, then the ability to find optimal conditions is improved, but the risk of abnormal operations and machine damage increases

Engineering Contradiction:
Improvesearch rangeVSAvoidabnormal operation risk
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent divides the reinforcement learning system into two separate agents: a first agent that performs exploration with a wide search range to find optimal conditions, and a second agent that performs exploitation with a limited search range to ensure safe operations. This segmentation allows each agent to specialize in one function, resolving the contradiction between comprehensive search and operational safety.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a condition adjustment unit as an intermediary that receives outputs from both agents and selectively applies them. This mediator coordinates the actions of the two agents, allowing the system to benefit from both wide exploration and safe exploitation without direct conflict between them.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If a wide search range is used in reinforcement learning, then the optimal molding condition can be found more effectively, but inappropriate settings may cause abnormal operations

Engineering Contradiction:
Improvemolding condition optimizationVSAvoidabnormal operation
Core Design Contradiction:
Manufacturing precisionVSObject-affected harmful factors

Solution Approach 1:

The patent dynamically adjusts the search range based on the agent being executed. When the first agent (exploration) is active, the search range is wide to enable comprehensive optimization. When the second agent (exploitation) is active, the search range is limited to ensure safety. This dynamic adjustment resolves the contradiction between thorough search and harm prevention.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter of search range width depending on which agent is performing the reinforcement learning. The condition adjustment unit modifies the effective search range parameter based on the current agent's needs, allowing the system to achieve both comprehensive optimization and safe operation under different conditions.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If the search range is limited to ensure safety, then abnormal operations are reduced, but the ability to find optimal conditions is restricted

Engineering Contradiction:
Improveoperation safetyVSAvoidcondition search capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent segments the reinforcement learning task into two distinct agents with different search range characteristics. The first agent handles wide-range exploration for optimal condition discovery, while the second agent handles narrow-range exploitation for safe operations. This segmentation allows the system to maintain both safety and optimization capability simultaneously.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system periodically alternates between the exploration phase (wide search) and exploitation phase (limited search) by selecting which agent to execute. This periodic switching allows the system to achieve comprehensive optimization over time while maintaining safety during exploitation phases, resolving the contradiction between search capability and operational safety.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS20240227266A9Reinforcement Learning Method, Non-Transitory Computer Readable Recording Medium, Reinforcement Learning Device and Molding Machine
Publication Date: 2024.07.11 THE JAPAN STEEL WORKS LTD
  • US20240227266A9 patent drawing
  • US20240227266A9 patent drawing
  • US20240227266A9 patent drawing

AI summary

A reinforcement learning method of a learning machine including a first agent adjusting a manufacture condition of a manufacturing device based on observation data obtained by observing a state of the manufacturing device and a second agent having a functional model or a functional approximator representing a relationship between the observation data and the manufacture condition in a different way from the first agent, comprises: adjusting the manufacture condition searched by the first agent that is performing reinforcement learning, using the observation data and the functional model or the functional approximator of the second agent; calculating reward data in accordance with a state of a product manufactured by the manufacturing device under the manufacture condition adjusted; and performing reinforcement learning on the first agent and the second agent based on the observation data and the reward data calculated.