Reinforcement Learning Agents for Molding Condition Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Reinforcement learning methods for adjusting molding conditions in manufacturing devices often result in inappropriate settings, leading to abnormal operations and disadvantages for both the machine and operators, due to unrestricted search ranges.
Innovation Solution
A reinforcement learning method involving a first agent that adjusts manufacturing conditions based on observation data and a second agent with a functional model representing the relationship between observation data and conditions, which calculates reward data and limits the search range to ensure safe optimization of molding conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If reinforcement learning is used to adjust molding conditions without limiting search range, then the ability to find optimal conditions is improved, but the risk of abnormal operations and machine damage increases
Solution Approach 1:
The patent divides the reinforcement learning system into two separate agents: a first agent that performs exploration with a wide search range to find optimal conditions, and a second agent that performs exploitation with a limited search range to ensure safe operations. This segmentation allows each agent to specialize in one function, resolving the contradiction between comprehensive search and operational safety.
Solution Approach 2:
The patent introduces a condition adjustment unit as an intermediary that receives outputs from both agents and selectively applies them. This mediator coordinates the actions of the two agents, allowing the system to benefit from both wide exploration and safe exploitation without direct conflict between them.
2Manufacturing precision
If a wide search range is used in reinforcement learning, then the optimal molding condition can be found more effectively, but inappropriate settings may cause abnormal operations
Solution Approach 1:
The patent dynamically adjusts the search range based on the agent being executed. When the first agent (exploration) is active, the search range is wide to enable comprehensive optimization. When the second agent (exploitation) is active, the search range is limited to ensure safety. This dynamic adjustment resolves the contradiction between thorough search and harm prevention.
Solution Approach 2:
The patent changes the parameter of search range width depending on which agent is performing the reinforcement learning. The condition adjustment unit modifies the effective search range parameter based on the current agent's needs, allowing the system to achieve both comprehensive optimization and safe operation under different conditions.
3Reliability
If the search range is limited to ensure safety, then abnormal operations are reduced, but the ability to find optimal conditions is restricted
Solution Approach 1:
The patent segments the reinforcement learning task into two distinct agents with different search range characteristics. The first agent handles wide-range exploration for optimal condition discovery, while the second agent handles narrow-range exploitation for safe operations. This segmentation allows the system to maintain both safety and optimization capability simultaneously.
Solution Approach 2:
The system periodically alternates between the exploration phase (wide search) and exploitation phase (limited search) by selecting which agent to execute. This periodic switching allows the system to achieve comprehensive optimization over time while maintaining safety during exploitation phases, resolving the contradiction between search capability and operational safety.
Data Source
AI summary
A reinforcement learning method of a learning machine including a first agent adjusting a manufacture condition of a manufacturing device based on observation data obtained by observing a state of the manufacturing device and a second agent having a functional model or a functional approximator representing a relationship between the observation data and the manufacture condition in a different way from the first agent, comprises: adjusting the manufacture condition searched by the first agent that is performing reinforcement learning, using the observation data and the functional model or the functional approximator of the second agent; calculating reward data in accordance with a state of a product manufactured by the manufacturing device under the manufacture condition adjusted; and performing reinforcement learning on the first agent and the second agent based on the observation data and the reward data calculated.


