Polymer Kneading Control Using Reinforcement Learning for Consistent Quality

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods rely on skilled operators' experience to decide kneading conditions, making it difficult to achieve appropriate kneaded products consistently.

Innovation Solution

A machine learning method and device that utilize reinforcement learning to determine optimal kneading conditions for polymer materials by evaluating state variables and updating decision functions based on rewards, incorporating parameters like physical properties and shape characteristics of the kneaded product.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If skilled operators decide kneading conditions based on years of experience, then appropriate kneaded products can be obtained, but it is difficult to be decided with ease and requires high operator expertise

Engineering Contradiction:
Improvekneaded product qualityVSAvoiddecision difficulty
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The system enables self-service by allowing the kneading device to automatically optimize its own operating conditions through reinforcement learning. The device learns from observed state variables and rewards to autonomously determine optimal kneading conditions without requiring external expert intervention, thus resolving the contradiction between product quality and operational ease.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical system of human expert decision-making with an automated machine learning system. The reinforcement learning model substitutes the need for skilled operators' experience-based judgments, using algorithms to process state variables and determine optimal kneading conditions, thereby eliminating the difficulty of operation while maintaining high product quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If machine learning is used to determine kneading conditions, then ease of operation is improved, but the system requires complex learning processes and state variable observations

Engineering Contradiction:
Improvedecision easeVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The controller serves multiple functions: it controls the rotors, monitors state variables, calculates rewards, and updates the reinforcement learning model. By making the controller multi-functional, the patent reduces the need for separate dedicated components for each function, thereby managing system complexity while maintaining ease of operation through automated machine learning.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If reinforcement learning is implemented to optimize kneading conditions, then productivity is improved through automated decision-making, but loss of time occurs during the learning process

Engineering Contradiction:
Improveautomated decision efficiencyVSAvoidlearning time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary learning actions by observing state variables and receiving rewards during the learning phase. This preliminary action allows the reinforcement learning model to build knowledge before actual production, enabling faster automated decision-making in subsequent operations. The learning process occurs in advance, reducing time loss during productive kneading operations.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12420452B2Machine learning method, machine learning device, machine learning program, communication method, and kneading device
Publication Date: 2025.09.23 KOBE STEEL LTD
  • US12420452B2 patent drawing
  • US12420452B2 patent drawing
  • US12420452B2 patent drawing

AI summary

A machine learning method includes: acquiring a state variable including at least one first evaluation parameter related to performance evaluation of a kneaded product and at least one kneading condition; calculating a reward for a decision result of the at least one kneading condition based on the state variable; updating a function for deciding the at least one kneading condition from the state variable based on the reward; and by repeating the update of the function, deciding a kneading condition under which the reward obtained becomes maximum, in which the at least one first evaluation parameter includes at least one of physical properties and shape characteristics related to the kneaded product.