Reinforcement Learning Control for Sheet Transport Jam Prevention
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Image forming devices face variations in use environments and conditions, leading to increased downtime due to sheet jams, as existing control methods are not optimized for unexpected conditions, requiring extensive man-hours for development to prevent jams under worst or typical conditions.
Innovation Solution
A machine learning device using reinforcement learning to acquire position information of transported objects, calculate rewards, and generate control information for the driving source to optimize transport control based on varying environmental and printing conditions, reducing downtime by adapting to specific use environments and conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If control is designed to prevent jams under worst or typical conditions, then reliability is improved, but device complexity and development time increase significantly
Solution Approach 1:
The system performs self-learning through reinforcement learning, where the machine learning device automatically acquires optimal control conditions by observing transport states and receiving reward signals. This eliminates the need for extensive manual control design for every possible use environment and condition, allowing the system to adapt autonomously to different scenarios including unexpected conditions.
Solution Approach 2:
The invention changes the control approach from fixed predetermined parameters to dynamically learned parameters. The machine learning device continuously updates control conditions based on observed transport states, environmental conditions, and printing conditions, allowing optimal control parameters to adapt to varying conditions without increasing structural complexity.
2Reliability
If extensive control design is performed to cover all use environments and conditions, then reliability is improved, but loss of time in development increases
Solution Approach 1:
The system performs preliminary learning actions by continuously acquiring transport states and calculating rewards during normal operation. This ongoing preliminary learning allows the system to build up knowledge of optimal control conditions across various use environments and conditions before actual jam situations occur, eliminating the need for extensive pre-development testing for every scenario.
Solution Approach 2:
The reinforcement learning process operates continuously during device operation, constantly acquiring new transport state data and updating control conditions. This continuous learning ensures that the system progressively improves its control coverage across all use environments and conditions without requiring separate development phases for each scenario.
3Ease of operation
If fixed control conditions are used for typical conditions, then ease of operation is maintained, but adaptability to unexpected conditions deteriorates
Solution Approach 1:
The control conditions transition from static fixed values to dynamic learned values. The machine learning device continuously updates control conditions based on real-time transport states, environmental conditions, and printing conditions, allowing the system to adapt to unexpected conditions while maintaining operational simplicity through automated adjustment without user intervention.
Solution Approach 2:
The system implements feedback through the reward calculation mechanism, where transport states are observed and compared against desired outcomes. This feedback loop allows the system to automatically adjust control conditions based on actual performance, enabling adaptation to unexpected conditions while maintaining ease of operation through self-correction rather than complex user control.
Data Source
AI summary
A machine learning device learns an action of a driving source in a transport device continuously transporting at least two transported objects along a transport path, and includes: a hardware processor that: acquires position information of the at least two transported objects on the transport path on the basis of a result of detection by a sensor provided in the transport path; calculates a reward on the basis of the position information acquired, according to a predetermined rule; learns an action by calculating an action value in reinforcement learning on the basis of the position information acquired and the reward calculated; and generates and outputs control information that causes the driving source to perform an action determined on the basis of a learning result.


