Conveyor Element Control Using Local RL for Piece-Good Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conveyor systems for singulating and aligning piece goods face challenges in efficiently controlling the separation and alignment of diverse goods with varying properties, due to high dimensionality and complexity in control processes, which limits adaptability to changing cargo flows and requires manual adjustments that are time-consuming and inefficient.
Innovation Solution
A computer-implemented method using reinforcement learning to control conveyor systems, where an agent selects actions based on state vectors and action vectors to adjust the speed of conveyor elements, reducing dimensionality by breaking down control problems into local action vectors and allowing for quick adaptation to changing conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional image processing and manual controller design are used, then the control system can be implemented with standard methods, but the system cannot adapt to changing cargo flow characteristics and requires time-consuming manual adjustments
Solution Approach 1:
The control system uses reinforcement learning agents that automatically adapt to changing cargo flow characteristics without manual intervention. The agents learn optimal control strategies through continuous interaction with the environment, enabling the system to self-adjust to different package types, weights, and flow patterns, eliminating the need for manual controller redesign
Solution Approach 2:
The system changes control parameters dynamically based on learned patterns from cargo flow data. The reinforcement learning agents continuously optimize conveyor element speeds and adjustments by learning from environmental feedback, allowing the system to adapt to varying package properties without manual parameter tuning
2Device complexity
If the control system focuses on the frontmost parcel only, then the control problem dimensionality is reduced, but the control quality decreases
Solution Approach 1:
The control problem is segmented by creating individual state vectors for each parcel rather than treating the entire cargo flow as a single control problem. Each parcel's state vector captures its specific position, velocity, and relevant properties, allowing the reinforcement learning agent to make targeted control decisions for each parcel while maintaining overall system coordination
Solution Approach 2:
The system transitions from controlling only the frontmost parcel to a multi-dimensional control approach where state vectors for multiple parcels are processed simultaneously. The reinforcement learning agent operates in an expanded state space that includes information about multiple parcels, enabling it to optimize control decisions that consider the positions and movements of several parcels rather than just one
3Productivity
If preset control processes are used for standard goods, then the system operates efficiently for similar piece goods, but cannot handle piece goods with different properties such as smooth plastic packages instead of grippy cardboard boxes
Solution Approach 1:
The reinforcement learning agents continuously learn from actual cargo flow data, automatically adapting control strategies to match the specific properties of different package types. When new package types with different friction properties are introduced, the agents learn the appropriate control adjustments through environmental feedback, maintaining efficient operation without manual reconfiguration
Solution Approach 2:
The system uses environmental feedback from sensor data about actual package positions and movements to continuously refine control decisions. The reinforcement learning agents receive rewards or penalties based on control quality metrics, allowing them to learn optimal control strategies for different package properties through continuous feedback loops
4Manufacturing precision
If manual optimization of control processes is performed, then control quality can be improved for specific conditions, but the process is very complex and time-consuming
Solution Approach 1:
The reinforcement learning agents perform automatic optimization of control processes by learning from environmental interactions. The agents continuously refine their control strategies through trial and error in the actual operating environment, eliminating the need for manual optimization processes while maintaining high control quality across varying conditions
Solution Approach 2:
The system performs preliminary learning and adaptation during normal operation rather than requiring separate optimization phases. The reinforcement learning agents continuously update their control strategies in real-time based on incoming cargo flow characteristics, so optimization occurs concurrently with operation rather than requiring downtime for manual tuning
Data Source
Figure 1~2C
Figure 3~4
Figure 5~6
AI summary
The use of reinforcement learning methods on conveyor systems (2) for piece goods (4) is quickly reaching its limit due to the high number of individual conveyor elements (12) which determine the dimensionality of the action vectors (a(t)). The invention relates to a computer-implemented method, to a device for processing data, and to a computer system for actuating a regulating device of a conveyor system (2) with individually actuatable conveyor elements (12) in order to achieve an alignment and/or a defined spacing of the piece goods (4), wherein the actuation of the regulating device (14) is determined by an agent which operates according to reinforcement learning methods. An individual local state vector sn(t) of a dimension which is ascertained in advance and which conforms to all of the piece goods (4) is generated for each piece good (4n) using an image. An action vector (an (t)) is selected individually for each piece good (4n) from an action space according to a strategy (policy) for the current state vector (sn (t)) of said piece good (4), said strategy being the same for all of the piece goods (4, 4n). The action vectors (an (t)) are projected onto the conveyor elements (12), and conflicts (for example multiple action vectors (an (t)) mapped to the same conveyor element (12)) are solved. After a cycle time (At) expires, state vectors (sn (t+Δt)) are generated again for each piece good (4n) and are evaluated using rewards, and the strategy is adapted.