Casting cleaning production line intelligent scheduling control method based on multi-agent reinforcement learning

By constructing the state space and reward function of the casting cleaning production line using a multi-agent reinforcement learning method, the coordination problem between workstations in the casting cleaning production line was solved, adaptive task allocation and resource optimization were realized, and the overall efficiency and flexibility of the production line were improved.

CN121934512APending Publication Date: 2026-04-28CRRC DALIAN INST CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CRRC DALIAN INST CO LTD
Filing Date
2026-01-27
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

The lack of coordination mechanism between workstations in the casting cleaning production line leads to an imbalance in cycle time, making it impossible to dynamically adjust according to equipment status and workload. Relying on manual experience makes it difficult to achieve overall optimized control.

Method used

A multi-agent reinforcement learning approach is adopted to construct the state space and reward function of the casting cleaning production line. Real-time scheduling instructions for the agents are obtained through reinforcement learning algorithms to achieve cycle time coordination and dynamic resource allocation between workstations.

Benefits of technology

Improve the overall utilization rate of the production line, reduce workpiece waiting time and equipment idle time, reduce the workload of system debugging and maintenance, and achieve adaptive task allocation and optimized control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121934512A_ABST
    Figure CN121934512A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a casting cleaning production line intelligent scheduling control method based on multi-agent reinforcement learning, and the method comprises the steps: enabling each processing station and a carrying execution unit to make a decision according to an environment state through introducing a multi-agent reinforcement learning frame based on an established state space for casting cleaning production line scheduling, and achieving the intelligent scheduling control of a casting cleaning production line. Through global value function optimization, beat coordination and resource dynamic allocation between stations are realized, workpiece waiting time and equipment idle time are reduced, fixed beat or manual intervention is eliminated, and the overall utilization rate of a production line is improved. Based on the setting of the reward function, the system can automatically optimize the task allocation logic according to the actual operation state, manual adjustment of beat parameters is not needed, the debugging and maintenance workload of the system is remarkably reduced, and manual dependence and debugging cost are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent manufacturing and industrial automation control technology, and in particular to an intelligent scheduling and control method for a casting cleaning production line based on multi-agent reinforcement learning. Background Technology

[0002] In the cleaning process of the foundry industry, the surface of castings typically requires multiple steps such as pretreatment, cutting, and grinding to remove excess structures such as risers, flash, and burrs. Traditional production lines often use fixed-time cycle control or rely on manual experience to set the operating logic of each station. While these methods can meet basic requirements when the process flow is relatively stable, they have significant drawbacks in actual production: First, there is a lack of coordination mechanisms between stations, and delays, idle periods, or blockages at a station can easily cause cycle imbalances, affecting the overall line efficiency. Second, the flow path of workpieces between different stations is fixed and cannot be dynamically adjusted according to equipment status, workload, or processing progress. Third, manual scheduling relies on experience-based rules and lacks data-driven and adaptive capabilities, making it difficult to achieve overall optimized control under complex operating conditions.

[0003] In recent years, reinforcement learning methods have attracted attention in the field of intelligent scheduling and control. These methods, through interactive learning between agents and the environment, can autonomously optimize decision-making strategies without relying on manual settings. However, existing research largely remains at the simulation level or focuses on single-device optimization. For complex scenarios such as casting cleaning production lines with multiple parallel workstations and collaborative material handling, effective multi-agent coordination and global optimization mechanisms are still lacking. Summary of the Invention

[0004] This invention discloses an intelligent scheduling and control method for a casting cleaning production line based on multi-agent reinforcement learning, in order to overcome the above-mentioned technical problems.

[0005] To achieve the above objectives, the technical solution of the present invention is as follows: A method for intelligent scheduling and control of a casting cleaning production line based on multi-agent reinforcement learning includes the following steps: S1: Establish a state space for scheduling the casting cleaning production line, including a global state vector, local observation vectors of the processing stations, and an action space for the agent oriented towards scheduling the casting cleaning production line. S2: Construct a reward function for scheduling the foundry cleaning production line; S3: Based on the state space and reward function for scheduling the casting cleaning production line, a reinforcement learning algorithm is used to obtain real-time scheduling instructions for the intelligent agent for scheduling the casting cleaning production line, so as to realize intelligent scheduling of the casting cleaning production line. Among them, the intelligent agents for scheduling the casting cleaning production line include intelligent agents corresponding to the material handling execution unit and intelligent agents corresponding to the processing station.

[0006] Furthermore, the global state vector is established as follows:

[0007] In the formula: In time step The global state vector under; Let be the decision vector for the machining station. ,in, Indicates the first The variable indicating whether a processing station is in a processing state, when the first processing station is in a processing state. While one processing station is processing... When the first When a processing station is idle ; Index to the processing station; This represents the total number of processing stations; This represents the remaining time vector for the processing station. ,in, For the current workpiece at the th The remaining processing time for each processing station; Let be the decision vector for whether the handling unit is in the workstation. ,in, To indicate whether the transport execution unit is in the first stage Variables for each processing station This is a variable used to indicate whether the transport execution unit is in the buffer. This is a variable used to indicate whether the transport execution unit is in the delivery area; A binary variable representing the load state of the transport execution unit. The time indicates that the workpiece is being clamped. The time indicates no load; Indicates the remaining execution time of the transport action of the transport execution unit; Indicates the batch task at the time step The number of remaining workpieces; A vector representing the current processing progress of the workpiece held by the handling unit. , Indicates that the transport execution unit is in the first stage The machining variables of a workpiece clamped at a machining station, when At this time, it indicates that the workpiece has completed the first step. The processes performed at each processing station; The local observation vector of the processing station is established as follows:

[0008] In the formula: For the first The observation vector of each processing station.

[0009] Furthermore, the action space of the intelligent agent for scheduling the casting cleaning production line is established as follows:

[0010] in: For the actions of the transport execution unit; For the first The actions of each processing station; For time step Joint actions;

[0011] in: This indicates that the transport execution unit is stationary and waiting for instructions; This indicates that the transport execution unit has moved from its current position to... Instructions, where, when hour, Indicates the processing station, when hour, Indicates the cache area, when hour, Indicates the delivery area; and These represent the loading and unloading operation instructions of the material handling unit, respectively.

[0012] Where: start indicates starting a task; process indicates that the current processing task is continuously executing; wait indicates that the processing station is idle or waiting to be moved. For the first The operating space of each processing station .

[0013] Furthermore, the reward function is established as follows:

[0014] In the formula: The value of the reward function; Rewards for output efficiency; For rhythm coordination; Penalty item for vacant workstations; Penalty items for non-conforming workpieces; This is a penalty item for sequential violation; in,

[0015] In the formula: This is the output reward coefficient; Indicates time step Internal output increment; Indicates at time step The number of workpieces that have not yet been cleaned, i.e., the batch task at the time step. The number of remaining workpieces;

[0016]

[0017] in: This is the beat coordination penalty coefficient; Indicates the current workpiece is at the [number]th [position]. The remaining processing time for each processing station; for The average remaining processing time for each processing station; It is the absolute value;

[0018] in, Indicates the first A variable indicating whether a processing station is in a processing state, where 1 indicates processing and 0 indicates idle; This represents the idle penalty coefficient.

[0019]

[0020] in, This is the penalty coefficient for non-conforming parts. A variable used to indicate whether the transport execution unit is in the delivery area. Indicates the workpiece at the 1st The first processing station, i.e. the... The completion status of each process step;

[0021] in, This refers to the penalty coefficient for violations of the process sequence. A binary variable representing the load state of the transport execution unit. The time indicates that the workpiece is being clamped. The time indicates no load.

[0022] Furthermore, the training objective function of the reinforcement learning algorithm is expressed as follows:

[0023] The target value is:

[0024] In the formula: For training the loss function; Representative of the experience replay pool The expected value of the transferred sample obtained from the sampling process is taken. For experience replay pool; The global action value function; For hybrid network parameters; Discount factor; The desired target Q value; For the next moment Candidate joint actions; The target hybrid network parameters are updated with a delay; For local Q-network parameter sets; For time step Instant rewards.

[0025] Beneficial Effects: This invention provides an intelligent scheduling and control method for a foundry cleaning production line based on multi-agent reinforcement learning. By introducing a multi-agent reinforcement learning framework and establishing a state space for scheduling the foundry cleaning production line, each processing station and handling unit can make autonomous decisions based on environmental conditions. Through global value function optimization, it achieves cycle time coordination and dynamic resource allocation between stations, reducing workpiece waiting time and equipment idle time, eliminating fixed cycle times or manual intervention, and improving the overall utilization rate of the production line. Based on the reward function setting of this invention, the system can automatically optimize the task allocation logic according to the actual operating status, without the need for manual adjustment of cycle time parameters, significantly reducing the workload of system debugging and maintenance, and reducing reliance on manual labor and debugging costs. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a flowchart of the intelligent scheduling and control method for a casting cleaning production line based on multi-agent reinforcement learning, as described in this invention. Figure 2 This is a schematic diagram of the multi-agent reinforcement learning framework structure in an embodiment of the present invention. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] This embodiment introduces an intelligent scheduling and control method for a casting cleaning production line based on multi-agent reinforcement learning, including the following steps: Figure 1 As shown: S1: Establish a state space for scheduling the casting cleaning production line, including a global state vector, local observation vectors of the processing stations, and an action space for the agents oriented towards scheduling the casting cleaning production line; wherein, the agents oriented towards scheduling the casting cleaning production line include the agents corresponding to the material handling execution units and the agents corresponding to the processing stations. The intelligent scheduling and control method for the casting cleaning production line in this embodiment is applicable to automated cleaning systems with multiple workstations and processes. The specific structure of the production line can be flexibly configured according to the type of casting, process requirements, and site layout, and is not limited to a single form. For ease of explanation, this embodiment uses a typical casting cleaning production line as an example: This production line mainly consists of multiple workstations and handling execution units. The workstations include processing workstations and operation workstations. In this embodiment, the processing workstations include a pre-processing workstation, a cutting workstation, and a grinding workstation. The operation workstations include a buffer area and a delivery area. Each workstation is connected to a host computer control system via an industrial communication network for status monitoring, task scheduling, and collaborative operation. The handling execution unit is a programmable handling robotic arm or a gantry-type gripper mechanism. Its picking and placing actions are executed by subroutines preset in the controller. The host computer calls the corresponding program through communication commands to achieve automatic transfer of workpieces between workstations. The pre-processing, cutting, and grinding workstations each undertake different cleaning processes, and these cleaning processes have a defined sequence. Workpieces must pass through the pre-processing workstation, cutting workstation, and grinding workstation sequentially to complete processing; skipping, reversing, or parallel execution is not allowed. The buffer area is used to temporarily store intermediate workpieces and buffer the differences in cycle time between processes. After completing all processes, the workpiece is transported to the delivery area by the handling execution unit. The host computer system collects the operating status information of each workstation and material handling unit in real time, and uses it as input to the reinforcement learning scheduling algorithm to achieve intelligent collaborative control of processing and material handling tasks at the workstation.

[0030] Specifically, to ensure that the reinforcement learning scheduling model accurately reflects the production line's operating status and makes decisions accordingly, this embodiment constructs a state space oriented towards scheduling optimization within a discrete-time step simulation environment. This state space centers on information "directly related to action feasibility and scheduling benefits," retaining only variables that substantially influence reinforcement learning decisions. Furthermore, it defines the global state and the local observations of each agent according to a centralized training and decentralized execution (CTDE) framework.

[0031] The global state vector is established as follows: Specifically, to achieve reinforcement learning scheduling control, this embodiment implements time step... The global state vector of the system is defined as follows:

[0032] In the formula: In time step The global state vector under; Let be the decision vector for the machining station. ,in, Indicates the first The variable indicating whether a processing station is in a processing state, when the first processing station is in a processing state. While one processing station is processing... When the first When a processing station is idle ; Index to the processing station; In this embodiment, there are 3 processing stations: a pre-processing station, a cutting station, and a grinding station. This represents the remaining time vector for the processing station. ,in, For the current workpiece at the th The remaining processing time of each processing station is used to describe the system load and cycle time occupancy. Let be the decision vector for whether the handling unit is in the workstation. ,in, To indicate whether the transport execution unit is in the first stage Variables for each processing station This is a variable used to indicate whether the transport execution unit is in the buffer. This is a variable used to indicate whether the transport execution unit is in the delivery area; specifically, when the transport unit is located at a certain workstation, the corresponding component is 1, and the components at other positions are 0. A binary variable representing the load state of the transport execution unit. The time indicates that the workpiece is being clamped. The time indicates no load; This indicates the remaining execution time of the transport action of the transport execution unit, used to determine whether the transport execution unit is in the execution phase; Indicates the batch task at the time step The remaining number of workpieces is the overall progress variable for the batch task, representing the number of workpieces in the current batch that have not yet been cleared. A vector representing the current processing progress of the workpiece held by the handling unit. , Indicates that the transport execution unit is in the first stage The machining variable for each workpiece held at a machining station is a variable that indicates whether the current workpiece has completed the current process. At this time, it indicates that the workpiece has completed the first step. The process performed at each processing station, when When the time is right, it means the workpiece has been cleaned.

[0033] Specifically, in actual implementation, time-related variables ( Normalization is used to eliminate dimensional differences in discrete variables. Input the reinforcement learning model in binary or one-hot encoded form.

[0034] The local observation vector of the processing station is established as follows: This embodiment employs a centralized training and distributed execution structure. During the training phase, each agent shares the global state vector. Joint optimization is performed through a centralized value function; during the execution phase, each agent optimizes based on its own local observation vector. Independent action selection enables distributed control and collaborative scheduling.

[0035] In this embodiment, the system's intelligent agents include the intelligent agent of the handling execution unit and the intelligent agents corresponding to the three processing stations, and its observation space is represented as follows: The meanings of the symbols are consistent with the definition of the global state space. Since the material handling execution unit undertakes the tasks of central scheduling and material flow, it needs to grasp the processing status and cycle time information of all workstations during the execution process. Therefore, its observation space is consistent with the global state space.

[0036] No. The observation vectors for each processing station are represented as follows:

[0037] in, Indicates the first A variable indicating whether a processing station is in a processing state; Indicates the current workpiece is at the [number]th [position]. The remaining processing time for each processing station; To indicate whether the transport execution unit is in the first stage The variables of each processing station (encoded by global position) (The corresponding components are given). A binary variable representing the load state of the transport execution unit. Indicates the number of remaining workpieces in the batch task. This variable indicates whether the current workpiece has completed the current process. The workstation agent can only acquire local information directly related to its own processing and handling interactions to ensure the independence and deployability of its decisions.

[0038] Specifically, the global state vector is used during the training phase. As input to the centralized value function, it enables joint optimization among multiple agents; during the execution phase, each agent optimizes according to its own observation vector. Independent decision-making enables collaborative control of workstation processing and transportation tasks without relying on global communication.

[0039] The action space of the agent oriented towards scheduling the casting cleaning production line is represented as follows: In the reinforcement learning scheduling and control model of this embodiment, the system at each time step The combined action is represented as:

[0040] in: For the actions of the transport execution unit; For the first The actions of each processing station; For time step Joint actions; The motion space of the transport execution unit is denoted as... This mainly includes two categories: position movement and pick-and-place operations. These actions can be formally represented as:

[0041] in: This indicates that the transport execution unit is stationary and waiting for instructions; This indicates that the transport execution unit has moved from its current position to... Instructions, where, when hour, Indicates the processing station, when hour, Indicates the cache area, when hour, Indicates the delivery area; and These represent the loading and unloading operation instructions of the material handling unit, respectively. Specifically, the executableness of an action is determined by state variables. , The signals from the workstation docking stations are jointly constrained.

[0042] No. The first workstation intelligent agent (i.e., the first...) The movement space of each processing station Defined as:

[0043] Where: start indicates that the task will be started when the workstation is ready for processing, specifically when the workstation is detected to be idle and has been loaded with materials; process indicates that the current processing task is continuously executing; wait indicates that the processing workstation is idle or waiting for material to be moved. For the first The operating space of each processing station .

[0044] This embodiment is based on the global state vector during the centralized training phase. With joint actions Optimize the global value function to achieve collaborative decision-making among agents and global convergence; during the distributed execution phase, each agent determines its own observation vector. Independent action selection enables parallel scheduling of workstation processing and material handling tasks.

[0045] S2: Construct a reward function for scheduling the foundry cleaning production line; In this embodiment, a comprehensive reward function is constructed during the reinforcement learning training process to guide the strategy to reduce workstation idle time while ensuring rhythm coordination. The reward function is expressed as follows:

[0046] In the formula: The value of the reward function; Rewards for output efficiency; For rhythm coordination; Penalty item for vacant workstations; Penalty items for non-conforming workpieces; This is a penalty item for sequential violation.

[0047] The output efficiency reward is used to measure the effective output completed by the system within the current time step, and is expressed as follows:

[0048] In the formula: This is the output reward coefficient; Indicates time step Internal output increment; Indicates at time step The number of workpieces that have not yet been cleaned, i.e., the batch task at the time step. The number of remaining workpieces; Specifically, output efficiency rewards encourage agents to improve their output per unit time through reasonable scheduling.

[0049] The cycle time coordination term is used to measure the balance of processing load between different workstations, and is expressed as follows:

[0050]

[0051] in: Indicates rhythmic coordination; This is the beat coordination penalty coefficient; Indicates the current workpiece is at the [number]th [position]. The remaining processing time for each processing station; for The average remaining processing time for each processing station; It is the absolute value; Specifically, the takt time coordination term is used to suppress takt time differences between workstations, promote synchronous operation of the system, and reduce waiting and bottlenecks.

[0052] The workstation idle penalty is used to penalize workstations that have the conditions for processing but are not performing tasks, as shown below:

[0053] in, This indicates a penalty for an idle workstation. Indicates the first A variable indicating whether a processing station is in a processing state, where 1 indicates processing and 0 indicates idle; This represents the idle penalty coefficient.

[0054] This algorithm encourages the reduction of unnecessary idle time and improves the utilization rate of workstation resources.

[0055] Specifically, when the handling unit arrives at the delivery area, the processing progress of the currently held workpiece is determined. If there are incomplete processes, penalties are imposed. The penalty items for defective workpieces are as follows:

[0056] in, This indicates the penalty for non-conforming workpieces; This is the penalty coefficient for non-conforming parts. A variable used to indicate whether the transport execution unit is in the delivery area (1 indicates yes, 0 indicates no). Indicates the workpiece at the 1st The first processing station, i.e. the... The completion status of each process step; The total number of processing stations is 3 in this embodiment, which are the pretreatment station, the cutting station, and the grinding station.

[0057] Specifically, when the handling unit has not reached the delivery area ( This item will not be triggered; when the transport execution unit arrives at the delivery area ( Furthermore, if there are incomplete processes, the system will incur penalties. This item is set to 0 only when the workpiece has completed all three processes, indicating no penalty.

[0058] Sequence violation penalties are used to punish actions by which an agent does not process data in a fixed order. They are represented as follows:

[0059] in, This refers to the penalty coefficient for violations of the process sequence. A binary variable representing the load state of the transport execution unit. The time indicates that the workpiece is being clamped. The time indicates no load.

[0060] Specifically, when the conveying execution unit clamps the workpiece ( If a workpiece has a situation where subsequent processes are completed but preceding processes are not, the item in parentheses will have a value of 1, triggering the sequence violation penalty, and the system will generate a penalty. When the handling unit is idle, or the workpiece's process completion status meets the predetermined processing sequence requirements, the sequence violation penalty item is set to 0, and no penalty is generated. By introducing this process sequence violation penalty item, the reinforcement learning scheduling strategy is guided to strictly follow the fixed processing sequence of the casting cleaning process during training, thereby avoiding scheduling and handling decisions that violate process specifications.

[0061] S3: Based on the state space and reward function for scheduling the foundry cleaning production line, a reinforcement learning algorithm is used to obtain real-time scheduling instructions for the agent scheduling the foundry cleaning production line, thereby achieving intelligent scheduling of the foundry cleaning production line; for example... Figure 2 As shown; The reinforcement learning environment in this embodiment is built upon the interaction between the workstation processing and handling execution units of the casting cleaning production line, and its operation follows the state transition law of discrete time steps. The system operates at time steps... The global state vector is denoted as Each agent selects a joint action based on its current state. via environment functions Transition to the next state and return instant rewards State transition function The position code is determined by both the physical logic and technological constraints of the production line when the material handling unit performs a movement action. Updated to the unique heat vector of the target station; when performing loading or unloading actions, the load status... Workstation processing markings It will change synchronously according to the docking conditions; if the first If a processing station is in the processing state, then its remaining processing time Decrease at each time step, when Automatically switches to idle state when the workpiece is in the first... After each processing station completes its corresponding process, the processing progress vector is... The amount The value was updated from 0 to 1, indicating that the current process has been completed; batch task quantity. The number of cycles decreases after the workpiece is cleaned until the task is completed.

[0062] To ensure that the state transition process conforms to safety and technological logic, this embodiment introduces an action space constraint mechanism for any processing agent. It is only allowed in the state The following set of possible actions The system makes decisions internally. Action constraints include, but are not limited to, the following: When a workstation is in a processing state, repeatedly triggering the "start processing" action is prohibited; when the transport execution unit is not holding a workpiece, performing the "unloading" operation is prohibited; when the target workstation has not completed its previous processing or is occupied, the transport execution unit is prohibited from performing the "loading" operation; when the transport execution unit is in motion, issuing new movement commands is prohibited. Through this mechanism, the reinforcement learning system avoids dangerous or illegal operations during both the training and execution phases, ensuring that the state transition process truly reflects the dynamic changes and safety constraints of the production line, thereby guaranteeing the stability of model learning and the deployability of the scheduling strategy.

[0063] In this embodiment, the reinforcement learning algorithm training process is completed in a program simulation environment. By constructing a digital simulation model corresponding to the casting and cleaning production line, the actions of the workstation processing, handling execution units, and buffer flow processes are simulated, enabling the algorithm to interactively learn and optimize parameters in a virtual environment. The training phase adopts a centralized training and distributed execution structure, with a value decomposition-based multi-agent reinforcement learning algorithm as the core, to achieve collaborative scheduling control among multiple workstations.

[0064] During the simulation, the environment changes at time steps. Output global state Each agent relies on local observations Select Action Together they form an action vector. And instant rewards calculated by the environment. With the next state All interactive samples Stored in the experience replay pool In this context, it is used for batch sampling during the training phase.

[0065] During the intensive training phase, each agent (including the agent in the transport execution unit and the agent corresponding to the processing station) has an independent local Q-network. The Q-network of the agent corresponding to the processing station is... This is used to estimate the value of local states and actions. In this embodiment, the training process employs the QMIX framework, a representative of mature value decomposition multi-agent reinforcement learning methods, for centralized training and distributed execution to achieve collaborative optimization among multiple agents. This framework uses a global value function with a hybrid network structure. To integrate the local Q values. The network input is the global state. and all local Q values Through nonlinear mapping Achieve a decomposable mapping of global value, where, The Q value represents the intelligent agent of the transport execution unit.

[0066] The training objective of this framework is to minimize the temporal difference error.

[0067] The target value is:

[0068] In the formula: The training loss function (temporal difference error loss); Representative of the experience replay pool Transfer samples obtained from sampling Take the expected value; As an experience replay pool, samples are decorrelated through random sampling; The global action value function; For hybrid network parameters; This is a discount factor used to balance current and future returns; The desired target Q value is used to guide parameter updates; For the next moment Candidate joint actions; The target hybrid network parameters are updated with a delay for stable training; For local Q-network parameter sets; This embodiment updates the information via backpropagation. and This enables joint optimization of the global value function and local strategies.

[0069] To improve training stability, this invention introduces an experience replay mechanism, a soft parameter update strategy, and... Greedy exploration ( The (-greedy) method is used to maintain exploratory nature in the early stages of training and improve convergence in the later stages; at the same time, constraint functions are used during the action sampling phase. Illegal actions are automatically screened to ensure that the training process always complies with the safety and process logic of the production line. After multiple rounds of simulation iterations, the local Q-networks of each agent converge to a stable policy. During the execution phase, it relies solely on its own observations and independent decision-making to achieve multi-station collaborative scheduling and global optimal control of the casting cleaning production line.

[0070] During the deployment phase, this embodiment downloads the trained model to the host computer of the production line control system and interfaces it with the status monitoring and command interface of the actual equipment. The deployed system retains only the local Q-networks of each agent for real-time decision-making, no longer relying on the global hybrid network from the centralized training phase, thus achieving distributed execution. Each agent collects workstation status signals and handling unit position data in real time, inputs them into the corresponding Q-network, selects the optimal action, and outputs control commands to the actuators. The host computer sends commands to each workstation control unit via an industrial communication bus, realizing automatic scheduling of actions such as workpiece loading, processing, and handling.

[0071] This embodiment optimizes the scheduling and control process of a casting cleaning production line based on the concept of multi-agent reinforcement learning. The system abstracts each functional unit in the production line (including pre-processing station, cutting station, grinding station, and handling execution unit, etc.) into an independent intelligent agent. Each intelligent agent autonomously decides to execute actions based on its own real-time status and environmental information to achieve the orderly flow of workpieces in the production line.

[0072] During the training phase, the system uses a procedural simulation environment to perform centralized training and distributed execution of strategies for each agent. Each agent outputs a local value function to evaluate the merits of its actions. All local values ​​are combined using a reinforcement learning method based on value decomposition to form a global value function describing the overall operating state of the production line. The system uses this global value as the optimization objective to continuously adjust the strategy parameters of each agent, gradually improving overall operating efficiency, workstation utilization, and cycle time coordination.

[0073] After training, the resulting model is deployed into the actual production line control system. The host computer, based on real-time collected workstation status information, invokes the corresponding strategy model and issues control commands to each workstation and handling unit. Each agent automatically completes workpiece processing, transfer, and cycle time matching based on the output of the reinforcement learning model, thereby achieving adaptive collaborative control of the production line without human intervention.

[0074] This embodiment presents an intelligent scheduling and control method for a casting cleaning production line based on multi-agent reinforcement learning, which enables adaptive collaborative control of each workstation and handling unit on the production line. The method constructs a typical simulation model of a casting cleaning production line, including a pre-processing workstation, a cutting workstation, a grinding workstation, and a handling unit, treating each functional unit (including the pre-processing workstation, cutting workstation, grinding workstation, and handling unit) as an independent agent. Each agent independently decides its current action based on its observed local state, and the host computer control system uses a reinforcement learning algorithm to centrally train the collaborative strategies of each agent.

[0075] During training, each agent outputs a local value function to describe the reward of its current action. The system employs a multi-agent reinforcement learning framework based on value decomposition (such as QMIX) to establish a global hybrid network, dynamically fusing the local value function with global state information to obtain the global value function. Through continuous iterative training, the global value is maximized, thereby obtaining the optimal production line scheduling strategy.

[0076] After training, the strategy models of each agent are deployed to the actual production line control system. The host computer collects the status information of the workstations and handling units in real time, and outputs control commands based on the trained strategy models to realize the automatic flow of workpieces and job scheduling. This method can dynamically adjust task allocation and job sequence under conditions of production cycle changes, uneven workstation loads, or equipment status fluctuations, thereby maintaining the overall high efficiency and stable operation of the production line.

[0077] The intelligent scheduling and control method in this embodiment has good flexibility and generalization ability. The strategy model obtained through simulation environment training can be transferred to production lines with different structures or configurations. Only a few parameters need to be fine-tuned to adapt to new working conditions, making it suitable for flexible production modes with multiple varieties and small batches.

[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for intelligent scheduling and control of a casting cleaning production line based on multi-agent reinforcement learning, characterized in that, Includes the following steps: S1: Establish a state space for scheduling the casting cleaning production line, including a global state vector, local observation vectors of the processing stations, and an action space for the agent oriented towards scheduling the casting cleaning production line. S2: Construct a reward function for scheduling the foundry cleaning production line; S3: Based on the state space and reward function for scheduling the casting cleaning production line, a reinforcement learning algorithm is used to obtain real-time scheduling instructions for the intelligent agent for scheduling the casting cleaning production line, so as to realize intelligent scheduling of the casting cleaning production line. Among them, the intelligent agents for scheduling the casting cleaning production line include intelligent agents corresponding to the material handling execution unit and intelligent agents corresponding to the processing station.

2. The intelligent scheduling and control method for a casting cleaning production line based on multi-agent reinforcement learning according to claim 1, characterized in that, The global state vector is established as follows: In the formula: In time step The global state vector under; Let be the decision vector for the machining station. ,in, Indicates the first The variable indicating whether a processing station is in a processing state, when the first processing station is in a processing state. While one processing station is processing... When the first When a processing station is idle ; Index to the processing station; This represents the total number of processing stations; This is the vector of remaining time at the processing station. ,in, For the current workpiece at the th The remaining processing time for each processing station; Let be the decision vector for whether the handling unit is in the workstation. ,in, To indicate whether the transport execution unit is in the first stage Variables for each processing station This is a variable used to indicate whether the transport execution unit is in the buffer. This is a variable used to indicate whether the transport execution unit is in the delivery area; A binary variable representing the load state of the transport execution unit. The time indicates that the workpiece is being clamped. The time indicates no load; Indicates the remaining execution time of the transport action of the transport execution unit; Indicates the batch task at the time step The number of remaining workpieces; A vector representing the current processing progress of the workpiece held by the handling unit. , Indicates that the transport execution unit is in the first stage The machining variables of a workpiece held at a machining station, when At this time, it indicates that the workpiece has completed the first step. The processes performed at each processing station; The local observation vector of the processing station is established as follows: In the formula: For the first The observation vector of each processing station.

3. The intelligent scheduling and control method for a casting cleaning production line based on multi-agent reinforcement learning according to claim 1, characterized in that, The action space of the intelligent agent for scheduling the casting cleaning production line is established as follows: in: For the actions of the transport execution unit; For the first The actions of each processing station; For time steps Joint actions; in: This indicates that the transport execution unit is stationary and waiting for instructions; This indicates that the transport execution unit has moved from its current position to... Instructions, where, when hour, Indicates the processing station, when hour, Indicates the cache area, when hour, Indicates the delivery area; and These represent the loading and unloading operation instructions of the material handling unit, respectively. Where: start indicates starting a task; process indicates that the current processing task is continuously executing; wait indicates that the processing station is idle or waiting to be moved. For the first The operating space of each processing station .

4. The intelligent scheduling and control method for a casting cleaning production line based on multi-agent reinforcement learning according to claim 1, characterized in that, The reward function is established as follows: In the formula: The value of the reward function; Rewards for output efficiency; For rhythm coordination; This is a penalty item for workstations being idle. Penalty items for non-conforming workpieces; This is a penalty item for sequential violation; in, In the formula: This is the output reward coefficient; Indicates time step Internal output increment; Indicates at time step The number of workpieces that have not yet been cleaned, i.e., the batch task at the time step. The number of remaining workpieces; in: This is the beat coordination penalty coefficient; Indicates the current workpiece is at the [number]th [position]. The remaining processing time for each processing station; for The average remaining processing time for each processing station; It is the absolute value; in, Indicates the first A variable indicating whether a processing station is in a processing state, where 1 indicates processing and 0 indicates idle; This is the idle penalty coefficient; in, This is the penalty coefficient for non-conforming parts. A variable used to indicate whether the transport execution unit is in the delivery area. Indicates the workpiece at the 1st The first processing station, i.e. the... The completion status of each process step; in, This refers to the penalty coefficient for violations of the process sequence. A binary variable representing the load state of the transport execution unit. The time indicates that the workpiece is being clamped. The time indicates no load.

5. The intelligent scheduling and control method for a casting cleaning production line based on multi-agent reinforcement learning according to claim 1, characterized in that, The training objective function of the reinforcement learning algorithm is expressed as follows: The target value is: In the formula: For training the loss function; Representative of the experience replay pool The expected value of the transferred sample obtained from the sampling process is taken. For experience replay pool; The global action value function; For hybrid network parameters; Discount factor; The desired target Q value; For the next moment Candidate joint actions; The target hybrid network parameters are updated with a delay; For local Q-network parameter sets; For time steps Instant rewards.