An Intelligent Dynamic Scheduling Method for AGVs in Automated Container Terminals

By applying the C-DQN multi-agent deep reinforcement learning algorithm in an automated container terminal, the scheduling problem of AGV in dynamic and complex environments was solved, and the AGV work efficiency was significantly improved, with the specific efficiency being increased to 77.6%.

CN118504866BActive Publication Date: 2025-06-20HARBIN INST OF TECH

Patent Information

Application Number
CN202410520911.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-28
Publication Date
2025-06-20
Estimated Expiration
2044-04-28

AI Technical Summary

Technical Problem

Existing intelligent algorithms are difficult to effectively solve the scheduling problem of AGV in dynamic and complex environments in automated container terminals, especially in multi-agent control, non-stationarity and partial observability.

Method used

The C-DQN multi-agent deep reinforcement learning algorithm is adopted, and the dock loading and unloading operations are divided into multiple modules to model separately, and the event logical relationship connects the modules to build a Markov decision-making process, train the multi-agent deep reinforcement learning model, and finally package it into an online real-time scheduling system.

Benefits of technology

The container operation sequence solution that minimizes AGV conflict waiting time is realized, and the working efficiency of AGV is improved, and the specific efficiency is increased to 77.6%, providing a solid theoretical foundation for the intelligent dynamic scheduling of AGV at an automated container terminal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118504866B_ABST
    Figure CN118504866B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of resource scheduling optimization for automated container terminals, and specifically relates to an intelligent dynamic scheduling method for AGVs in automated container terminals, which includes the following steps: Step 1: Conduct simulation modeling by dividing modules according to different functions of the automated container terminal; Step 2: Connect various different functional modules through event logical relationships; Step 3: Build a C-DQN multi-agent deep reinforcement learning algorithm framework; Step 4: Construct a Markov decision process; Step 5: Train the C-DQN multi-agent deep reinforcement learning model; Step 6: Package the trained model into an online real-time scheduling system, and the devices participating in the work include quay front handling equipment, horizontal transportation equipment, and yard operation equipment. The present invention divides the terminal handling operation into multiple modules for separate modeling, controls multiple AGVs, and realizes the solution of the container operation sequence that minimizes the conflict waiting time of AGVs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of resource scheduling optimization for automated container terminals, and particularly relates to an intelligent dynamic scheduling method for AGVs in automated container terminals. Background Art

[0002] With the development of the transportation industry, container terminals are also transforming towards the direction of intelligence and automation. To promote the construction of new automated container terminals and the unmanned system construction of bulk dry bulk terminals, and to accelerate the comprehensive applications such as intelligent dispatching of ports and stations and remote control of equipment, China aims to build a global intelligent shipping development center and form a high-quality intelligent shipping system. Currently, there are 27 automated container terminals under construction and 18 completed automated container terminals in China, with the total number ranking first in the world.

[0003] Compared with traditional container terminals, automated container terminals have greater transportation capacity, more complex environments, and more difficult scheduling. Automated container terminals adopt automatic control technology, requiring the realization of full or partial unmanned operation of terminal handling equipment, dispatching control systems, etc., in order to improve work efficiency and reduce labor costs. Therefore, an intelligent algorithm that can improve scheduling efficiency is crucial. Existing intelligent algorithms can control each AGV along the Internet of Things system to carry out operations such as loading, unloading, and transportation. As a container horizontal transportation equipment connecting the terminal berth and the yard, AGV has become the key to affecting the handling efficiency of automated container terminals.

[0004] Currently, the application of traditional operations research methods is more common for scheduling the handling equipment of automated container terminals. However, this method has great limitations and is mostly applicable to static deterministic problems, and it is difficult to cope with dynamic complex environments. Intelligent algorithms based on deep reinforcement learning have natural advantages for dynamic scheduling problems, but there are also problems such as non-stationarity, partial observability, game and cooperation in the control of multi-agent systems. Summary of the Invention

[0005] The purpose of the present invention is to provide an intelligent dynamic scheduling method for AGVs in automated container terminals, which can divide the terminal handling operations into multiple modules for separate modeling, control multiple AGVs, and solve the container operation sequence that minimizes the conflict waiting time of AGVs.

[0006] The technical solution adopted by the present invention is specifically as follows:

[0007] An intelligent dynamic scheduling method for AGVs in automated container terminals, comprising the following steps:

[0008] Step 1: Conduct simulation modeling by dividing modules according to different functions of the automated container terminal;

[0009] Step 2: Connect various different functional modules through event logical relationships;

[0010] Step 3: Build a C-DQN multi-agent deep reinforcement learning algorithm framework;

[0011] Step 4: Construct a Markov decision process;

[0012] Step 5: Train the C-DQN multi-agent deep reinforcement learning model;

[0013] Step 6: Package the trained model into an online real-time scheduling system.

[0014] When the automated container terminal performs loading and unloading operations in Step 1, the equipment involved in the work includes quay front loading and unloading equipment, horizontal transportation equipment, and yard operation equipment;

[0015] Among them, according to the working processes of different loading and unloading equipment in the automated container terminal, they can be respectively modeled as:

[0016] Quay crane module, responsible for simulating the loading and unloading of containers between the ship and the AGV;

[0017] Yard crane module, responsible for simulating the loading and unloading of containers between the yard and the AGV;

[0018] AGV module, responsible for simulating the container transportation work of the AGV;

[0019] AGV queue manager, responsible for managing the working sequence of multiple AGVs at the ship berth and the yard;

[0020] AGV action selector, responsible for communicating with external algorithms and issuing work instructions to the AGV;

[0021] Task reporting module, responsible for evaluating the decisions of the AGV and reporting to the operator.

[0022] In Step 1, the container loading and unloading tasks of the automated container terminal are also modeled, and different container tasks are abstracted into a five-tuple:

[0023] Task_information = ['Container', 'Loading or unloading', 'Quay cranes number', 'Field cranes number', 'Container tonnage'] (1)

[0024] Among them, 'Container' represents the number of the container task, 'Loading or unloading' represents the loading and unloading type of the container task, 'Quay cranes number' represents the quay crane number related to the container task, 'Field cranes number' represents the yard crane number related to the container task, and 'Container tonnage' represents the weight level of the container.

[0025] In step 2, analyze the working relationships between the various modules of the automated container terminal simulation model in step 1, find the interactions and logical connections between the various modules, and describe these logical relationships with events; realize the communication between the various modules through the activation and capture of events, simulate the working process of each module through the delay function, and the various modules run in parallel to simulate the various activities of the terminal, thus realizing the simulation operation of the automated container terminal.

[0026] Step 3 includes the following steps:

[0027] First, construct a C-DQN multi-agent deep reinforcement learning algorithm framework. The C-DQN multi-agent deep reinforcement learning algorithm consists of a prediction network and a target network;

[0028] Among them, the prediction network is used to output the value Q of each action in the current state of the agent, and the target network is used to output the maximum value Q of the action in the next state of the agent target ;

[0029] Then, use a mask to process the action value Q output by the prediction network and the target network, and mask out the actions that do not conform to the rules; among them, the mask is a string of 0-1 arrays, and the dimension size is the same as the action output dimension of the neural network model; in the mask, the mask position corresponding to the action that conforms to the rules is 1, and the mask position corresponding to the action that does not conform to the rules is 0. Process the output of the neural network with the mask, and the action value at the mask position of 1 remains unchanged, and the action value at the mask position of 0 is replaced with a minimum value;

[0030] Next, use a greedy strategy with a fixed step size and variable greed rate to select the action of the AGV and generate a random number;

[0031] When the random number is less than the greed rate ε, use the Softmax function to select the action with the largest output value in the prediction network;

[0032] When the random number is greater than the greed rate ε, the AGV randomly selects an action that conforms to the rules;

[0033] After the C-DQN multi-agent deep reinforcement learning algorithm model is trained a certain number of times, gradually increase the value of the greed rate ε, so that the probability of the neural network randomly selecting an action becomes smaller, and the model starts to output a definite action;

[0034] Then, store the information of each task completed by the AGV as samples in the experience replay buffer, and use the method of centralized learning and distributed execution for training; when training, train on the states and environmental information of all agents, and when outputting decisions, each agent makes a separate judgment to output a decision action that conforms to the current environmental state. The samples in the experience replay buffer take the form of a mixture of the state information of multiple agents, as follows:

[0035] (s i ,a i ,r i ,s′1,s′2,...,s′ n ) (2)

[0036] Among them, i represents the serial number of the agent, s i represents the state before the agent i executes the action, a i represents the action executed by the agent i, r i represents the reward obtained by the agent i after executing the action, s′1, s′2,..., s′ n respectively represent the states of the remaining agents after they execute their respective current actions when the agent i starts to execute the action a i ;

[0037] Finally, whenever the neural network model needs to be trained, randomly extract a batch of samples from the experience replay buffer for learning. According to the sample information, the value information of the predicted network outputting the current action is: Q(s i ,a i ; θ);

[0038] Among them, θ represents the parameter information of the prediction network. The target network synthesizes the information of multiple agents and outputs the target value as:

[0039]

[0040] Among them, γ represents the discount factor, θ′ represents the parameter information of the target network. According to the action values obtained from the two neural networks, use the mean square loss function to measure the degree of difference, and then perform gradient backpropagation to update the network parameters. The mean square loss function is as follows:

[0041] loss=E[(Q target -Q(s i ,a i ; θ)) 2 (4)

[0042] Among them, E represents the average value of a preset random variable.

[0043] The fourth step includes the following steps:

[0044] First, construct the state space of the C-DQN multi-agent deep reinforcement learning. The state space includes:

[0045] Container task status information, where 1 indicates not completed and 0 indicates completed;

[0046] AGV task information, the container task selected by the AGV;

[0047] AGV progress information, where 0 indicates that the AGV is in a waiting state, 0.5 indicates that the AGV is in the process of picking up a container, and 1 indicates that the AGV is in the process of placing a container;

[0048] AGV time information, the time the AGV has experienced in the current task state;

[0049] Handling equipment working information, where 0 indicates that the quay crane or yard crane is not working and 1 indicates that the quay crane or yard crane is working;

[0050] Handling equipment time information, the time the quay crane or yard crane has experienced in the current working state;

[0051] Tonnage information, the weight of the last container placed by the quay crane or yard crane;

[0052] Among them, when the handling equipment in the automated container terminal increases, the dimension of the state space increases accordingly;

[0053] Then, construct the action space of the deep reinforcement learning. The action space includes all container tasks and a waiting action. The dimension of the action space is N + 1, where N is the number of container tasks;

[0054] Next, in order to reflect the game cooperation relationship among multiple agents in the deep reinforcement learning, the waiting time of the AGV during the working process is used as the optimization goal, as follows:

[0055] When the AGV is performing an operation, the reward calculation basis is 20 - t waiting / 10, where t waiting is the waiting duration of the AGV in one task;

[0056] When the AGV performs a waiting action, a penalty of -1.5 is given;

[0057] When the action selected by the AGV conforms to the container flipping rule, the reward is +3 on the original basis;

[0058] When the action selected by the AGV does not conform to the rules of turning over the box, the reward is -1 on the original basis;

[0059] Among them, the reward function of the optimization target is expressed as:

[0060]

[0061] Finally, construct the parameter structure of the neural network; both the prediction network and the target network adopt a three-layer network structure. The first layer network contains 128 neurons, the second layer network contains 64 neurons, and the networks are fully connected. The optimizer uses the Adam optimizer, and the linear rectifier function ReLU is used as the activation function, as follows:

[0062]

[0063] Moreover, the discount factor γ is obtained by an empirical formula, where L is set as the number of steps that the agent is expected to consider next, and it is expressed as:

[0064]

[0065] The fifth step includes the following steps:

[0066] First, initialize the initial state of the automated container terminal simulation model, input the initial state into the C-DQN multi-agent deep reinforcement learning algorithm model, and obtain the task selection of the current AGV after mask processing;

[0067] Secondly, input this task into the automated container terminal simulation model through the AGV action selector. The AGV executes this action, and the AGV module starts to work, interacting with the quay crane module, the yard crane module, and the AGV queue manager;

[0068] Then, when the AGV finishes executing the current action, interact with the external algorithm through the AGV action selector, record the current state, the executed task, the obtained reward, and the next state, store them in a temporary tuple, and increment the tuple counter by one.

[0069] The fifth step also includes the following steps:

[0070] First, judge whether the data in the tuple meets all the conditions for forming a sample. If so, store this sample in the experience replay buffer, delete the current sample from the tuple, and decrement the tuple counter by one;

[0071] Or judge that the data in the tuple does not meet all the conditions for forming a sample. Execute the next task of this agent, and at the same time wait for other agents to finish executing their current actions, record their information in the tuple, and combine it with the tuple to form a sample;

[0072] Next, it is judged that the experience replay buffer of the C-DQN multi-agent deep reinforcement learning algorithm is full, then the C-DQN multi-agent deep reinforcement learning algorithm model starts training. A batch of samples is randomly selected from the experience replay buffer and input into the prediction network and the target network respectively. The obtained Q(s i , a i ; θ) and Q target The degree of difference Loss is obtained through the mean square loss function and the gradient is backpropagated. The Adam optimizer is used to optimize the parameters of the prediction network;

[0073] Or it is judged that the experience replay buffer of the C-DQN multi-agent deep reinforcement learning algorithm is not full, then the automated container terminal simulation model continues to run.

[0074] The following steps are also included in step five:

[0075] First, it is judged that the number of training times of the C-DQN multi-agent deep reinforcement learning algorithm reaches the set value, then the model parameters of the prediction network are copied to the target network;

[0076] Or it is judged that the number of training times of the C-DQN multi-agent deep reinforcement learning algorithm does not reach the set value, then the automated container terminal simulation model continues to run;

[0077] Then, it is judged that there are unfinished container tasks, then the current state is input into the C-DQN multi-agent deep reinforcement learning algorithm model, and the next container task is obtained after masking processing and handed over to the AGV by the AGV action selector for work;

[0078] Or it is judged that there are no unfinished container tasks, then this round of container tasks is all completed, and the task reporting module outputs the task scheduling information of this round of AGV and the total reward obtained;

[0079] Finally, it is judged the number of training times of the C-DQN multi-agent deep reinforcement learning algorithm model; after reaching a certain number of training times, the greed rate ε is gradually increased; when the number of training times reaches the set maximum value, the training is stopped, the final AGV scheduling scheme and reward value are output, the model parameters of the prediction network are saved and output.

[0080] The following steps are included in step six:

[0081] After obtaining the trained C-DQN multi-agent deep reinforcement learning algorithm model using step five, the model is applied to the online real-time scheduling system of the AGV;

[0082] When the AGV needs to make a decision during operation, the current environmental state is input into the C-DQN multi-agent deep reinforcement learning algorithm model, and the actions that the AGV should execute are output by the algorithm model;

[0083] After the AGV completes the task, the environmental state information is updated in real time, and the algorithm model outputs the next action again, so as to realize the intelligent dynamic scheduling of the AGV.

[0084] The technical effects achieved by the present invention are as follows:

[0085] An intelligent dynamic scheduling method for AGVs in an automated container terminal according to the present invention can divide the terminal loading and unloading operations into multiple modules for separate modeling according to different work contents of the automated container terminal, has good scalability and versatility, and can control multiple AGVs to solve the container operation sequence with the minimum AGV conflict waiting time.

[0086] An intelligent dynamic scheduling method for AGVs in an automated container terminal according to the present invention connects different modules according to the event logical relationship to obtain an automated container terminal simulation model.

[0087] An intelligent dynamic scheduling method for AGVs in an automated container terminal according to the present invention proposes a C-DQN multi-agent deep reinforcement learning algorithm and constructs a network model, and trains the algorithm model on the basis of the automated container terminal simulation model. Finally, the trained algorithm model is exported to realize the intelligent dynamic scheduling of AGVs in the automated container terminal.

[0088] An intelligent dynamic scheduling method for AGVs in an automated container terminal according to the present invention, compared with the AGV scheduling method based on the traditional DQN deep reinforcement learning algorithm, uses the C-DQN multi-agent deep reinforcement learning algorithm to schedule the AGV, which can improve the working efficiency of the AGV by 77.6%, laying a solid theoretical foundation for the intelligent dynamic scheduling of AGVs in the automated container terminal. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] Figure 1 is a model framework diagram of the C-DQN multi-agent deep reinforcement learning algorithm of the present invention;

[0090] Figure 2 is a schematic diagram of the pseudocode of the C-DQN multi-agent deep reinforcement learning algorithm of the present invention;

[0091] Figure 3 is a training flow chart of the C-DQN multi-agent deep reinforcement learning algorithm of the present invention;

[0092] Figure 4 is a schematic diagram of the automated container terminal simulation model in Embodiment 1 of the present invention;

[0093] Figure 5 It is the operation diagram of the automated container terminal simulation model in the first embodiment of the present invention;

[0094] Figure 6 It is the method flow chart of an AGV intelligent dynamic scheduling method for an automated container terminal in the first embodiment of the present invention;

[0095] Figure 7 It is the reward curve diagram of the C-DQN multi-agent deep reinforcement learning algorithm in the first embodiment of the present invention;

[0096] Figure 8 It is the AGV scheduling Gantt chart solved by the C-DQN multi-agent deep reinforcement learning algorithm in the first embodiment of the present invention;

[0097] Figure 9 It is the reward curve diagram of the DQN algorithm in the second embodiment of the present invention;

[0098] Figure 10 It is the AGV scheduling Gantt chart solved by the DQN algorithm in the second embodiment of the present invention. Detailed implementation manners

[0099] In order to make the purpose and advantages of the present invention clearer, the present invention will be specifically described below in conjunction with embodiments. It should be understood that the following text is only used to describe one or several specific implementation manners of the present invention, and does not strictly limit the scope of protection of the specific requests of the present invention.

[0100] Embodiment 1:

[0101] As Figure 1-8 shown, an AGV intelligent dynamic scheduling method for an automated container terminal includes the following steps:

[0102] Step 1: Perform simulation modeling by dividing modules according to different functions of the automated container terminal;

[0103] First of all, this experiment simulated an automated container terminal simulation environment. The quay berths of the terminal are arranged along the shore, and the yard forms a 90° angle with the quay berths of the terminal. The distance position information in the terminal is shown in the following table, with the unit of m;

[0104] Table 1 Distance position information in the terminal

[0105] Quayside Crane 1 Quayside Crane 2 Quayside Crane 3 Yard Crane 1 Yard Crane 2 Yard Crane 3 Quayside Crane 1 0 200 400 200 150 100 Quayside Crane 2 200 0 200 282.84 250 223.61 Quayside Crane 3 400 200 0 447.21 427.20 412.31 Yard Crane 1 200 282.84 447.21 0 50 100 Yard Crane 2 150 250 427.20 50 0 50 Yard Crane 3 100 223.61 412.31 100 50 0

[0106] When an automated container terminal conducts loading and unloading operations, the main equipment involved in the work includes quay front loading and unloading equipment, horizontal transportation equipment, and yard operation equipment. According to the work processes of different loading and unloading equipment in the automated container terminal, they can be respectively modeled as: quay crane module, yard crane module, AGV module, AGV queuing manager, AGV action selector, and task reporting module, as follows:

[0107] The quay crane module is responsible for simulating the loading and unloading of containers between the ship and the AGV;

[0108] The yard crane module is responsible for simulating the loading and unloading of containers between the yard and the AGV;

[0109] The AGV module is responsible for simulating the container transportation work of the AGV;

[0110] The AGV queuing manager is responsible for managing the working sequence of multiple AGVs at the ship berth and the yard;

[0111] The AGV action selector is responsible for communicating with external algorithms and issuing work instructions to the AGV;

[0112] The task reporting module is responsible for evaluating the decisions of the AGV and reporting to the operator;

[0113] The relevant parameters of the loading and unloading equipment in the automated container terminal simulation model are shown in the following table:

[0114] Table 2 Relevant Parameters of Loading and Unloading Equipment

[0115]

[0116]

[0117] Then, model the container loading and unloading tasks of the automated container terminal, and abstract different container tasks into a five-tuple:

[0118] Task_information = ['Container', 'Loading or unloading', 'Quay cranes number', 'Field cranes number', 'Container tonnage'] (1)

[0119] Among them, 'Container' represents the number of the container task, 'Loading or unloading' represents the loading and unloading type of the container task, 'Quay cranes number' represents the quay crane number related to the container task, 'Field cranes number' represents the yard crane number related to the container task, and 'Container tonnage' represents the weight level of the container;

[0120] A total of 30 container tasks are set for the AGV to schedule. The container tasks include stuffing tasks and discharging tasks. The container task information is integrated into the form of a five-tuple for the algorithm model to read and identify. The relevant task information is shown in the following table:

[0121] Table 3 Container Task Information

[0122]

[0123]

[0124] Step 2: Connect each different functional module through event logical relationships;

[0125] Analyze the working relationships between the various modules of the automated container terminal simulation model, find the interactions and logical connections between the various modules, and describe these logical relationships with events; realize the communication between the various modules through the activation and capture of events, simulate the working process of each module through the delay function, and the various modules run in parallel to simulate the various activities of the terminal, realizing the simulation operation of the automated container terminal;

[0126] Step 3: Build a C-DQN multi-agent deep reinforcement learning algorithm framework;

[0127] First, build a C-DQN multi-agent deep reinforcement learning algorithm framework. The C-DQN multi-agent deep reinforcement learning algorithm consists of a prediction network and a target network;

[0128] Among them, the prediction network is used to output the value Q of each action of the agent in the current state, and the target network is used to output the maximum value Q of the action of the agent in the next state target ;

[0129] Then, the action values Q output by the prediction network and the target network are processed with a mask to mask out the actions that do not conform to the rules; the mask is a string of 0-1 arrays with the same dimensionality as the action output dimensionality of the neural network model; the mask contains information about whether the actions conform to the rules, where the mask positions corresponding to the actions that conform to the rules are 1, and the mask positions corresponding to the actions that do not conform to the rules are 0. The output of the neural network is processed with the mask, and the action values at the mask positions of 1 remain unchanged, while the action values at the mask positions of 0 are replaced with extremely small values;

[0130] Next, a greedy strategy with a fixed step size and variable greed rate is used to select the actions of the AGV, and a random number is generated;

[0131] When the random number is less than the greed rate ε, the Softmax function is used to select the action with the maximum output value in the prediction network;

[0132] When the random number is greater than the greed rate ε, the AGV randomly selects a conforming action;

[0133] After the C-DQN multi-agent deep reinforcement learning algorithm model is trained a certain number of times, the value of the greed rate ε is gradually increased, so that the probability of the neural network randomly selecting actions becomes smaller, and the model begins to output definite actions;

[0134] Then, the information of each item for the AGV to complete the task is stored as a sample in the experience replay buffer. To solve the complex problems of multi-agent reinforcement learning, a centralized learning and distributed execution method is used for training; during training, the states and environmental information of all agents are trained, and when outputting decisions, each agent makes a separate judgment to output a decision action that conforms to the current environmental state. The samples in the experience replay buffer take the form of mixing the state information of multiple agents, as follows:

[0135] (s i , a i , r i , s′1, s′2,..., s′ n ) (2)

[0136] where i represents the serial number of the agent, s i represents the state of agent i before performing the action, a i represents the action performed by agent i, r i represents the reward obtained by agent i after performing the action, and s′1, s′2,..., s′ n respectively represent the states of the remaining agents after they have performed their respective current actions when agent i starts to perform action a i ;

[0137] Finally, whenever the neural network model needs to be trained, a small batch of samples is randomly drawn from the experience replay buffer for learning. According to the sample information, the value information of the current action predicted by the prediction network is: Q(s i , a i ; θ);

[0138] where θ represents the parameter information of the prediction network. The target network synthesizes the information of multiple agents and sets the target value as:

[0139]

[0140] where γ represents the discount factor, θ′ represents the parameter information of the target network. According to the action values obtained from the two neural networks, the mean squared loss function is used to measure the degree of difference, and then gradient backpropagation is performed to update the network parameters. The mean squared loss function is as follows:

[0141] loss = E[(Q target - Q(s i , a i ; θ)) 2 (4)

[0142] where E represents the average value of a preset random variable;

[0143] Step 4: Construct a Markov decision process;

[0144] First, construct the state space of the C-DQN multi-agent deep reinforcement learning. The state space includes:

[0145] Container task status information, 1 indicates not completed, 0 indicates completed;

[0146] AGV task information, the container task selected by the AGV;

[0147] AGV progress information, 0 indicates that the AGV is in a waiting state, 0.5 indicates that the AGV is in the process of picking up the container, 1 indicates that the AGV is in the process of placing the container;

[0148] AGV time information, the time the AGV has experienced in the current task state;

[0149] Handling equipment working information, 0 indicates that the quay crane or yard crane is not working, 1 indicates that the quay crane or yard crane is working;

[0150] Handling equipment time information, the time the quay crane or yard crane has experienced in the current working state;

[0151] Tonnage information, the weight of the last container placed by the quay crane or yard crane;

[0152] When the number of loading and unloading equipment in an automated container terminal increases, the dimension of the state space increases accordingly;

[0153] Then, construct the action space of deep reinforcement learning. The action space includes all container tasks and a waiting action. The dimension of the action space is N + 1, where N is the number of container tasks;

[0154] Next, in order to reflect the game cooperation relationship among multiple agents in deep reinforcement learning, the waiting time of the AGV during the working process is taken as the optimization goal. The longer the waiting time of the AGV, the less reward it gets;

[0155] When the AGV performs an operation, the reward basis is calculated as 20 - t waiting / 10, where t waiting is the waiting duration of the AGV in one task;

[0156] When the AGV performs a waiting action, a penalty of -1.5 is given;

[0157] When the action selected by the AGV conforms to the container - turning rule, the reward is +3 on the original basis;

[0158] When the action selected by the AGV does not conform to the container - turning rule, the reward is -1 on the original basis;

[0159] Among them, the reward function of the optimization goal is expressed as:

[0160]

[0161] Finally, construct the parameter structure of the neural network; both the prediction network and the target network adopt a three - layer network structure. The first - layer network contains 128 neurons, the second - layer network contains 64 neurons. The networks are fully connected. The optimizer adopts the Adam optimizer, and the rectified linear unit ReLU is used as the activation function, as follows:

[0162]

[0163] Moreover, the discount factor γ is obtained by an empirical formula. Let L be the number of steps that the agent is expected to consider next, which is expressed as:

[0164]

[0165] Step 5: Train the C - DQN multi - agent deep reinforcement learning model;

[0166] First, initialize the initial state of the automated container terminal simulation model, input the initial state into the C - DQN multi - agent deep reinforcement learning algorithm model, and obtain the task selection of the current AGV after masking processing;

[0167] Secondly, input the task into the automated container terminal simulation model through the AGV action selector. The AGV executes this action, and the AGV module starts to work, interacting with the quay crane module, yard crane module, and AGV queuing manager;

[0168] Then, after the AGV finishes executing the current action, it interacts with the external algorithm through the AGV action selector, records the current state, executed task, obtained reward, and next state, stores them in a temporary tuple, and increments the tuple counter by one;

[0169] Judge whether the data in the tuple meets all the conditions for forming a sample. If it meets, store the sample in the experience replay buffer, delete the current sample from the tuple, and decrement the tuple counter by one; if it does not meet, execute the next step, and at the same time wait for other agents to finish their current actions, record their information in the tuple, and combine it with the current sample to form a sample;

[0170] Next, judge whether the experience replay buffer of the C-DQN multi-agent deep reinforcement learning algorithm is full; if the experience replay buffer is full, the C-DQN multi-agent deep reinforcement learning algorithm model starts training. Randomly select a small batch of samples from the experience replay buffer and input them into the prediction network and the target network respectively. Obtain Q(s i ,a i ; θ) and Q target Calculate the degree of difference Loss through the mean square loss function and backpropagate the gradient. Use the Adam optimizer to optimize the parameters of the prediction network; if the experience replay buffer is not full, continue to run the automated container terminal simulation model;

[0171] Then, judge whether the number of training times of the C-DQN multi-agent deep reinforcement learning algorithm reaches the set value; if it reaches the set value, copy the model parameters of the prediction network to the target network; if it does not reach the set value, continue to run the automated container terminal simulation model;

[0172] Then, judge whether there are still unfinished container tasks; if there are, input the current state into the C-DQN multi-agent deep reinforcement learning algorithm model, obtain the next container task after masking, and hand it over to the AGV by the AGV action selector for work; if not, all container tasks in this round are completed, and the task reporting module outputs the task scheduling information of the AGV in this round and the total reward obtained;

[0173] Finally, determine the number of training times of the C-DQN multi-agent deep reinforcement learning algorithm model; when a certain number of training times is reached, gradually increase the greed rate ε; when the number of training times reaches the set maximum value, stop training, output the final AGV scheduling scheme and reward value, save the model parameters of the prediction network and output them;

[0174] Step Six: Package the trained model into an online real-time scheduling system;

[0175] After obtaining the trained C-DQN multi-agent deep reinforcement learning algorithm model using Step Five, apply this model to the online real-time scheduling system of AGV;

[0176] When an AGV needs to make a decision during operation, input the current environmental state into the C-DQN multi-agent deep reinforcement learning algorithm model, and the algorithm model outputs the action that the AGV should execute;

[0177] When the AGV completes a task, update the environmental state information in real time, and the algorithm model re-outputs the next action, thereby realizing the intelligent dynamic scheduling of the AGV. The best scheduling scheme of the AGV in this example is shown in the following table:

[0178] Table 4 Best Scheduling Scheme of AGV

[0179]

[0180]

[0181] Finally, it can be analyzed from the AGV scheduling results that after a period of learning, the AGV has learned the best container task scheduling method and can coordinate with each other to achieve the goal of the shortest conflict waiting time for the AGV. The task selection strategy of each AGV tends to select a container that is close to the destination of the previous task and does not interfere with the work of other AGVs for operation. Among them, AGV1 waited for 60.00 seconds during the work process, AGV2 waited for 62.58 seconds, AGV3 waited for 62.13 seconds, and the three AGVs waited for 184.71 seconds in total. Finally, the task was completed in 1939.87 seconds.

[0182] According to the automated container terminal simulation environment, at the beginning of the task, the AGV has reached the position of the first container to be retrieved and is waiting for the quay crane or the yard crane to finish working. Therefore, there will inevitably be a waiting time for the AGV when performing the first container task. From the final scheduling plan, it can be seen that the first container task of the AGV always selects the yard crane with the least working time. And except for the inevitable waiting time of the first task, the total waiting time of the AGV is 4.71 seconds. Thus, it can be seen that this scheduling plan is already close to the ideal optimal scheduling plan, proving that the method of the present invention has good effects.

[0183] Embodiment 2:

[0184] As Figure 9-10 shown, to verify the effectiveness of the method of the present invention, the method of the present invention is experimentally compared with the AGV scheduling method based on the traditional DQN deep reinforcement learning algorithm. The AGV scheduling method based on the traditional DQN deep reinforcement learning algorithm also uses the automated container terminal simulation model described above and schedules the same container tasks. The experimental results are shown in the following table:

[0185] Table 5 Experimental Results

[0186]

[0187]

[0188] From the results, it can be analyzed that among the AGV scheduling methods based on the traditional DQN deep reinforcement learning algorithm, AGV1 waited for a total of 294.77 seconds during the work process, AGV2 waited for a total of 152.73 seconds, AGV3 waited for a total of 377.63 seconds, and the three AGVs waited for a total of 825.13 seconds. The task completion time was 2124.38 seconds.

[0189] Through comparison, it can be proved that compared with the AGV scheduling method based on the traditional DQN deep reinforcement learning algorithm, using the C-DQN multi-agent deep reinforcement learning algorithm to schedule the AGV can improve the working efficiency of the AGV by 77.6%. The method of the present invention has good effects.

[0190] In summary, the present invention can divide the terminal handling operations into multiple modules and model them separately according to different working contents of the automated container terminal, and has good scalability and versatility. Connect different modules according to the event logical relationship to obtain an automated container terminal simulation model. And a C-DQN multi-agent deep reinforcement learning algorithm is proposed and a network model is constructed, and the algorithm model is trained on the basis of the automated container terminal simulation model. Finally, the trained algorithm model is exported to realize the intelligent dynamic scheduling of the AGV in the automated container terminal.

[0191] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. The structures, devices, and operation methods not specifically described and explained in the present invention are implemented according to the conventional means in the art without special instructions and limitations.

Claims

1. An intelligent dynamic scheduling method for AGVs in an automated container terminal, characterized in that: The following steps are involved: Step 1: Conduct simulation modeling based on the different functions of the automated container terminal; When the automated container terminal in step 1 performs loading and unloading operations, the equipment involved in the work includes terminal front loading and unloading equipment, horizontal transportation equipment, and yard operation equipment; Among them, according to the workflow of different loading and unloading equipment in the automated container terminal, they can be modeled as follows: The quay crane module is responsible for simulating the loading and unloading of containers between ships and AGVs; The yard bridge module is responsible for simulating the loading and unloading of containers between the yard and the AGV; AGV module, responsible for simulating AGV's container transportation work; AGV queue manager, responsible for managing the working sequence of multiple AGVs at ship berths and yards; AGV action selector, responsible for communicating with external algorithms and issuing work instructions to AGV; The task reporting module is responsible for evaluating the AGV's decision and reporting it to the operator; In step 1, the container loading and unloading tasks of the automated container terminal are also modeled, and different container tasks are abstracted into a five-tuple: Task_information=['Container', 'Loading or unloading', 'Quay cranesnumber', 'Field cranes number', 'Container tonnage'] (1) Wherein, 'Container' indicates the number of the container task, 'Loading or unloading' indicates the loading and unloading type of the container task, 'Quay cranes number' indicates the quay crane number related to the container task, 'Fieldcranes number' indicates the field crane number related to the container task, and 'Container tonnage' indicates the weight level of the container; Step 2: Connect different functional modules through event logic relationships; In the step 2, the working relationship between the modules of the automated container terminal simulation model of step 1 is analyzed, the interaction and logical connection between the modules are found, and these logical relationships are described by events; the communication between the modules is realized by activating and capturing events, the working process of each module is simulated by a delay function, and the modules are run in parallel to simulate various activities of the terminal, thereby realizing the simulation operation of the automated container terminal; Step 3: Build the C-DQN multi-agent deep reinforcement learning algorithm framework; The step three comprises the following steps: First, we build the C-DQN multi-agent deep reinforcement learning algorithm framework. The C-DQN multi-agent deep reinforcement learning algorithm consists of a prediction network and a target network. Among them, the prediction network is used to output the value Q of each action in the current state of the intelligent agent, and the target network is used to output the maximum value Q of the action in the next state of the intelligent agent target ; Then, the action value Q output by the prediction network and the target network is processed with a mask to block out actions that do not conform to the rules; the mask is a string of 0-1 arrays, and the dimension size is the same as the dimension of the action output of the neural network model; in the mask, the mask position corresponding to the action that conforms to the rules is 1, and the mask position corresponding to the action that does not conform to the rules is 0. The output of the neural network is processed with a mask, and the action value at the mask position 1 remains unchanged, and the action value at the mask position 0 is replaced with a minimum value; Next, a greedy strategy with a fixed step size and variable greed rate is used to select the action of the AGV and generate a random number; When the random number is less than the greedy rate ε, the Softmax function is used to select the action with the largest output value in the prediction network; When the random number is greater than the greedy rate ε, AGV randomly selects an action that complies with the rules; After the C-DQN multi-agent deep reinforcement learning algorithm model is trained a certain number of times, the value of the greed rate ε is gradually increased, so that the probability of the neural network randomly selecting an action becomes smaller, and the model begins to output a certain action; Then, the information of the AGV completing the task is stored as a sample in the experience replay buffer, and the training is carried out in a centralized learning distributed execution mode; during the training, the state and environmental information of all agents are trained, and when outputting decisions, each agent makes a separate judgment and outputs a decision action that conforms to the current environmental state. The samples in the experience replay buffer are in the form of a mixture of multiple agent state information, as follows: (s i ,a i ,r i ,s′1,s′2,...,s′ n ) (2) Among them, i represents the serial number of the agent, s i represents the state of agent i before executing the action, a i represents the action performed by agent i, r i represents the reward obtained by agent i after performing the action, s′1, s′2, ..., s′ n They represent when agent i starts to perform action a i , the state of the remaining agents after they have completed their respective current actions; Finally, whenever the neural network model needs to be trained, a batch of samples are randomly selected from the experience replay buffer for learning. Based on the sample information, the value information of the current action output by the network is predicted to be: Q(s i , a i ;θ); Among them, θ represents the parameter information of the prediction network, the target network integrates the information of multiple agents, and the output target value is set as: Among them, γ represents the discount factor, θ′ represents the parameter information of the target network, and the mean square loss function is used to measure the degree of difference according to the action values ​​obtained by the two neural networks, so as to perform gradient back propagation to update the network parameters. The mean square loss function is as follows: loss=E[(Q target -Q(s i ,a i ;θ)) 2 ] (4) Where E represents the average value of a preset random variable; Step 4: Construct the Markov decision process; The step 4 comprises the following steps: First, construct the state space of C-DQN multi-agent deep reinforcement learning, which includes: Container task status information, 1 means not completed, 0 means completed; AGV task information, the container task selected by the AGV; AGV progress information, 0 means the AGV is in the waiting state, 0.5 means the AGV is in the process of picking up boxes, and 1 means the AGV is in the process of placing boxes; AGV time information: the time the AGV has been in the current task state; Loading and unloading equipment working information, 0 means the quay crane or yard crane is not working, 1 means the quay crane or yard crane is working; Time information of loading and unloading equipment, the time that the quay crane or yard crane has been in the current working state; Tonnage information, the weight of the last container placed by the quay crane or yard crane; Among them, when the loading and unloading equipment of the automated container terminal increases, the dimension of the state space increases accordingly; Then, the action space of deep reinforcement learning is constructed. The action space includes all container tasks and a waiting action. The dimension of the action space is N+1, where N is the number of container tasks. Next, in order to reflect the game cooperation relationship between multiple agents in deep reinforcement learning, the waiting time of AGV in the working process is taken as the optimization target, as follows: When AGV is working, the calculation basis of reward is 20-t waiting / 10, where waiting The waiting time of AGV in a task; When the AGV is in a waiting action, a penalty of -1.5 is given; When the action selected by the AGV meets the box-turning rules, the reward is +3 on the original basis; When the action selected by the AGV does not meet the box-turning rules, the reward is reduced by 1. Among them, the reward function of the optimization objective is expressed as: Finally, the parameter structure of the neural network is constructed; both the prediction network and the target network adopt a three-layer network structure. The first layer of the network contains 128 neurons, and the second layer of the network contains 64 neurons. The networks are fully connected. The optimizer uses the Adam optimizer and the linear rectification function ReLU is used as the activation function, as follows: In addition, the discount factor γ is obtained using an empirical formula, where L is set to be the number of steps that the agent is expected to consider next, expressed as: Step 5: Train the C-DQN multi-agent deep reinforcement learning model; The step five includes the following steps: First, the initial state of the automated container terminal simulation model is initialized, and the initial state is input into the C-DQN multi-agent deep reinforcement learning algorithm model, and the current AGV task selection is obtained after mask processing; Secondly, the task is input into the automated container terminal simulation model through the AGV action selector. The AGV executes the action, and the AGV module starts working, interacting with the quay crane module, the yard crane module, and the AGV queue manager. Then, when the AGV completes the current action, it interacts with the external algorithm through the AGV action selector, records the current state, executed task, reward, and next state, stores them in a temporary tuple, and the tuple counter increases by one; The step five also includes the following steps: First, determine whether the data in the tuple meets all the conditions for forming a sample, store the sample in the experience playback buffer, delete the current sample from the tuple, and reduce the tuple counter by one; Or it is determined that the data in the tuple does not meet all the conditions for forming a sample, and the agent executes the next task, while waiting for other agents to complete their current actions, recording their information in the tuple, and combining them with the tuple to form a sample; Next, if the experience replay buffer of the C-DQN multi-agent deep reinforcement learning algorithm is full, the C-DQN multi-agent deep reinforcement learning algorithm model starts training, and a batch of samples are randomly selected from the experience replay buffer and input into the prediction network and the target network respectively. i , a i ;θ) and Q target The difference degree Loss is obtained through the mean square loss function and the gradient is back-propagated, and the parameters of the prediction network are optimized using the Adam optimizer; Or if it is determined that the experience replay buffer of the C-DQN multi-agent deep reinforcement learning algorithm is not filled, the automated container terminal simulation model continues to run; The step five also includes the following steps: First, determine whether the number of training times of the C-DQN multi-agent deep reinforcement learning algorithm reaches the set value, and then copy the model parameters of the prediction network to the target network; Or if it is determined that the number of training times of the C-DQN multi-agent deep reinforcement learning algorithm has not reached the set value, the automated container terminal simulation model continues to run; Next, if it is determined that there are unfinished container tasks, the current state is input into the C-DQN multi-agent deep reinforcement learning algorithm model, and the next container task is obtained after mask processing, and then handed over to the AGV by the AGV action selector; Or if it is determined that there is no unfinished container task, then all container tasks in this round are completed, and the task reporting module outputs the task scheduling information and total reward obtained for this round of AGV; Finally, determine the number of training times of the C-DQN multi-agent deep reinforcement learning algorithm model; when a certain number of training times is reached, start to gradually increase the greed rate ε; when the number of training times reaches the set maximum value, stop training, output the final AGV scheduling plan and reward value, save the model parameters of the prediction network and output; Step 6: Package the trained model into an online real-time scheduling system.

2. The method for intelligent dynamic scheduling of AGVs in an automated container terminal according to claim 1, characterized in that: The step six includes the following steps: After obtaining the trained C-DQN multi-agent deep reinforcement learning algorithm model in step 5, the model is applied to the online real-time scheduling system of AGV; When the AGV needs to make a decision during operation, the current environment state is input into the C-DQN multi-agent deep reinforcement learning algorithm model, and the algorithm model outputs the action that the AGV should perform; When the AGV completes the task, the environmental status information is updated in real time, and the algorithm model re-outputs the next action, thereby realizing the intelligent dynamic scheduling of the AGV.

Citation Information

Patent Citations

  • DDQN-based automatic container terminal AGV scheduling method

    CN114912809A

Cited By

  • Multi-agent-based OHT dynamic scheduling and traffic control optimization method and system

    CN121921977A