A Deep Reinforcement Learning Method for Conflict-Free Path Planning of Multiple AGVs in a Warehousing Environment

Through deep reinforcement learning and dynamic grouping mechanisms, a multi-AGV path planning method is built, which solves the problems of low path planning efficiency and dimensional disasters in multi-AGV systems, and achieves efficient conflict-free path planning.

CN119879967BActive Publication Date: 2025-07-08BEIJING WUZI UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411893075.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-07-08
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

The prior art is difficult to effectively carry out multi-AGV path planning in a complex environment with dynamic changes, and traditional methods cannot fully utilize the system dynamic data, resulting in poor path planning effect, and the single agent reinforcement learning method fails to effectively consider the interaction between agents, which is prone to network dimension explosion problems.

Method used

Deep reinforcement learning method is adopted, combined with dynamic grouping mechanism, and multiple AGVs are dynamically grouped and communicated through evaluation networks, attention networks and communication networks, and distributed conflict-free path planners are built, and multi-agent asynchronous advantage actor critic algorithms are used for training to optimize path planning.

Benefits of technology

The dimensional disaster problem in multi-AGV conflict-free path planning is solved, the efficiency and accuracy of path planning is improved, and efficient collaborative decision-making of multi-AGV systems is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119879967B_ABST
    Figure CN119879967B_ABST
Patent Text Reader

Abstract

The present invention provides a method for conflict-free path planning of multiple AGVs in a warehouse environment based on deep reinforcement learning, including: creating a virtual warehouse simulation training environment, and extracting virtual warehouse simulation training environment information from the virtual warehouse simulation training environment; determining the state observation value encoding of each AGV according to the virtual warehouse simulation training environment information; setting the action space of the AGV and constructing a reward function for deep reinforcement learning; constructing a distributed conflict-free multi-AGV path planner incorporating a dynamic grouping mechanism, and through a deep reinforcement learning network, iteratively training the distributed conflict-free multi-AGV path planner according to the reward function, the state observation value encoding and the action space of each AGV, updating the network parameters of the distributed conflict-free multi-AGV path planner to obtain a trained distributed conflict-free multi-AGV path planner; using the trained distributed conflict-free multi-AGV path planner to perform conflict-free path planning for multiple AGVs in the actual warehouse environment to obtain a conflict-free path for each AGV.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of path planning for logistics robots, and particularly to a method for conflict-free path planning of multiple AGVs in a warehousing environment based on deep reinforcement learning. Background Art

[0002] In recent years, new technologies and new business models have emerged on a large scale, and the traditional logistics industry is evolving and upgrading from manual operation to automation. With the rapid development of logistics robots and Internet of Things technologies, more and more enterprises have started strategic layout in the field of intelligent logistics. In order to improve work efficiency and reduce labor intensity, e-commerce enterprises have begun to use logistics robots (such as AGVs, Automated Guided Vehicles) to achieve automated picking, and multi-robot path planning is a key link in the scheduling of automated picking systems.

[0003] In the process of implementing the present invention, the applicant found that there are at least the following problems in the prior art: Most of the traditional research on robot path planning targets static environments or simple situations where the system environment is known, constructs models and methods based on the idea of operations research optimization, and it is difficult to meet the ever-changing complex environments in the current research background, and it is impossible to make full use of the dynamic data of the system, resulting in poor actual application effects of the models and methods. The research on the path planning problem of logistics robots based on reinforcement learning that has been carried out mostly uses the method of single-agent reinforcement learning, emphasizes the interaction and feedback between a single agent and the environment, and rarely considers the impact of the interaction between agents on the results. Moreover, the current methods mostly use centralized training, and as the scale of agents increases, the problem of network dimensionality explosion often occurs during the training process. Summary of the Invention

[0004] An embodiment of the present invention provides a method for conflict-free path planning of multiple AGVs (Automated Guided Vehicles) in a warehousing environment based on deep reinforcement learning to solve the problem of dimensionality disaster existing in the joint decision-making of multiple AGVs and achieve conflict-free path planning of multiple AGVs.

[0005] To achieve the above object, an embodiment of the present invention provides a method for conflict-free path planning of multiple AGVs in a warehousing environment based on deep reinforcement learning, including:

[0006] Create a virtual warehousing simulation training environment, and extract virtual warehousing simulation training environment information from the virtual warehousing simulation training environment;

[0007] According to the virtual warehousing simulation training environment information, determine the state observation value encoding of each AGV among multiple AGVs in the virtual warehousing simulation training environment;

[0008] Set the action space of the AGV, and construct a reward function for deep reinforcement learning;

[0009] Construct a distributed conflict-free multi-AGV path planner with a dynamic grouping mechanism. Encode and action space according to the reward function and the state observation values of each AGV, and iteratively train the distributed conflict-free multi-AGV path planner to update the network parameters of the distributed conflict-free multi-AGV path planner, obtaining a trained distributed conflict-free multi-AGV path planner;

[0010] Use the trained distributed conflict-free multi-AGV path planner to perform conflict-free path planning for multiple AGVs in the actual warehousing environment, obtaining conflict-free paths for each AGV;

[0011] Among them, the virtual warehousing simulation training environment information includes the initial positions, obstacle positions, shelf positions, and road information of multiple AGVs in the virtual warehousing simulation training environment; the state observation value encoding of the AGV includes: the position information and movement direction of the AGV, the target shelf position, whether it is carrying goods, and the position information and movement direction of obstacles and other AGVs within the sensing radius; the action space includes: turning left, turning right, moving forward, loading, and unloading; the sensing radius is used to define the observation range of the AGV;

[0012] The distributed conflict-free multi-AGV path planner includes a global network model and a local network model for each AGV; the global network model and the local network model for each AGV have the same network structure and both correspondingly include an evaluation network, an attention network, a communication network, and a deep reinforcement learning network;

[0013] The evaluation network is used to fuse the features of other AGVs within the observation range of each AGV into the features of the AGV for each AGV, obtaining a feature vector of the AGV, and determining the adjacency matrix of the AGV according to the feature vectors of all AGVs within the observation range of the AGV;

[0014] The attention network is used to calculate the communication group of the AGV centered on the AGV for each AGV based on the self-attention mechanism according to the adjacency matrix of the AGV and the state observation value encoding of the AGV included in the adjacency matrix of the AGV;

[0015] The communication network is used to perform network communication within the communication group centered on each AGV and update the state observation value encoding of each AGV within the communication group;

[0016] The deep reinforcement learning network is constructed based on the multi-agent asynchronous advantage actor-critic algorithm. For the AGVs after the communication groups are divided by the dynamic grouping mechanism, action decisions are made in the policy network according to the state observation values encoded by each AGV, path planning is completed, and rewards are obtained according to the reward function. With the goal of maximizing the rewards, the model parameters of each AGV's local network model are continuously iteratively updated, and the model parameters of the global network model are updated using the model parameters of each AGV's local network model, where the deep reinforcement learning network includes a policy network and a critic network.

[0017] The above technical solution has the following beneficial effects: The dynamic grouping mechanism is introduced. Multiple AGVs are dynamically grouped through the evaluation network and the attention network, and communication between AGVs is carried out through the communication network. Combining deep reinforcement learning to train the network parameters solves the curse of dimensionality problem that occurs in the multi-AGV joint decision-making during the conflict-free path planning process of multiple AGVs, and improves the efficiency of the conflict-free path planning of multiple AGVs. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a flowchart of a method for conflict-free path planning of multiple AGVs in a warehousing environment based on deep reinforcement learning according to one of the embodiments of the present invention;

[0020] Figure 2 It is an algorithm flowchart of an evaluation network of a method for conflict-free path planning of multiple AGVs in a warehousing environment based on deep reinforcement learning according to one of the embodiments of the present invention;

[0021] Figure 3 It is a schematic diagram of the orientation area of an AGV according to one of the embodiments of the present invention;

[0022] Figure 4 It is a schematic structural diagram of an attention network of a method for conflict-free path planning of multiple AGVs in a warehousing environment based on deep reinforcement learning according to one of the embodiments of the present invention;

[0023] Figure 5 It is a schematic structural diagram of a communication network of a method for conflict-free path planning of multiple AGVs in a warehousing environment based on deep reinforcement learning according to one of the embodiments of the present invention;

[0024] Figure 6Schematic diagram of the deep reinforcement learning network of a multi-AGV conflict-free path planning method for a warehousing environment according to one embodiment of the present invention;

[0025] Figure 7 Flowchart of training the deep reinforcement learning network of a multi-AGV conflict-free path planning method for a warehousing environment according to one embodiment of the present invention;

[0026] Figure 8 Schematic diagram of the network structure of a distributed conflict-free multi-AGV path planner of a multi-AGV conflict-free path planning method for a warehousing environment according to one embodiment of the present invention;

[0027] Figure 9 Reward iteration graph generated during training in the simulation experiment of the multi-AGV conflict-free path planning task according to one embodiment of the present invention;

[0028] Figure 10 Schematic example diagram of the virtual warehousing simulation training environment according to one embodiment of the present invention;

[0029] Figure 11 Schematic diagram of the virtual warehousing simulation training environment of Warehouse 1 used in the simulation experiment according to one embodiment of the present invention;

[0030] Figure 12 Schematic diagram of the virtual warehousing simulation training environment of Warehouse 2 used in the simulation experiment according to one embodiment of the present invention;

[0031] Figure 13 Schematic diagram of the virtual warehousing simulation training environment of Warehouse 3 used in the simulation experiment according to one embodiment of the present invention;

[0032] Figure 14 Schematic diagram of quantization coding of the virtual warehousing simulation training environment according to one embodiment of the present invention. Detailed implementation manners

[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0034] On the one hand, as Figure 1 shown, an embodiment of the present invention provides a multi-AGV conflict-free path planning method for a warehousing environment based on deep reinforcement learning, including:

[0035] Step S10, create a virtual warehousing simulation training environment, and extract virtual warehousing simulation training environment information from the virtual warehousing simulation training environment;

[0036] Step S11, according to the virtual warehousing simulation training environment information, determine the state observation value encoding of each Automated Guided Vehicle (AGV) among the multiple AGVs in the virtual warehousing simulation training environment;

[0037] Step S12, set the action space of the AGV, and construct a reward function for deep reinforcement learning;

[0038] Step S13, construct a distributed conflict-free multi-AGV path planner incorporating a dynamic grouping mechanism, and perform iterative training on the distributed conflict-free multi-AGV path planner according to the reward function, the state observation value encoding of each AGV, and the action space, update the network parameters of the distributed conflict-free multi-AGV path planner, and obtain a trained distributed conflict-free multi-AGV path planner;

[0039] Step S14, use the trained distributed conflict-free multi-AGV path planner to perform conflict-free path planning for multiple AGVs in the actual warehousing environment, and obtain a conflict-free path for each AGV;

[0040] Among them, the virtual warehousing simulation training environment information includes the initial positions, obstacle positions, shelf positions, and road information of multiple AGVs in the virtual warehousing simulation training environment; the state observation value encoding of the AGV includes: the position information and movement direction of the AGV, the target shelf position, whether it is carrying goods, and the position information and movement direction of obstacles and other AGVs within the sensing radius; the action space includes: turning left, turning right, moving forward, loading, and unloading; the sensing radius is used to define the observation range of the AGV;

[0041] The distributed conflict-free multi-AGV path planner includes a global network model and a local network model for each AGV; the global network model and the local network model for each AGV have the same network structure and both correspondingly include an evaluation network, an attention network, a communication network, and a deep reinforcement learning network;

[0042] The evaluation network is used to, for each AGV, fuse the features of other AGVs within the observation range of the AGV into the features of the AGV to obtain a feature vector of the AGV, and determine the adjacency matrix of the AGV according to the feature vectors of all AGVs within the observation range of the AGV;

[0043] The attention network is used to calculate, for each AGV, based on the self-attention mechanism, the communication group centered on the AGV according to the adjacency matrix of the AGV and the state observation values of the AGVs included in the adjacency matrix of the AGV for encoding.

[0044] The communication network is used to perform network communication within the communication group centered on the AGV for each AGV, and update the state observation value encoding of each AGV within the communication group.

[0045] The deep reinforcement learning network is constructed based on the multi-agent asynchronous advantage actor-critic algorithm, and is used to make action decisions in the policy network for the AGVs after the communication groups are divided by the dynamic grouping mechanism according to the state observation value encoding of each AGV, complete path planning, and obtain rewards according to the reward function, with the goal of maximizing the rewards, continuously iteratively updating the model parameters of each AGV's local network model, and using the model parameters of each AGV's local network model to update the model parameters of the global network model, where the deep reinforcement learning network includes a policy network and a critic network.

[0046] The embodiments of the present invention have the following technical effects: introducing a dynamic grouping mechanism, dynamically grouping multiple AGVs through an evaluation network and an attention network, and performing communication between AGVs through a communication network, combined with deep reinforcement learning to train network parameters, solving the dimensionality disaster problem that occurs in the multi-AGV joint decision-making during the conflict-free path planning process of multiple AGVs, and improving the efficiency of the conflict-free path planning of multiple AGVs.

[0047] Further, a reward function for deep reinforcement learning is constructed, including:

[0048] The reward function is represented by the following formula (1):

[0049] R(s,a) = R1 + R2 + R3 (1)

[0050] where R1 represents the reward value obtained according to the distance between the AGV and the nearest obstacle, R2 represents the reward value obtained when the AGV reaches the target shelf position, and R3 represents the penalty value set for the time spent in the path planning process of the AGV, which is used to measure the cost paid by the AGV in the process of avoiding other agents and non-target shelves;

[0051] wherein, during the period without collision, R1 is determined by formula (2):

[0052] R1 = u × obs_distance (2)

[0053] and according to the following collision penalty correction formula (3), R1 is punished and corrected when a collision is identified:

[0054]

[0055] Among them, u represents the reward and punishment factor set for the distance between the AGV and the obstacle, which is a positive value; obs_distance represents the distance between the current AGV and the obstacle closest to it; the collision penalty correction formula means that when the distance obs_distance from the AGV to the obstacle is less than or equal to the preset minimum collision distance dis crash it is recognized that a collision occurs, and 100 is subtracted from R1 to punish and correct the value of R1;

[0056] The initial value of R2 is 0, and according to the following arrival reward formula (4), R2 is rewarded and corrected when it is recognized that the target shelf is reached:

[0057]

[0058] Among them, tar_distance represents the distance between the current AGV and the target shelf, and dis reach is the set arrival distance. When the distance tar_distance between the robot and the target shelf is less than the default arrival distance dis reach it means that the current AGV has completed path planning and reached the target shelf, and it is recognized that the target shelf is reached;

[0059] R2' is a continuous arrival reward function, specifically formula (5):

[0060]

[0061] Among them, ep_count represents the cumulative value of the number of consecutive target arrival rounds of the AGV. If the AGV has a conflict collision during the simulation process or reaches the maximum training time step but does not reach the target shelf in a round, then ep_count is cleared and accumulated again; k' represents the reward coefficient for the number of consecutive target arrival shelves ep_count of the AGV;

[0062] Determine R3 according to formula (6):

[0063] R3 = q × step_count (6)

[0064] Among them, q is the penalty coefficient based on the arrival time, and step_count is the cumulative step value.

[0065] Furthermore, the construction of the distributed conflict-free multi-AGV path planner introducing a dynamic grouping mechanism, encodes the reward function and the state observation values and action space of each AGV, and iteratively trains the distributed conflict-free multi-AGV path planner to update the network parameters of the distributed conflict-free multi-AGV path planner, obtaining the trained distributed conflict-free multi-AGV path planner, including:

[0066] Define a global iteration count and set the initial value of the global iteration count to 0;

[0067] Increment the global iteration count by one;

[0068] Determine whether the global iteration count is less than a preset global iteration count threshold;

[0069] If it is determined that the global iteration count is less than the preset global iteration count threshold, then perform the following steps:

[0070] For each AGV in the virtual warehouse simulation training environment, synchronize the model parameters of the local network model corresponding to the AGV using the model parameters of the global network model, and set the time step count corresponding to the AGV to 0;

[0071] According to the virtual warehouse simulation training environment information and the preset perception radius, determine the state observation value encoding of each AGV in the virtual warehouse simulation training environment;

[0072] For each AGV in the virtual warehouse simulation training environment, based on the evaluation network of the local network model of the AGV according to the state observation value encoding of the AGV, determine the feature vector of the AGV; the feature vector of the AGV characterizes the environmental information within the perception radius of the AGV;

[0073] Generate the adjacency matrix corresponding to the AGV according to the feature vector of the AGV; the adjacency matrix characterizes the adjacent relationship between all AGVs within the perception radius of the AGV;

[0074] Based on the attention network of the local network model of the AGV, obtain the similarity matrix corresponding to the AGV according to the adjacency matrix corresponding to the AGV and the state observation value encodings of all AGVs in the adjacency matrix; according to the similarity matrix corresponding to the AGV, determine other AGVs with similar movement paths to the AGV, and form a communication group centered on the AGV with the AGV and the other AGVs obtained with similar movement paths to the AGV; wherein, each row element in the similarity matrix body characterizes the corresponding relationship between the AGV corresponding to the row and other AGVs with similar running paths;

[0075] Encode the state observation values corresponding to all AGVs in the communication group centered on the AGV and input them into the communication network of the local network model corresponding to the AGV for prediction, to obtain the predicted state observation value encoding of the AGV;

[0076] Input the predicted state observation value encoding of the AGV into the policy network of the deep reinforcement learning network to determine the action to be executed by the AGV;

[0077] Execute the action to be executed by the AGV, so that the AGV interacts with the virtual warehousing simulation training environment, updates the position information in the state observation value encoding of the AGV, and updates the virtual warehousing simulation training environment information according to the updated position information of the AGV;

[0078] Based on the preset reward value calculation function, calculate the reward value corresponding to the current time step count of the AGV according to the state observation value encoding of the AGV before executing the action to be executed and the updated state observation value encoding of the AGV;

[0079] Increment the time step count corresponding to the AGV by one, and determine whether the position information in the updated state observation value encoding of the AGV is equal to the target shelf position in the state observation value encoding of the AGV, and whether the time step count corresponding to the AGV is greater than the preset local time step count threshold; if it is determined that the position information in the updated state observation value encoding of the AGV is not equal to the target shelf position in the state observation value encoding of the AGV, and the time step count corresponding to the AGV is less than the preset local time step count threshold, then return to the step of determining the state observation value encoding of each AGV in the virtual warehousing simulation training environment according to the virtual warehousing simulation training environment information and the preset sensing radius to execute;

[0080] Based on the critic network in the local network model of the AGV, determine the expected return of the AGV at the current time step count according to the position information and reward value of the AGV at the current time step count;

[0081] According to the expected return of the AGV at the current time step count, determine the expected return corresponding to each time step count before the current time step count, and perform gradient update on the model parameters in the local network model of the AGV according to the expected return corresponding to each time step count before the current time step count; the model parameters in the local network model of the AGV include the model parameters in the evaluation network, attention network, communication network and deep reinforcement learning network in the local network model;

[0082] Update the model parameters of the corresponding model in the global network model according to the model parameters of the local network model of the AGV after gradient update, and return to execute the step of incrementing the global iteration count by one; wherein, the model parameters of the global network model include: the model parameters in the evaluation network, attention network, communication network, and deep reinforcement learning network in the global network model;

[0083] If it is determined that the global iteration count is greater than or equal to the preset global iteration number threshold, save the model parameters of the corresponding model in the global network model.

[0084] Further, use the trained distributed conflict-free multi-AGV path planner to perform conflict-free path planning for multiple AGVs in the actual warehousing environment to obtain the conflict-free paths of each AGV, including:

[0085] Obtain the actual warehousing environment information, where the actual warehousing environment information includes the positions of multiple AGVs, obstacle positions, shelf positions, and road information in the actual warehousing environment;

[0086] Initialize the model parameters of the local network model of each AGV among the multiple AGVs in the actual warehousing environment using the model parameters of the global network model in the trained distributed conflict-free multi-AGV path planner;

[0087] At each time step at a preset time step interval, periodically construct a communication group centered on the AGV through the evaluation network and attention network of the local network model of each AGV in the actual warehousing environment, and update the state observation value encoding of each AGV in the communication group centered on the AGV through the communication network in the local network model of the AGV, input the updated state observation value encoding of the AGV into the policy network in the deep reinforcement learning network of the local network model of the AGV to obtain the to-be-executed action of the AGV corresponding to the time step, execute the to-be-executed action, and determine the position information of the AGV corresponding to the time step;

[0088] Arrange the AGV position information of each AGV in the actual warehousing environment at each time step in chronological order to obtain the conflict-free planned path of each AGV.

[0089] Further, determine the state observation value encoding of each AGV in the virtual warehousing simulation training environment according to the virtual warehousing simulation training environment information and the preset perception radius, including:

[0090] For each AGV in the virtual warehousing simulation training environment, based on the virtual warehousing simulation training environment information, determine the obstacles and other AGVs within the sensing radius centered on the AGV, and construct the state observation value encoding of the AGV according to the obstacle positions corresponding to the determined obstacles, the position information and movement directions of other AGVs.

[0091] Further, for each AGV in the virtual warehousing simulation training environment, based on the evaluation network of the local network model of the AGV and according to the state observation value encoding of the AGV, determine the feature vector of the AGV, including:

[0092] According to the state observation value encoding of the AGV, determine all AGVs within the neighborhood of the AGV, and vectorize the state observation value encodings of all AGVs within the AGV into corresponding initial feature values; the neighborhood of the AGV represents the range covered by the preset sensing radius of the AGV.

[0093] Save the initial feature values of all AGVs within the neighborhood of the AGV separately as the previous updated feature values.

[0094] Set the initial value of the iteration count.

[0095] Judge whether the iteration count is greater than the preset iteration count. If the iteration count is greater than the preset iteration count, end the iteration loop, and use the initial feature value of the AGV after the iteration loop ends as the feature vector of the AGV; otherwise, continue to execute the following iteration loop:

[0096] Calculate the average of the initial feature value of the AGV and the previous updated feature values corresponding to other AGVs within the neighborhood of the AGV, and then multiply it by the weight corresponding to the current iteration count to obtain an aggregated value. Nonlinearize the aggregated value through a preset nonlinear function to obtain the non - linearized aggregated value, and use the non - linearized aggregated value to update the initial feature value of the AGV.

[0097] Perform normalization processing on the initial feature value of the AGV, and use the normalized initial feature value of the AGV to update the initial feature value of the AGV.

[0098] For each other AGV within the neighborhood of the AGV, judge the position relationship of the other VGA relative to the AGV.

[0099] In the case where it is judged that the other AGV is in the forward area of the AGV, update the previous updated feature value corresponding to the other AGV with the value obtained by multiplying the initial feature value of the other AGV by the first coefficient.

[0100] When it is determined that the other AGV is in the backward area of the AGV, update the previous updated eigenvalue corresponding to the other AGV with the value obtained by multiplying the initial eigenvalue of the other AGV by a second coefficient;

[0101] Return to the step of determining whether the number of judgment iterations is greater than the preset number of iterations and execute;

[0102] Among them, determining the positional relationship between the other VGA and the AGV includes:

[0103] According to the state observation value coding of the AGV, obtain the position information of other AGVs within the sensing radius of the AGV, and divide the other AGVs within the sensing radius of the AGV into AGVs within the forward area of the AGV and AGVs within the backward area of the AGV; among them, the forward area of the AGV is defined as the first quadrant and the second quadrant of the coordinate system with the position information in the state observation value coding of the AGV as the origin; the backward area of the AGV is defined as the third quadrant and the fourth quadrant of the coordinate system with the position information in the state observation value coding of the AGV as the origin; the Y-axis of the coordinate system with the position information in the state observation value coding of the AGV as the origin is defined as the straight line pointing from the position information in the state observation value coding of the AGV to the target shelf position in the state observation value coding of the AGV, and the positive direction of the Y-axis faces the target shelf position in the state observation value coding of the AGV, and the X-axis of the coordinate system with the position information in the state observation value coding of the AGV as the origin is perpendicular to the Y-axis;

[0104] Among them, the first coefficient is greater than the second coefficient, both the first coefficient and the second coefficient are greater than 0 and the sum of the first coefficient and the second coefficient is 1; the first coefficient, the second coefficient and the weights corresponding to all the number of iterations are model parameters of the evaluation network.

[0105] Furthermore, generating an adjacency matrix corresponding to the AGV according to the eigenvector of the AGV includes:

[0106] Calculate the similarity values of the eigenvectors of all AGVs within the sensing radius of the AGV, and set the element positions corresponding to the two AGVs with the obtained similarity values greater than the preset adjacency similarity threshold to 1 in the adjacency matrix, otherwise set to 0.

[0107] Further, according to the adjacency matrix corresponding to the AGV and the state observation values encoding of all AGVs in the adjacency matrix, based on the attention network of the local network model of the AGV, a similarity matrix corresponding to the AGV is obtained. According to the similarity matrix corresponding to the AGV, other AGVs with similar movement paths to the AGV are determined, and the AGV and the other AGVs with similar movement paths obtained together form a communication group centered on the AGV, including:

[0108] Vectorize the state observation values encoding of all AGVs in the adjacency matrix corresponding to the AGV into corresponding encoding vectors, and arrange all the obtained encoding vectors in the order of AGVs in the adjacency matrix to obtain an observation value encoding matrix;

[0109] Use the observation value encoding matrix to initialize the query matrix, key matrix, and value matrix in the attention network of the local network model of the AGV;

[0110] Obtain the similarity matrix at the current time step according to the following formula (7)

[0111]

[0112] where M is the adjacency matrix at the current time step; Q is the query matrix, K is the key matrix, d k is the dimension of the key matrix, and T represents the transpose matrix of the current matrix;

[0113] Calculate the normalized similarity matrix according to the following formula (8):

[0114]

[0115] where the softmax function is the normalization function; represents the interaction relationship between AGVs before normalization, represents the sum of the exponential operations on the relationship between the i-th AGV and other k-th AGVs;

[0116] Partition the other AGVs corresponding to the columns or rows where the values of the elements corresponding to the AGV in the row or column of Attention(Q, K, V) are 1 into the communication group centered on the AGV.

[0117] Further, at preset time step intervals, periodically at each time step, construct a communication group centered on the AGV through the evaluation network and the attention network of the local network model of each AGV in the actual warehousing environment, and update the state observation value encoding of each AGV in the communication group centered on the AGV through the communication network in the local network model of the AGV. Input the updated state observation value encoding of the AGV into the policy network in the deep reinforcement learning network of the local network model of the AGV to obtain the to-be-executed action of the AGV corresponding to the time step, and execute the to-be-executed action to determine the position information of the AGV corresponding to the time step, including:

[0118] For each time step, determine the state observation value encoding of each AGV in the actual warehousing environment according to the actual warehousing environment information and the preset perception radius; the actual warehousing environment information includes: the position information of multiple AGVs, the obstacle positions, the shelf positions, and the road information;

[0119] For each AGV in the actual warehousing environment, determine the feature vector of the AGV based on the evaluation network in the local network model of the AGV according to the state observation value encoding of the AGV;

[0120] Generate the adjacency matrix corresponding to the AGV according to the feature vector of the AGV;

[0121] Based on the attention network in the local network model of the AGV, obtain the similarity matrix corresponding to the AGV according to the adjacency matrix corresponding to the AGV and the state observation value encodings of all AGVs in the adjacency matrix;

[0122] According to the similarity matrix corresponding to the AGV, determine other AGVs with similar movement paths to the AGV, and form a communication group centered on the AGV by combining the AGV and the other AGVs with similar movement paths obtained;

[0123] Input the state observation value encodings of all AGVs in the communication group centered on the AGV into the communication network in the local network model of the AGV for prediction to obtain the predicted state observation value encoding of the AGV;

[0124] Input the predicted state observation value encoding of the AGV into the policy network of the deep reinforcement learning network of the local network model of the AGV to determine the to-be-executed action of the AGV;

[0125] Execute the to-be-executed actions of the AGV to enable the AGV to interact with the actual warehousing environment, update the position information in the state observation value encoding of the AGV, and update the actual warehousing environment information according to the updated position information in the state observation value encoding of the AGV;

[0126] Determine whether the position information in the state observation value encoding of the AGV is equal to the target shelf position in the state observation value encoding of the AGV;

[0127] If it is determined that the position information in the state observation value encoding of the AGV is not equal to the target shelf position in the state observation value encoding of the AGV, then return to the step of determining the state observation value encoding of each AGV in the actual warehousing environment according to the actual warehousing environment information and the preset sensing radius for each time step.

[0128] The above technical solutions of the embodiments of the present invention will be described in detail below in conjunction with specific application examples. For technical details not introduced during the implementation process, reference can be made to the relevant descriptions above.

[0129] The embodiments of the present invention provide a multi-AGV conflict-free path planning method based on deep reinforcement learning and dynamic grouping, which simulates the actual warehousing picking scenario, constructs a deep reinforcement learning multi-AGV conflict-free path planner introducing a dynamic grouping mechanism, and solves the dimensionality disaster problem of multi-AGV joint decision-making through the dynamic grouping mechanism. It includes:

[0130] Select all orders in a certain period that have been processed in the warehouse. After task allocation, through sensing technology, the system monitors the physical environment in the warehouse in real time, including the current position of the AGV, the distribution of goods, obstacles, etc. The system needs to collect and manage various task requirements, including task type, task urgency, starting point and target position of the task, etc. Build a multi-AGV path planning model, considering multiple factors such as the number of AGVs, speed, task priority, physical environment, etc., for path planning. Path planning includes path generation and real-time path adjustment. Path generation is to generate the best path for each AGV, ensure the minimum path length, avoid conflicts between AGVs, so that tasks can be executed efficiently; real-time path adjustment is to monitor the position and environmental changes of the AGV during task execution, and if obstacles or emergencies occur, the path can be adjusted in time to respond.

[0131] Furthermore, the method further includes: using a multi-AGV conflict-free path planning model that introduces a dynamic grouping mechanism for deep reinforcement learning to solve the problem. The overall network architecture of the model includes input, dynamic grouping, collaborative communication, and output. The input information of the model includes environmental information, the real-time position information of the AGVs, and the sensing radius of the AGVs. A dynamic grouping mechanism is introduced to complete the information interaction between agents. To achieve the dynamic grouping of AGVs, the model takes each AGV as the central agent and dynamically determines the communication group centered on this AGV. First, through the evaluation network, the label of the central agent is extracted from the features of the agents existing within the observation range of the central agent, and at the same time, an adjacency matrix is constructed. Suppose there are N associated agent nodes in the system, and each agent node has its own features. The feature vectors of the agent nodes can form an N×D-dimensional matrix X. Through an N×N-dimensional adjacency matrix, the association relationship of the agent nodes can be clearly shown. The model uses a self-attention unit to extract the feature vectors of the agents to calculate the similarity of the observed value encoding between the agents, and uses the adjacency matrix to filter out the agents lacking mutual association. Finally, the attention weight values between the agents are calculated through the softmax function, so as to obtain the final AGV grouping information.

[0132] To achieve the collaborative communication of AGVs, the model designs a communication network and a deep reinforcement learning network based on a centralized training and decentralized execution architecture. The central agent selects the agents within the group to communicate in the communication network and outputs the shared communication information, so that the agents within the group can obtain more comprehensive observation information. The agents within the group understand and infer the behaviors of other AGVs within the group based on the comprehensive observation information, cooperate with other AGVs to make decisions and help each other to reach the target shelf, output the action sequence of the AGVs, obtain rewards, and maximize the cumulative sum of the rewards of all AGVs in the system.

[0133] As Figure 8 shown, a model network structure including an evaluation network, an attention network, a communication network, and a deep reinforcement learning network is designed, specifically including:

[0134] The evaluation network mainly completes the feature extraction of the central agent and its neighboring agents, predicts the label of the central agent through the feature vectors of the agents within the observation range of the central agent, and constructs an adjacency matrix at the same time. In the cooperative planning problem of multi-agents, assume that the task scenario is a two-dimensional environment (W×H) with obstacles. Assume that V={v1, v2,..., v N} is the set of N agents in the environment, and N(v i ) is the direct neighborhood of the agent {v i , i = 1, 2,..., N}. The sensing radius of each agent is r ob , and the observation range matrix of the agent on the map at time t is Among them, W ob and H ob respectively represent the width and height of the agent's observation range. The communication radius of the agent is r com , and the communication network is C(V, ε t ), where is a V×V matrix. If (v i , v j ) ∈ ε t , it means that v i and v j can communicate with each other. At the same time, if two agents can communicate with each other at time t, it means that ||P i - P j || ≤ r com . Among them, P i , P j ∈R 2 , and P i , P j represent the position vectors of agents v i , v j . The goal of the model is to effectively navigate the agents to their respective target positions without conflicts. The target path is connected by the action sequence of each agent at each moment. The evaluation network needs to create a mapping function F to convert the observation range matrix of the agent into a decision , that is C t represents the communication network C at time t. To effectively aggregate agents in a two-dimensional space, the model uses a mean aggregation function for information aggregation and outputs the feature vector of the agent (AGV).

[0135] As Figure 2 shown, the main idea of the evaluation network is that in each iteration, the agent continuously aggregates the state information of the agents within its observation range. During the continuous iteration process, the agent can continuously accumulate more and more feature vectors. The forward propagation algorithm can be divided into three steps. First, randomly sample the agents adjacent to the central agent; second, summarize the environmental information of the sampled agents from the outside to the inside in turn. Finally, use the aggregated information as the input of the fully connected layer to obtain the label of the central agent. The model obtains the weighted cognition of the agent to the environment and the similarity between the observation value encodings by referring to the self-attention mechanism.

[0136] As Figure 3As shown, the environmental information of the agents in the system changes dynamically in each training cycle. When implementing dynamic grouping in the model, it is necessary to introduce the observable features of the target agent itself, that is, the specific positional relationship of the agent to screen the neighboring agents of the target agent. First, the evaluation network defines the "forward area" and the "backward area" according to the movement direction of the agent. Secondly, a rectangular coordinate system is constructed based on the above areas, with the central agent as the origin, and the line connecting the central agent and the direction of its target position point as the vertical axis. During the movement of the agent, the direction facing the target location is defined as the positive direction of the vertical axis of the coordinate axis. In this coordinate system, the first and second quadrants are the forward areas, and the third and fourth quadrants are the backward areas. The pseudocode of the evaluation network algorithm in the model is shown in Table 1. The feature vector of the central agent is determined through the following steps:

[0137] According to the state observation value encoding of the AGV (equivalent to the central agent here), determine all AGVs (equivalent to all agents) within the domain of the AGV (equivalent to the central agent here), and vectorize the state observation value encodings of all AGVs within the AGV (equivalent to the central agent here) into corresponding initial feature values; the neighborhood of the AGV (equivalent to the central agent here) represents the range covered by the preset perception radius of the AGV (equivalent to the central agent here);

[0138] Save the initial feature values of all AGVs (equivalent to all agents) within the neighborhood of the AGV (equivalent to the central agent here) separately as the previous update feature values;

[0139] Set the initial value of the iteration count; it can be set to 0 or 1, which is used to control the number of iteration loops; the preset number of iterations can be used as a hyperparameter and debugged during the actual training process;

[0140] Judge whether the iteration count is greater than the preset number of iterations. If the iteration count is greater than the preset number of iterations, end the iteration loop, and use the initial feature value of the AGV (equivalent to the central agent here) after the iteration loop ends as the feature vector of the AGV (equivalent to the central agent here). Otherwise, continue to execute the following iteration loop:

[0141] Take the mean of the initial feature value of the AGV (equivalent to the central agent here) and the previous update feature values corresponding to other AGVs (equivalent to other agents) within the neighborhood of the AGV (equivalent to the central agent here), and then multiply it by the weight corresponding to the current iteration count to obtain an aggregated value. Nonlinearize the aggregated value through a preset nonlinear function to obtain a nonlinearized aggregated value, and use the nonlinearized aggregated value to update the initial feature value of the AGV (equivalent to the central agent here);

[0142] Normalize the initial eigenvalue of the AGV (equivalent to the central agent here), and use the normalized initial eigenvalue of the AGV (equivalent to the central agent here) to update the initial eigenvalue of the AGV (equivalent to the central agent here);

[0143] For each other AGV (equivalent to other agents) within the domain of the AGV (equivalent to the central agent here), determine the positional relationship of the other VGA (equivalent to other agents) relative to the AGV (equivalent to the central agent here);

[0144] When it is determined that the other AGV (equivalent to other agents) is in the forward region of the AGV (equivalent to the central agent here), use the value obtained by multiplying the initial eigenvalue of the other AGV (equivalent to other agents) by the first coefficient to update the previous updated eigenvalue corresponding to the other AGV (equivalent to other agents);

[0145] When it is determined that the other AGV (equivalent to other agents) is in the backward region of the AGV (equivalent to the central agent here), use the value obtained by multiplying the initial eigenvalue of the other AGV (other agents) by the second coefficient to update the previous updated eigenvalue corresponding to the other AGV (equivalent to other agents);

[0146] Return to the step of determining whether the number of judgment iterations is greater than the preset number of iterations and execute;

[0147] Among them, determining the positional relationship of the other VGA (equivalent to other agents) relative to the AGV (equivalent to the central agent here) includes:

[0148] Encode according to the state observation value of the AGV to obtain the position information of other AGVs within the sensing radius of the AGV, and divide the other AGVs within the sensing radius of the AGV into the AGVs within the forward region of the AGV and the AGVs within the backward region of the AGV; wherein, the forward region of the AGV is defined as the first quadrant and the second quadrant of the coordinate system with the position information in the state observation value encoding of the AGV as the origin; the backward region of the AGV is defined as the third quadrant and the fourth quadrant of the coordinate system with the position information in the state observation value encoding of the AGV as the origin; the Y-axis of the coordinate system with the position information in the state observation value encoding of the AGV as the origin is defined as the straight line pointing from the position information in the state observation value encoding of the AGV to the target shelf position in the state observation value encoding of the AGV, and the positive direction of the Y-axis faces the target shelf position in the state observation value encoding of the AGV, and the X-axis of the coordinate system with the position information in the state observation value encoding of the AGV as the origin is perpendicular to the Y-axis; wherein, the first coefficient is greater than the second coefficient, both the first coefficient and the second coefficient are greater than 0 and the sum of the first coefficient and the second coefficient is 1; the first coefficient, the second coefficient and the weights corresponding to all the iteration times are the model parameters of the evaluation network. The non-linear function includes but is not limited to the ReLu function, the Sigmoid function or the Tanh function.

[0149]

[0150] Table 1 Evaluation Network Algorithm

[0151] The attention network module determines the similarities and differences between different agents by analyzing the agent feature information. It focuses on the key features of the agents, such as position, speed, status, etc., to identify patterns and trends related to the interaction with the environment, enabling the system to accurately perceive the status and behavior of the agents, so as to make appropriate preparations for subsequent cooperative communication. Through dynamic grouping, the attention network module can classify agents with similar features or behaviors into the same group to create a more organized agent network, making the cooperative communication more efficient. For example, dividing agents with similar tasks or goals into the same group can promote information exchange and cooperation between them. Figure 4 Shows the attention network structure. M is the adjacency matrix extracted by the evaluation network from the feature vectors of the agents (AGVs). The matrix M shows that AGV No. 1 can observe AGV No. 3. Similarly, AGV No. 3 can also observe AGV No. 1; AGV No. 1 cannot obtain the feature information of AGV No. 2.

[0152]

[0153] The model obtains the weighted cognition of the agent towards the environment and the similarity between the observation value encodings through the self-attention mechanism. Among them, the matrix X represents the feature vector transformed from the observation value encoding matrix of the agent. In the initial state, the query matrix Q, the key matrix K, and the value matrix V in the self-attention mechanism are the same as the observation value encoding matrix X. The similarity values between any two AGV observation value encoding matrices can be obtained.

[0154] The similarity matrix at the current time step is obtained according to formula (7), where M is the adjacency matrix at the current time step; Q is the query matrix, K is the key matrix, and d k is the dimension of the key matrix, and T represents the transpose matrix of the current matrix;

[0155] The normalized similarity matrix is calculated according to formula (8), where the softmax function is the normalization function; represents the interaction relationship between AGVs before normalization, represents the sum after exponentiating the relationship between the i-th AGV and other k-th AGVs;

[0156] The AGVs corresponding to the columns or rows where the elements with a value of 1 are located in the row or column corresponding to the AGV in Attention(Q, K, V) are divided into the communication group centered on the AGV.

[0157] Based on Attention(Q, K, V), when the element value at the i-th row and j-th column is equal to 0, the i-th AGV cannot obtain the feature vector of the j-th AGV, such as information like position and distance; when the element value at the i-th row and j-th column is not equal to 0, the i-th AGV can obtain the information of the j-th AGV.

[0158] The communication network module is a key tool for information exchange between agents. It can assist in establishing effective communication channels between agents, promoting cooperation and collaborative work. And as Figure 5 shown, the bidirectional LSTM unit embedded in it is a neural network unit that can capture the characteristics of sequence data. Through its bidirectional recursive structure, it can comprehensively utilize past and future information. It has the ability to selectively output information that promotes cooperation, thus helping agents make more intelligent decisions based on the understanding and prediction of the behaviors of other agents. Agents will select group members from the surrounding agents and establish communication groups. When multiple groups simultaneously select agent i, the model allows information to flow between different groups. Formulas (9) and (10) indicate that agent i first updates the information of this agent once in group P, then agent i joins communication group Q and updates its own parameters again. The finally updated parameter information will directly affect the parameter updates of other agents in group Q.

[0159]

[0160] In the above formula, C represents the communication network. The communication network uses a bidirectional LSTM network, and the bidirectional LSTM units can selectively output information that promotes cooperation, helping the agents make decisions based on understanding and predicting the behaviors of other agents.

[0161] In the deep reinforcement learning network, as Figure 6 shown, the global network is a grid model shared by AGVs in the warehousing environment, including two parts: the Actor network (policy network) and the Critic network (critic network). There are n independent sub-threads under the global network, each sub-thread corresponds to an AGV, and the network structure in each sub-thread is the same as that of the common neural network. During the training process, each sub-thread independently interacts with the environment to obtain experience values and does not affect each other.

[0162] Since the strategies of relevant agents in the multi-agent system will affect the final optimal strategy of the current agent, when the traditional method applies the reinforcement learning algorithm to solve problems, it exceeds the solving ability of the single-agent Markov storage process. The model combines the characteristics of the multi-AGV system, considers the dynamic environmental changes, and uses deep reinforcement learning with an introduced dynamic grouping mechanism to construct a multi-AGV path planning framework. In this architecture, the AGV can make decisions multiple times at different time periods. In the same time period, the AGV's spatial positions are used for dynamic grouping, and the AGVs within the group are expected to achieve a balance of relevance so that all AGVs can obtain an efficient and conflict-free path planning, thus greatly reducing the decision space and solving the problem of the dimensional explosion of path decisions in the multi-AGV system. In the deep reinforcement learning network, each AGV will establish a centralized discriminator Critic, which can obtain all the global states within the group and the action sets of all AGVs, and output the corresponding value function Q(x,a1,..,a n) to alleviate the problem of unstable internal environment in multi-agent systems. At the same time, the Actor network (policy network) of each AGV only makes decisions based on local observation information, realizing the distributed management of multi-agent systems. The centralized Q-value learning and distributed policy execution are used to solve problems. The Q-value can obtain the observation information o and action a of all AGVs in the group, and the policy π outputs the actions of AGVs according to the environmental observation information of each AGV. The sub-thread uses the interaction data with the environment to estimate the loss function of its own neural network, and uses the estimation result to update the parameters of the common neural network model, continuously improving the reliability and accuracy of the model training results. As the number of training times increases, the sub-thread continuously adjusts the parameters of its internal neural network to keep in sync with other neural networks. Through interaction with the external environment, the network model in the sub-thread can effectively help itself and other common network models to achieve optimization. During the training process, different agents learn and obtain parameterized experiences in the interaction with the environment, and each sub-thread copies the training parameters of the global network before working. After the AGV in the warehousing environment interacts with the environment, the algorithm calculates the gradient and transmits the gradient back to the central control center to further update the parameters of the global network. During the execution of the algorithm, each actor obtains parameters from the central control center for training and then transmits the parameters back to the central control network after training.

[0163] Furthermore, the multi-agent asynchronous advantage actor-critic algorithm is used to solve the multi-AGV conflict-free path planning model of deep reinforcement learning with a dynamic grouping mechanism, and a simulation experiment platform is developed to verify the effectiveness of the model and algorithm. Validity experiments are carried out for multi-AGV warehousing environments of different scales; comparative experiments are carried out in two dimensions: whether to dynamically group the AGVs in the simulation environment and various deep reinforcement learning solution algorithms.

[0164] 1.1 Algorithm Design: In the case of multi-agent games, the state space and action space of the environment often grow exponentially. When solving the equilibrium solution, if traditional value function-based training algorithms are used, the training time of the algorithm is often very long. To solve this problem, the embodiment of the present invention creates a policy-based approximate solution algorithm, the MAA3C algorithm, and each agent in the system can obtain its own optimal policy.

[0165] 1.1.1 Multi-Agent Asynchronous Advantage Actor-Critic Algorithm: The Actor-Critic Algorithm combines the temporal difference algorithm and the policy gradient algorithm to solve the path planning problem of multi-agent systems. The Multi-Agent Asynchronous Advantage Actor-Critic (MAA3C) algorithm extends the theoretical results of the traditional single-agent Actor-Critic algorithm to the process of multi-agent games. During the training process, different agents learn and obtain parameterized experiences by interacting with the environment. Each sub-thread copies the training parameters of the global network before working. After the AGV in the warehouse environment interacts with the environment, the algorithm calculates the gradient and transmits the gradient back to the central control center to further update the parameters of the global network. During the execution of the algorithm, each actor obtains parameters from the central control center for training and then transmits the parameters back to the central control network after training.

[0166] 1.1.2 Solution Process: The embodiment of the present invention uses the MAA3C algorithm to solve the deep reinforcement learning multi-AGV conflict-free path planning model with a dynamic grouping mechanism. The input of the MAA3C algorithm is the observation information of the agent, including environmental information and the feature information of other agents; the output of the algorithm is the optimal policy π * (s), that is, the action set corresponding to the optimal action value function Q * (s,a). During the training process of the algorithm, the AGV can interact with the environment to obtain changing and unknown environmental information in the surrounding environment. The algorithm solution process is as Figure 7 shown.

[0167] First, the system takes the current state of the AGV as input, and the Actor network outputs the parameters corresponding to the normal distribution. After sampling, the action a of the AGV in the current state is obtained. t . Second, the AGV executes the action a in step 1 t , and the environment feedbacks the state s at the next moment. t+1 . The Critic network gives an immediate reward r according to the action and state taken by the agent. t . Third, the Actor network updates the parameters according to the immediate reward, and at the same time combines s t+1 to output new parameters σ and μ, and samples to obtain the current action a. t+1 . Fourth, the AGV executes the action a t+1 , and continuously iterates in the manner of step 2 and step 3, repeating in a cycle until the conflict-free path planning is completed.

[0168] 1.2 Experimental Design and Result Analysis: The research conducts effectiveness experiments for multi-AGV application environments of different scales; comparative experiments are carried out in two dimensions: whether to dynamically group the AGVs in the simulation environment and various deep reinforcement learning solution algorithms.

[0169] 1.2.1 Experimental Design: Based on the above experimental operating conditions, the embodiments of the present invention developed a set of experimental simulation environment - multi-AGV warehouse simulation environment using Python code to simulate the warehouse environment. The experimental environment developed in the embodiments of the present invention simulates the real-world warehouse environment, providing a guarantee for the implementation of the model and algorithm. In the simulation environment, each AGV is regarded as a particle. A certain number of training rounds are set in the experiment. The AGV starts from the initial point, and in each round, after processing the data, it is input into the model, and during the continuous iteration of the algorithm, it conducts conflict-free path planning to reach the target shelf position. If the AGV conflicts with other AGVs during the training process or completes the path optimization and reaches the target shelf, the round ends. The code loops continuously according to such logic until the maximum value of the training rounds set by the system is reached. To prevent the problem of ineffective training caused by the algorithm not converging, a maximum time step threshold is set during the training process to avoid ineffective training time. In a multi-robot warehouse environment, the observation range of the robot contains all relevant information of the cells adjacent to the robot and is partially observable. By default, the observation range of each agent is the information within a 3×3 square centered on the agent, and the visual range of the robot can be modified as an environmental parameter. The agent can obtain its own position, rotation, and load, as well as the position information of other agents within the visual range and the shelves around the robot. There are four discrete actions available in the action space of the multi-robot warehouse environment: In this environment simulation, the robot has the following discrete action space: A = {Turn Left, Turn Right, Forward, Load}. The first three actions allow the robot to rotate and move forward. Loading and unloading are only effective when the robot is located below the target shelf. When the AGV does not carry any goods, it can freely shuttle under the shelf; but when the AGV carries the shelf, the AGV must travel through the visible channels in the figure. The AGV cannot move to an occupied cell when moving forward or rotating in the grid world. The AGV loads the shelf when it is located at the position of the target shelf. Figure 10 It is a schematic diagram of the warehouse simulation training environment, where the square P1 is the picking station, P2 is the shelf, the total number of AGV cars is 8, 1 (rectangular) AGV is not carrying a shelf, 7 (hexagonal) AGVs are carrying shelves, and the number of requested target shelves P3 is 3. The AGV needs to start from the picking station, complete conflict-free path planning to reach the target shelf, and obtain rewards. In this paper, the total cumulative return of all agents is selected as the performance index. The following is a specific description of this warehouse environment:

[0170] (1) Observation space: In a multi-robot warehouse environment, the observation range of a robot contains all relevant information of the cells adjacent to the robot and is partially observable. By default, the observation range of each agent is the information within a 3×3 square centered on the agent, and the visible range of the robot can be modified as an environmental parameter. The agent can obtain its own position, rotation, and load, as well as the position information of other agents within the visible range and the shelves around the robot. The following array details the encoded state observations of two AGVs:

[0171] [[8,3,0,0,1,0,0,0,0,1,0,0,0,1,0,0,1,0,0,0,1,1,0,1,0,0,0,0,0,0,1,0,0,0,1,0,1,0,1,0,0,1,0,0,1,0,0,0,0,0,1,1,0,0,0,1,0,0,1,0,0,0,1,0,0,1,0,0,0,0,0],

[0172] [4,9,0,0,0,1,0,1,0,1,0,0,0,0,0,0,1,0,0,0,0,0,0,1,0,0,0,0,0,0,1,0,0,0,0,0,1,0,0,1,0,0,0,0,1,0,0,0,0,0,0,1,0,0,0,0,0,0,1,0,0,0,0,0,0,1,0,0,0,0,0],...]

[0173] Each element in the array corresponds to the observation value of each agent. Among them, the first two values correspond to the x and y coordinates of the AGV itself. The third is encoded as "1" or "0", depending on whether the AGV is currently carrying a shelf. If it is carrying a shelf, it is 1; if not, it is 0. The next four values are the one-hot code encoding of the current direction the AGV is facing, including the four directions of up, down, left, and right. For example, (1,0,0,0) can represent up, (0,1,0,0) represents down, (0,0,1,0) represents left, and (0,0,0,1) represents right. Then, a single value "1" or "0" indicates whether the AGV is in a channel position. If it is a channel, it is not allowed to place a shelf at this position. The remaining values in the encoding can be divided into 9 groups of 7 elements, each group corresponding to a square in the observation radius (within the 3x3 observation range centered on the agent, which can be customized according to parameters). In each group of elements, the first element is 1 or 0, representing whether there is an AGV at this position. If there is an AGV, it then represents the one-hot code of the driving direction of the adjacent AGV. The sixth element indicates whether there is a shelf at this position. If there is a shelf, the seventh element indicates whether the shelf is the target shelf.

[0174] (2) Action Space: In the multi-robot warehouse environment, there are four discrete actions available in the action space. In this environment simulation, the robot has the following discrete action space: A = {Turn Left, Turn Right, Forward, Load}. The first three actions allow the robot to rotate and move forward. Loading and unloading are only valid when the robot is located under the target shelf. When the AGV is not carrying any goods, it can freely shuttle under the shelf. However, when the AGV is carrying a shelf, the AGV must travel through the visible passage in the figure. The AGV cannot move to an occupied cell when moving forward or rotating in the grid world. The AGV loads the shelf when it is located at the position of the target shelf.

[0175] (3) Reward function: If the AGV successfully reaches the target shelf, it will receive a reward. A major challenge in the environment is how the AGV completes path planning to reach the target shelf efficiently and without conflicts. The training objective of reinforcement learning is to maximize the cumulative reward. In reinforcement learning, the reward function can quantitatively evaluate the learning effect of the AGV. Therefore, the design of the reward function is an important factor affecting the effect of reinforcement learning and plays an extremely important role in the final training result. The reward function R(s,a) can be represented by formula (1); the R1 part of formula (1) describes the shortest distance between the AGV and the obstacle. In formula (2) of R1, obs_distance represents the distance between the current AGV and the nearest obstacle, and u is the reward and punishment factor set for the distance between the AGV and the obstacle. To enable the agent to successfully avoid obstacles, a safety distance needs to be defined. In the experiment, u is set to a positive value. From the formula, it can be seen that the greater the distance between the obstacle agent and the current AGV, the more rewards generated by the return function, and the better the effect. R2 represents the reward for the AGV to reach the target shelf position. In the simulation environment, a small arrival distance is set. If the distance between the AGV and the target shelf is less than the set value, it means that the AGV has completed collision-free path planning and reached the target shelf at this time. If the AGV can reach the target shelf continuously several times, it means that during the training process, the AGV is using the learned experience. The experiment can optimize the reward function to increase the reward level for the AGV so that the AGV can complete path planning more excellently. In the simulation experiment, in order to make the agent's path planning more efficient and minimize the cost as much as possible, an experiment sets a reward function to punish the number of steps the agent spends to avoid obstacles. The third part R3 of formula (1) represents a penalty function set for the time spent in the AGV's path planning process. R3 is defined as q×step_count, where q is the penalty coefficient based on the arrival time, and step_count is the cumulative step value, which is used to measure the cost paid by the AGV in the process of avoiding other agents and non-target shelves. Therefore, the reward function sets the penalty factor q as a negative number. The larger the cumulative time step step_count during the training process, the greater the penalty imposed by the reward function on the AGV. Based on whether the robot can successfully avoid other agents on non-target shelves, whether it can reach the target shelf, and whether it spends less time, etc., the reward function is further set so that the robot can finally achieve collision-free path planning with a stable path and less time overhead in a complex environment. Regarding the level of avoiding obstacles, if the distance between robots or between a robot and an obstacle is always within the safety distance range, the AGV will receive a reward. If the distance between robots or between a robot and a shelf is less than the safety distance, it is defaulted that a collision occurs, and a penalty will be received. According to the collision penalty correction formula (3), when a collision is identified, R1 is penalized and corrected, and dis is definedcrash is the safety distance. If the distance between two AGVs is less than the safety distance, it can be regarded as a conflict occurring between the AGVs. According to the setting criteria of the reward function, the agent is punished. Regarding the rewards and punishments for whether the AGV can complete conflict-free path planning during the training process, see formula (4). Among them, tar_distance represents the distance between the current AGV and the target shelf, and dis reach is the set arrival distance. When the distance tar_distance between the robot and the target shelf is less than the default arrival distance dis reach it means that the current AGV has completed path planning and reached the target shelf. In order to further enable the AGV to reach the target shelf position more quickly and continuously, a continuous arrival reward function R2' is adopted, and the specific logic is shown in formula (5). Among them, ep_count represents the cumulative value of the number of consecutive target rounds reached by the AGV. If the AGV has a conflict collision during the simulation process or reaches the maximum training time step but does not reach the target shelf in a round, ep_count is cleared and the accumulation starts again. k' in formula (5) represents the reward coefficient for the number of consecutive times ep_count that the AGV reaches the target shelf in the simulation environment.

[0176] In the simulation experiment, the network structure of the policy network is three layers. The number of neurons in the middle layer neural network is set to 128, and the last layer neural network outputs the action probabilities that the AGV can choose. The hyperparameters required during the experiment are shown in Table 2.

[0177] Parameter Name Description Value hiddendimension Hidden layer dimension 64 / 128 learningrate Learning rate 0.0005 rewardStandardisation Reward standardisation true networktype Network type Fully connected network (FC) entropycoefficient Entropy coefficient 0.01 targetupdate Target function update 0.01 (soft update soft) n-step Maximum number of steps per episode 5 γ Discount factor 1.0 epochsize Number of iterative updates in an epoch 10 batchsize Number of steps taken for an update 500

[0178] Table 2 Hyperparameters in the algorithm

[0179] The experiment uses a random number seed for replication and calculates the return with a 95% confidence interval. The parameters of different algorithms are optimized according to the different environments used, and remain constant for the remaining tasks in the same environment. For each combination of hyperparameters, different seeds are evaluated to obtain the optimal combination of hyperparameters for generating the results to be evaluated. Generally, all algorithms are evaluated when the number of hyperparameter combinations in each environment is roughly the same to ensure consistency.

[0180] Simulation experiment environment settings: The simulation experiment can customize the settings of the warehouse size, the number of AGVs in the system, the number of shelves, and the position of the target shelf. By default, the number of AGVs in the system is equal to the number of target shelves. The embodiments of the present invention verify the effectiveness of the MAA3C algorithm through experiments in environments of three different scales of warehouses (Warehouse 1, Warehouse 2, Warehouse 3). The three warehouses contain 9, 25, and 49 AGVs (agents) respectively, and the environmental parameters are shown in Table 3.

[0181] Warehouse parameter description Parameter Warehouse 1 Warehouse 2 Warehouse 3 Warehouse code Grid_Code warehouse1 warehouse2 warehouse3 Warehouse scale Grid_Sacle small medium large Number of columns in the warehouse Shelf_Colums 3 5 7 Column height in the warehouse Column_Height 3 5 7 Number of rows in the warehouse Shelf_Rows 3 5 7 Number of agents Agent_Number 9 25 49 Number of target shelves Request_Sizes 9 25 49 Agent observation range Sensor_Range 3*3 3*3 3*3 Grid world height Grid_Height 14 32 56 Grid world width Grid_Width 10 16 22 Grid world size Grid_Size 140 512 1232

[0182] Table 3 Warehouse environmental parameter table for different scales

[0183] Among them, the height formula of the grid world is shown in Formula (11):

[0184] Grid_Weight = (2 + 1) * Shelf_Columns + 1 (11)

[0185] The width formula of the grid world is shown in Formula (12):

[0186] Grid_Height = (Column_Height + 1) * Shelf_Rows + 2 (12)

[0187] According to the parameter input algorithm in the above table, the simulation maps of the virtual warehousing simulation training environments of Warehouse 1, Warehouse 2, and Warehouse 3 can be obtained, as shown in Figure 11 , Figure 12 and Figure 13 . Among them, the slant-filled squares represent shelves, the blank squares represent roads, the plus-filled squares represent target shelves, and the hexagons represent AGVs.

[0188] 1.2.2 Experimental results and analysis: Taking the experimental environment of Warehouse 1 as an example, during the experiment, each grid in the simulation environment is encoded to generate the quantization encoding map of the simulation environment of Warehouse 1, as shown in Figure 14 .

[0189] In the simulation environment of Warehouse 1, 9 AGVs correspond to 9 target tasks. Based on the multi-AGV conflict-free path planning model and the MAA3C algorithm, the running trajectories of each AGV under different task numbers can be obtained, as shown in Table 4.

[0190]

[0191] Table 4 AGV path trajectory table of simulation Warehouse 1

[0192] Figure 9Describes the reward iteration graph generated during training in the multi-AGV conflict-free path planning task simulation experiment, showing the variation trend of the overall reward of the warehousing system with the number of training steps in simulation environments of different scales. Among them, the horizontal axis represents the time steps during the training process, and the vertical axis is the reward value corresponding to each time step. The overall reward curves of Warehouse 1, Warehouse 2, and Warehouse 3 correspond to the curves warehouse1-MAA3C, curve warehouse2-MAA3C, and curve warehouse3-MAA3C in the figure respectively. It can be seen from the convergence values returned by the agents that in simulation experiment environments of different scales, the overall system reward can achieve final convergence as the number of training time steps increases. In the initial rounds, the reward value is relatively low. As the number of training steps increases, the rewards obtained by the agents gradually increase during the process of reaching the target shelves, and the overall system reward value gradually increases until it finally converges. Due to the setting of the reward function, as the number of AGVs in the simulation environment increases and the environment scale increases, the reward value of the system fluctuates; however, the increase in the simulation environment scale has no impact on the policy convergence of the AGVs, and the global policy still remains stable, verifying the effectiveness of the model and algorithm. Based on the model structure analysis, the input of the centralized Critic network includes the features of all AGVs themselves and the local observation features of all AGVs on the environment. During the policy evaluation process, the model effectively solves the instability problem of the system in the multi-agent warehousing environment according to the above input information. The centralized model network structure improves the policies of each agent in the system and can achieve the convergence of the overall model.

[0193] In the multi-AGV conflict-free path planning model based on deep reinforcement learning with a dynamic grouping mechanism proposed in the embodiments of the present invention, the idea of dynamic grouping is adopted to realize information interaction among relevant AGVs in the warehousing environment, and the MAA3C algorithm is designed for model solving. Therefore, the comparative experiments are carried out from two dimensions. One is whether to introduce the dynamic grouping mechanism of AGVs in the experimental simulation environment, and the other is based on the above-provided algorithm dimension, using different deep reinforcement learning algorithms for model solving and comparison. The experiment will use the above two dimensions to solve the conflict-free path planning problem of AGVs for different scales of warehousing environments. Tables 5 and 6 respectively show the average return values obtained during the code training process of seven algorithms, namely VDN (Value Decomposition Network), QMIX (Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning), COMA (Counterfactual Multi-Agent Policy Gradients), MADDPG (Multi-Agent Deep Deterministic Policy Gradient), MAPPO (Multi-Agent Proximal Policy Optimization, a deep reinforcement learning algorithm), MAA2C (Multi-Agent Advantage Actor Critic), and MAA3C (Multi-Agent Asynchronous Advantage Actor-Critic), without considering dynamic grouping and considering dynamic grouping. By comparing the training process data of agents in different algorithms under the conditions of non-dynamic grouping and dynamic grouping, it can be found that in most cases, the average return of all algorithms considering dynamic grouping is higher than that without considering dynamic grouping, and the difference is more obvious with the increase of scale. This shows that introducing dynamic grouping is very effective in solving the large-scale AGV path planning problem in the warehousing environment. Whether dynamic grouping is adopted or not, the MAA3C algorithm obtains the highest average return value compared with other algorithms; at the same time, with the increase of scale, the effect is more significant. This is because the MAA3C algorithm uses the joint action value function to train the agents. Compared with the local action value function, the centralized training fully considers the behaviors of other agents and realizes the joint policy evaluation among agents.The experimental results further illustrate the effectiveness and feasibility of using the MAA3C algorithm with dynamic grouping for model solution.

[0194]

[0195] Table 5 Average returns and 95% confidence intervals of different algorithms without dynamic grouping

[0196]

[0197] Table 6 Average returns and 95% confidence intervals of different algorithms with dynamic grouping

[0198] The embodiments of the present invention have the following technical effects: In the multi-AGV path planning in a warehouse environment, a dynamic grouping mechanism is introduced, multiple AGVs are dynamically grouped through an evaluation network and an attention network, and communication between AGVs is carried out through a communication network. Combining deep reinforcement learning to train the network parameters solves the curse of dimensionality problem that occurs in the joint decision-making process of multi-AGV conflict-free path planning, and improves the efficiency of multi-AGV conflict-free path planning.

[0199] It should be understood that the specific order or hierarchy of steps in the disclosed process is an example of an exemplary method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process can be rearranged without departing from the scope of the present disclosure. The appended method claims present the elements of the various steps in an exemplary order and are not intended to be limited to the specific order or hierarchy described.

[0200] In the above detailed description, various features are combined in a single embodiment to simplify the present disclosure. This method of disclosure should not be interpreted as reflecting an intention that the embodiments of the claimed subject matter require more features than are clearly recited in each claim. On the contrary, as reflected in the appended claims, the present invention lies in a state with fewer features than all the features of the disclosed single embodiment. Therefore, the appended claims are hereby expressly incorporated into the detailed description, where each claim stands alone as a separate preferred embodiment of the present invention.

[0201] In order to enable any person skilled in the art to implement or use the present invention, the above-described disclosed embodiments have been described. For those skilled in the art; various modification methods of these embodiments are obvious, and the general principles defined herein can also be applied to other embodiments without departing from the spirit and scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments given herein, but is consistent with the broadest scope of the principles and novel features disclosed in this application.

[0202] The foregoing description includes examples of one or more embodiments. Of course, it is not possible to describe all possible combinations of components or methods for the purpose of describing the above embodiments, but those of ordinary skill in the art should recognize that the various embodiments can be further combined and arranged. Accordingly, the embodiments described herein are intended to cover all such changes, modifications, and variations that fall within the scope of the appended claims. Further, with respect to the term "comprising" as used in the specification or claims, that term is intended to be inclusive in a manner similar to the term "including". Additionally, any use of the term "or" in the specification or claims is intended to mean "non-exclusive or".

[0203] The specific embodiments described above further elaborate on the object, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for conflict-free path planning of multiple AGVs in a warehousing environment based on deep reinforcement learning, characterized in that Including: Create a virtual warehousing simulation training environment, and extract virtual warehousing simulation training environment information from the virtual warehousing simulation training environment; According to the virtual warehousing simulation training environment information, determine the state observation value encoding of each Automated Guided Vehicle (AGV) among multiple AGVs in the virtual warehousing simulation training environment; Set the action space of the AGV and construct a reward function for deep reinforcement learning; Construct a distributed conflict-free multi-AGV path planner incorporating a dynamic grouping mechanism. According to the reward function, the state observation value encoding and the action space of each AGV, perform iterative training on the distributed conflict-free multi-AGV path planner, update the network parameters of the distributed conflict-free multi-AGV path planner, and obtain a trained distributed conflict-free multi-AGV path planner; Use the trained distributed conflict-free multi-AGV path planner to perform conflict-free path planning for multiple AGVs in the actual warehousing environment, and obtain a conflict-free path for each AGV; Wherein, the virtual warehousing simulation training environment information includes the initial positions, obstacle positions, shelf positions, and road information of multiple AGVs in the virtual warehousing simulation training environment; the state observation value encoding of the AGV includes: the position information and movement direction of the AGV, the target shelf position, whether it is carrying goods, and the position information and movement direction of obstacles and other AGVs within the sensing radius; the action space includes: turning left, turning right, moving forward, loading, and unloading; the sensing radius is used to define the observation range of the AGV; The distributed conflict-free multi-AGV path planner includes a global network model and a local network model for each AGV; the global network model and the local network model for each AGV have the same network structure and both correspondingly include an evaluation network, an attention network, a communication network, and a deep reinforcement learning network; The evaluation network is used to, for each AGV, fuse the features of other AGVs within the observation range of the AGV into the features of the AGV to obtain a feature vector of the AGV, and determine the adjacency matrix of the AGV according to the feature vectors of all AGVs within the observation range of the AGV; The attention network is used to, for each AGV, based on the self-attention mechanism, calculate a communication group centered on the AGV according to the adjacency matrix of the AGV and the state observation value encoding of the AGV included in the adjacency matrix of the AGV; The communication network is used to, for each AGV, perform network communication within the communication group centered on the AGV and update the state observation value encoding of each AGV within the communication group; The deep reinforcement learning network is constructed based on the multi-agent asynchronous advantage actor-critic algorithm. For the AGVs after the communication groups are divided by the dynamic grouping mechanism, action decisions are made in the policy network according to the state observation values encoded by each AGV, path planning is completed, and rewards are obtained according to the reward function. With the goal of maximizing the rewards, the model parameters of each AGV's local network model are continuously iteratively updated, and the model parameters of the global network model are updated using the model parameters of each AGV's local network model, where the deep reinforcement learning network includes a policy network and a critic network.

2. The multi-AGV conflict-free path planning method for the warehousing environment based on deep reinforcement learning according to claim 1, characterized in that Construct the reward function of deep reinforcement learning, including: The reward function is represented by the following formula: R(s,a) = R1 + R2 + R3 Among them, R1 represents the reward value obtained according to the distance between the AGV and the nearest obstacle, R2 represents the reward value obtained when the AGV reaches the target shelf position, and R3 represents the penalty value set for the time spent in the path planning process of the AGV, which is used to measure the cost paid by the AGV in the process of avoiding other agents and non-target shelves; Among them, R1 is determined by the following formula during the non-collision period: R1 = u × obs_distance And according to the following collision penalty correction formula, R1 is penalty-corrected when a collision is identified: Where, u represents the reward and punishment factor set for the distance between the AGV and the obstacle, which is a positive value; obs_distance represents the distance between the current AGV and the nearest obstacle; the collision penalty correction formula means that when the distance obs_distance from the AGV to the obstacle is less than or equal to the preset minimum collision distance dis crash it is recognized that a collision has occurred, and 100 is subtracted from R1 to punish and correct the value of R1; The initial value of R2 is 0, and according to the following arrival reward formula, R2 is reward-corrected when it is identified that the target shelf is reached: Among them, tar_distance represents the distance between the current AGV and the target shelf, and dis reach is the set arrival distance. When the distance tar_distance between the robot and the target shelf is less than the default arrival distance dis reach it means that the current AGV has completed path planning and reached the target shelf, and is recognized as reaching the target shelf; R2' is a continuous arrival reward function, specifically the formula: Among them, ep_count represents the cumulative value of the number of consecutive target arrival rounds of the AGV. If the AGV has a collision or reaches the maximum training time step but does not reach the target shelf during the simulation process, then ep_count is cleared and the accumulation is restarted; k' represents the reward coefficient for the number of consecutive target shelf arrivals ep_count of the AGV; Determine R3 according to the following formula: R3 = q × step_count Among them, q is the penalty coefficient based on the arrival time, and step_count is the cumulative step value.

3. The multi-AGV conflict-free path planning method for warehouse environment based on deep reinforcement learning according to claim 1, characterized in that, The constructed distributed conflict-free multi-AGV path planner introducing the dynamic grouping mechanism is iteratively trained according to the reward function and the state observation value encoding and action space of each AGV, and the network parameters of the distributed conflict-free multi-AGV path planner are updated to obtain the trained distributed conflict-free multi-AGV path planner, including: Define the global iteration count and set the initial value of the global iteration count to 0; Increment the global iteration count by one; Judge whether the global iteration count is less than the preset global iteration number threshold; If it is judged that the global iteration count is less than the preset global iteration number threshold, then execute the following steps: For each AGV in the virtual warehousing simulation training environment, synchronize the model parameters of the local network model corresponding to the AGV using the model parameters of the global network model, and set the time step count corresponding to the AGV to 0; Determine the state observation value encoding of each AGV in the virtual warehousing simulation training environment according to the virtual warehousing simulation training environment information and the preset perception radius; For each AGV in the virtual warehousing simulation training environment, based on the evaluation network of the local network model of the AGV, determine the feature vector of the AGV according to the state observation value encoding of the AGV; the feature vector of the AGV characterizes the environmental information within the perception radius of the AGV; Generate the adjacency matrix corresponding to the AGV according to the feature vector of the AGV; the adjacency matrix characterizes the adjacent relationship between all AGVs within the perception radius of the AGV; Based on the attention network of the local network model of the AGV, obtain the similarity matrix corresponding to the AGV according to the adjacency matrix corresponding to the AGV and the state observation value encodings of all AGVs in the adjacency matrix. According to the similarity matrix corresponding to the AGV, determine other AGVs having similar movement paths with the AGV, and form a communication group centered on the AGV together with the AGV and the other AGVs having similar movement paths obtained; wherein, each row element in the similarity matrix body characterizes the correspondence between the AGV corresponding to the row and other AGVs having similar running paths; Input the state observation value encodings of all AGVs in the communication group centered on the AGV into the communication network of the local network model corresponding to the AGV for prediction, and obtain the predicted state observation value encoding of the AGV; Input the predicted state observation value encoding of the AGV into the policy network of the deep reinforcement learning network to determine the to-be-executed action of the AGV; Execute the to-be-executed action of the AGV to enable the AGV to interact with the virtual warehousing simulation training environment, update the position information in the state observation value encoding of the AGV, and update the virtual warehousing simulation training environment information according to the updated position information of the AGV; Based on a preset reward value calculation function, calculate the reward value corresponding to the current time step count of the AGV according to the state observation value encoding of the AGV before executing the to-be-executed action and the updated state observation value encoding of the AGV; Increment the time step count corresponding to the AGV by one, and determine whether the position information in the updated state observation value encoding of the AGV is equal to the target shelf position in the state observation value encoding of the AGV, and whether the time step count corresponding to the AGV is greater than a preset local time step count threshold; if it is determined that the position information in the updated state observation value encoding of the AGV is not equal to the target shelf position in the state observation value encoding of the AGV, and the time step count corresponding to the AGV is less than the preset local time step count threshold, then return to the step of determining the state observation value encoding of each AGV in the virtual warehousing simulation training environment according to the virtual warehousing simulation training environment information and the preset perception radius for execution; Based on the position information and reward value of the AGV at the current time step count, determine the expected return of the AGV at the current time step count based on the critic network in the local network model of the AGV; Based on the expected return of the AGV at the current time step count, determine the expected return corresponding to each time step count before the current time step count, and perform gradient update on the model parameters in the local network model of the AGV according to the expected return corresponding to each time step count before the current time step count; the model parameters in the local network model of the AGV include the model parameters in the evaluation network, attention network, communication network, and deep reinforcement learning network in the local network model; According to the model parameters of the local network model of the AGV after gradient update, update the model parameters of the corresponding model in the global network model, and return to the step of incrementing the global iteration count by one; wherein, the model parameters of the global network model include: the model parameters in the evaluation network, attention network, communication network, and deep reinforcement learning network in the global network model; If it is determined that the global iteration count is greater than or equal to the preset global iteration number threshold, save the model parameters of the corresponding model in the global network model.

4. The multi-AGV conflict-free path planning method for the warehousing environment based on deep reinforcement learning according to claim 1, wherein Use the trained distributed conflict-free multi-AGV path planner to perform conflict-free path planning for multiple AGVs in the actual warehousing environment, and obtain the conflict-free paths of each AGV, including: Obtain the actual warehousing environment information, where the actual warehousing environment information includes the positions of multiple AGVs, obstacle positions, shelf positions, and road information in the actual warehousing environment; Initialize the model parameters of the local network model of each AGV among the multiple AGVs in the actual warehousing environment using the model parameters of the global network model in the trained distributed conflict-free multi-AGV path planner; At preset time step intervals, periodically construct a communication group centered on the AGV through the evaluation network and attention network of the local network model of each AGV in the actual warehousing environment at each time step, and update the state observation value encoding of each AGV in the communication group centered on the AGV through the communication network in the local network model of the AGV, input the updated state observation value encoding of the AGV into the policy network in the deep reinforcement learning network of the local network model of the AGV to obtain the action to be executed by the AGV corresponding to the time step, execute the action to be executed, and determine the position information of the AGV corresponding to the time step; Arrange the AGV position information of each AGV in the actual warehousing environment at each time step in chronological order to obtain the conflict-free planned path of each AGV.

5. The multi-AGV conflict-free path planning method for the warehousing environment based on deep reinforcement learning according to claim 3, wherein Determine the state observation value encoding of each AGV in the virtual warehousing simulation training environment according to the virtual warehousing simulation training environment information and the preset perception radius, including: For each AGV in the virtual warehousing simulation training environment, based on the virtual warehousing simulation training environment information, determine the obstacles and other AGVs within the sensing radius centered on the AGV, and construct the state observation value encoding of the AGV according to the obstacle positions corresponding to the determined obstacles, the position information and movement directions of other AGVs.

6. The multi-AGV conflict-free path planning method for the warehousing environment based on deep reinforcement learning according to claim 3, wherein For each AGV in the virtual warehousing simulation training environment, based on the state observation value encoding of the AGV, determine the feature vector of the AGV based on the evaluation network of the local network model of the AGV, including: According to the state observation value encoding of the AGV, determine all AGVs within the neighborhood of the AGV, and vectorize the state observation value encodings of all AGVs within the AGV into corresponding initial feature values; the neighborhood of the AGV represents the range covered by the preset sensing radius of the AGV. Save the initial feature values of all AGVs within the neighborhood of the AGV separately as the previous updated feature values. Set the initial value of the iteration count. Judge whether the iteration count is greater than the preset iteration count. If the iteration count is greater than the preset iteration count, end the iteration loop, and use the initial feature value of the AGV after the iteration loop ends as the feature vector of the AGV. Otherwise, continue to execute the following iteration loop: Take the average of the initial feature value of the AGV and the previous updated feature values corresponding to other AGVs within the neighborhood of the AGV, and then multiply it by the weight corresponding to the current iteration count to obtain an aggregated value. Nonlinearize the aggregated value through a preset nonlinear function to obtain the non - linearized aggregated value, and use the non - linearized aggregated value to update the initial feature value of the AGV. Perform normalization processing on the initial feature value of the AGV, and use the normalized initial feature value of the AGV to update the initial feature value of the AGV. For each other AGV within the neighborhood of the AGV, judge the position relationship of the other AGV relative to the AGV. In the case where it is judged that the other AGV is in the forward area of the AGV, update the previous updated feature value corresponding to the other AGV with the value obtained by multiplying the initial feature value of the other AGV by the first coefficient. In the case where it is judged that the other AGV is in the backward area of the AGV, update the previous updated feature value corresponding to the other AGV with the value obtained by multiplying the initial feature value of the other AGV by the second coefficient. Return to the step of judging whether the iteration count is greater than the preset iteration count for execution. Among them, judging the position relationship of the other AGV relative to the AGV includes: Encoding based on the state observation value of the AGV, obtaining the position information of other AGVs within the sensing radius of the AGV, and dividing the other AGVs within the sensing radius of the AGV into the AGVs within the facing area of the AGV and the AGVs within the backward area of the AGV; wherein, the facing area of the AGV is defined as the first quadrant and the second quadrant of the coordinate system with the position information in the state observation value encoding of the AGV as the origin; the backward area of the AGV is defined as the third quadrant and the fourth quadrant of the coordinate system with the position information in the state observation value encoding of the AGV as the origin; the Y-axis of the coordinate system with the position information in the state observation value encoding of the AGV as the origin is defined as the straight line pointing from the position information in the state observation value encoding of the AGV to the target shelf position in the state observation value encoding of the AGV, and the positive direction of the Y-axis faces the target shelf position in the state observation value encoding of the AGV, and the X-axis of the coordinate system with the position information in the state observation value encoding of the AGV as the origin is perpendicular to the Y-axis; Wherein, the first coefficient is greater than the second coefficient, both the first coefficient and the second coefficient are greater than 0 and the sum of the first coefficient and the second coefficient is 1; the first coefficient, the second coefficient and the weights corresponding to all the iteration times are the model parameters of the evaluation network.

7. The multi-AGV conflict-free path planning method for warehouse environment deep reinforcement learning according to claim 3, wherein Generating an adjacency matrix corresponding to the AGV according to the feature vector of the AGV, including: Calculating the similarity values of the feature vectors of all AGVs within the sensing radius of the AGV, and setting the element positions corresponding to the two AGVs with the obtained similarity values greater than the preset adjacency similarity threshold in the adjacency matrix to 1, otherwise setting them to 0.

8. The multi-AGV conflict-free path planning method for warehouse environment based on deep reinforcement learning according to claim 3, wherein Based on the attention network of the local network model of the AGV, obtaining the similarity matrix corresponding to the AGV according to the adjacency matrix corresponding to the AGV and the state observation value encodings of all AGVs in the adjacency matrix, and determining other AGVs with similar movement paths to the AGV according to the similarity matrix corresponding to the AGV, and forming a communication group centered on the AGV by the AGV and the other AGVs with similar movement paths obtained, including: Vectorizing the state observation value encodings of all AGVs in the adjacency matrix corresponding to the AGV into corresponding encoding vectors, and arranging the obtained all encoding vectors in the order of the AGVs in the adjacency matrix to obtain an observation value encoding matrix; Initializing the query matrix, key matrix and value matrix in the attention network in the local network model of the AGV with the observation value encoding matrix; Obtain the similarity matrix for the current time step according to the following formula Among them, M is the adjacency matrix at the current time step; Q is the query matrix, K is the key matrix, and d k is the dimension of the key matrix, and T represents the transpose matrix of the current matrix; Calculating the normalized similarity matrix according to the following formula: Among them, the softmax function is a normalization function; It represents the interaction relationship between AGVs before normalization, It represents the sum of the exponential operations on the relationship between the i-th AGV and other k-th AGVs; Dividing the other AGVs corresponding to the columns or rows with the value of 1 on the row or column corresponding to the AGV in Attention(Q, K, V) into the communication group centered on the AGV.

9. The multi-AGV conflict-free path planning method for warehouse environment based on deep reinforcement learning according to claim 4, wherein At preset time step intervals, periodically at each time step, through the evaluation network and attention network of the local network model of each AGV in the actual warehousing environment, construct a communication group centered on the AGV, and update the state observation value encoding of each AGV in the communication group centered on the AGV through the communication network in the local network model of the AGV. Input the updated state observation value encoding of the AGV into the policy network in the deep reinforcement learning network of the local network model of the AGV to obtain the action to be executed by the AGV at the corresponding time step, execute the action to be executed, and determine the position information corresponding to the time step of the AGV, including: For each time step, determine the state observation value encoding of each AGV in the actual warehousing environment according to the actual warehousing environment information and the preset perception radius; the actual warehousing environment information includes: the position information of multiple AGVs, the obstacle positions, the shelf positions, and the road information; For each AGV in the actual warehousing environment, determine the feature vector of the AGV based on the evaluation network in the local network model of the AGV according to the state observation value encoding of the AGV; Generate the adjacency matrix corresponding to the AGV according to the feature vector of the AGV; Based on the attention network in the local network model of the AGV, obtain the similarity matrix corresponding to the AGV according to the adjacency matrix corresponding to the AGV and the state observation value encodings of all AGVs in the adjacency matrix; According to the similarity matrix corresponding to the AGV, determine other AGVs with similar movement paths to the AGV, and form a communication group centered on the AGV by the AGV and the other AGVs with similar movement paths obtained; Input the state observation value encodings corresponding to all AGVs in the communication group centered on the AGV into the communication network in the local network model of the AGV for prediction to obtain the predicted state observation value encoding of the AGV; Input the predicted state observation value encoding of the AGV into the policy network of the deep reinforcement learning network of the local network model of the AGV to determine the action to be executed by the AGV; Execute the action to be executed by the AGV to enable the AGV to interact with the actual warehousing environment, update the position information in the state observation value encoding of the AGV, and update the actual warehousing environment information according to the position information in the updated state observation value encoding of the AGV; Judge whether the position information in the state observation value encoding of the AGV is equal to the target shelf position in the state observation value encoding of the AGV; If it is judged that the position information in the state observation value encoding of the AGV is not equal to the target shelf position in the state observation value encoding of the AGV, then return to the step of determining the state observation value encoding of each AGV in the actual warehousing environment according to the actual warehousing environment information and the preset perception radius for execution.

Citation Information

Patent Citations

  • Cluster track automatic planning method based on multi-agent safety reinforcement learning

    CN116661503A

  • Unmanned vehicle adaptive path planning method based on dynamic window method and near-end strategy

    CN116679719A