Multi-intelligent Communication Reinforcement Learning Agent Path Planning Method and System Based on Warehouse Environment

By introducing deep reinforcement learning and multi-agent communication in multi-agent path planning, combined with greed priority and deadlock detection mechanisms, the existing methods have solved the problem of long calculation time and easy to fall into deadlock in large-scale complex environments, and efficient and real-time path planning and deadlock processing are achieved.

CN115496287BActive Publication Date: 2025-05-30HARBIN INST OF TECH SHENZHEN GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211179911.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-15
Publication Date
2025-05-30
Estimated Expiration
2042-11-15

AI Technical Summary

Technical Problem

The existing multi-agent path planning methods have a long calculation time and poor scalability in large-scale complex environments. Distributed learning methods are difficult to ensure optimal solutions and are prone to deadlocks. The information provided by existing reinforcement learning methods in MAPF problems is too redundant, which affects agent decision-making.

Method used

A multi-intelligent communication reinforcement learning body path planning method based on the warehousing environment is proposed. Through deep reinforcement learning combined with multi-intelligent communication, greedy priority allocation and communication topology based on the adjacency matrix are adopted to generate planning paths, and a deadlock detection mechanism is designed to handle abnormal states.

Benefits of technology

It improves the distributed decoupling and real-time nature of path planning, alleviates the unstability of reinforcement learning algorithms, enhances the success rate of path planning, and effectively deals with deadlock situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115496287B_ABST
    Figure CN115496287B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-intelligent communication reinforcement learning body path planning method and system based on a warehousing environment. The method includes: generating a map, obtaining the starting point, target point and obstacle information of the intelligent body and inputting them into a neural network, obtaining the self-feature of the intelligent body through an observation value processing module, allocating the intelligent body by using a greedy-based priority, selecting neighbor intelligent bodies for each intelligent body based on an adjacency matrix and according to the allocated priority, each intelligent body receiving the communication messages of the neighbor intelligent bodies it selects and forming neighbor features, forming final features according to the neighbor features and its own features, and inputting the final features into a decision network module to generate a planned path. The present invention introduces communication to alleviate the environmental instability caused by reinforcement learning, improves the effectiveness by selecting communication intelligent bodies according to priorities, and introduces a new deadlock detection mechanism to enable the intelligent body to break out of the deadlock.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi-intelligent communication reinforcement learning body path planning method and system based on a warehousing environment, belonging to the field of artificial intelligence. Background Technique

[0002] The multi-agent path finding problem (abbreviated as MAPF) has many applications in real life. For example, path finding is a very important part of games, and most of the motion planning problems of game programs are path planning problems. In the vehicle scheduling problem, how to reasonably schedule trains so that they can reach their respective destinations without conflicts can be solved by modeling it into a MAPF problem. Aircraft route planning needs to plan non-conflicting routes for all flights, which is essentially also a path planning problem. The warehousing robot system is essentially to plan paths for hundreds or thousands of intelligent warehousing robots at the same time, ensuring that there are no conflicts between the robots and they can reach the destination quickly. Thus, the solution of the multi-agent path finding problem has great practical significance.

[0003] Currently, there are mainly two types of multi-agent path planning algorithms, namely centralized planning methods and distributed path planning methods, and the learning-based method belongs to the distributed planning method. Although the centralized solution method has optimality, in a complex environment with a large number of agents, the calculation time of this method also increases with the increase of the scale, and may even exceed the specified time, resulting in poor scalability. The distributed learning-based method has good scalability and real-time performance, but it cannot guarantee to obtain the optimal solution. The learning-based method has been proven to be an effective way to solve the MAPF problem. For large-scale warehousing systems, the distributed strategy obtained by the learning-based method can effectively improve efficiency and scalability, but compared with the traditional centralized planner, the learning-based method is more likely to fall into deadlock. In recent years, communication learning has also developed greatly in the field of multi-agents, and now it has also begun to explore introducing communication into the MAPF problem or combining it with the reinforcement learning method. However, for reinforcement learning, the current communication method will provide too redundant information to interfere with the decision-making of the agent.

[0004] At the same time, since most of the existing reinforcement learning methods applied to the MAPF problem use Independent RL, in this case each agent regards other agents around it as part of the environment, which will inevitably lead to an unstable environment, thus affecting the decision-making process of the agent. Moreover, imitation learning also has its own limitations. The error of imitation learning comes from the fitting error of the observed values that have been seen and the generalization error of the observed values that have not been seen. Therefore, there are still certain defects in the planning of imitation learning in different training environments. In addition, due to the fact that the learning-based method only uses local information for decision-making, it will inevitably fall into a deadlock state, and currently existing multi-agent path planning methods do not have a good detection and handling mechanism for deadlocks. Summary of the Invention

[0005] The present invention provides a multi-intelligent communication reinforcement learning body path planning method and system based on a warehousing environment, aiming to solve at least one of the technical problems existing in the prior art.

[0006] The technical solution of the present invention relates to a multi-intelligent communication reinforcement learning body path planning method based on a warehousing environment. The method according to the present invention includes the following steps:

[0007] S10. Generate a warehousing environment map, obtain the starting point, target point and obstacle information of each agent and input them into a neural network based on deep reinforcement learning; the observation value processing module obtains the own characteristics of each agent according to the input observation values;

[0008] S20. Allocate agents according to the own characteristics of each agent and using a greedy-based priority; select neighbor agents for each agent based on the adjacency matrix and according to the allocated priority. Each agent receives the communication messages of the neighbor agents it selects and forms neighbor characteristics;

[0009] S30. Form final characteristics according to the neighbor characteristics and the own characteristics, and input the final characteristics into the decision network module to generate a planned path.

[0010] Further, in the step S10:

[0011] The input channels of the observation values of the observation value processing module include nine, and the nine input channels are channels 0 to 8 respectively; among them, channel 0 contains information about other agents within the field of view of the agent, channel 1 contains information about the obstacle map within the field of view of the agent, channel 2 contains information about the target point map within the field of view of the agent, channels 3 to 6 all contain information about the surrounding agent maps of the agent in the past multiple time steps, channel 7 contains information about the target point deviation map, and channel 8 contains information about whether the agent is in an abnormal state.

[0012] Further, for the step S10, the neural network includes:

[0013] Four convolutional layers for processing the input information of channels 0 to 6;

[0014] A first fully connected layer for processing the input information of channel 7;

[0015] A second fully connected layer for processing the input information of channel 8;

[0016] A third fully connected layer for splicing the output information of the four convolutional layers, the first fully connected layer and the second fully connected layer, and the third fully connected layer outputs the final features;

[0017] A fourth fully connected layer and a fifth fully connected layer, wherein the final features are split into the fourth fully connected layer and the fifth fully connected layer;

[0018] A sixth fully connected layer for merging the output information of the fourth fully connected layer and the fifth fully connected layer, and the sixth fully connected layer outputs the planned path.

[0019] Further, the step S20 includes:

[0020] S21. Obtain the Manhattan distance from the current position to the target position of each agent, and set the reciprocal of the Manhattan distance as the priority of the agent;

[0021] S22. When the Manhattan distances of two agents are equal, set the priority according to the congestion degree of the two agents; wherein, the congestion degree represents the number of obstacles within the field of view of the agent;

[0022] S23. Obtain the communication range of the current agent, and calculate the current position distance between the current agent and other agents; set the other agents whose current position distance is within the communication range as neighbor agents;

[0023] S24. According to the set communication connection upper limit, select the agent with the highest priority from other agents whose current position distance exceeds the communication range as a neighbor agent;

[0024] S24. Determine whether the neighbor agent is an agent that has reached the target point; if so, disconnect the communication connection between the neighbor agent that has reached the target point and the current agent.

[0025] Further, the priority of the agent is represented by the following formula:

[0026]

[0027] In the formula, vi represents an agent, represents the number of other agents within the field of view of the agent; is the Manhattan distance from the current position of the agent to the target position, represents the priority of the agent.

[0028] Furthermore, in the step S30:

[0029] The output of the fourth fully connected layer is a state value function that is only related to the state and not related to the action;

[0030] The output of the fifth fully connected layer is an advantage function that is related to both the state and the action;

[0031] The Q-value function of the neural network is obtained based on the state value function and the advantage function.

[0032] Furthermore, the input value of the Q-value function is obtained through the following calculation:

[0033]

[0034] where, represents the output of the own characteristics of each agent of the observation value processing module, then represents the feature set of the entire multi-agent system; represents the neighbor agent list of the i-th agent, then represents the graph transfer operator of the adjacency matrix network, where 1 means receiving the communication message of this neighbor agent, and 0 means not receiving the communication message of this neighbor agent; w i and b i respectively represent the weights and biases of the cross network; represents the final feature output, x represents the input value of the Q network.

[0035] Furthermore, the training of the neural network includes a deadlock detection module, and the deadlock detection module includes the following steps:

[0036] Obtain the own position situation of the agent in the past multiple time steps and detect it through the abnormal state detection module;

[0037] If it is detected that the current agent stays at a non-target point position for more than three time steps, then determine that the current agent is a stagnant abnormal agent; judge whether the congestion degree of the stagnant agent is greater than or equal to the congestion threshold, if so, pre-label the stagnant abnormal agent;

[0038] If it is detected that the current agent has been moving back and forth between two positions in the past four time steps, it is determined that the current agent is a wandering abnormal agent; it is judged whether the congestion degree of the wandering agent is greater than or equal to the congestion threshold, and if so, the wandering abnormal agent is pre-labeled;

[0039] It is judged whether there are two or more pre-labeled agents within the communication range. If so, it is determined that the pre-labeled agents in the adjacency matrix are in a deadlock state;

[0040] Adjust the parameters of the neural network according to the information of the agents in the deadlock state.

[0041] The technical solution of the present invention also relates to a computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a processor, the above-mentioned method is implemented.

[0042] The technical solution of the present invention also relates to a multi-agent communication reinforcement learning body path planning system based on a warehousing environment. The system includes a computer device, and the computer device includes the above-mentioned computer-readable storage medium.

[0043] The beneficial effects of the present invention are as follows:

[0044] The multi-agent communication reinforcement learning body path planning method and system based on a warehousing environment of the present invention are modeled as a multi-agent path planning problem for solution. By combining reinforcement learning with multi-agent communication, reinforcement learning is used to achieve distributed decoupling of path planning, and communication is used to alleviate the instability of the reinforcement learning algorithm applied to the MAPF problem; and a communication topology based on priority is designed. Before communication, agents are selected for communication through priority, and unnecessary communication contacts are erased, which is beneficial to ensuring the effectiveness of communication objects; a new deadlock detection mechanism is also designed. By using the outstanding performance of auxiliary information during the learning process, agents learn a method to break out of the deadlock, realizing the detection and handling of abnormal states during the planning process. Brief Description of the Drawings

[0045] Figure 1 is a basic flowchart of the multi-agent communication path planning method according to the present invention.

[0046] Figure 2 is a schematic diagram of the neural network algorithm according to the method of the present invention.

[0047] Figure 3 is a schematic diagram of the neural network structure according to the method of the present invention.

[0048] Figure 4 is a schematic diagram of the structure of the observation value input channel according to the method of the present invention.

[0049] Figure 5Schematic diagram of the first embodiment of the communication module according to the method of the present invention.

[0050] Figure 6 Basic flowchart of the deadlock detection module according to the method of the present invention.

[0051] Figure 7 Schematic diagram of the second embodiment of the deadlock detection module according to the method of the present invention.

[0052] Figure 8 Simulation map of the storage environment with a simple structure according to an embodiment of the present invention.

[0053] Figure 9 Simulation map of the storage environment with a complex structure according to an embodiment of the present invention.

[0054] Figure 10a and Figure 10b Comparison diagram of the test results of the algorithm of the present invention and the DHC algorithm.

[0055] Figure 11a and Figure 11b Comparison diagram of the test results of different variants of the method of the present invention. Detailed implementation manners

[0056] The concept, specific structure and technical effects of the present invention will be clearly and completely described below in conjunction with the embodiments and the drawings, so as to fully understand the purpose, solution and effects of the present invention.

[0057] It should be noted that, unless otherwise specified, when a certain feature is referred to as "fixed" or "connected" to another feature, it can be directly fixed or connected to the other feature, or indirectly fixed or connected to the other feature. As used herein, the singular forms "a", "the" and "said" are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in the description of this specification herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term "and / or" used herein includes any combination of one or more of the related listed items.

[0058] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish elements of the same type from each other. For example, without departing from the scope of this disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element. The use of any and all examples or exemplary language ("for example", "such as", etc.) provided herein is only intended to better illustrate the embodiments of the present invention and will not impose a limitation on the scope of the present invention unless otherwise required.

[0059] Referring to Figures 1 to 3 , in some embodiments, the multi-intelligent communication reinforcement learning body path planning method based on the warehousing environment according to the present invention at least includes the following steps:

[0060] S10. Generate a warehousing environment map, obtain the starting point, target point, and obstacle information of each intelligent body and input them into a neural network based on deep reinforcement learning; the observation value processing module obtains the self-characteristics of each intelligent body according to the input observation values;

[0061] S20. Allocate intelligent bodies according to the self-characteristics of each intelligent body and using a greedy-based priority; select neighbor intelligent bodies for each intelligent body based on the adjacency matrix and according to the allocated priority, and each intelligent body receives the communication messages of the neighbor intelligent bodies it selects and forms neighbor characteristics;

[0062] S30. Form final characteristics according to the neighbor characteristics and self-characteristics, and input the final characteristics into the decision network module to generate a planned path.

[0063] Specific implementation of step S10

[0064] The reinforcement learning framework of the present invention is IQL based on the D3QN (Dueling Double Deep Q Network) algorithm, and its network architecture mainly includes four modules (see Figure 2 ), namely the input channel module (Agent), the observation value processing module (Observation Module), the communication message generation and communication module (Communication Moudle), and the decision network module (Q-Nerwork).

[0065] The method of the embodiment of the present invention uses a neural network based on deep reinforcement learning (see Figure 3) Its neural network includes four convolutional layers for processing the input information of channels 0 to 6; a first fully connected layer for processing the input information of channel 7; a second fully connected layer for processing the input information of channel 8; a third fully connected layer for splicing the output information of the four convolutional layers, the first fully connected layer and the second fully connected layer, and the third fully connected layer outputs the final features; the final features are split into the fourth fully connected layer and the fifth fully connected layer; the sixth fully connected layer combines the output information of the fourth fully connected layer and the fifth fully connected layer, and the sixth fully connected layer outputs to the Q network.

[0066] Specifically, a warehouse environment map is generated, and the starting point, target point, and obstacle information of each agent are obtained and input into a neural network based on deep reinforcement learning to obtain observations. Then, an observation processing module obtains the self-features of each agent according to the observations.

[0067] Among them, for the input channels of the observations of the observation processing module, see Figure 3 As shown, the observations of the observation processing module are divided into nine channels from 0 to 8 for input. Channel 0 is other agents within the field of view. Channel 1 is the obstacle map within the field of view. Channel 2 is the target point map within the field of view, and when the target point is outside the field of view, it is projected onto the boundary of the field of view to guide the agent's forward direction. Channels 3 to 6 are the maps of surrounding agents in the past four steps. Channel 7 belongs to an auxiliary channel, containing the target point deviation map information, which is used to represent the differences in the horizontal and vertical coordinates of the agent from the target point in the current and the past three time steps, and the distance between the agent and the target point and whether it is in a wandering state can be learned from it. Channel 8 is an auxiliary channel, which is used to represent whether the agent is in an abnormal state.

[0068] The observation processing module consists of four convolutional layers, and the middle two convolutional layers are combined into a residual block (see Figure 2 the right part) to solve the degradation problem of the deep neural network. Channels 0 to 6 all enter the convolutional layer first. Channels 7 and 8 are divided into two auxiliary vectors. One vector (channel 7) represents the deviation of the actions in the past few steps from the horizontal and vertical coordinates of the target point, and the other vector (channel 8) represents whether the current agent has stopped, wandered, and deadlocked. See Figure 4 . These two vectors are respectively output through the first fully connected layer and the second fully connected layer, then spliced together with the output of the convolutional layer, and then passed through the third fully connected layer to obtain the output of the observation processing module.

[0069] Specific implementation manner of step S20

[0070] The output of the observation value processing module enters the communication module for communication. Before communication, the communication objects (neighboring agents) are selected first. First, a greedy-based priority is assigned to each agent. Then, the adjacency matrix and the priority are used to select the respective communication objects for each agent. After the communication objects are selected, in the communication module, the communication messages received from the neighboring agents are aggregated for each agent to obtain neighbor features. The neighbor features are concatenated with the original state value (self-feature) of the agent to form the final state tensor (final feature). The final feature is input into the final decision network to generate the planned path points.

[0071] Specifically, refer to Figure 5 , in order to improve the effectiveness of the neighboring agents selected by the agent, it is necessary to assign priorities to each agent first. The method of the embodiment of the present invention adopts a greedy-based priority assignment strategy. It should be noted that in the MAPF setting of the embodiment of the present invention, the agents and their tasks are homogeneous, and each task does not have a pre-assigned priority. The priority strategy is distance-based, that is, in each time step, the priority of each agent is set to the reciprocal of the Manhattan distance from the current position of the agent to the target point. Thus, the closer the Manhattan distance, the higher the priority of the agent.

[0072] At the same time, the number of obstacles within the field of view of each agent is set as the congestion degree, where the above-mentioned obstacles include static obstacles and other agents. For two agents with equal Manhattan distances, it is set that the higher the congestion degree of the agent, the higher its priority. The above-mentioned set priorities will change dynamically during the path planning process of the agent, which has better adaptability than fixed priorities.

[0073] Specifically, the priority of the agent is represented by the following formula:

[0074]

[0075] In the formula, v i represents the agent, represents the number of other agents within the field of view of the agent; is the Manhattan distance from the current position of the agent to the target position, represents the priority of the agent.

[0076] After the priority assignment of the agent is completed, it enters the communication object (neighboring agent) selection stage. Refer to the adjacency matrix Figure 5, first, obtain the communication range of the agent, then calculate the current position distance between each agent and other agents, then set the position distances of all other agents beyond the communication range to 0, and set the other agents with the current position distance within the communication range as neighbor agents. Select the agent with the highest priority among the other agents beyond the communication range as a neighbor agent. Specifically, according to the set communication connection upper limit n, and the number of neighbor agents within the communication range is m, select n - m agents with the highest priority from the other agents beyond the communication range as neighbor agents. Compared with the method of communicating with all agents within the communication range, the above method of the embodiment of the present invention helps to avoid unnecessary computational consumption and redundancy of messages.

[0077] Furthermore, determine whether the selected neighbor agent is an agent that has reached the target point. If so, disconnect the communication connection between the neighbor agent that has reached the target point and the current agent, thereby improving the reception effectiveness of the agent, helping to eliminate the influence of the agent that has reached the target point on the decision-making of the current agent, and avoiding the possibility that the current agent adopts an avoidance strategy when approaching the agent that has reached the target point.

[0078] Here, a specific embodiment is used for illustration. Refer to Figure 5 , sphere 1 is the current agent. The communication range (observation range) of the current agent is shown as the bold box in Figure 1. The communication connection upper limit is set to three. In the first step, select all agents within the communication range (refer to the bold box in Figure 2) as neighbor agents. The selected neighbor agents include sphere 2 agent and sphere 3 agent that has reached the target point directly above. In the second step, according to the communication connection upper limit of three and the fact that two neighbor agents have been selected, then select an agent with the highest priority from the agents outside the bold box in Figure 2 as a neighbor agent. Specifically, for sphere 4 agent and sphere 5 agent, according to the priority definition of the embodiment of the present invention, the congestion degree of sphere 5 agent is greater and the Manhattan distance to the target point is smaller. Therefore, the priority of sphere 5 agent is higher, and sphere 5 agent is selected as a neighbor agent. In the last step, disconnect the communication connection of sphere 2 agent that has reached the target point. Therefore, the communication objects of sphere 1 agent are sphere 2 agent and sphere 5 agent.

[0079] Specific implementation manner of step S30

[0080] The neural network based on the D3QN framework in the embodiments of the present invention separates the main network and the target network through Double DQN, and performs action selection and value estimation respectively. Meanwhile, the finally extracted feature is a feature tensor that combines the state of the agent itself (self feature) and communication information (neighbor feature). After the final feature extraction by the Dueling DQN algorithm, the information (feature tensor) is split into the fourth fully connected layer and the fifth fully connected layer. Among them, the fourth fully connected layer represents the state value function V(s; θ, β) that is only related to the state and not related to the action, and the value function V is used to make a long-term judgment on the current state; the fifth fully connected layer represents the advantage function A(s, a; θ, α) that is related to both the state and the action, and the advantage function A is used to measure the relative goodness or badness of each action. Finally, the outputs of the two fully connected layers are combined to obtain the final Q-value function Q(s, a; θ, α, β), and the action corresponding to the maximum Q value is selected. The decision network module generates a planned path point according to the output of the Q-value function.

[0081] Specifically, the selected neighbor agents passed through Figure 5 The final graph transfer operator of the network Where represents the communication object table (neighbor agent table) of the i-th agent, 1 means receiving the message transmitted by this agent, and 0 means not receiving the message transmitted by this agent. Set the state output of each agent in the previous observation value processing module to Then the state set of the entire multi-agent system is represented as The previously calculated GSO operator is represented as And there is:

[0082]

[0083] In the formula, where represents a communication connection relationship between the i-th agent and other agents, Then for the i-th agent, the value x input to the Q network is calculated as follows:

[0084] First, set represents the final feature, Then the obtained are respectively input into the weight generation fully connected layer (the fourth fully connected layer) and the bias fully connected layer (the fifth fully connected layer) of the hyper network to obtain the weight w i and the bias b i , and the output of the final feature interaction process is It should be noted that if there are multiple hops in the communication module, multiple feature interactions are required to spread the state information beyond the communication range, and the state features after the final feature processing will be input into the Q network through the sixth fully connected layer.

[0085] Specific implementation manner of the deadlock detection mechanism

[0086] In the method of the embodiment of the present invention, during the training process of the neural network, a brand-new deadlock detection mechanism is introduced (see Figure 6 ), to implement embedding the abnormal situation existing in the current intelligent agent in the auxiliary input channel. Specifically, the deadlock is divided into two situations. One is the deadlock in the stagnant abnormal state, that is, multiple intelligent agents are locked and stagnated on the road; the other is the deadlock in the wandering abnormal state, that is, due to the learning characteristics, two intelligent agents make avoidance decisions for each other at the same time. Before judging whether the intelligent agent has a deadlock, first, according to the position of each intelligent agent recorded in the past multiple time steps, the abnormal state detection module is used to detect whether the intelligent agent is in the stagnant abnormal state or the wandering abnormal state, and the first-round detection and the second-round detection of the intelligent agent are carried out (see Figure 7 ).

[0087] Among them, the first-round detection includes: for the detection of the stagnant abnormal state, if the intelligent agent stays at a non-target point position for more than three time steps, it is determined that the intelligent agent is in the stagnant abnormal state, and at this time, the stagnant abnormal state flag bit corresponding to the intelligent agent is set to 1. For the wandering abnormal state, since the path planning intelligent agent trained by the reinforcement learning method may wander on the map, or some work processing makes the stationary intelligent agent return to the previous position, by detecting whether the intelligent agent has wandered back and forth between two positions in the past four time steps, it is determined whether the intelligent agent is in the wandering abnormal state. For example, if the position of the intelligent agent at time t-1 is the same as the position at time t-3, and the position at time t is the same as the position at time t-2, then the intelligent agent is in a wandering state on the map, and the corresponding wandering abnormal state flag bit is set to 1. Then, the congestion degree of the intelligent agent in the past multiple time steps is obtained to obtain the congestion state of the area where the intelligent agent is located in the past few time steps. When it is detected that the intelligent agent is in the stagnant abnormal state or the wandering abnormal state and the congestion degree is greater than or equal to 1 and remains unchanged or increases for multiple steps, this intelligent agent is pre-labeled as deadlocked. It should be noted that setting the congestion degree determination threshold to 1 is beneficial to preventing the state where a single intelligent agent is identified as deadlocked due to being in the stagnant and wandering states in the early stage of training.

[0088] However, this does not guarantee that the agent must be in a deadlock state. Since a deadlock state must be caused by multiple agents competing for channel resources, a single agent entering this abnormal state of suspected deadlock may be due to the error of the training network or the agent being in a crowded channel and being forced to take evasive actions to allow different agents to pass. Therefore, after obtaining the pre-labeled agents in the first round of deadlock detection, a second round of detection is required.

[0089] The second round of detection includes: If two or more pre-labeled deadlock agents appear within the communication range, it is confirmed that these agents are in a deadlock state. If a pre-labeled agent has no other pre-labeled agents within the communication range, it is determined that this agent is in a non-deadlock state.

[0090] Finally, information on the deadlock state is added to the auxiliary channel (input channel 8), and relevant penalties are imposed on the deadlock state to encourage the agent to learn to break out of the deadlock state. Based on the fact that a deadlock state must be caused by multiple agents competing for channel resources, and a single agent entering this abnormal state of suspected deadlock may be due to the error of the training network or the agent being in a crowded channel and being forced to take evasive actions to allow different agents to pass, the above-mentioned deadlock detection mechanism is set up in the embodiments of the present invention to detect and handle abnormal states during the planning process. At the same time, a special part for recording the agents within the field of view of each agent is provided in the input channel, so the calculation of the congestion degree does not consume too much computing resources.

[0091] Here, a specific embodiment is used for illustration. Refer to Figure 7 , at time t, it is detected that agent 1 has been moving back and forth between two adjacent coordinate points in the last four time steps, being in a wandering abnormal state. At the same time, the congestion degree of agent 1 in the past four moments is 1, so it is pre-labeled as a suspected deadlock agent. After all possible pre-labeled deadlock agents are detected in the above-mentioned first round. Then, in the second round of detection, it is checked whether there are other pre-labeled agents within the communication range of a pre-labeled deadlock agent. Refer to Figure 7 There is another suspected deadlock agent No. 2 within 2 grids (communication range) of the field of view of agent 1 in

[0092] The present invention conducts actual tests on the multi-agent path planning method based on communication reinforcement learning, and tests the simulation map of a simple-structured warehouse environment (refer to Figure 8 ). The test results are shown in Figure 9 . By adding a communication module and a deadlock detection module, the success rate of the path planning algorithm has been greatly improved, proving the effectiveness of these two modules in the method proposed by the present invention.

[0093] Furthermore, the method of the present invention is used to test the simulation map of the warehouse environment in a complex environment. Refer to Figure 9 the two complex environments shown. In the complex environment, the obstacle density is increased, and a single channel is introduced to make it easier to appear in a suspected deadlock state. Each test randomly generates Figure 9 one of the two environments, and at the same time, the test results of the method of the present invention (OURS) and the current optimal method (DHC) in the complex environment are compared. The test results are shown in Figure 10a and Figure 10b . From this, it can be obtained that the success rate of the algorithm of the present invention in the complex environment has been improved. In addition, the method of the present invention is modified, for example, deleting the communication module (DD), or deleting both the communication module (DD) and the deadlock detection module (CM) at the same time. Refer to Figure 11a and Figure 11b . As shown, the success rate of the algorithm of the present invention will drop significantly, especially when the number of agents is large, it is even difficult to plan a collision-free path for all agents to the target point. From this, it can be obtained that the improvement of the algorithm of the present invention not only comes from the efficiency of communication, but also from the fact that the deadlock detection mechanism provides additional abnormal state information for the agents so that the agents have a greater possibility of getting out of the deadlock state. This can prove the generalization performance of the method proposed by the present invention in different environments, as well as the solving ability for this kind of single-channel environment that is extremely prone to deadlocks.

[0094] It should be recognized that the method steps in the embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or computer instructions stored in a non-transitory computer-readable memory. The method can use standard programming techniques. Each program can be implemented in a high-level procedure or object-oriented programming language to communicate with the computer system. However, if necessary, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, for this purpose, the program can run on a dedicated integrated circuit programmed for this purpose.

[0095] In addition, the operations of the processes described herein can be performed in any suitable order, unless otherwise indicated herein or otherwise clearly contradicted by the context. The processes described herein (or variations and / or combinations thereof) can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed commonly on one or more processors, by hardware, or a combination thereof. The computer program includes a plurality of instructions executable by one or more processors.

[0096] Further, the method can be implemented in any type of computing platform operatively connected to a suitable one, including but not limited to personal computers, minicomputers, mainframes, workstations, network or distributed computing environments, separate or integrated computer platforms, or communicating with charged particle tools or other imaging devices, etc. Aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into the computing platform, such as a hard disk, optical read and / or write storage medium, RS1M, ROM, etc., such that it can be read by a programmable computer and can be used to configure and operate the computer to perform the processes described herein when the storage medium or device is read by the computer. In addition, the machine-readable code, or portions thereof, can be transmitted via wired or wireless networks. When such media include instructions or programs that implement the above-described steps in conjunction with a microprocessor or other data processor, the invention described herein includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention can also include the computer itself.

[0097] A computer program can be applied to input data to perform the functions described herein, thereby transforming the input data to generate output data stored in non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the transformed data represents physical and tangible objects, including a specific visual depiction of the physical and tangible objects generated on the display.

[0098] As described above, it is only a preferred embodiment of the present invention, and the present invention is not limited to the above-described embodiments. As long as the same means are used to achieve the technical effects of the present invention, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Within the scope of protection of the present invention, various different modifications and changes can be made to its technical solutions and / or implementation manners.

Claims

1. A multi-intelligent communication reinforcement learning agent path planning method based on warehousing environment, It is characterized in that The method comprises the following steps: S10, generating a warehouse environment map, obtaining the starting point, target point and obstacle information of each intelligent agent and inputting them into a neural network based on deep reinforcement learning; the observation value processing module obtains the characteristics of each intelligent agent according to the input observation value; S20, assigning the agents according to their own characteristics and using a greedy priority; selecting neighboring agents for each agent based on the adjacency matrix and according to the assigned priority, each agent receiving communication messages from the selected neighboring agents and forming neighbor characteristics; S30, forming a final feature according to the neighbor feature and the own feature, and inputting the final feature into a decision network module to generate a planned path; Wherein, the step S20 includes: S21, obtain the Manhattan distance from the current position of each agent to the target position, and set the reciprocal of the Manhattan distance as the priority of the agent; S22. When the Manhattan distances of the two agents are equal, setting the priority of the two agents according to the congestion degree; wherein the congestion degree represents the number of obstacles within the field of vision of the agents; S23, obtaining the communication range of the current agent, calculating the current position distance between the current agent and other agents; setting the other agents within the communication range of the current position as neighboring agents; S24, according to the set communication connection upper limit, select the agent with the highest priority from other agents whose current location distance exceeds the communication range as the neighbor agent; S24, determining whether the neighboring agent is an agent that has reached the target point; if so, disconnecting the communication connection between the neighboring agent that has reached the target point and the current agent; Wherein, the training of the neural network includes a deadlock detection module, and the deadlock detection module includes the following steps: Obtain the agent's own position in the past multiple time steps and detect it through the abnormal state detection module; If it is detected that the current agent stays at a non-target point for more than three time steps, the current agent is determined to be a stagnant abnormal agent; whether the congestion degree of the stagnant abnormal agent is greater than or equal to the congestion threshold is determined, and if so, the stagnant abnormal agent is pre-marked; If it is detected that the current agent has been moving back and forth between two positions for the past four time steps, the current agent is determined to be an abnormal wandering agent; whether the crowding degree of the abnormal wandering agent is greater than or equal to the crowding threshold is determined, and if so, the abnormal wandering agent is pre-marked; Determine whether there are two or more pre-marked agents within the communication range, and if so, determine that the pre-marked agents in the adjacency matrix are in a deadlock state; The parameters of the neural network are adjusted based on the information of the agent in the deadlock state.

2. The method according to claim 1, It is characterized in that In step S10: The input channels of the observed value processing module for the observed values include nine, and the nine input channels are channels 0 to 8 respectively; among them, channel 0 contains information of other agents within the field of view of the agent, channel 1 contains information of the obstacle map within the field of view of the agent, channel 2 contains information of the target point map within the field of view of the agent, channels 3 to 6 all contain information of the surrounding agent maps of the agent in the past multiple time steps, channel 7 contains information of the target point deviation map, and channel 8 contains information on whether the agent is in an abnormal state.

3. The method according to claim 2, wherein, for the step S10, the neural network includes: Four convolutional layers for processing the input information of channels 0 to 6; A first fully connected layer for processing the input information of channel 7; A second fully connected layer for processing the input information of channel 8; A third fully connected layer for splicing the output information of the four convolutional layers, the first fully connected layer and the second fully connected layer, and the third fully connected layer outputs the final features; A fourth fully connected layer and a fifth fully connected layer, wherein the final features are split into the fourth fully connected layer and the fifth fully connected layer; A sixth fully connected layer for combining the output information of the fourth fully connected layer and the fifth fully connected layer.

4. The method according to claim 3, wherein, The priority of the agent is represented by the following formula: ; In the formula, represents the number of other agents within the field of view of the agent; represents the Manhattan distance from the current position of the agent to the target position, represents the priority of the agent.

5. The method according to claim 3, wherein, In the step S30: The output of the fourth fully connected layer is a state value function that is only related to the state and not related to the action; The output of the fifth fully connected layer is an advantage function that is related to both the state and the action; The Q-value function of the neural network is obtained based on the state value function and the advantage function.

6. A computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a processor, the method according to any one of claims 1 to 5 is implemented.

7. A multi-agent communication reinforcement learning body path planning system based on a warehousing environment, wherein, it includes: A computer device, and the computer device includes the computer-readable storage medium according to claim 6.

Citation Information

Patent Citations

  • Moving state monitoring method and device based on Manhattan distance and computer equipment

    CN114722581A

  • Systems and methods for path planning with latent state inference and graphical relationships

    US20220147051A1