Multi-agent path planning method based on congestion perception and cache communication
By introducing congestion perception and cache communication mechanisms in multi-agent path planning, combined with deep reinforcement learning and distributed path planning methods, the low communication efficiency and congestion collision problems in large-scale tasks are solved, and efficient and stable multi-agent path planning is achieved.
Patent Information
- Application Number
- CN202510303433.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-20
AI Technical Summary
The existing multi-agent path planning methods are difficult to effectively perceive congestion in the environment in large-scale tasks, resulting in inefficient communication between agents and the method based on deep reinforcement learning lacks effective congestion collision solutions.
The multi-agent path planning method based on congestion perception and cache communication is adopted to build static and dynamic congestion information through local environmental information collection and processing, and combine observation encoder, collection module and Q network to realize distributed path planning. Each agent has an independent cache area for information storage and interaction, and trains the network model through reinforcement learning to optimize the agent's behavioral decisions.
It effectively reduces computing costs, improves the scalability of multi-agent path planning, reduces the occurrence of congestion, and ensures the stability of multi-agent collaboration in large-scale tasks.
Smart Images

Figure CN120178875A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of artificial intelligence, multi-agent systems, path planning, etc., and particularly relates to a multi-agent path planning method based on congestion perception and cached communication. Background Art
[0002] The multi-agent path planning task is an important part of large-scale robot systems. Currently, multi-agent path planning based on grid maps has been widely applied to intelligent warehousing and office robots, aiming to find non-conflicting paths for a group of agents from their starting points to their end points, while minimizing the time cost for each agent to complete the task. Solving the optimal solution to the multi-agent path planning task problem has been proven to be an NP-hard problem. Traditional search-based centralized solvers such as CBS and PCBS find the shortest path for each agent under constraints by establishing a high-dimensional constraint space. The search-based centralized solver will cause the computational amount to increase exponentially as the scale of agents expands, and it is difficult to scale to larger-scale tasks. At present, a large number of bounded sub-optimal solvers such as ECBS have emerged, which alleviate the problem of rapid growth of the computational amount by reducing the search space and computational time.
[0003] With the development of multi-agent systems, agents can independently configure environmental perception devices such as infrared scanners and vision sensors, and each agent can independently make appropriate behavior decisions based on the environmental information it perceives. The further development of deep learning has promoted the application of deep reinforcement learning in multi-agent systems. Currently, there are some distributed multi-agent path planning solvers based on reinforcement learning, which make independent behavior decisions within a limited local field of view by having each agent share the same policy. This learning-based solution method can be effectively extended to any number of agents. For multi-agent systems, agent communication in a dynamic environment is a major factor affecting the system effect. How to determine effective communication objects and reduce unnecessary communication frequencies as the task scale expands is a key problem. At the same time, there has been a lack of effective solutions to the congestion and collision problems among multi-agents based on deep reinforcement learning. Based on this, we need a multi-agent path planning method that can effectively perceive the congestion situation in the environment and reasonably communicate information between agents. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the present invention provides a multi-agent path planning method based on congestion perception and cached communication, which solves the above problems and effectively realizes multi-agent path planning in a grid map.
[0005] To achieve the above object, the present invention proposes the following technical solutions:
[0006] A multi-agent path planning method based on congestion awareness and caching communication, comprising the following steps:
[0007] Step 1, collection and processing of multi-agent local environment information;
[0008] Step 2, constructing congestion information according to the task environment;
[0009] Step 3, constructing a network model suitable for multi-agent path planning, the network model including an observation encoder, a pooling module, and a Q network;
[0010] Step 4, constructing a caching communication mechanism suitable for multi-agent systems, where each agent in the system has an independent caching area for storing communication information;
[0011] Step 5, training the network model established in Step 3 using a reinforcement learning strategy to obtain trained weights, and each agent in the system distributes and deploys the model and shares weight parameters;
[0012] Step 6: Obtain the local observable information of each agent according to Step 1 and combine it with the caching communication information in Step 4 to send it into their respective models for action decision-making, and obtain the next movement action of each agent.
[0013] Further, in the said Step 1, the local observable environment of the agent is a two-dimensional grid map, and the environmental information collected by it includes: the positions of obstacles, the positions of other agents, and the heuristic information in the four directions of up, down, left, and right within the range of ξ×ξ centered on the agent; the calculation method of the heuristic information is as follows:
[0014] Regard the information of each movement behavior as an independent channel of the same size as ξ×ξ. When the behavior associated with the channel that the agent takes can get closer to the target, the current position is marked as 1, otherwise it is marked as 0. At this time, the positions of the behavior channels associated with the optimal behavior should be as many as possible as 1 except for obstacles.
[0015] Still further, in the said Step 2, the construction of congestion information includes static congestion information and dynamic congestion information;
[0016] The static congestion information is based on the target position initialized by the task. Through breadth-first search, the average shortest distance of each grid cell relative to the reachable end point is obtained, and then the static congestion information of each cell in the map is obtained through standardized processing of the average shortest distance, that is, the initial potential congestion possibility;
[0017] The dynamic congestion information is based on the position information, collision situation, and potential movement direction of each agent during the task operation. By collecting the number Count(x,y) of each grid cell selected as the next potential position and the collision diffusion loss collion of each cellc ost(x, y) can be used to obtain the overall loss cost(x, y) = Count(x, y) + collion c ost(x, y), and the overall loss is normalized to obtain the dynamic congestion value.
[0018] Furthermore, in step 2, the calculation process of the collision diffusion loss of the dynamic congestion value in the congestion information is as follows: starting from the collision point, the collision congestion value is diffused outward, and the diffusion rules are as follows:
[0019] 1) If the diffused unit is an obstacle, stop diffusion;
[0020] 2) If the diffused unit is an agent, continue to diffuse K units to the four neighborhoods according to this agent;
[0021] 3) If the diffused unit is a narrow passage where only one agent can pass in one direction at a time, diffuse along the narrow passage to the end of the narrow passage;
[0022] 4) If the diffused unit is an ordinary empty unit, and the current relative to the collision point has diffused E units, then continue to diffuse K - E units to the four neighborhoods;
[0023] If the maximum diffusion times is K m , then this collision congestion value will continue for a number of time steps;
[0024]
[0025] where collion c ost(x, y) represents the collision congestion penalty value of the cell (x, y), V diffused represents the set of grid cells penalized due to collision, and S represents how many steps the current cell has diffused relative to the collision point.
[0026] In step 3, the observation encoder is a network module composed of four consecutive 3×3 convolutions, LeakyReLU activation functions, and a GRU unit. After receiving the environmental information and congestion information collected in steps 1 and 2, it finally outputs the observation encoding of the entire scene abstraction;
[0027] The aggregation module consists of an information encoder and an attention mechanism. The information encoder consists of a fully connected layer and a GRU unit. The information encoding output by it passes through the attention mechanism part and goes through a layer normalization, a multi-head attention calculation unit, a GRU unit, a layer normalization, a multi-layer perceptron, and a GRU unit in sequence, and finally outputs the decision vector for path planning;
[0028] The Q network contains two branches, a state-value function and an advantage function. Substantially, these two functions are two fully connected layers. The input decision features are respectively fed into the two functions to obtain the estimated value of the decision and the advantage value. The final Q value of each behavior can be calculated according to the following formula:
[0029]
[0030] where η is the network parameter shared by the state-value function and the advantage function, and α and β are the network parameters of the state-value function and the advantage function respectively. is the action space of the agent, s represents the current state, and w represents the currently taken behavior. The entire multi-agent path planning task is trained through the training network and the target network based on the Q value.
[0031] In step 4, the communication process of the cache communication mechanism is as follows:
[0032] Step 4.1: At time step t, agent a i will encode its own observation to obtain a message encoding through the message encoder
[0033] Step 4.2: Calculate the cosine similarity between and the previous broadcast message . If the similarity is less than δ, then will be broadcast to other agents within the FOV;
[0034] Step 4.4 Agent a i will receive the broadcast message from other agents a j within the FOV. After receiving the broadcast message j of a , a i will store it in its own receive cache and update the information validity value Val j = C. If the broadcast message of a j is not received, the information validity value of a j in the cache will be decremented by 1;
[0035] Step 4.5: When making a behavior decision, a i will take out all the cache information within the FOV with a validity value greater than 0 from the message receive cache and perform information fusion with its own message encoding through the message aggregation module to achieve multi-agent communication and cooperation;
[0036] Step 4.6: If a i does not receive the broadcast message of a j within the FOV at time t, and the information validity value of a j in the cache is 0, then it will actively send a message to aj Send a broadcast request, a j After receiving the request signal, re- Broadcast to other agents within the FOV.
[0037] In step 3, the calculation process of the final decision features obtained by the observation encoder and the aggregation module is as follows:
[0038] Step 3.1: The local observation information of the agent Pass through the observation encoder to obtain
[0039] Step 3.2: Observation encoding Convert to message encoding through the message encoder
[0040] Step 3.3: After layer normalization, retrieve the message encoding c broadcast by other agents within the local field of view received from the cache j , and aggregate the feature information scattered on h attention heads. The calculation formula is as follows:
[0041]
[0042] where d k is the dimension of the key, respectively represent the weight parameters for mapping the query, key, and value, represents the agent a at time t i the set of agents within the observable range, represents all attention heads;
[0043] Step 3.4: Aggregate features Combine with the original information Obtain through the GRU unit After layer normalization and multi-layer perceptron, and then combine with the original information through the GRU unit to obtain the final decision features The calculation formula is as follows:
[0044]
[0045] In step 5, the training process is as follows:
[0046] Step 5.1: Send the decision features into the state value function and the advantage function respectively, and calculate the Q value according to the following formula:
[0047]
[0048] Among them, Val is the state value function, Adv is the advantage function, η is the network parameter shared by the state value function and the advantage function, α and β are the network parameters of the state value function and the advantage function respectively. is the action space of the agent, s represents the current state, and w represents the current behavior taken.
[0049] Step 5.2: Based on the Q value calculated by the training network and the labeled reward estimated by the target network, calculate the n-step temporal difference error based on the mean square error to train this method. The loss calculation formula is as follows:
[0050]
[0051]
[0052] Among them is the reward obtained by agent a i at time step t, θ represents the parameters of the training network, and θ t represents the parameters of the target network, and γ is the discount factor.
[0053] In the said Step 6, the model constructed in Step 3 is deployed distributively in each independent agent sharing the same weight parameters trained in Step 5; each agent obtains the environmental information within the local observable field of view according to Step 1, constructs the congestion information using Step 2, conducts information interaction based on the communication mechanism described in Step 4, aggregates the information and sends it into the model to calculate the action Q value, and selects the action with the highest Q value as the next movement operation, effectively avoiding collisions and greatly improving the solution success rate of the multi-agent path planning task.
[0054] The beneficial effects of the present invention are as follows:
[0055] 1) The present invention utilizes the collection of local observation information, and realizes a distributed multi-agent path planning algorithm through the design of a neural network sharing weights combined with the communication process. Each agent can independently deploy a network model with the same weight parameters, make effective behavior decisions according to their respective observation information, effectively reduces the computational cost compared with the centralized planner, and has good scalability.
[0056] 2) Based on the distributed planner of mainstream deep reinforcement learning, the present invention incorporates congestion perception technology to encourage agents to avoid congested areas and reduce the occurrence of congestion; to avoid the problem of selecting effective communication objects for agents in large-scale tasks, the present invention proposes a caching communication mechanism, which circumvents the problem of selecting communication objects, and while effectively conducting multi-agent information interaction, ensures the solution stability of multi-agent cooperation tasks on a large scale. Description of the Drawings
[0057] Figure 1It is the flowchart of the method of the present invention.
[0058] Figure 2 It is the network structure diagram of the multi-agent path planning method based on congestion awareness and cache communication.
[0059] Figure 3 It is the schematic diagram of the map scene of the multi-agent path planning method based on congestion awareness and cache communication.
[0060] Figure 4 It is the schematic diagram of the local observable information of the multi-agent path planning method based on congestion awareness and cache communication. Among them, (a) represents the heuristic information (upper); (b) represents the heuristic information (lower), (c) represents the heuristic information (left), (d) represents the heuristic information (right), (e) represents the obstacle information, (f) represents the agent information, (g) represents the congestion information, and (h) represents the visited information.
[0061] Figure 5 It is the cache communication flowchart of the multi-agent path planning method based on congestion awareness and cache communication. Specific implementation manner
[0062] The present invention will be further described below in conjunction with the accompanying drawings of the specification, but the protection scope of the present invention is not limited thereto.
[0063] Reference Figures 1 to 5 , a multi-agent path planning method based on congestion awareness and cache communication, the method includes the following steps:
[0064] Step 1, acquisition and processing of multi-agent local environment information;
[0065] The local observable range of the agent is a rectangular area of 9×9. The observable environmental information includes eight parts, namely obstacle position information, agent position information, heuristic information of the four behaviors of up, down, left, and right, congestion information, and visited information. Figure 3 and Figure 4 respectively represent the map environment and observable environment information of the multi-agent path planning task; Figure 4 In the subgraphs (a) to (f) of, it shows the local information observable by the agent, which are the heuristic guidance of the four behaviors of up, down, left, and right, the relative positions of surrounding obstacles and other agents. Figure 4 The subgraph (h) of is used to represent the area that the agent has traveled through.
[0066] The process of the said Step 1 is as follows:
[0067] Step 1.1, starting from each initial target end point, traverse all reachable cells in the map using the breadth-first search algorithm, and record the shortest step lengths from each cell to all reachable end points;
[0068] Step 1.2, the calculation methods of the heuristic information corresponding to the four behaviors are as follows:
[0069] Regard the information of each movement behavior as an independent channel (matrix) of size 9×9. When the behavior associated with the channel enables the agent to get closer to the target, that is, when taking the action to move in the direction where the shortest step length to the current target point decreases, the current position is marked as 1, otherwise it is marked as 0. At this time, for the behavior channel associated with the optimal behavior, as many positions as possible should be 1 except for obstacles. An example can be seen in Figure 4 subgraphs (a) to (d) in
[0070] Step 2, construct congestion information according to the task environment;
[0071] As Figure 4 shown in subgraph (g), after determining the number of task agents and the task scenario, the congestion status of the abstracted environment can be obtained according to the known fixed information, such as the shortest step lengths from each unit to reachable end points obtained in Step 1.1;
[0072] The process of the above Step 2 is as follows:
[0073] Step 2.1, calculate the average shortest step length from each cell to all its reachable end points;
[0074] Step 2.2, standardize the average shortest step length as static congestion information;
[0075] Step 2.3, calculate the number of agents (the number of potential next positions selected) Count(x, y) in the four-neighborhood of each grid cell at each time step;
[0076] Step 2.4, at each time step, spread the collision loss collion_cost(x, y) from all colliding agents starting from the collision point. The spreading rules are as follows:
[0077] 1) If the cell to be spread is an obstacle, stop spreading;
[0078] 2) If the cell to be spread is an agent, continue to spread K cells to the four-neighborhood according to this agent;
[0079] 3) If the cell to be spread is a narrow passage (only one agent can pass through in one direction at a time), spread along the narrow passage to the end of the narrow passage;
[0080] 4) If the diffused unit is an ordinary empty unit and the current relative collision point has diffused E units, continue to diffuse K - E units to the four neighboring areas;
[0081]
[0082] where K m is the maximum diffusion times, and this collision congestion value will last for a time step of V diffused represents the set of grid cells to which the penalty value due to collision has spread, and S represents how many steps the current cell has diffused relative to the collision point. Calculate the total loss cost(x, y) = Count(x, y) + collion_cost(x, y), and the normalized total loss is regarded as the dynamic congestion value;
[0083] Step 2.5, weighted sum the static congestion value and the dynamic congestion value as the observable local congestion information of each cell in the environment congestion(x, y) = α·congestion st (x, y) + (1 - α)·congestion dyn (x, y);
[0084] Step 3, construct a network model suitable for multi-agent path planning, Figure 2 which represents the overall structure of the network model;
[0085] In the said Step 3, the calculation process of the final decision-making features obtained by the observation encoder and the pooling module is as follows:
[0086] Step 3.1, input the local observable environment information of the agent into the observation encoder, and the observation encoder is a network module composed of four consecutive 3×3 convolutions, LeakyReLU activation functions, and a GRU unit, which abstracts the environment information into an observation code;
[0087] Step 3.2, further convert the observation code into a message code through the message encoder, and the message code is composed of a fully connected layer and a GRU unit;
[0088] Step 3.3, send the message code into layer normalization, and then map the message code into Query through the weight matrix W q Subsequently, retrieve all communication information with valid values greater than 0 and corresponding agent objects still within the local observable field of view from the local cache, and map them into Key and Value through the weight matrices W K 、W V Perform attention calculation on the three of them on h attention heads:
[0089]
[0090]
[0091] The multi-head attention mechanism enables the model to fuse multi-agent information and perform task collaboration for auxiliary decision-making, where c j represents the communication information of the corresponding agent j in the cache, and d k is the dimension of Key, H represents the number of attention heads, is the output of the attention;
[0092] Step 3.4, the attention and the output of the message encoder are jointly passed through a GRU unit for context feature fusion, and its output continues to pass through layer normalization and a multi-layer perceptron, and then the output is combined with to jointly pass through a GRU unit to obtain the behavior decision vector
[0093]
[0094] In the above formula, LN represents layer normalization, and MLP represents a multi-layer perceptron, which consists of two fully connected layers;
[0095] Step 4, construct a cache communication mechanism suitable for multi-agent systems. Each agent in the system has an independent cache area to store communication information, Figure 5 showing the process of cache communication;
[0096] The process of Step 4 is as follows:
[0097] Step 4.1: At time step t, agent a i will encode its own observation to obtain the message encoding through the message encoder
[0098] Step 4.2: Calculate the cosine similarity between and the previous broadcast information . If the similarity is less than δ, then will be broadcast to other agents within the FOV;
[0099] Step 4.4 Agent a i will receive the broadcast information from other agents a j within the FOV. After receiving the broadcast information j of a a i will store it in its own receive cache and update the information validity value Val j = C. If the broadcast information of a j is not received, the information validity value of a j in the cache will be decremented by 1;
[0100] Step 4.5: In a i When making a behavior decision, all cached information within the FOV and with a valid value greater than 0 and its own message encoding will be taken out from the message receiving cache. Information fusion is performed through the message aggregation module to achieve multi-agent communication and cooperation;
[0101] Step 4.6: If at time t a i Not received within FOVa j Broadcast information, and a in the cache j If the effective value of the information is 0, it will automatically send a j Send a broadcast request, a j After receiving the request signal, Broadcast to other agents in FOV;
[0102] Step 5: The network model established in step 3 is trained using a reinforcement learning strategy to obtain trained weights. Each agent in the system deploys the model in a distributed manner and shares weight parameters.
[0103] Based on D3QN, the model training is performed, and the hidden layer output is input into the state value function Val and the advantage function Adv respectively. The training network is used to estimate the behavior and the target network is used to estimate the target Q value to improve the overestimation of the behavior value. The training steps are as follows:
[0104] Step 5.1: Set the decision features The state value function and advantage function are input respectively, and the Q value is calculated according to the following formula:
[0105]
[0106] Where Val is the state value function, Adv is the advantage function, η is the network parameter shared by the state value function and the advantage function, α and β are the network parameters of the state value function and the advantage function respectively. is the action space of the agent, s represents the current state, and w represents the current action.
[0107] Step 5.2, based on the Q value calculated by the training network and the label reward estimated by the target network, the n-step temporal difference error is calculated based on the mean square error to train this method. The loss calculation formula is as follows:
[0108]
[0109] in is an intelligent agent i The reward obtained at time step t, θ represents the parameters of the training network, θ trepresents the parameters of the target network and γ is the discount factor.
[0110] In step 5.3, the parameters of the training network are updated in the target network at a predefined number of iterations.
[0111] The D3QN-based deep reinforcement learning training method adopts a priority experience replay mechanism to accelerate the convergence of training and uses curriculum learning to stabilize the learning process. The curriculum learning starts from training a single agent on a 10×10 map to training 16 agents on a 40×40 map. The difficulty of the environment is increased (the map size is expanded and the number of agents is increased) only when the task success rate in the current environment reaches 0.9.
[0112] In step 6, each agent deploys a model with the same weight, and each agent makes an action decision based on the local observation information and cached communication information sent to the model to obtain the next movement action of each agent. During the training process, this method uses the curriculum learning proposed in step 5 to train in a map of a random environment, and at the same time combines the congestion information in step 2 to give the agent the ability to perceive congestion in the environment, significantly improving the agent's obstacle avoidance ability and task solving success rate. The cache communication method described in step 4 is used to effectively avoid the difficulty of selecting communication objects in multi-agent systems in large-scale tasks, so that this method still maintains a good solution success rate as the task scale expands.
[0113] The contents described in the embodiments of this specification are merely enumerations of implementation forms of the inventive concept and are for illustrative purposes only. The protection scope of the present invention should not be considered to be limited to the specific forms described in this embodiment, and the protection scope of the present invention also extends to equivalent technical means that can be thought of by ordinary technicians in this field based on the inventive concept.
Claims
1. A multi-agent path planning method based on congestion awareness and cache communication, characterized in that: The method comprises the following steps: Step 1: Collection and processing of multi-agent local environment information; Step 2, construct congestion information according to the task environment; Step 3, construct a network model suitable for multi-agent path planning; Step 4: construct a cache communication mechanism suitable for multi-agent systems; Step 5: Use reinforcement learning strategy to train the network model and distribute the shared model weight parameters; In step 6, the agent inputs the model output behavior decisions based on local observable information and cached communication content.
2. A multi-agent path planning method based on congestion perception and cache communication as claimed in claim 1, characterized in that: In step 1, the local observable environment of the agent is a two-dimensional grid map, and the collected environmental information includes: the position of obstacles within the range of ξ×ξ centered on the agent, the positions of other agents, and four heuristic information of up, down, left, and right. The calculation method of the heuristic information is as follows: The information of each moving behavior is regarded as an independent channel of the same size as ξ×ξ. When the agent takes the channel-associated behavior to get closer to the goal, the current position is marked as 1, otherwise it is marked as 0. At this time, the behavior channel associated with the optimal behavior should have as many positions as possible as 1 except for obstacles.
3. A multi-agent path planning method based on congestion perception and cache communication as described in claim 1 or 2, characterized in that: In step 2, the construction of congestion information includes static congestion information and dynamic congestion information; The static congestion information is based on the target position of the task initialization, and the average shortest distance of each grid cell relative to the achievable end point is obtained through breadth-first search, and then the static congestion information of each cell of the map is obtained by normalizing the average shortest distance, that is, the initial potential congestion possibility; The dynamic congestion information is based on the position information, collision situation and potential movement direction of each agent during the task operation. By collecting the number of each grid cell selected as the next potential position Count(x,y) and the collision diffusion loss collion of each cell c ost(x,y), we can get the total loss cost(x,y)=Count(x,y)+collion c ost(x,y), the overall loss is normalized to obtain the dynamic congestion value.
4. A multi-agent path planning method based on congestion perception and cache communication as claimed in claim 3, characterized in that: In step 2, the calculation process of the collision diffusion loss of the dynamic congestion value in the congestion information is as follows: the collision congestion value is diffused outward from the collision point, and the diffusion rule is as follows: 1) If the diffused unit is an obstacle, stop diffusing; 2) If the diffused unit is an agent, continue to diffuse K units to the four neighborhoods according to the agent; 3) If the diffused unit is a narrow road, agents can only pass in one direction at a time and diffuse along the narrow road to the end of the narrow road; 4) If the diffused cell is an ordinary empty cell, and the current relative collision point has diffused E cells, then continue to diffuse KE cells to the four neighborhoods; If the maximum diffusion number is K m , then the collision congestion value will continue time steps; Among them c ost(x,y) represents the collision congestion penalty of cell (x,y), V diffused It represents the set of grid cells to which the penalty value is diffused due to collision, and S represents how many steps the current cell has diffused relative to the collision point.
5. A multi-agent path planning method based on congestion perception and cache communication as claimed in claim 1 or 2, characterized in that: In step 3, the observation encoder is a network module composed of four consecutive 3×3 convolutions and LeakyReLU activation functions and a GRU unit, which finally outputs an abstract observation code of the entire scene after inputting the environmental information and congestion information collected in steps 1 and 2; The aggregation module is composed of an information encoder and an attention mechanism. The information encoder is composed of a fully connected layer and a GRU unit. The information encoding output by the information encoder passes through a layer normalization, a multi-head attention calculation unit, a GRU unit, a layer normalization, a multi-layer perceptron, and a GRU unit in the attention mechanism part, and finally outputs a decision vector for path planning. The Q network contains two branch state value functions and an advantage function. The two functions are essentially two fully connected layers. The input decision features are sent to the two functions to obtain the estimated value and advantage value of the decision. The final Q value of each behavior can be calculated according to the following formula: Where η is the network parameter shared by the state value function and the advantage function, α and β are the network parameters of the state value function and the advantage function respectively, It is the action space of the agent, s represents the current state, w represents the current action, and the entire multi-agent path planning task is trained through the training network and the target network according to the Q value.
6. A multi-agent path planning method based on congestion perception and cache communication as claimed in claim 1 or 2, characterized in that: In step 4, the communication process of the cache communication mechanism is: Step 4.1: At time step t, agent a i It will encode its own observations Get the message code through the message encoder Step 4.2: and last broadcast information Calculate the cosine similarity. If the similarity is less than δ, Broadcast to other agents within the FOV; Step 4.4 Agent a i Will receive messages from other agents in the FOV j The broadcast information of j Broadcast information After that, a i It will store it in its own receiving buffer and update the effective value of the information Val j =C, if a is not received j The broadcast information of j The effective value of the information is reduced by 1; Step 4.5: In a i When making a behavior decision, all cached information within the FOV and with a valid value greater than 0 and its own message encoding will be taken out from the message receiving cache. Information fusion is performed through the message aggregation module to achieve multi-agent communication and cooperation; Step 4.6: If at time t a i Not received within FOVa j Broadcast information, and a in the cache j If the effective value of the information is 0, it will automatically send a j Send a broadcast request, a j After receiving the request signal, Broadcast to other agents in FOV.
7. A multi-agent path planning method based on congestion perception and cache communication as claimed in claim 5, characterized in that: In step 3, the calculation process of the final decision feature obtained by the observation encoder and the aggregation module is: Step 3.1: Local observation information of the agent By observing the encoder Step 3.2: Observation Coding Convert to message encoding through message encoder Step 3.3: After layer normalization, the message encoding c broadcast by other agents in the local field of view is retrieved from the cache j , feature information is collected and distributed on h attention heads. The calculation formula is as follows: where d k is the dimension of the key, Represents the weight parameters for mapping query, key and value respectively, represents agent a at time t i The set of agents within the observable range, represents all attention heads; Step 3.4: Aggregate features With original information Obtained through the GRU unit After layer normalization and multi-layer perceptron, the final decision feature is obtained by combining the original information with the GRU unit. The calculation formula is as follows:
8. A multi-agent path planning method based on congestion perception and cache communication as claimed in claim 1 or 2, characterized in that: In step 5, the training process is as follows: Step 5.1: Decision features The state value function and advantage function are input respectively, and the Q value is calculated according to the following formula: Where Val is the state value function, Adv is the advantage function, η is the network parameter shared by the state value function and the advantage function, α and β are the network parameters of the state value function and the advantage function respectively. is the action space of the agent, s represents the current state, and w represents the current action; Step 5.2: Based on the Q value calculated by the training network and the label reward estimated by the target network, the n-step temporal difference error is calculated based on the mean square error to train this method. The loss calculation formula is as follows: in is an intelligent agent i The reward obtained at time step t, θ represents the parameters of the training network, θ t represents the parameters of the target network and γ is the discount factor.
9. A multi-agent path planning method based on congestion perception and cache communication as claimed in claim 1 or 2, characterized in that: In step 6, the model constructed in step 3 shares the same weight parameters trained in step 5 and is distributedly deployed in each independent agent; each agent obtains the environmental information within the local observable field of view according to step 1, constructs congestion information using step 2, and interacts with information based on the communication mechanism described in step 4. The information is collected and sent to the model to calculate the Q value of the behavior, and the behavior with the highest Q value is selected as the next moving operation, which effectively avoids collisions and greatly improves the success rate of solving multi-agent path planning tasks.