Distributed multi-agent path planning method based on reinforcement learning reward molding

By designing a reward shaping mechanism in multi-agent path planning, quantifying the impact of agent behavior on neighbors and training the policy network, the path conflict problem between agents is solved, and efficient collaborative path planning is achieved, suitable for large-scale and communication-constrained environments.

CN120274781APending Publication Date: 2025-07-08ZHEJIANG UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510420334.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the multi-agent path planning, it is difficult to design reasonable reward functions to promote cooperation among agents to avoid collisions, resulting in path conflicts and congestion problems, and the calculation complexity is high, making it difficult to apply to large-scale agent systems.

Method used

A reward shaping mechanism based on reinforcement learning is designed. By quantifying the impact of each agent's behavior on neighbors and integrating it into the reward function, a distributed reinforcement learning algorithm is used to train a path planning strategy network, and the agent's strategies are optimized using local observation data to realize collaborative path planning.

Benefits of technology

Without increasing computing overhead, the collaboration capability and path planning efficiency of multi-agent systems are improved, and are suitable for large-scale and communication-constrained environments, significantly reduce computing costs, and improve task success rate and path planning efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120274781A_ABST
    Figure CN120274781A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed multi-agent path planning method based on reinforcement learning reward molding, and belongs to the field of collaborative path planning. The method comprises the following steps: modeling a multi-agent path planning problem; by designing a reward shaping mechanism, the influence of agent behaviors on neighbors is quantified, and the influence is fused into a reward function, so that cooperative collision avoidance is realized while the agents are guided to maximize self accumulated rewards; a distributed reinforcement learning algorithm is adopted to train an intelligent agent, so that the intelligent agent can perform efficient path planning based on local observation data. According to the method, the conflict problem caused by local observation among the multiple agents is solved through reward modeling, the success rate and efficiency of the multi-agent path planning task are effectively improved, and low calculation overhead in the reasoning stage is kept. The method is suitable for multi-agent path planning tasks in a large-scale multi-agent scene, and is widely applied to the fields of transportation, logistics scheduling and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of multi-robot collaborative path planning, and mainly relates to a distributed multi-agent path planning method based on reinforcement learning reward shaping. Background Art

[0002] With the wide application of multi-agent systems, especially in the fields of traffic management, logistics distribution, automated production, etc., multi-agent path planning technology has become an important research direction. The core goal of multi-agent path planning is to plan a set of efficient and conflict-free paths for multiple agents (such as robots, autonomous driving vehicles, etc.), so that all agents can reach their respective target positions and avoid collisions with each other.

[0003] Traditional multi-agent path planning methods are generally divided into two categories: centralized and distributed. Centralized methods usually require a central controller to coordinate the path planning of all agents, and plan paths for each agent through global information to ensure the conflict-freeness of the paths. However, the computational complexity of centralized methods is relatively high, and with the increase in the number of agents, their solution efficiency decreases significantly, making it difficult to apply to scenarios with a large number of agents. Distributed methods avoid path conflicts by allowing each agent to make independent decisions. This method does not rely on a central controller and can effectively reduce the computational burden and improve the scalability of the system. In recent years, distributed path planning methods based on reinforcement learning have been widely studied. Especially in the decision-making of multi-agent systems, reinforcement learning can adapt to complex environments through autonomous learning and effectively handle the challenges brought by local observations. However, the application of reinforcement learning in multi-agent path planning still faces some problems. Especially in the design of the reward function, how to design a reasonable reward function to promote cooperation and collision avoidance among agents is a difficult point in current research.

[0004] Existing reinforcement learning methods usually train by designing independent reward functions for each agent, but this often leads agents to pursue the maximization of individual rewards, thus ignoring cooperation with other agents, and then problems such as path congestion and collisions occur. To solve this problem, the reward shaping method emerged. Reward shaping guides agents to not only focus on their own rewards but also consider cooperation and conflicts with other agents by designing a new type of reward function, thereby promoting the coordination and cooperation of multi-agent systems.

[0005] However, there is currently no reward shaping method designed specifically for the multi-agent path planning task. The existing reward shaping methods mainly include the following two categories: (1) Reward shaping based on the rewards of neighboring agents: For example, by introducing a "cooperation coefficient", the average reward of the agent's neighbors is weighted and fused with its own reward, thus guiding the agent to choose behaviors that are more conducive to cooperation. Although such methods are computationally simple, they introduce a large variance into the shaped reward function, which has an adverse effect on reinforcement learning. (2) Reward shaping based on game theory: For example, the Shapley value is used to calculate the impact of the agent's behavior on cooperation, and rewards are allocated according to the cooperation value. Although such methods can achieve stable cooperation in some applications, their computational complexity is relatively high. Especially in the multi-agent path planning problem, due to the large number of agents and the dynamic changes in neighbor relationships, directly applying such methods will significantly increase the computational overhead. Summary of the Invention

[0006] The object of the present invention is to provide a distributed multi-agent path planning method based on reinforcement learning reward shaping, aiming at the deficiencies of the existing technology, to solve the path conflict problem of agents under local observations, and at the same time improve the path planning efficiency and the collaborative ability of the system.

[0007] To solve the above problems, the present invention provides the following technical solutions:

[0008] The present invention provides a distributed multi-agent path planning method based on reinforcement learning reward shaping. The method is used for path planning of multi-agents from their respective starting points to their respective ending points and realizes collaborative collision avoidance. The method includes the following steps:

[0009] Step 1: Based on the scenario where the multi-agents are located, model the multi-agent path planning problem;

[0010] Step 2: Each agent calculates the current action a based on its own policy network and the current state s i and interacts with the environment using a. j i

[0011] Step 3: Design a reward shaping mechanism to quantify the impact of each agent's behavior on its neighbors and incorporate this impact into the reward function. The reward function is:

[0012]

[0013] where is the basic reward given by the environment to agent A i , is an index for evaluating the impact of the agent's behavior on its neighbors, and α is a weight coefficient;

[0014] Step 4: Adopt a distributed reinforcement learning algorithm to train the path planning policy network of the agent through local observation data;

[0015] Step 5: Evaluate the policy network trained in Step 4; Calculate the gradient of α according to the performance of the policy, and update α; Repeat Steps 4 and 5 until the update step size of α is less than the preset value;

[0016] Step 6: Apply the trained path planning policy network to the actual multi-agent path planning task, and achieve collaborative path planning of multi-agents while ensuring no conflict.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0018] (1) The reward shaping mechanism proposed by the present invention is only introduced in the training stage, and no additional computational cost and communication volume are required in the inference stage, which is applicable to real-time path planning of large-scale agent systems and multi-agent path planning in communication-constrained environments.

[0019] (2) The reward shaping mechanism proposed by the present invention is simple to calculate and will not significantly increase the computational overhead in the training process. Compared with the existing reward shaping methods, the present invention significantly reduces the computational cost while maintaining high cooperation efficiency.

[0020] (3) The reward function design proposed by the present invention is universal, and the agent behavior (such as obstacle avoidance priority, neighbor influence range, etc.) can be adjusted according to actual needs, which is applicable to the task planning requirements of various scenarios. Description of the Drawings

[0021] Figure 1 It is a schematic diagram of the principle of the distributed reinforcement learning algorithm adopted by the present invention.

[0022] Figure 2 It is a schematic diagram of a multi-agent environment in an embodiment of the method of the present invention.

[0023] Figure 3 It is a schematic diagram of the task success rate under different numbers of agents in an embodiment of the method of the present invention.

[0024] Figure 4 It is a schematic diagram of the average number of steps to complete the task under different numbers of agents in an embodiment of the method of the present invention. Detailed Embodiments

[0025] The present invention will be further described and illustrated below in conjunction with the detailed embodiments. The embodiments are only demonstrations of the present disclosure content and do not delimit the scope of limitation. The technical features of each embodiment of the present invention can be combined correspondingly without conflict.

[0026] The present invention proposes a distributed multi-agent path planning method based on reinforcement learning reward shaping. First, the multi-agent path planning problem is modeled to construct a reinforcement learning environment, clarifying the state, action, and state update method of the agent. Then, the agent calculates and selects the current action based on its own policy network and local observation data, and interacts with the environment. Next, a reward shaping mechanism is designed to quantify the impact of the agent's behavior on its neighbors and incorporate it into the reward function. After that, a distributed reinforcement learning algorithm is used to train the path planning strategy of the agent using local observation data. Finally, the trained strategy is applied to the actual task to achieve cooperative path planning of the agents without conflicts.

[0027] A distributed multi-agent path planning method based on reinforcement learning reward shaping proposed by the present invention specifically includes the following steps:

[0028] Step 1: Model the multi-agent path planning problem based on the scenario where the multi-agents are located;

[0029] In a specific embodiment of the present invention, Step 1 specifically includes: Agent A i The state at time t consists of the positions of the agents, the positions of the obstacles within the surrounding W×H range, and the distances from all grids within the W×H range to the end point of the agent. The action space of the agent contains 5 actions: "up", "down", "left", "right", and "stop". At each moment t, agent i selects an action a from the action space according to the locally observed state information and the policy i and executes it. After executing the action, the agent moves to a new position and obtains a reward value given by the environment

[0030] The reward function of the environment adopted by the present invention (i.e., the reward value ) can be in the form of the environment reward function in any existing path planning method, and the design of the environment reward function can be flexibly selected from the prior art according to the actual situation. These environment reward functions generally mainly include the following parts:

[0031] (1) Goal-reaching reward: When all agents successfully reach the target position, a significant positive reward R goal is given to each agent to encourage all agents to complete the task.

[0032] (2) Collision penalty: When an agent collides with a neighbor or an obstacle, a large negative reward R collision is given to guide the agent to avoid conflicts.

[0033] (3) Action penalty: For each action executed, a small negative reward R is given step , to encourage the agent to choose an efficient path to the goal as much as possible. If the agent stays at its goal point, then R step = 0.

[0034] Step 2: Each agent calculates the current action a i based on its own policy network and the current state s i , and interacts with the environment using a i ;

[0035] Step 3: Calculate the shaped reward function for agent A i according to the weight α, and the calculation method is:

[0036]

[0037] where is the base reward given by the environment to agent A i , is an index for evaluating the impact of the agent's behavior on its neighbors, and α is a weight coefficient used to balance the environmental base reward and the neighbor impact index

[0038] Calculate the impact on its neighbors for each agent A i after taking the action respectively

[0039] In this embodiment, is obtained by the following method:

[0040] For each agent A i within the neighbor range of agent A j , traverse all its possible actions a j ; In each possible action case, calculate the reward r j that agent A can obtain in the current state i and the action of A j , and select the maximum value as the evaluation index for this neighbor:

[0041]

[0042] where is the reward r j that agent A can obtain under the actions j and a j ;

[0043] Average the maximum reward values of all neighbors to obtain Agent A i The influence index of the current action on its neighbors

[0044]

[0045] Wherein, is the size of the neighbor set.

[0046] Step 4: As Figure 1 shown, adopt a distributed reinforcement learning algorithm to train the path planning policy network of the agent through local observation data;

[0047] The local observation data includes the position information of other agents and obstacles within a certain range around, and the distance from the surrounding grids to the target point of the agent; the agent infers the optimal path planning policy through the local observation data, and finally avoids conflicts between agents in path planning.

[0048] Furthermore, the distributed reinforcement learning algorithm is trained using a Deep Q-Network (DQN) or its variant algorithms, and the algorithm optimizes the policy by maximizing the Q i (s i ,a i ) value function:

[0049]

[0050] Wherein, the trajectory is the state-action sequence obtained by Agent A i when using the policy π i , represents the mathematical expectation over all possible trajectories τ i obtained by the policy π i , and γ is the discount factor.

[0051] The optimization objective of training is:

[0052]

[0053] Wherein, s i ′ is the next state, and a i ′ is the next action.

[0054] In a specific embodiment of the present invention, Step 4 specifically includes:

[0055] Step 4.1: Combine the shaped reward function with the state action of the agent and the state at the next moment into a set of experiences Store it into the global experience pool B.

[0056] Step 4.2: Sample b n groups of experiences from the experience pool B, and calculate the loss of the current network using these sampled experiences through the temporal difference method. The specific calculation method is as follows:

[0057]

[0058] Where:

[0059]

[0060] Step 4.3: Calculate the gradient of the loss function with respect to the network parameter θ through the backpropagation algorithm and update the network parameter using the gradient descent method:

[0061]

[0062] Where η is the learning rate.

[0063] Step 5: Evaluate the policy network trained in Step 4; calculate the gradient of α according to the performance of the policy, and update α; repeat Steps 4 and 5 until the update step of α is less than the preset value;

[0064] Specifically, with the fixed reward shaping weight parameter α, conduct several rounds of training and record the evaluation metrics of the policy (such as task success rate, average path length, etc.). Estimate the gradient of α through the following formula:

[0065]

[0066] Where J(α) represents the performance of the current policy (such as cumulative reward), and ∈ is a small perturbation value used to approximately calculate the gradient. Subsequently, update α according to the learning rate η α :

[0067]

[0068] Repeat the process from Step 4 to Step 5 to gradually optimize the policy of the agent and the reward shaping weight parameter α. When the update amplitude of α is less than the preset value, stop the training. At this time, the policy of the agent has converged, and the reward shaping parameter has also reached the optimization goal, and the training process is completed.

[0069] Step 6: Apply the trained path planning policy network to the actual multi-agent path planning task, and achieve collaborative path planning of multi-agents while ensuring no conflicts.

[0070] The mobile robots and sensors used in the method of the present invention are all conventional model devices; those skilled in the art can implement the method of the present invention through programming.

[0071] The effectiveness of the present invention is verified below in combination with a specific embodiment.

[0072] Consider a Figure 2 two-dimensional grid environment as shown, where 30% of the grids are randomly placed with obstacles (i.e., Figure 2 the gray grids in i ). There are multiple agents simultaneously performing tasks in the environment, and each agent has a different starting point start i and ending point goal Figure 2 in i A j and A i represent two agents, and their positions are the starting points at the current time step, and g j and g

[0073] Figure 3 show the comparison of success rates in 200 experiments before and after applying the method of the present invention. It can be clearly seen from the figure that when the method of the present invention is not used (directly using the reward function provided by the environment), the success rate of the strategy is relatively low, and in a two-dimensional grid environment of 40×40 size and an environment of 64 agents, the success rate can only reach about 60%. As Figure 3 can be seen, when the number of agents reaches a certain number or the map size is small, the success rate of the existing path planning method without using the method of the present invention will be significantly reduced, indicating that the existing environment reward function cannot handle the path collaborative planning task in the scenario of a relatively large number of multi-agents. After applying the path planning method based on reinforcement learning reward shaping of the present invention, the success rate is significantly improved, reaching about 90%. Under different experimental conditions, the success rate generally increases by about 10%, and in some scenarios, it even increases by more than 25%. This shows that the present invention can effectively solve the conflict problem of path planning of multi-agents in a complex environment and greatly enhance the ability of agents to cooperate to complete tasks.

[0074] Figure 4Shows the comparison of the average step lengths to complete tasks in 200 experiments before and after applying the method of the present invention. It can be clearly seen from the figure that when the method of the present invention is not used, the average step length for multiple agents to complete tasks is relatively long. After applying the path planning method based on reinforcement learning reward shaping of the present invention, the average step length is significantly optimized. Under different experimental conditions, the average step length is generally shortened by about 20 steps, and in some scenarios, it is even shortened by more than 30 steps. This shows that through the reward shaping mechanism of the present invention, agents are effectively guided to learn more efficient path planning strategies. When facing a complex environment, agents can better consider the cooperation relationship between themselves and other agents, avoid unnecessary path choices, and thus quickly reach the target position with fewer steps, greatly improving the overall task execution efficiency.

[0075] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the claims of the present invention. All of these fall within the protection scope of the present invention.

Claims

1. A distributed multi-agent path planning method based on reinforcement learning reward shaping, which is used for multi-agent path planning from their respective starting points to their respective ending points and realizes cooperative collision avoidance, and is characterized in that, The method includes the following steps: Step 1: Model the multi-agent path planning problem based on the scenario where the multi-agents are located; Step 2: Each agent calculates the current action a based on its policy network and the current state s i and interacts with the environment using a i , i ​ Step 3: Design a reward shaping mechanism to quantify the impact of each agent's behavior on its neighbors and incorporate this impact into the reward function, where the reward function is: Among them, is the basic reward given by the environment to agent A i , is an index for evaluating the impact of the agent's behavior on its neighbors, and α is the weight coefficient; Step 4: Adopt a distributed reinforcement learning algorithm to train the path planning policy network of the agents through local observation data; Step 5: Evaluate the policy network trained in Step 4; calculate the gradient of α according to the performance of the policy network and update α; repeat Steps 4 and 5 until the update step size of α is less than the preset value; Step 6: Apply the trained path planning policy network to the actual multi-agent path planning task to achieve collaborative path planning of multi-agents while ensuring no conflicts.

2. The method according to claim 1, wherein The said Step 1 includes: Build a reinforcement learning environment for multi-agent path planning in a two-dimensional grid map. Among them, the state s of agent i i consists of the positions of all agents, the positions of obstacles, and the distances from all grids within the range of W×H to the end point of agent i within the surrounding range of W×H; The action space of the agent includes 5 actions: "up", "down", "left", "right", and "stop"; At each moment t, agent i selects an action a from the action space according to the locally observed state information and the policy ; after executing the action, the agent moves to a new position and obtains a reward value given by the environment i ​ 3. The method according to claim 1, wherein The metric of the influence of the agent's behavior on its neighbors is obtained by the following method: For each agent A i within the neighbor range at time t j , traverse all its possible actions a j ; in each possible action case, calculate the reward r j that agent A can obtain in the current state i and the action of A j , and select the maximum value as the evaluation index for this neighbor: Among them, is the agent A j in action and a j the reward that can be obtained Average the maximum reward values of all neighbors to obtain Agent A i The influence index of the current action on its neighbors Among them, is the size of the neighbor set of agent A i at time t.

4. The method according to claim 1, wherein In step 4, the distributed reinforcement learning algorithm is trained using a Deep Q-Network (DQN) or its variant algorithm, and the algorithm optimizes the policy by maximizing the Q i (s i ,a i ) value function: Among them, the trajectory is the state-action sequence obtained by the agent A i when using the policy π i , and denotes the mathematical expectation over all possible trajectories τ i obtained by the policy π i , where γ is the discount factor.

5. The method according to claim 4, characterized in that, In Step 4, the optimization objective for training is: where s i ' is the next state, a i ' is the next action.

6. The method according to claim 4, wherein In Step 4, the local observation data includes the position information of other agents and obstacles within a certain range around, as well as the distance from the surrounding grids to the agent's target point; the agent infers the optimal path planning strategy through the local observation data and finally avoids conflicts between agents in path planning.

7. The method according to claim 1, characterized in that The said Step 4 specifically includes: Step 4.1: Combine the shaped reward function with the state of the agent action and the state at the next moment to form a set of experiences and store them in the global experience pool B; Step 4.2: Sample b n sets of experiences from the experience pool B, and calculate the loss of the current network using these sampled experiences. The specific calculation method is as follows: Where: Step 4.3: Calculate the gradient of the loss function with respect to the network parameter θ using the backpropagation algorithm and update the network parameters using the gradient descent method: Where η is the learning rate.

8. The method according to claim 1, wherein The said Step 5 specifically includes: With the reward function shaping weight parameter α fixed, conduct several rounds of training and record the evaluation metrics of the policy; the evaluation metrics are the task success rate or the average path length; Estimate the gradient of α through the following formula: Among them, I(α) represents the performance of the current strategy, ∈ is a small perturbation value used for approximate gradient calculation; subsequently, according to the learning rate η α Update α: When the update amplitude of α is less than the preset value, stop training; At this time, the policy of the agent has converged, the reward shaping parameter has also reached the optimization objective, and the training process is completed.

Citation Information

Cited By

  • Automatic collaborative decision-making method, system and equipment for enterprise resource planning, medium and product

    CN120494762A

  • Unmanned aerial vehicle route automatic planning system based on AI identification

    CN121115888A

  • Model training method, electronic equipment, readable storage medium and program product

    CN121279350A

  • Model training method, electronic device, readable storage medium, and program product

    CN121279350B