A Design Method for Path Planning Reward Function Based on Deep Reinforcement Learning
By planning the reward function based on deep reinforcement learning, the cost of escape speed obstacles of the agent is calculated and the reward value is weighted to adjust the reward value, which solves the problem of insufficient collision avoidance performance and stability in multi-agent systems, and achieves efficient collision avoidance and rule compliance in dynamic environments.
Patent Information
- Application Number
- CN202411662518.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-20
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-11-20
AI Technical Summary
When planning paths in multi-agent systems, especially in dynamic environments and multi-agent collaboration scenarios, it is difficult to ensure high stability and good collision avoidance performance, and the existing reward mechanisms are insufficient in generalization capabilities in different tasks or environments.
A path planning reward function based on deep reinforcement learning is designed. By calculating the Euclidean distance between the agent's expected speed and the current speed, distinguishing dynamic and static obstacles, calculating the cost of escape speed obstacles, and using importance factor weighting, adjusting the reward value to affect the agent's collision avoidance behavior, complying with relevant operating rules such as COLREGs.
The convergence speed and generalization ability of deep reinforcement learning algorithms are improved, ensuring that the agent maintains good collision avoidance performance and stability in complex environments, adapt to diverse tasks and environmental changes, and avoid overfitting and local optimization problems.
Smart Images

Figure CN119575965B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of path planning, and particularly to a design method for a path planning reward function based on deep reinforcement learning. Background Art
[0002] With the rapid development of artificial intelligence and automation technologies, unmanned technologies have been widely applied in many fields such as marine monitoring, aerial patrol, and logistics transportation. Among them, multi-agent systems (MAS) have become an important research direction in intelligent applications due to their characteristics of collaborative operation, flexible task allocation, and high degree of automation. A multi-agent system can include various types of autonomous devices such as unmanned surface vehicles (USVs), unmanned aerial vehicles (UAVs), and unmanned vehicles, and has shown great potential when performing tasks such as environmental monitoring, search and rescue missions, and cargo transportation.
[0003] International related research shows that in practical applications, a key problem faced by the agent system is how to perform effective path planning and collision avoidance operations. Especially in the scenario of multi-agent collaborative work, the complexity of path planning is further increased. Planning a path not only needs to consider the task efficiency of each agent, but also needs to ensure that they avoid interfering with and colliding with each other during task execution, and at the same time follow the operation rules of a specific field. For example, a multi-USV system applied at sea should strictly comply with the International Regulations for Preventing Collisions at Sea (COLREGs).
[0004] The Chinese patent "CN 111880549 B" proposes an optimization method for the reward function of deep reinforcement learning for the path planning of unmanned surface vessels (USVs). This method preprocesses the environmental information and dynamically adjusts the navigation strategy using the distances between the USV, obstacles, and the target point. The reward mechanism includes setting a "reward domain" near the target point and a "hazard domain" near the obstacles, and dynamically adjusts the reward or penalty values according to the number of times the USV reaches the target point and the number of times it collides with obstacles. This method aims to accelerate the convergence speed of the deep reinforcement learning model, help the unmanned surface vessel avoid obstacles more effectively, and optimize the path planning. The patent improves the convergence speed of the learning algorithm through complex parameter and threshold adjustments, but this may lead to overfitting and stability problems. Its reward and penalty mechanism is set based on the distances between the USV, the target point, and the obstacles, but these distance thresholds may not be applicable in different tasks or environments. The optimization method of this patent mainly targets specific simulation scenarios, and its generalization ability still needs to be further verified for environmental changes in practical applications, especially scenarios with dynamic obstacles or multi-agent cooperation. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a design method for a path planning reward function based on deep reinforcement learning in view of the deficiencies of the above-mentioned prior art, which can ensure that the intelligent agent can still make effective collision avoidance decisions in a complex dynamic environment, can ensure high stability, good collision avoidance performance, and comply with relevant operating rules.
[0006] A design method for a path planning reward function based on deep reinforcement learning includes the following steps:
[0007] Step 1: Calculate the expected speed of the intelligent agent;
[0008] The expected speed of the intelligent agent is shown in the following formula:
[0009]
[0010] where v des is the expected speed of the intelligent agent, (x1, y1) are the coordinates of the current position of the intelligent agent, (x2, y2) are the coordinates of the target point, is the unit vector from the current position (x1, y1) of the intelligent agent to the target point (x2, y2), v max is the maximum speed of the intelligent agent's travel, ||v max || is the modulus of the maximum speed of the intelligent agent's travel;
[0011] Step 2: Calculate the Euclidean distance between the current travel speed and the expected speed of the intelligent agent, define the reward formula, and assign the calculation result obtained from the reward formula as the reward value to the intelligent agent. The reward formula is shown in the following formula:
[0012]
[0013] where r diff is the calculation result, v is the current driving speed of the agent, and d(v, v des ) is the Euclidean distance between the current driving speed of the agent and the desired speed;
[0014] Step 3: Classify the obstacles into dynamic obstacles and static obstacles by determining whether the obstacles have a speed;
[0015] Step 4: Define the speed obstacles generated by dynamic obstacles and static obstacles within a certain range around the agent respectively;
[0016] Step 5: Calculate the cost values of the lowest escape speed obstacles when the agent faces two types of obstacles respectively, and take the negative of the cost values as the reward values to affect the collision avoidance behavior of the agent;
[0017] Set the function This function calculates the positional relationship of point p with respect to the directed line by calculating the cross product; where is the directed line, (a x , a y ) is the starting point of the directed line , (u x , u y ) is the direction of the directed line , and (p x , p y ) is the position coordinate of point p;
[0018] The calculation formulas for the cost values of the lowest escape speed obstacles when the agent faces two types of obstacles are shown as follows:
[0019]
[0020] where r voca represents the cost value of the lowest escape speed obstacle when the agent faces dynamic obstacles, and r vocs represents the cost value of the lowest escape speed obstacle when the agent faces static obstacles; represents the two sides of the speed obstacle, represents the cut-off line, r c is the collision radius of the agent, O c is the center of the circle at the top of the speed obstacle, and M is the point on the static obstacle closest to the current driving speed of the agent;
[0021] Step 6: Use the importance factor to weight the cost value of the minimum escape speed obstacle when the agent faces dynamic obstacles, obtaining the weighted cost;
[0022] The calculation formula of the importance factor is shown as follows:
[0023]
[0024] where β is the importance factor, and ||d|| represents the relative distance between agents;
[0025] Weight the importance factor to the cost value of the minimum escape speed obstacle when the agent faces dynamic obstacles, dynamically adjusting the influence degree of speed obstacles at different distances on the agent's reward value;
[0026] The weighted cost is shown as follows:
[0027] r voca-β = βr voca
[0028] where r voca-β represents the weighted cost;
[0029] Step 7: Confirm the safest speed adjustment direction of the agent in the current state through the gradient of the weighted cost, and give the agent rewards and punishments to guide the agent to choose a suitable direction when avoiding obstacles;
[0030] The magnitude of the outer product of the safest speed adjustment direction e of the agent in the current state and the agent's current driving speed v can reflect the direction of e relative to v, that is, when v×e > 0, it means e is to the left of v, and when v×e < 0, e is to the right of v; when hoping that the agent learns a strategy of avoiding collisions to the right, when v×e < 0, a reward is given at this time to encourage the agent to avoid collisions from the right side, and when v×e > 0, a punishment is given to punish the agent for avoiding collisions to the left side;
[0031] Step 8: Give different amounts of punishment according to the type of obstacle when the agent collides;
[0032] Step 9: Give corresponding rewards according to the number of agents reaching the end point;
[0033] The reward value is defined as follows:
[0034] r task =(2·N arrive -N total )·5
[0035] where r task is the reward value, N arrive is the number of agents reaching the target point in an episode, Ntotal is the total number of agents in the current scenario.
[0036] The beneficial effects of adopting the above technical solutions are as follows: By calculating the cost of escaping the speed obstacle, the present invention provides a multi-agent path planning reward mechanism with high stability and strong collision avoidance performance. Compared with the overfitting and stability problems existing in the prior art, the present invention can, to a certain extent, meet specific collision avoidance rules similar to COLREGs through parameter adjustment, thereby enhancing the adaptability of the system in complex scenarios. Based on the calculation of the cost of escaping the speed obstacle, this reward mechanism effectively influences the collision avoidance decision of the agent, enabling it to maintain good collision avoidance performance in a complex dynamic environment and avoid falling into local optima. This design not only improves the convergence speed of the deep reinforcement learning algorithm but also enhances the generalization ability of the algorithm, especially in multi-agent cooperation and dynamic obstacle environments. In this way, the present invention demonstrates stronger stability and flexibility in practical applications, can adapt to diverse task requirements and environmental changes, and overcomes the limitations of traditional methods. Brief Description of the Drawings
[0037] Figure 1 is a flowchart of a design method for a path planning reward function based on deep reinforcement learning provided by an embodiment of the present invention;
[0038] Figure 2 A geometric relationship diagram of speed obstacles provided by an embodiment of the present invention, where (a) is a schematic diagram of agents A and B, (b) is a schematic diagram of the speed obstacle of agent B with respect to A, (c) is a schematic diagram of agent A and static obstacle O, and (d) is a schematic diagram of the speed obstacle of static obstacle O with respect to A;
[0039] Figure 3 is a schematic diagram of collision avoidance principles in a meeting situation stipulated by COLREGs provided by an embodiment of the present invention, where (a) is a schematic diagram of collision avoidance principles in an overtaking scenario, (b) is a schematic diagram of collision avoidance principles in a head-on scenario, (c) is a schematic diagram of collision avoidance principles in a left crossing scenario, and (d) is a schematic diagram of collision avoidance principles in a right crossing scenario;
[0040] Figure 4 is a schematic diagram of an unmanned ship overtaking a target ship in an overtaking scenario provided by an embodiment of the present invention, where (a) is a trajectory diagram of the overtaking scenario at time t1, (b) is a trajectory diagram of the overtaking scenario at time t2, and (c) is a trajectory diagram of the overtaking scenario at time t3;
[0041] Figure 5 is a schematic diagram of collision avoidance between USV1 and USV2 in a head-on scenario provided by an embodiment of the present invention, where (a) is a trajectory diagram of the head-on scenario at time t1, (b) is a trajectory diagram of the head-on scenario at time t2, and (c) is a trajectory diagram of the head-on scenario at time t3;
[0042] Figure 6 Collision avoidance schematic diagrams of USV1 and USV2 in the left crossing scenario provided by the embodiments of the present invention, where (a) is the trajectory diagram of the left crossing scenario at time t1, (b) is the trajectory diagram of the left crossing scenario at time t2, and (c) is the trajectory diagram of the left crossing scenario at time t3;
[0043] Figure 7 Collision avoidance schematic diagrams of USV1 and USV2 in the right crossing scenario provided by the embodiments of the present invention, where (a) is the trajectory diagram of the right crossing scenario at time t1, (b) is the trajectory diagram of the right crossing scenario at time t2, and (c) is the trajectory diagram of the right crossing scenario at time t3;
[0044] Figure 8 Collision avoidance schematic diagram of an on - ship unmanned aerial vehicle in the on - ship unmanned aerial vehicle test scenario I provided by the embodiments of the present invention;
[0045] Figure 9 Collision avoidance schematic diagram of an on - ship unmanned aerial vehicle in the on - ship unmanned aerial vehicle test scenario II provided by the embodiments of the present invention. Detailed implementation manners
[0046] The following combines the accompanying drawings and embodiments to further describe in detail the detailed implementation manners of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0047] In this embodiment, a design method for a path - planning reward function based on deep reinforcement learning includes the following steps:
[0048] Step 1: Calculate the expected speed of the agent; the calculation is as shown in formula (1):
[0049]
[0050] where v des is the expected speed of the agent, (x1, y1) are the coordinates of the current position of the agent, (x2, y2) are the coordinates of the target point, is the unit vector from the current position (x1, y1) of the agent to the target point (x2, y2), v max is the maximum speed of the agent's travel, ||v max || is the modulus of the maximum speed of the agent's travel;
[0051] Step 2: Calculate the Euclidean distance between the current travel speed and the expected speed of the agent, define the reward formula, and assign the calculation result obtained from the reward formula as the reward value to the agent. The reward formula is as shown in formula (2):
[0052]
[0053] Among them, r diff is the calculation result, v is the current driving speed of the agent, and d(v, v des ) is the Euclidean distance between the current driving speed of the agent and the desired speed;
[0054] Step 3: Classify the obstacles into dynamic obstacles and static obstacles by judging whether the obstacles have speed;
[0055] Step 4: Use the optimal interactive collision avoidance algorithm to define the velocity obstacles generated by dynamic obstacles and static obstacles within a certain range around the agent respectively;
[0056] When the obstacle is a dynamic obstacle, as shown in Figure 2 (a), assume there are agent A and agent B, and the central position coordinates of the agents are p A and p B respectively, and the collision radii are r A and r B respectively. Within the time window τ, the geometric relationship of the velocity obstacle of agent B to agent A is shown in Figure 2 (b). Figure 2 In (b), assume the desired speeds of A and B are and respectively. Then u is the shortest vector from the relative desired speed to . is the optimal interactive collision avoidance speed set of A to B, and the formula for the velocity obstacle of agent B to agent A is defined as shown in formula (3):
[0057]
[0058] Among them, D(p B -p A , r A +r B ) represents a circle with p B -p A as the center and r A +r B as the radius, and t represents time;
[0059] When the obstacle is a static obstacle, as shown in Figure 2 (c), assume there is a static obstacle O and an agent A with a collision radius of r A . The velocity obstacle is shown in the figure as . Within the time window τ, the geometric relationship of the velocity obstacle of the static obstacle O to the agent A is shown in Figure 2 (d). is the optimal interactive collision avoidance speed set of A to O, and the formula for the velocity obstacle of the static obstacle O to the agent A is defined as shown in formula (4):
[0060]
[0061] Among them, D(p A , r A ) represents a circle centered at p A with radius r A ;
[0062] Step 5: Calculate the cost values of the lowest escape speed obstacles when the agent faces two types of collision obstacles respectively, and take the negative of the cost values as the reward values to affect the collision avoidance behavior of the agent;
[0063] Set the function This function calculates the positional relationship of point p relative to the directed line by calculating the cross product; among them, is the directed line, (a x , a y ) is the starting point of the directed line , (u x , u y ) is the directed vector used to determine the direction of the directed line , and (p x , p y ) are the position coordinates of point p;
[0064] The calculation of the cost values of the lowest escape speed obstacles when the agent faces two types of collision obstacles is shown in formulas (5) and (6):
[0065]
[0066] Among them, r voca represents the cost value of the lowest escape speed obstacle when the agent faces a dynamic obstacle, and r vocs represents the cost value of the lowest escape speed obstacle when the agent faces a static obstacle. As Figure 1 shows, r voca and r vocs respectively consider the interaction situation between agents and the situation of the agent avoiding static obstacles; represents that the two sides (VOlegs) of the velocity obstacle VO include leg1 and leg2, represents the cutoff line (cutoffline), r c is the collision radius of the agent, O c is the center of the circle at the top of the velocity obstacle VO, that is, (p A - p B ) / τ, and M is the point on the static obstacle closest to the current driving speed v of the agent;
[0067] The cost values r of the lowest escape speed obstacle when the agent faces two types of collision obstacles voca and r vocs are made negative and given as reward values to the agent.
[0068] Step 6: Use the importance factor to weight the cost value of the lowest escape speed obstacle when the agent faces dynamic obstacles to obtain the weighted cost;
[0069] The calculation formula of the importance factor is shown in formula (7):
[0070]
[0071] where β is the importance factor, calculated using the exponential function, ||d|| represents the relative distance between agents, used to adjust the size of β, the smaller ||d|| is, the larger β will be;
[0072] The importance factor β is weighted to the cost value r of the lowest escape speed obstacle when the agent faces dynamic obstacles voca to dynamically adjust the influence degree of speed obstacles at different distances on the agent's reward value;
[0073] The weighted cost is shown in formula (8):
[0074] r voca-β = βr voca (8)
[0075] where r voca-β represents the weighted cost;
[0076] Step 7: Confirm the safest speed adjustment direction of the agent in the current state through the gradient of the weighted cost, and give rewards and punishments to the agent to guide the agent to choose the appropriate direction when avoiding collisions;
[0077] The magnitude of the outer product of the safest speed adjustment direction e of the agent in the current state and the agent's current driving speed v can reflect the direction of e relative to v, that is, when v×e > 0, it means e is to the left relative to v, and when v×e < 0, e is to the right relative to v; if it is desired that the agent learns a strategy to avoid collisions to the right, then when v×e < 0, a reward of 0.2 is given at this time, and when v×e > 0, a punishment of -0.2 is given to punish the agent for avoiding collisions to the left; this can be flexibly adjusted according to specific collision avoidance rules;
[0078] Step 8: Give different amounts of punishment according to the type of obstacle when the agent collides;
[0079] When colliding with other dynamic obstacles, a punishment of -8 is given, and when colliding with static obstacles, a punishment of -5 is given.
[0080] Step 9: Give corresponding rewards according to the number of agents reaching the end point;
[0081] The reward value is defined as shown in formula (9):
[0082] r task = (2·N arrive - N total )·5 (9)
[0083] where r task is the reward value, N arrive is the number of agents reaching the target point in an episode, and N total is the total number of agents in the current scenario.
[0084] Next, in combination with a specific application, policy training will be carried out through the multi-agent deep reinforcement learning framework MA-POCA to illustrate the method for designing a multi-agent path planning reward function based on deep reinforcement learning in the embodiments of the present invention.
[0085] Embodiment 1: Collision avoidance of multiple unmanned ships satisfying the International Regulations for Preventing Collisions at Sea (COLREGs)
[0086] COLREGs altogether contain 38 rules, but not all rules can be transformed into mathematical problems. Therefore, in this embodiment, a subset of COLREGs is considered: Rule 8, Rules 13-17. For the convenience of subsequent discussion, the ship encounter situations in Rules 13-15 in this subset are summarized, as shown in Table 1.
[0087] OS represents the own ship, and TS represents the target ship. In the overtaking situation, as Figure 3 (a) shows, the own ship can choose to turn right or left, and the overtaken target ship can sail normally; when the ships meet head-on, as Figure 3 (b), both the own ship and the target ship should adjust their headings to the right to avoid collision; in the crossing situation, as Figure 3 (c)(d) show, if there is a target ship on the left side of the own ship, the own ship has the right of way; if there is a target ship on the right side of the own ship, the own ship does not have the right of way. The ship without the right of way should give way to the ship with the right of way.
[0088] Table 1 Collision avoidance behaviors in encounter situations stipulated by COLREGs
[0089]
[0090] In addition, Rule 8 requires that a ship, in any situation, once it detects a risk of collision, must take appropriate and sufficient collision avoidance measures, including changing course or speed; Rule 16 stipulates that when a ship has the obligation to keep out of the way, it needs to take timely collision avoidance actions; Rule 17 stipulates that if the ship with the right of way observes that the give-way ship has not taken effective collision avoidance measures, then that ship needs to take actions to avoid collision.
[0091] In Test Scenario I of Embodiment 1, it is an overtaking scenario. As Figure 4 shown, the unmanned ship (USV) and the target ship are sailing on the same route. In order to reach the destination faster, the USV needs to overtake the other party. According to the provisions of the COLREGS, under the condition of ensuring safety, the USV needs to complete the overtaking behavior from the left or right side of the target ship.
[0092] In Test Scenario II of Embodiment 1, it is a head-on scenario. As Figure 5 shown, USV1 and USV2 form a head-on situation. At this time, both parties are give-way ships and both need to take right-turning measures to ensure safe navigation.
[0093] In Test Scenario III of Embodiment 1, it is a crossing scenario. As Figure 6 shown, from the perspective of USV2, USV1 and the ship itself form a situation of left crossing. The ship itself has the right of way and in principle only needs to maintain the speed and course. However, considering safety, USV2 also takes certain collision avoidance actions, which comply with the provisions of the COLREGS. Figure 7 In [the figure], USV2 and the target ship USV1 form a situation of right crossing. In this case, USV1 has the right of way and USV2 needs to actively avoid USV1.
[0094] Through the above three test scenarios, it can be seen that the present invention can meet the ship collision avoidance rules stipulated by the COLREGs.
[0095] Embodiment 2: Autonomous taxiing during the take-off and landing stage of an on-board unmanned aerial vehicle
[0096] In Embodiment 2, simulation experiments were carried out in the application scenario of an on-board unmanned aerial vehicle. As Figure 8 and Figure 9 shown, there are a large number of static obstacles in both Test Scenarios I and II. In Scenario I, three unmanned aerial vehicles taxi autonomously at the same time, and in Scenario II, four unmanned aerial vehicles taxi autonomously at the same time. 50 tests were carried out in each scenario. The test results are shown in Table 2. It can be seen that the method for designing the path planning reward function based on deep reinforcement learning proposed by the present invention can train a strategy with good collision avoidance performance and can improve the task execution efficiency and safety in the application scenario of multi-agent collaborative tasks.
[0097] Table 2 Results of Collision Avoidance Performance Test
[0098]
[0099] The object of the present invention is to provide a method for optimizing the path planning reward function applicable to multi-agent systems. By calculating the cost for an agent to escape from the VO and using it as the reward value to influence the collision avoidance behavior of the agent, it is ensured that the agent can still make effective collision avoidance decisions in a complex dynamic environment, capable of ensuring high stability, good collision avoidance performance, and compliance with relevant operation rules such as COLREGs, etc. In this way, overfitting and stability problems can be avoided, and the generalization ability of the algorithm in multi-agent cooperation scenarios can be improved, making this method more adaptable to diverse tasks and environmental changes in practical applications. The method proposed by the present invention is not only applicable to USVs, but can also be applied to other types of agents, such as the scheduling of on-board drones on rescue ships and scientific research ships, improving the task execution efficiency and safety in a wide range of application scenarios of multi-agent collaborative tasks.
[0100] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope defined by the claims of the present invention.
Claims
1. A design method for a path planning reward function based on deep reinforcement learning, characterized in that: Including the following steps: Step 1: Calculate the expected speed of the agent; Step 2: Calculate the Euclidean distance between the current driving speed and the expected speed of the agent, define a reward formula, and assign the calculation result obtained from the reward formula as the reward value to the agent; Step 3: Classify the obstacles into two categories, dynamic obstacles and static obstacles, by judging whether the obstacles have a speed; Step 4: Use the optimal interaction collision avoidance algorithm to define the speed obstacles generated by dynamic obstacles and static obstacles within a certain range around the agent respectively; Step 5: Calculate the cost values of the lowest escape speed obstacles when the agent faces the two types of obstacles respectively, and take the negative of the cost values as the reward values to affect the collision avoidance behavior of the agent; Step 6: Use the importance factor to weight the cost value of the lowest escape speed obstacle when the agent faces dynamic obstacles to obtain the weighted cost; Step 7: Confirm the safest speed adjustment direction in the current state of the agent through the gradient of the weighted cost, and give rewards and punishments to the agent to guide the agent to choose a suitable direction when avoiding collisions; Step 8: Give different amounts of punishment according to the type of obstacle when the agent collides; Step 9: Give corresponding rewards according to the number of agents reaching the end point; Among them, the specific method of the said step 5 is: set a function This function calculates the positional relationship of point p with respect to the directed line by calculating the cross product; where is the directed line, (a x , a y ) is the starting point of the directed line , (u x , u y ) is the direction of the directed line , (p x , p y ) is the position coordinates of point p; The calculation formula for the cost value of the lowest escape speed obstacle when the agent faces the two types of obstacles is shown as follows: Among them, r voca represents the cost value of the lowest escape speed obstacle when the agent faces a dynamic obstacle, and r vocs represents the cost value of the lowest escape speed obstacle when the agent faces a static obstacle; represents the two sides of the speed obstacle, represents the cut-off line, and r c is the collision radius of the agent, O c is the center of the circle at the top of the speed obstacle, and M is the point on the static obstacle closest to the current driving speed of the agent; Among them, the calculation formula for the importance factor in Step 6 is shown as follows: Among them, β is the importance factor, and ||d|| represents the relative distance between agents; Weight the importance factor to the cost value of the lowest escape speed obstacle when the agent faces dynamic obstacles, and dynamically adjust the influence degree of speed obstacles at different distances on the reward value of the agent; The weighted cost is shown as follows: r voca-β =βr voca Among them, r voca-β represents the weighted cost; Among them, the specific method of Step 7 is: The magnitude of the outer product of the safest speed adjustment direction e in the current state of the agent and the current driving speed v of the agent can reflect the direction of e relative to v, that is, when v×e>0, it means that e is to the left relative to v, and when v×e<0, e is to the right relative to v; When the agent learns the strategy of avoiding collisions to the right, when v×e<0, a reward is given at this time to encourage the agent to avoid collisions from the right side, and when v×e>0, a punishment is given to punish the agent for avoiding collisions to the left side.
2. The design method of a path planning reward function based on deep reinforcement learning according to claim 1, characterized in that: The expected speed of the agent described in Step 1 is shown in the following formula: Among them, v des is the desired speed of the agent, (x1, y1) are the coordinates of the agent's current position, and (x2, y2) are the position coordinates of the target point, is the unit vector from the agent's current position (x1, y1) to the target point (x2, y2), and v max is the maximum speed at which the agent travels, ||v max || is the magnitude of the maximum speed at which the agent travels.
3. The design method of a path planning reward function based on deep reinforcement learning according to claim 2, characterized in that: The reward formula in Step 2 is shown as follows: where r diff is the calculation result, v is the current driving speed of the agent, and d(v, v des ) is the Euclidean distance between the current driving speed and the desired speed of the agent.
4. The design method of a path planning reward function based on deep reinforcement learning according to claim 1, characterized in that: The reward value in Step 9 is defined as follows: r task = (2·N arrive - N total )·5 where r task is the reward value, N arrive is the number of agents that reach the target point in an episode, and N total is the total number of agents in the current scenario.
Citation Information
Patent Citations
Deep reinforcement learning reward function optimization method for unmanned ship path planning
CN111880549B
Deep reinforcement learning reward function optimization method for unmanned ship path planning
CN111880549A
Intelligent collision avoidance method for a swarm of unmanned surface vehicles based on deep reinforcement learning
US20220189312A1