Obstacle-avoiding pursuit game method and system for multiple unmanned vehicles

By constructing the cost function of the pursuit and obstacle avoidance of multiple unmanned vehicles, and using the gradient heavy center self-coordinated obstacle function to convert the safe distance constraint into the obstacle penalty term, the problem that unmanned vehicles cannot effectively avoid obstacles in the pursuit and fugitive game is solved, and the safety and stability of the system are improved.

CN120370949APending Publication Date: 2025-07-25JIANGNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510504988.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the existing game methods for avoiding obstacles and chasing games for many unmanned vehicles, the speed of pursuit is limited, and the penalty item only works when the vehicle is very close to an obstacle, resulting in the unmanned vehicles being unable to effectively avoid obstacles, affecting the safety and stability of the system.

Method used

The cost function and the cost function of the escaped and pursuer unmanned vehicles are constructed. The safety distance constraint is converted into obstacle penalty terms through the gradient recenter self-coordinated obstacle function, and the target strategy is obtained using the value iteration algorithm to ensure that the unmanned vehicles maintain a safe distance at each time step.

Benefits of technology

It realizes a flexible switching strategy for unmanned vehicles in complex environments, accurately avoids obstacles, and improves the safety and stability of the system and the efficiency of pursuit.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120370949A_ABST
    Figure CN120370949A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multi-agent systems, in particular to a multi-unmanned-vehicle obstacle avoidance pursuit game method and a multi-unmanned-vehicle obstacle avoidance pursuit game system. The method comprises the following steps: constructing a pursuit cost function and an obstacle avoidance cost function of an escaper unmanned vehicle and each pursuit unmanned vehicle, wherein the pursuit cost function is constructed based on a local error function and a pursuit strategy; an obstacle avoidance problem is converted into an unconstrained optimization problem through a gradient weight center self-coordination obstacle function, the obstacle avoidance problem is weighted and introduced into an obstacle avoidance cost function, the obstacle avoidance cost function is constructed by using a relative safety function, an obstacle avoidance strategy and an obstacle penalty term converted by a safety distance constraint, and the obstacle avoidance strategy and a pursuit strategy are optimized based on a value iteration algorithm. And the Nash equilibrium of pursuit game in the obstacle environment is realized. According to the invention, it can be effectively ensured that the unmanned vehicle always keeps a safe distance from the obstacle in the process of pursuit game, and the stability and safety of the unmanned vehicle are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-agent systems, and in particular, to a multi-unmanned vehicle obstacle avoidance pursuit-evasion game method and system. Background Art

[0002] Consensus control, as the core issue of cooperative control in multi-agent systems, aims to make the states of all agents asymptotically converge to the same value or track a given reference trajectory through distributed communication and local decision-making. Game theory provides a new theoretical framework for achieving consensus control in multi-agent systems. In multi-agent systems, due to the existence of mutual communication among individuals, from the perspective of game theory, the relationship between agents can be described as cooperation or competition. When all agents optimize to obtain individual optimal controllers, the game reaches a Nash equilibrium, and at the same time, the system as a whole reaches a Nash equilibrium. Cooperative control of multi-agent systems is widely applied in pursuit-evasion game problems and is commonly used in search and rescue tasks and computer game fields. For example, in a personnel search and rescue operation in a forest fire, the people lost in the fire area are the "evaders", and the unmanned vehicles carrying rescue supplies and equipment are the "pursuers". The unmanned vehicles flexibly avoid obstacles such as areas where the fire spreads, high-temperature zones, and fallen trees according to the path planned by the consensus control algorithm, and achieve efficient search and rescue in a complex environment. In a pursuit-evasion game, the game objective of the pursuer unmanned vehicle agent is to capture the evader as soon as possible, while the objective of the evader agent is to avoid being captured as much as possible. When the game reaches a Nash equilibrium, the pursuer unmanned vehicle agent can still achieve the capture purpose when the evader agent is maximally far away.

[0003] In the application of unmanned vehicle confrontation, the pursuer unmanned vehicle expects to stay around the evader unmanned vehicle in a formation, spray a capture net to capture it, and then take it to a safe area. However, there will also be obstacle factors in the pursuit-evasion game scenario. Therefore, it is of great significance to set appropriate obstacle avoidance strategies for agents to achieve the pursuit-evasion goal while maintaining a safe distance from obstacles. Currently, for the obstacle avoidance pursuit-evasion game of multi-unmanned vehicles, the common method is to set a cost function that takes into account both the pursuit-evasion and obstacle avoidance tasks for each pursuer and evader vehicle. When constructing this cost function, mainly the distance between the agent and the obstacle is used as a penalty term and incorporated into the cost function. Its operating mechanism is to try to prompt the unmanned vehicle to actively increase the distance from the obstacle during the execution of the pursuit-evasion task by minimizing this cost function. In principle, when the unmanned vehicle approaches the obstacle, the value of the penalty term increases, resulting in an increase in the value of the entire cost function. In order to minimize the cost function, the decision-making system of the unmanned vehicle will adjust the vehicle's actions to increase the distance from the obstacle. The penalty term in the cost function can help the unmanned vehicle avoid obstacles in the environment to a certain extent and ensure the smooth progress of the pursuit-evasion operation.

[0004] However, there are obvious flaws in the cost function setting that takes the distance between the intelligent agent and the obstacle as a penalty term. First, during the execution of the pursuit mission, in order to minimize the cost function, the decision-making system of the unmanned vehicle will adjust the vehicle's actions to increase the distance from the obstacle, making it difficult for the unmanned vehicle to focus on the pursuit target. For example, when the pursuing vehicle finds the fugitive in a relatively open area with a small number of obstacles, it should normally accelerate the pursuit. However, because of the need to consider the penalty terms brought by distant obstacles, the pursuit speed is ultimately limited, making it difficult to accurately plan the optimal action plan to approach the fugitive as quickly as possible. Moreover, the penalty item has a strong lag and will only take effect when the unmanned vehicle is very close to an obstacle. During the driving process, when the vehicle is still a certain distance away from the obstacle but has shown a dangerous approaching trend, the penalty item cannot take effect in time and cannot send an adjustment signal to the vehicle's decision-making system, resulting in the vehicle being unable to change its action strategy in advance and effectively avoid obstacles. This makes it difficult to accurately and efficiently prevent unmanned vehicles from approaching dangerous situations when performing obstacle avoidance and pursuit game tasks, which has a serious impact on the safety and stability of the system operation. Summary of the invention

[0005] To this end, the technical problem to be solved by the present invention is to overcome the defects of the existing multi-unmanned vehicle obstacle avoidance and pursuit game method that the pursuit speed is limited and its penalty item only takes effect when the vehicle is very close to the obstacle, resulting in the unmanned vehicle being unable to effectively avoid the obstacle.

[0006] In order to solve the above technical problems, the present invention provides a multi-unmanned vehicle obstacle avoidance and pursuit game method, comprising:

[0007] Construct the pursuit cost function and obstacle avoidance cost function of the escaper unmanned vehicle and each pursuer unmanned vehicle at each time step, including:

[0008] Based on the local error function of the current unmanned vehicle at each time step and the control input generated by the pursuit and escape strategy of the current unmanned vehicle at each time step, a pursuit and escape cost function of the current unmanned vehicle at each time step is constructed;

[0009] Based on the safety radius of the current unmanned vehicle, the safety radius of the obstacle, and the distance between the current unmanned vehicle and the obstacle at each time step, a relative safety function of the current unmanned vehicle at each time step is constructed; a safety distance constraint is set: the relative safety function of the current unmanned vehicle at each time step is not greater than the safety threshold;

[0010] Through the gradient re-centering self-coordinating obstacle function, the safety distance constraint is converted into the obstacle penalty term of the current unmanned vehicle at each time step;

[0011] Based on the relative safety function of the current driverless vehicle at each time step, the control input generated by the obstacle avoidance strategy of the current driverless vehicle at each time step, and the obstacle penalty term of the current driverless vehicle at each time step, construct the obstacle avoidance cost function of the current driverless vehicle at each time step;

[0012] Based on the pursuit-evasion cost function between the evader driverless vehicle and each pursuer driverless vehicle at each time step, and the Q function corresponding to the obstacle avoidance cost function, use the value iteration algorithm to obtain the target pursuit-evasion strategy and obstacle avoidance strategy of each driverless vehicle at each time step;

[0013] Based on the target pursuit-evasion strategy and obstacle avoidance strategy of each driverless vehicle at each time step, generate the control input of the target pursuit-evasion strategy and obstacle avoidance strategy of each driverless vehicle at each time step.

[0014] Preferably, the relative safety function of the evader driverless vehicle at each time step is:

[0015] η e (k)=(r e +r o ) / ξ e (k),

[0016] The relative safety function of each pursuer driverless vehicle at each time step is:

[0017]

[0018] where η e (k) is the relative safety function of the evader driverless vehicle at the k-th time step, r e is the safety radius of the evader driverless vehicle, r o is the safety radius of the obstacle, ξ e (k) is the distance between the evader driverless vehicle and the obstacle at the k-th time step, is the dynamic model of the evader driverless vehicle at the k-th time step, x e (k) is the horizontal coordinate of the evader driverless vehicle at the k-th time step, y e (k) is the vertical coordinate of the evader driverless vehicle at the k-th time step, is the two-dimensional coordinate of the obstacle, is the relative safety function of the i-th pursuer driverless vehicle at the k-th time step, is the safety radius of the i-th pursuer driverless vehicle, r o is the safety radius of the obstacle, is the distance between the i-th pursuer driverless vehicle and the obstacle at the k-th time step, is the dynamic model of the $i$-th pursuer unmanned vehicle at the $k$-th time step, is the horizontal coordinate of the $i$-th pursuer unmanned vehicle at the $k$-th time step, is the vertical coordinate of the $i$-th pursuer unmanned vehicle at the $k$-th time step.

[0019] Preferably, the obstacle penalty term of the evader unmanned vehicle at each time step is: where, is the gradient re-centered self-concordant barrier function, and the gradient re-centered self-concordant barrier function converts the safety distance constraint $\eta$ of the evader unmanned vehicle e $(k + n)\leq v$ e into the obstacle penalty term of the evader unmanned vehicle at the $k$-th time step $v$ e is the safety threshold of the evader unmanned vehicle, and $\eta$ e $(k + n)$ is the relative safety function of the evader unmanned vehicle at the $(k + n)$-th time step;

[0020] The obstacle penalty term of the $i$-th pursuer unmanned vehicle at each time step is: where, the gradient re-centered self-concordant barrier function converts the safety distance constraint of the $i$-th pursuer unmanned vehicle into the obstacle penalty term of the $i$-th pursuer unmanned vehicle at the $k$-th time step is the safety threshold of the $i$-th pursuer unmanned vehicle, is the relative safety function of the $i$-th pursuer unmanned vehicle at the $(k + n)$-th time step.

[0021] Preferably, the pursuit-evasion cost function of the evader unmanned vehicle at each time step is:

[0022]

[0023] The pursuit-evasion cost function of each pursuer unmanned vehicle at each time step is:

[0024]

[0025] where, $J$ e $(\delta$ e $(k), u$ e $(k))$ is the pursuit-evasion cost function of the evader unmanned vehicle at the $k$-th time step, $\delta$ e $(k)$ is the local error function of the evader unmanned vehicle at the $k$-th time step, $\delta$ e $(k + n)$ is the local error function of the evader unmanned vehicle at the $(k + n)$-th time step, $u$ e $(k)=[V$ e $(k),\theta$ e $(k)]$T , (·) T is the transpose, u e (k) is the control input generated by the evader's unmanned vehicle's pursuit and evasion strategy at the k-th time step, u e (k + n) is the control input generated by the evader's unmanned vehicle's pursuit and evasion strategy at the k + n-th time step, V e (k) is the speed of the evader's unmanned vehicle at the k-th time step, θ e (k) is the heading angle of the evader's unmanned vehicle at the k-th time step, U e (·) is the single-step cost of the evader's unmanned vehicle regarding pursuit and evasion, P e is the first symmetric positive definite matrix of the evader's unmanned vehicle, R e is the second symmetric positive definite matrix of the evader's unmanned vehicle, is the pursuit and evasion cost function of the i-th pursuer's unmanned vehicle at the k-th time step, is the local error function of the i-th pursuer's unmanned vehicle at the k-th time step, The local error function of the i-th pursuer's unmanned vehicle at the k + n-th time step, is the control input generated by the i-th pursuer's unmanned vehicle's pursuit and evasion strategy at the k-th time step, is the control input generated by the i-th pursuer's unmanned vehicle's pursuit and evasion strategy at the k + n-th time step, is the speed of the i-th pursuer's unmanned vehicle at the k-th time step, is the heading angle of the i-th pursuer's unmanned vehicle at the k-th time step, is the single-step cost of the i-th pursuer's unmanned vehicle regarding pursuit and evasion, is the first symmetric positive definite matrix of the i-th pursuer's unmanned vehicle, is the second symmetric positive definite matrix of the i-th pursuer's unmanned vehicle, k is the time step, and n is the time step offset index.

[0026] Preferably, the obstacle avoidance cost function of the evader's unmanned vehicle at each time step is:

[0027]

[0028] The obstacle avoidance cost function of each pursuer's unmanned vehicle at each time step is:

[0029]

[0030]

[0031] Among them, is the obstacle avoidance cost function of the evader's unmanned vehicle at the k-th time step, ηe $(k)$ is the relative safety function of the escaping unmanned vehicle at the $k$-th time step, $\eta$ e $(k + n)$ is the relative safety function of the escaping unmanned vehicle at the $(k + n)$-th time step, (·) T is the transpose, is the control input generated by the obstacle avoidance strategy of the escaping unmanned vehicle at the $k$-th time step, is the control input generated by the pursuit-evasion strategy of the escaping unmanned vehicle at the $(k + n)$-th time step, $V$ e (k) is the speed of the escaping unmanned vehicle at the $k$-th time step, $\theta$ e (k) is the heading angle of the escaping unmanned vehicle at the $k$-th time step, $\beta$ e is the obstacle avoidance parameter of the escaping unmanned vehicle, $\beta$ e > 0, $L$ e (·) is the single-step cost of the escaping unmanned vehicle regarding obstacle avoidance, is the third symmetric positive definite matrix of the escaping unmanned vehicle, is the fourth symmetric positive definite matrix of the escaping unmanned vehicle, is the obstacle avoidance cost function of the $i$-th pursuer unmanned vehicle at the $k$-th time step, is the relative safety function of the $i$-th unmanned vehicle at the $k$-th time step, is the relative safety function of the $i$-th unmanned vehicle at the $(k + n)$-th time step, is the control input generated by the obstacle avoidance strategy of the $i$-th pursuer unmanned vehicle at the $k$-th time step, is the control input generated by the pursuit-evasion strategy of the $i$-th pursuer unmanned vehicle at the $(k + n)$-th time step, is the speed of the $i$-th pursuer unmanned vehicle at the $k$-th time step, is the heading angle of the $i$-th pursuer unmanned vehicle at the $k$-th time step, is the obstacle avoidance parameter of the $i$-th pursuer unmanned vehicle, is the single-step cost of the $i$-th pursuer unmanned vehicle regarding obstacle avoidance, is the third symmetric positive definite matrix of the $i$-th pursuer unmanned vehicle, is the fourth symmetric positive definite matrix of the $i$-th pursuer unmanned vehicle, $k$ is the time step, and $n$ is the time step offset index.

[0032] Preferably, the value range of the safety threshold of the escaping unmanned vehicle is The value range of the safety threshold of the $i$-th pursuer unmanned vehicle is where, $r$ e is the safety radius of the escaping unmanned vehicle, $\tau$e is the detection radius of the escaping unmanned vehicle, is the safety radius of the i-th pursuer unmanned vehicle, r o is the safety radius of the obstacle, is the detection radius of the i-th pursuer unmanned vehicle.

[0033] Preferably, the local error function of the escaping unmanned vehicle at each time step is:

[0034]

[0035] The local error function of each pursuer unmanned vehicle at each time step is:

[0036]

[0037] where δ e (k) is the local error function of the escaping unmanned vehicle at the k-th time step, is the local error function of the i-th pursuer unmanned vehicle at the k-th time step, is the set of neighbor pursuer unmanned vehicles of the escaping unmanned vehicle, a ei is the connection gain coefficient between the escaping unmanned vehicle and the i-th pursuer unmanned vehicle, is the set of neighbor pursuer unmanned vehicles of the i-th pursuer unmanned vehicle, is the dynamic model of the escaping unmanned vehicle at the k-th time step, is the dynamic model of the i-th pursuer unmanned vehicle at the k-th time step, is the expected displacement between the i-th pursuer unmanned vehicle and the j-th neighbor pursuer unmanned vehicle, is the expected displacement between the i-th pursuer unmanned vehicle and the escaping unmanned vehicle, i is the pursuer unmanned vehicle index, and j is the neighbor pursuer unmanned vehicle index.

[0038] Preferably, the dynamic model of the escaping unmanned vehicle at the k-th time step is x e (k), y e (k) update iteration formula is:

[0039]

[0040] The dynamic model of each pursuer unmanned vehicle at the k-th time step is The update iteration formula of is:

[0041]

[0042] where x e(k + 1) is the horizontal coordinate of the evader unmanned vehicle at the (k + 1)-th time step, y e (k + 1) is the vertical coordinate of the evader unmanned vehicle at the (k + 1)-th time step, x e (k) is the horizontal coordinate of the evader unmanned vehicle at the k-th time step, y e (k) is the vertical coordinate of the evader unmanned vehicle at the k-th time step, T is the sampling interval, V e (k) is the speed of the evader unmanned vehicle at the k-th time step, θ e (k) is the heading angle of the evader unmanned vehicle at the k-th time step, is the horizontal coordinate of the i-th pursuer unmanned vehicle at the (k + 1)-th time step, y e (k + 1) is the vertical coordinate of the i-th pursuer unmanned vehicle at the (k + 1)-th time step, is the horizontal coordinate of the i-th pursuer unmanned vehicle at the k-th time step, is the vertical coordinate of the i-th pursuer unmanned vehicle at the k-th time step, is the speed of the i-th pursuer unmanned vehicle at the k-th time step, is the heading angle of the i-th pursuer unmanned vehicle at the k-th time step, k represents the time step, [·] T is the transpose.

[0043] Preferably, based on the pursuit-evasion cost function and the Q function corresponding to the obstacle avoidance cost function of the evader unmanned vehicle and each pursuer unmanned vehicle at each time step, the value iteration algorithm is used to obtain the target pursuit-evasion strategy and obstacle avoidance strategy of each unmanned vehicle at each time step, including:

[0044] For each time step, it is judged whether the distance between each unmanned vehicle and the obstacle at the current time step is less than or equal to the detection radius of the unmanned vehicle. If it is less than or equal to, the target obstacle avoidance strategy of the unmanned vehicle at the current time step is obtained through the Q function corresponding to the obstacle avoidance cost function of the unmanned vehicle at the current time step. If it is greater than, the target pursuit-evasion strategy of the unmanned vehicle at the current time step is obtained through the Q function corresponding to the pursuit-evasion cost function of the unmanned vehicle at the current time step, and the target pursuit-evasion strategy and obstacle avoidance strategy of each unmanned vehicle at each time step are obtained.

[0045] The present invention also provides a multi-unmanned vehicle obstacle avoidance and pursuit-evasion game system, including:

[0046] A cost function construction module for constructing the pursuit-evasion cost function and the obstacle avoidance cost function of the evader unmanned vehicle and each pursuer unmanned vehicle at each time step, including:

[0047] Based on the local error function of the current driverless vehicle at each time step and the control input generated by the pursuit-evasion strategy for the current driverless vehicle at each time step, construct the pursuit-evasion cost function of the current driverless vehicle at each time step;

[0048] Based on the safety radius of the current driverless vehicle, the safety radius of the obstacle, and the distance between the current driverless vehicle and the obstacle at each time step, construct the relative safety function of the current driverless vehicle at each time step; Set the safety distance constraint: the relative safety function of the current driverless vehicle at each time step is not greater than the safety threshold;

[0049] Through the gradient re-centered self-coordinated barrier function, convert the safety distance constraint into the barrier penalty term of the current driverless vehicle at each time step;

[0050] Based on the relative safety function of the current driverless vehicle at each time step, the control input generated by the obstacle avoidance strategy for the current driverless vehicle at each time step, and the barrier penalty term of the current driverless vehicle at each time step, construct the obstacle avoidance cost function of the current driverless vehicle at each time step;

[0051] The policy solving module is used to obtain the target pursuit-evasion strategy and obstacle avoidance strategy of each driverless vehicle at each time step by using the value iteration algorithm based on the Q functions corresponding to the pursuit-evasion cost function and the obstacle avoidance cost function of the evader driverless vehicle and each pursuer driverless vehicle at each time step;

[0052] The control generation module is used to generate the control inputs of the target pursuit-evasion strategy and the obstacle avoidance strategy of each driverless vehicle at each time step based on the target pursuit-evasion strategy and the obstacle avoidance strategy of each driverless vehicle at each time step.

[0053] The above technical solution of the present invention has the following beneficial effects compared with the prior art:

[0054] A multi-unmanned vehicle obstacle avoidance pursuit and escape game method and system according to the present invention. The present invention constructs a cost function for unmanned vehicle pursuit and escape and obstacle avoidance. The pursuit and escape cost function is constructed around the local error function of the current unmanned vehicle and the control input generated by the pursuit and escape strategy. Its goal is clearly directed at efficiently completing the pursuit and escape task. It can accurately calculate the best action plan that the unmanned vehicle needs to make to approach or move away from the target during the pursuit or escape process. The obstacle avoidance cost function is constructed based on the safety radius of the unmanned vehicle, the safety radius of the obstacle, and the real-time distance between the two, focusing on ensuring that the unmanned vehicle effectively avoids obstacles on the driving path. Compared with the existing single cost function that tries to take into account all tasks but often fails to do both well, it can more accurately meet the core needs of different task links. By judging whether the distance between the unmanned vehicle and the obstacle is less than the detection radius of the unmanned vehicle in each iteration, the dominant strategy of the unmanned vehicle in each time step is determined. In a relatively empty area with few obstacles, the pursuit and escape cost function dominates, and the unmanned vehicle can fully execute the pursuit and escape strategy to approach or move away from the target at the fastest speed. Once entering a complex area with dense obstacles, the obstacle avoidance cost function immediately comes into play, guiding the vehicle to adjust its actions and giving priority to ensuring driving safety. This mechanism of flexibly switching the dominant cost function according to the scenario enables the unmanned vehicle to quickly adapt to environmental changes and formulate the most appropriate action strategy, greatly improving the flexibility of the unmanned vehicle to cope with different scenarios.

[0055] The present invention comprehensively considers the safety radius of the unmanned vehicle itself, the safety radius of the obstacle, and the real-time distance between the two, and constructs a relative safety function; the safety radius of the unmanned vehicle reflects the minimum space range required to ensure its own safe operation in the normal driving state, and the safety radius of the obstacle represents the size of the area around the obstacle that may affect the driving safety of the vehicle. Through the comprehensive operation of multi-dimensional information, it can more accurately and comprehensively describe the safety situation between the unmanned vehicle and the obstacle. Through the gradient re-centered self-coordinated obstacle function, the safety distance constraint where the relative safety function is less than or equal to the safety threshold is converted into an obstacle penalty term for the unmanned vehicle, setting a clear constraint boundary for the obstacle avoidance problem. When the unmanned vehicle reaches the preset safety distance from the obstacle, the system will immediately impose a large penalty, thus effectively ensuring that the unmanned vehicle always maintains a safe distance from the obstacle and ensuring that the unmanned vehicle can perceive potential dangers in advance through the change of the obstacle avoidance cost function. Compared with the existing method that only sets a penalty term according to the real-time distance between the two, the present invention effectively prevents the dangerous situation of the unmanned vehicle approaching the obstacle and improves the safety and stability of the unmanned vehicle system. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to the specific embodiments of the present invention in conjunction with the drawings, wherein:

[0057] Figure 1 It is a schematic flow chart of a multi-unmanned vehicle obstacle avoidance pursuit and escape game method of the present invention.

[0058] Figure 2 It is a step flow chart for constructing the pursuit and escape cost function and the obstacle avoidance cost function.

[0059] Figure 3 It is a schematic diagram of multi-unmanned vehicles in a two-dimensional coordinate system.

[0060] Figure 4 It is a communication topology diagram of multi-unmanned vehicles.

[0061] Figure 5 It is a coordinate change trajectory diagram of multi-unmanned vehicles.

[0062] Figure 6 It is a local error trajectory diagram of multi-unmanned vehicles.

[0063] Figure 7 It is a coordinate change trajectory diagram of multi-unmanned vehicles in a two-dimensional coordinate system. Specific Embodiment

[0064] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the specific embodiments cited are not intended to limit the present invention.

[0065] Refer to Figure 1 As shown, Embodiment 1 of the present invention provides a multi-unmanned vehicle obstacle avoidance pursuit and escape game method, including the following steps:

[0066] As Figure 2 shown, Figure 2 It is a step flow chart for constructing the pursuit and escape cost function and the obstacle avoidance cost function.

[0067] Step S1: Construct the pursuit and escape cost function and the obstacle avoidance cost function of the evader unmanned vehicle and each pursuer unmanned vehicle at each time step. The construction process includes:

[0068] Step S11: Based on the local error function of the current unmanned vehicle at each time step and the control input generated by the pursuit and escape strategy of the current unmanned vehicle at each time step, construct the pursuit and escape cost function of the current unmanned vehicle at each time step;

[0069] In this embodiment, specifically, the construction process of the local error function of the evader unmanned vehicle and each pursuer unmanned vehicle includes:

[0070] Obtain the dynamic models of the evader unmanned vehicle and each pursuer unmanned vehicle, and based on the dynamic models of the evader unmanned vehicle and each pursuer unmanned vehicle, construct the local error functions of the evader unmanned vehicle and each pursuer unmanned vehicle.

[0071] As shown Figure 3 in Figure 3 the figure, it is a schematic diagram of multiple unmanned vehicles in a two-dimensional coordinate system. The coordinates of each unmanned vehicle can be represented by (x m , y m ), where x m is the horizontal coordinate (x-coordinate) of the m-th unmanned vehicle, and y m is the vertical coordinate (y-coordinate) of the m-th unmanned vehicle. And each unmanned vehicle will adjust its position according to its own speed and heading angle. The multi-unmanned vehicle scenario of the present invention consists of N pursuer unmanned vehicles and one escapee unmanned vehicle. The dynamic characteristics of the pursuer unmanned vehicles can be described by the following non-linear model.

[0072] In this embodiment, specifically, the dynamic model of the escapee unmanned vehicle at the k-th time step is The update iteration formulas for x e (k) and y e (k) are:

[0073]

[0074] The dynamic model of each pursuer unmanned vehicle at the k-th time step is The update iteration formula for

[0075]

[0076] is: e where x e (k + 1) is the horizontal coordinate of the escapee unmanned vehicle at the (k + 1)-th time step, y e (k + 1) is the vertical coordinate of the escapee unmanned vehicle at the (k + 1)-th time step, x e (k) is the horizontal coordinate of the escapee unmanned vehicle at the k-th time step, y e (k) is the vertical coordinate of the escapee unmanned vehicle at the k-th time step, T is the sampling interval. In this embodiment, T = 0.5s, V e (k) is the speed of the escapee unmanned vehicle at the k-th time step, θ is the horizontal coordinate of the i-th pursuer unmanned vehicle at the (k + 1)-th time step, y e (k + 1) is the vertical coordinate of the i-th pursuer unmanned vehicle at the (k + 1)-th time step, is the horizontal coordinate of the i-th pursuer unmanned vehicle at the k-th time step, is the vertical coordinate of the i-th pursuer unmanned vehicle at the k-th time step, is the speed of the i-th pursuer unmanned vehicle at the k-th time step, $\theta_{i}(k)$ is the heading angle of the $i$-th pursuer unmanned vehicle at the $k$-th time step, where $k$ represents the time step, [·] T is the transpose, and $i = 1, 2, \ldots, N$.

[0077] In this embodiment, specifically, the local error function of the evader unmanned vehicle at each time step is:

[0078]

[0079] For the pursuer unmanned vehicle, its goal is to capture the evader unmanned vehicle at the minimum cost, while the goal of the evader unmanned vehicle is to stay away from the pursuer unmanned vehicle at the minimum cost. Therefore, the local error function of each pursuer unmanned vehicle at each time step is:

[0080]

[0081] where $\delta$ e (k) is the local error function of the evader unmanned vehicle at the $k$-th time step, is the local error function of the $i$-th pursuer unmanned vehicle at the $k$-th time step, is the set of neighboring pursuer unmanned vehicles of the evader unmanned vehicle, $a$ ei is the connection gain coefficient between the evader unmanned vehicle and the $i$-th pursuer unmanned vehicle, is the set of neighboring pursuer unmanned vehicles of the $i$-th pursuer unmanned vehicle, is the dynamic model of the evader unmanned vehicle at the $k$-th time step, is the dynamic model of the $i$-th pursuer unmanned vehicle at the $k$-th time step, is the expected displacement between the $i$-th pursuer unmanned vehicle and the $j$-th neighboring pursuer unmanned vehicle, is the expected displacement between the $i$-th pursuer unmanned vehicle and the evader unmanned vehicle, where $i$ is the pursuer unmanned vehicle index and $j$ is the neighboring pursuer unmanned vehicle index.

[0082] In this embodiment, preferably, the pursuit and evasion cost function of the evader unmanned vehicle at each time step is:

[0083]

[0084]

[0085] where $J$ e (\delta e (k), u e (k)) is the pursuit and evasion cost function of the evader unmanned vehicle at the $k$-th time step, $\delta$ e (k) is the local error function of the evader unmanned vehicle at the $k$-th time step, $\delta$ e(k + n) is the local error function of the evader unmanned vehicle at the (k + n)-th time step, u e (k) = [V e (k), θ e (k)] T , (·) T is the transpose, u e (k) is the control input generated by the pursuit-evasion strategy of the evader unmanned vehicle at the k-th time step, u e (k + n) is the control input generated by the pursuit-evasion strategy of the evader unmanned vehicle at the (k + n)-th time step, V e (k) is the speed of the evader unmanned vehicle at the k-th time step, θ e (k) is the heading angle of the evader unmanned vehicle at the k-th time step, U e (·) is the single-step cost of the evader unmanned vehicle regarding pursuit-evasion, P e is the first symmetric positive definite matrix of the evader unmanned vehicle, R e is the second symmetric positive definite matrix of the evader unmanned vehicle,

[0086]

[0087] The pursuit-evasion cost function of each pursuer unmanned vehicle at each time step is:

[0088]

[0089] Among them, is the pursuit-evasion cost function of the i-th pursuer unmanned vehicle at the k-th time step, is the local error function of the i-th pursuer unmanned vehicle at the k-th time step, The local error function of the i-th pursuer unmanned vehicle at the (k + n)-th time step, is the control input generated by the pursuit-evasion strategy of the i-th pursuer unmanned vehicle at the k-th time step, is the control input generated by the pursuit-evasion strategy of the i-th pursuer unmanned vehicle at the (k + n)-th time step, is the speed of the i-th pursuer unmanned vehicle at the k-th time step, is the heading angle of the i-th pursuer unmanned vehicle at the k-th time step, is the single-step cost of the i-th pursuer unmanned vehicle regarding pursuit-evasion, P i p is the first symmetric positive definite matrix of the i-th pursuer unmanned vehicle, is the second symmetric positive definite matrix of the i-th pursuer unmanned vehicle, k is the time step, and n is the time step offset index.

[0090] The present invention comprehensively considers the error functions and control inputs of each unmanned vehicle, and accumulates and sums them over multiple future time steps. The local error function reflects the gap between the current state and the desired state of the escapee / pursuer, and the control input reflects the actions taken by the escapee / pursuer to change the state. This comprehensive consideration enables the pursuit-evasion cost function to comprehensively and accurately evaluate the performance of the escapee / pursuer during the entire pursuit-evasion process, providing an accurate quantitative basis for optimizing the escape strategy.

[0091] Step S12: Based on the safety radius of the current unmanned vehicle, the safety radius of the obstacle, and the distance between the current unmanned vehicle and the obstacle at each time step, construct the relative safety function of the current unmanned vehicle at each time step; set the safety distance constraint: the relative safety function of the current unmanned vehicle at each time step is not greater than the safety threshold.

[0092] In this embodiment, preferably, the relative safety function of the escapee unmanned vehicle at each time step is:

[0093] η e (k) = (r e + r o ) / ξ e (k),

[0094] where η e (k) is the relative safety function of the escapee unmanned vehicle at the k-th time step, r e is the safety radius of the escapee unmanned vehicle, r o is the safety radius of the obstacle, ξ e (k) is the distance between the escapee unmanned vehicle and the obstacle at the k-th time step, x e (k) is the dynamic model of the escapee unmanned vehicle at the k-th time step, x e (k) is the horizontal coordinate of the escapee unmanned vehicle at the k-th time step, y e (k) is the vertical coordinate of the escapee unmanned vehicle at the k-th time step, is the two-dimensional coordinate of the obstacle.

[0095] The relative safety function of each pursuer unmanned vehicle at each time step is:

[0096]

[0097] where, is the relative safety function of the i-th unmanned vehicle at the k-th time step, is the safety radius of the i-th pursuer unmanned vehicle, r o is the safety radius of the obstacle, is the distance between the i-th pursuer unmanned vehicle and the obstacle at the k-th time step, is the dynamic model of the i-th pursuer unmanned vehicle at the k-th time step, is the horizontal coordinate of the i-th pursuer unmanned vehicle at the k-th time step, is the vertical coordinate of the i-th pursuer unmanned vehicle at the k-th time step, is the two-dimensional coordinate of the obstacle.

[0098] The present invention comprehensively considers the safety radius of the unmanned vehicle itself, the safety radius of the obstacle, and the real-time distance between the two, and constructs a relative safety function; by taking into account the safety radius of the unmanned vehicle itself, it fully ensures that when the unmanned vehicle takes actions, enough safe operation space is reserved for itself to prevent damage to its own structure or system failure caused by being too close to the obstacle. The consideration of the safety radius of the obstacle avoids the influence of sudden situations of the obstacle on the unmanned vehicle at a seemingly safe distance due to ignoring the potential dangerous area around the obstacle. Through the comprehensive operation of multi-dimensional information, the safety situation between the unmanned vehicle and the obstacle can be described more accurately and comprehensively.

[0099] Step S13: Convert the safety distance constraint into an obstacle penalty term for the current unmanned vehicle at each time step through the gradient re-centered self-concordant barrier function;

[0100] In this embodiment, preferably, the obstacle penalty term for the evader unmanned vehicle at each time step is: where, is the gradient re-centered self-concordant barrier function, and the gradient re-centered self-concordant barrier function converts the safety distance constraint η e (k + n) ≤ v e into the obstacle penalty term for the evader unmanned vehicle at the k-th time step v e is the safety threshold of the evader unmanned vehicle, and η e (k + n) is the relative safety function of the evader unmanned vehicle at the k + n-th time step;

[0101] The obstacle penalty term for the i-th pursuer unmanned vehicle at each time step is: where the gradient re-centered self-concordant barrier function converts the safety distance constraint of the i-th pursuer unmanned vehicle into the obstacle penalty term for the i-th pursuer unmanned vehicle at the k-th time step is the safety threshold of the i-th pursuer unmanned vehicle, is the relative safety function of the i-th pursuer unmanned vehicle at the k + n-th time step.

[0102] In this embodiment, specifically, the gradient re-centering self-coordination barrier function can be expressed as:

[0103]

[0104] in, D m (z)≤0 represents the w-th constraint, φ(z) is the original barrier function, z is the independent variable of the gradient re-centering self-coordination barrier function, φ(0) is the value of φ(z) when z=0, is the gradient operator about z, W is the sorting of constraints, w is the constraint index, D m (z) is a function of z.

[0105] The present invention innovatively transforms the obstacle avoidance problem into a constraint problem, and further transforms it into an unconstrained optimization problem by cleverly combining the obstacle function. This design shows many extremely significant beneficial effects in the unmanned vehicle obstacle avoidance and pursuit game. In the stage of transforming the obstacle avoidance problem into a constraint problem, a strict safety distance constraint is formed by clearly defining the relationship between the unmanned vehicle's own safety radius, the obstacle's safety radius, and the real-time distance between the two. Compared with the traditional fuzzy obstacle avoidance consideration method, this constraint form enables the unmanned vehicle system to accurately judge when obstacle avoidance actions need to be taken in different scenarios, greatly improving the accuracy and reliability of obstacle avoidance decisions. The gradient re-centering self-coordinating obstacle function is introduced to cleverly transform the above safety distance constraint into an obstacle penalty term, setting a clear constraint boundary for the obstacle avoidance problem, so that when the unmanned vehicle reaches a preset safety distance with the obstacle, the system will immediately impose a larger penalty, thereby effectively ensuring that the unmanned vehicle always maintains a safe distance with the obstacle, and ensuring that the unmanned vehicle can perceive potential dangers in advance through changes in the obstacle avoidance cost function.

[0106] Step S14: constructing an obstacle avoidance cost function of the current unmanned vehicle at each time step based on the relative safety function of the current unmanned vehicle at each time step, the control input generated by the obstacle avoidance strategy of the current unmanned vehicle at each time step, and the obstacle penalty term of the current unmanned vehicle at each time step;

[0107] In this embodiment, preferably, the obstacle avoidance cost function of the escaper unmanned vehicle at each time step is:

[0108]

[0109] in, is the obstacle avoidance cost function of the escaper unmanned vehicle at the kth time step, η e (k) is the relative safety function of the escaped unmanned vehicle at the kth time step, η e (k+n) is the relative safety function of the escaped unmanned vehicle at the k+nth time step, (·) T is the transpose, is the control input generated by the obstacle avoidance strategy of the evader UGV at the k-th time step, is the control input generated by the pursuit-evasion strategy of the evader UGV at the (k + n)-th time step, V e v(k) is the speed of the evader UGV at the k-th time step, θ e θ(k) is the heading angle of the evader UGV at the k-th time step, β e is the obstacle avoidance parameter of the evader UGV, β e > 0, L e (·) is the single-step cost of the evader UGV regarding obstacle avoidance, is the third symmetric positive definite matrix of the evader UGV, is the fourth symmetric positive definite matrix of the evader UGV,

[0110] The obstacle avoidance cost function of each pursuer UGV at each time step is:

[0111]

[0112] where, is the obstacle avoidance cost function of the i-th pursuer UGV at the k-th time step, is the relative safety function of the i-th UGV at the k-th time step, is the relative safety function of the i-th UGV at the (k + n)-th time step, is the control input generated by the obstacle avoidance strategy of the i-th pursuer UGV at the k-th time step, is the control input generated by the pursuit-evasion strategy of the i-th pursuer UGV at the (k + n)-th time step, is the speed of the i-th pursuer UGV at the k-th time step, is the heading angle of the i-th pursuer UGV at the k-th time step, is the obstacle avoidance parameter of the i-th pursuer UGV, is the single-step cost of the i-th pursuer UGV regarding obstacle avoidance, is the third symmetric positive definite matrix of the i-th pursuer UGV, is the fourth symmetric positive definite matrix of the i-th pursuer UGV, k is the time step, and n is the time step offset index.

[0113] In the present invention, by separately constructing the cost functions for the pursuit-evasion and obstacle avoidance of the unmanned vehicle, the pursuit-evasion cost function is closely related to the local error function of the unmanned vehicle at each time step and the control input generated by the pursuit-evasion strategy. This construction method highly focuses the goal of the pursuit-evasion task and can precisely plan the optimal action plan to approach or move away from the target according to the current state of the unmanned vehicle and the pursuit-evasion target in great detail. Taking the pursuit scenario as an example, the pursuit-evasion cost function comprehensively considers local error factors such as the position deviation and relative speed between the pursuer and the evader, and combines the control inputs such as speed and steering generated by the preset pursuit-evasion strategy to accurately calculate the best action that the pursuer should take at each moment, thereby efficiently promoting the progress of the pursuit-evasion task. For the construction of the obstacle avoidance cost function, it focuses on the safety radius of the unmanned vehicle itself, the safety radius of the obstacle, and their real-time dynamic distance. This enables full consideration of the core elements of the obstacle avoidance task and can comprehensively and real-time evaluate the safety state of the unmanned vehicle with surrounding obstacles during driving. When the unmanned vehicle approaches an obstacle, the obstacle avoidance cost function will timely and accurately calculate the potential risk according to the changes in the safety radius and real-time distance, and prompt the unmanned vehicle to adjust the driving path to effectively avoid the obstacle, escorting the safe driving of the unmanned vehicle.

[0114] Different from the single cost function of traditional multi-unmanned vehicles, the dual cost function system of the present invention has significant task adaptability advantages. In the actual complex and changeable environment, the requirements for pursuit-evasion and obstacle avoidance in different regions are significantly different. In an open and obstacle-scarce scenario, the pursuit-evasion cost function, with its high pertinence to the pursuit-evasion task, dominates the decision-making of the unmanned vehicle, enabling it to wholeheartedly execute the pursuit-evasion strategy and achieve the pursuit-evasion goal at the fastest speed. However, once the unmanned vehicle enters a complex area with dense obstacles, the obstacle avoidance cost function quickly plays a dominant role, guiding the unmanned vehicle to prioritize attention to and ensure driving safety. This mechanism of intelligently and flexibly switching the dominant cost function according to the scenario endows the unmanned vehicle with strong environmental adaptability, enabling it to quickly and accurately formulate the most suitable action strategy for the current situation in various scenarios, and greatly improving the comprehensive performance and reliability of the unmanned vehicle in dealing with complex environments.

[0115] Step S2: Based on the Q functions corresponding to the pursuit-evasion cost function and the obstacle avoidance cost function of the evader unmanned vehicle and each pursuer unmanned vehicle at each time step, use the value iteration algorithm to obtain the target pursuit-evasion strategy and obstacle avoidance strategy of each unmanned vehicle at each time step;

[0116] In this embodiment, specifically, the Q function corresponding to the pursuit-evasion cost function of the evader unmanned vehicle at each time step is:

[0117] Q e (δ e (k),u e (δ eU((k))) = U e (δ e (k), u e (δ e (k)) + Q e (δ e (k + 1), u e (δ e (k + 1))),

[0118] where Q e (δ e (k), u e (k)) is the Q - function corresponding to the pursuit - evasion cost function of the evader UGV at the k - th time step, δ e (k) is the local error function of the evader UGV at the k - th time step, u e (δ e (k)) is the pursuit - evasion strategy of the evader UGV at the k - th time step, U e (·) is the single - step cost of the evader UGV regarding pursuit - evasion, Q e (δ e (k + 1), u e (δ e (k + 1))) is the Q - function corresponding to the pursuit - evasion cost function of the evader UGV at the (k + 1) - th time step, δ e (k + 1) is the local error function of the evader UGV at the (k + 1) - th time step, u e (δ e (k + 1)) is the pursuit - evasion strategy of the evader UGV at the (k + 1) - th time step.

[0119] The Q - function corresponding to the obstacle - avoidance cost function of the evader UGV at each time step is:

[0120]

[0121] where is the Q - function corresponding to the obstacle - avoidance cost function of the evader UGV at the k - th time step, η e (k) is the relative safety function of the evader UGV at the k - th time step, is the obstacle - avoidance strategy of the evader UGV at the k - th time step, L e (·) is the single - step cost of the evader UGV regarding obstacle - avoidance, is the Q - function corresponding to the obstacle - avoidance cost function of the evader UGV at the (k + 1) - th time step, η e (k + 1) is the relative safety function of the evader UGV at the (k + 1) - th time step, is the obstacle - avoidance strategy of the evader UGV at the (k + 1) - th time step.

[0122] The Q function corresponding to the pursuit and evasion cost function of each pursuer unmanned vehicle at each time step is as follows:

[0123]

[0124] Among them, is the Q function corresponding to the pursuit and evasion cost function of the i-th pursuer unmanned vehicle at the k-th time step, is the local error function of the i-th pursuer unmanned vehicle at the k-th time step, is the pursuit and evasion strategy of the i-th pursuer unmanned vehicle at the k-th time step, is the single-step cost of the i-th pursuer unmanned vehicle regarding pursuit and evasion, is the Q function corresponding to the pursuit and evasion cost function of the i-th pursuer unmanned vehicle at the (k + 1)-th time step, is the local error function of the i-th pursuer unmanned vehicle at the k-th time step, is the pursuit and evasion strategy of the i-th pursuer unmanned vehicle at the (k + 1)-th time step,

[0125] The Q function corresponding to the obstacle avoidance cost function of each pursuer unmanned vehicle at each time step is as follows:

[0126]

[0127] Among them, is the Q function corresponding to the obstacle avoidance cost function of the i-th pursuer unmanned vehicle, is the relative safety function of the i-th pursuer unmanned vehicle at the k-th time step, is the obstacle avoidance strategy of the i-th pursuer unmanned vehicle at the k-th time step, is the single-step cost of the i-th pursuer unmanned vehicle regarding obstacle avoidance, is the Q function corresponding to the pursuit and evasion cost function of the i-th pursuer unmanned vehicle at the (k + 1)-th time step, is the relative safety function of the i-th pursuer unmanned vehicle at the (k + 1)-th time step, is the obstacle avoidance strategy of the i-th pursuer unmanned vehicle at the (k + 1)-th time step.

[0128] Taking the pursuit and evasion cost function of each pursuer unmanned vehicle at each time step as an example, the derivation process of the Q function corresponding to the pursuit and evasion cost function of each pursuer unmanned vehicle at each time step is as follows:

[0129]

[0130] The optimal Q functions corresponding to the pursuit and evasion cost functions and obstacle avoidance cost functions of the evader unmanned vehicle and each pursuer unmanned vehicle at each time step satisfy:

[0131]

[0132] At this time, the optimal pursuit-evasion strategy and the optimal obstacle avoidance strategy of all unmanned vehicles can be obtained by solving the above equations as follows:

[0133]

[0134] Among them, is the optimal pursuit-evasion strategy of the evader unmanned vehicle at the k-th time step, is the optimal obstacle avoidance strategy of the evader unmanned vehicle at the k-th time step, is the optimal pursuit-evasion strategy of the i-th pursuer unmanned vehicle at the k-th time step, is the optimal obstacle avoidance strategy of the i-th pursuer unmanned vehicle at the k-th time step, is the optimal Q function corresponding to the pursuit-evasion cost function of the evader unmanned vehicle at the k-th time step, is the optimal Q function corresponding to the obstacle avoidance cost function of the evader unmanned vehicle at the k-th time step, is the optimal Q function corresponding to the pursuit-evasion cost function of the i-th pursuer unmanned vehicle at the k-th time step, The optimal Q function corresponding to the obstacle avoidance cost function of the i-th pursuer unmanned vehicle at the k-th time step, is the optimal Q function corresponding to the pursuit-evasion cost function of the evader unmanned vehicle at the (k + 1)-th time step, is the optimal Q function corresponding to the obstacle avoidance cost function of the evader unmanned vehicle at the (k + 1)-th time step, is the optimal Q function corresponding to the pursuit-evasion cost function of the i-th pursuer unmanned vehicle at the (k + 1)-th time step, The optimal Q function corresponding to the obstacle avoidance cost function of the i-th pursuer unmanned vehicle at the (k + 1)-th time step.

[0135] In this embodiment, specifically, the Nash equilibrium solutions of the evader unmanned vehicle and the pursuer unmanned vehicle are:

[0136]

[0137] Among them, J e (.) is the pursuit-evasion cost function of the evader unmanned vehicle at the k-th time step, is the pursuit-evasion cost function of the i-th pursuer unmanned vehicle at the k-th time step, δ e (k) is the local error function of the evader unmanned vehicle at the k-th time step, is the optimal control input generated by the pursuit-evasion strategy of the evader unmanned vehicle at the k-th time step, is the optimal control input generated by the pursuit-evasion strategy of the neighboring pursuers of the evader unmanned vehicle at the k-th time step, ue (k) is the control input generated by the evader unmanned vehicle's pursuit and escape strategy at the k-th time step, is the optimal control input generated by the pursuit and escape strategy of the i-th pursuer unmanned vehicle at the k-th time step, is the optimal control input generated by the neighboring pursuers of the i-th pursuer unmanned vehicle's pursuit and escape strategy at the k-th time step, is the local error function of the i-th pursuer unmanned vehicle at the k-th time step, is the control input generated by the pursuit and escape strategy of the i-th pursuer unmanned vehicle at the k-th time step.

[0138] In this embodiment, preferably, based on the pursuit and escape cost function and the Q function corresponding to the obstacle avoidance cost function of the evader unmanned vehicle and each pursuer unmanned vehicle at each time step, the value iteration algorithm is used to obtain the target pursuit and escape strategy and obstacle avoidance strategy of each unmanned vehicle at each time step, including:

[0139] For each time step, it is judged whether the distance between each unmanned vehicle and the obstacle at the current time step is less than or equal to the detection radius of the unmanned vehicle. If it is less than or equal to, the target obstacle avoidance strategy of the unmanned vehicle at the current time step is obtained through the Q function corresponding to the obstacle avoidance cost function of the unmanned vehicle at the current time step. If it is greater than, the target pursuit and escape strategy of the unmanned vehicle at the current time step is obtained through the Q function corresponding to the pursuit and escape cost function of the unmanned vehicle at the current time step, and the target pursuit and escape strategy and obstacle avoidance strategy of each unmanned vehicle at each time step are obtained.

[0140] When the distance between the unmanned vehicle and the obstacle is less than or equal to the detection radius, the Q function corresponding to the obstacle avoidance cost function is immediately enabled to obtain the obstacle avoidance strategy, which can effectively avoid collision accidents in time and ensure the safety of the unmanned vehicle itself and the surrounding environment. When the distance is greater than the detection radius, the unmanned vehicle focuses on the pursuit and escape task, and determines the pursuit and escape strategy by means of the Q function corresponding to the pursuit and escape cost function, so that the unmanned vehicle can execute the pursuit and escape action with higher efficiency without the interference of obstacles, reduce the time and energy consumed by unnecessary obstacle avoidance actions, and improve the pursuit and escape efficiency of the unmanned vehicle.

[0141] In this embodiment, specifically, taking the i-th pursuer unmanned vehicle as an example, the value iteration algorithm is used to obtain the target pursuit and escape strategy and obstacle avoidance strategy of the i-th pursuer unmanned vehicle at the k-th time step, including:

[0142] Step S2101: Set the initial iteration step l = 0 and the maximum iteration step l max , and arbitrarily give the pursuit strategy of the i-th pursuer unmanned vehicle at the initial iteration step and the initial obstacle avoidance strategy Set a sufficiently small positive real number and the pursuit learning rate Obstacle avoidance learning rate Let the Q function corresponding to the pursuit - evasion cost function of the i - th pursuer unmanned vehicle at the initial iteration step The Q function corresponding to the obstacle - avoidance cost function of the i - th pursuer unmanned vehicle at the initial iteration step

[0143] Step S2102: Determine whether the distance between the i - th pursuer unmanned vehicle and the obstacle at the k - th time step is less than the detection radius of the unmanned vehicle. If it is less than or equal, go to step S2106; if it is greater, go to step S2103;

[0144] Step S2103: Update the pursuit Q function by solving the Q function corresponding to the pursuit - evasion cost function of the i - th pursuer unmanned vehicle at the l - th iteration at the k - th time step. The formula is:

[0145]

[0146] Where, is the pursuit strategy of the i - th pursuer unmanned vehicle at the l - th iteration at the k - th time step, is the pursuit strategy of the i - th pursuer unmanned vehicle at the (l - 1) - th iteration at the k - th time step, is the Q function corresponding to the pursuit - evasion cost function of the i - th pursuer unmanned vehicle at the (l + 1) - th iteration at the k - th time step, is the Q function corresponding to the pursuit - evasion cost function of the i - th pursuer unmanned vehicle at the l - th iteration at the k - th time step;

[0147] Step S2104: Improve the pursuit strategy of the i - th pursuer unmanned vehicle at the k - th time step along the gradient - descent direction of the updated pursuit - evasion Q function. The formula is:

[0148]

[0149] Where, represents the first - order partial derivative

[0150] Step S2105: Determine whether l = l max holds. If it holds, stop the iteration and output the target pursuit - evasion strategy of the i - th pursuer unmanned vehicle at the l max - th iteration at the k - th time step; otherwise, let l = l + 1 and go to step 3103;

[0151] Step S2106: Update the obstacle - avoidance Q function by solving the Q function corresponding to the obstacle - avoidance cost function of the i - th pursuer unmanned vehicle at the l - th iteration at the k - th time step. The formula is:

[0152]

[0153] Among them, is the obstacle avoidance strategy of the $i$-th pursuer unmanned vehicle at the $k$-th time step and the $l$-th iteration, is the obstacle avoidance strategy of the $i$-th pursuer unmanned vehicle at the $k$-th time step and the $(l - 1)$-th iteration, is the Q function corresponding to the obstacle avoidance cost function of the $i$-th pursuer unmanned vehicle at the $(l + 1)$-th iteration at the $k$-th time step, is the Q function corresponding to the obstacle avoidance cost function of the $i$-th pursuer unmanned vehicle at the $l$-th iteration at the $k$-th time step;

[0154] Step S2107: Improve the obstacle avoidance strategy of the $i$-th pursuer unmanned vehicle at the $k$-th time step along the gradient descent direction of the updated obstacle avoidance Q function. The formula is:

[0155]

[0156] Among them, represents the first-order partial derivative

[0157] Step S2108: Determine whether $l = l$ max holds. If it holds, stop the iteration and output the target obstacle avoidance strategy of the $i$-th pursuer unmanned vehicle at the $k$-th time step and the $l$ max -th iteration. Otherwise, let $l = l + 1$ and go to Step S2106.

[0158] In this embodiment, optionally, the condition for stopping the iteration in Step S2105 is:

[0159]

[0160] The condition for stopping the iteration in Step S2108 is:

[0161]

[0162] Among them, $\|\cdot\|$ is the norm.

[0163] For the evader unmanned vehicle, its value iteration Q-learning algorithm is similar to the above algorithm. It will also judge whether an obstacle enters the detection range of the evader unmanned vehicle at each time step. If it enters, update and execute the target obstacle avoidance strategy of the evader unmanned vehicle. Otherwise, update and execute the target evasion strategy of the evader unmanned vehicle.

[0164] When solving optimization problems using traditional methods, they usually rely on an initial admissible policy. However, it is not easy to determine a suitable initial admissible policy. In the actual environment, various factors are intertwined, and it is very difficult to accurately judge in advance which initial policy can effectively guide multiple unmanned vehicles to the optimal solution. The present invention uses a value iteration Q-learning algorithm to solve the optimization problem, avoiding the need for an initial admissible policy. It does not need to preset a specific initial policy, and directly explores the optimal policy step by step according to the environmental feedback in the continuous iterative learning process, greatly reducing the difficulty of policy acquisition, and better fitting the complexity and uncertainty of the actual engineering environment, and being more in line with the actual engineering.

[0165] In this embodiment, taking the i-th pursuer unmanned vehicle as an example, preferably, a value iteration algorithm is implemented based on an execution-evaluation network to obtain the target pursuit-evasion strategy and obstacle avoidance strategy of the i-th pursuer unmanned vehicle, including:

[0166] Step S2201: Let the initial iteration step l = 0 and the initial time step k = 0, and select the maximum iteration number l max and the maximum time step k max , let the initial first evaluation weight second evaluation weight of the evaluation network corresponding to the i-th pursuer unmanned vehicle be a zero vector, and set the value of the first evaluation weight of the i-th pursuer unmanned vehicle corresponding to l = 1 the value of the second evaluation weight of l = 1 Arbitrarily give the initial first execution weight and Set the first learning rate of the evaluation network corresponding to the i-th pursuer unmanned vehicle the second learning rate of the evaluation network the first learning rate of the execution network the second learning rate of the execution network Select a sufficiently small positive real number

[0167] Step S2202: Judge whether the distance between the i-th pursuer unmanned vehicle and the obstacle at the k-th time step is less than the detection radius of the unmanned vehicle. If it is less than or equal to, go to step S2208. If it is greater, go to step S2203;

[0168] Step S2203: Calculate the approximate Q-function corresponding to the pursuit-evasion cost function of the \(i\)-th pursuer unmanned vehicle at the \((l + 1)\)-th iteration in the \(k\)-th time step based on the first evaluation weight of the evaluation network corresponding to the \(i\)-th pursuer unmanned vehicle at the \((l + 1)\)-th iteration in the \(k\)-th time step; and calculate the pursuit strategy of the \(i\)-th pursuer unmanned vehicle at the \(l\)-th iteration in the \(k\)-th time step based on the first execution weight of the execution network corresponding to the \(i\)-th pursuer unmanned vehicle at the \(l\)-th iteration in the \(k\)-th time step

[0169] The formula for calculating the approximate Q-function corresponding to the pursuit-evasion cost function of the \(i\)-th pursuer unmanned vehicle at the \((l + 1)\)-th iteration in the \(k\)-th time step based on the first evaluation weight of the evaluation network corresponding to the \(i\)-th pursuer unmanned vehicle at the \((l + 1)\)-th iteration in the \(k\)-th time step is as follows:

[0170]

[0171] where represents the first evaluation weight of the evaluation network corresponding to the \(i\)-th pursuer unmanned vehicle at the \((l + 1)\)-th iteration in the \(k\)-th time step, is the approximate Q-function corresponding to the pursuit-evasion cost function at the \((l + 1)\)-th iteration in the \(k\)-th time step, is the pursuit-evasion activation function of the evaluation network corresponding to the \(i\)-th pursuer unmanned vehicle.

[0172] The role of the execution network is to approximate the control strategy and then generate a control action to be applied to the multi-unmanned vehicle scenario. Calculate the pursuit strategy of the \(i\)-th pursuer unmanned vehicle at the \(l\)-th iteration in the \(k\)-th time step based on the first execution weight of the execution network corresponding to the \(i\)-th pursuer unmanned vehicle at the \(l\)-th iteration in the \(k\)-th time step The formula is:

[0173]

[0174] where is the pursuit strategy of the \(i\)-th pursuer unmanned vehicle at the \(l\)-th iteration in the \(k\)-th time step, is the first execution weight of the execution network at the \(l\)-th iteration in the \(k\)-th time step, is the pursuit-evasion activation function of the execution network.

[0175] Step S2204: Calculate the pursuit-evasion estimation error of the evaluation network corresponding to the \(i\)-th pursuer unmanned vehicle at the \((l + 1)\)-th iteration in the \(k\)-th time step based on the approximate Q-function corresponding to the pursuit-evasion cost function of the \(i\)-th pursuer unmanned vehicle at the \((l + 1)\)-th iteration in the \(k\)-th time step. The formula is:

[0176]

[0177] wherein, is the pursuit-evasion estimation error of the evaluation network at the (l + 1)-th iteration at the k-th time step corresponding to the i-th pursuer unmanned vehicle;

[0178] To accelerate the network learning process, multiple evaluation networks are used to approximate the Q function at each iteration. Specifically, N c weight vectors are selected, and the corresponding network outputs are calculated. After updating each weight vector the evaluation network with the smallest pursuit-evasion estimation error is selected from them to improve the control strategy.

[0179] Step S2205: Update the first evaluation weight of the evaluation network corresponding to the i-th pursuer unmanned vehicle at the (l + 1)-th iteration at the k-th time step by using the gradient descent method, and set the N c weight bias terms of the first execution weight of the evaluation network corresponding to the i-th pursuer unmanned vehicle at the (l + 1)-th iteration at the k-th time step After calculating the pursuit-evasion estimation error of the evaluation network at the (l + 1)-th iteration at the k-th time step corresponding to each weight bias term, all weight bias terms are updated;

[0180] wherein, the formula for updating the first evaluation weight of the evaluation network at the (l + 1)-th iteration at the k-th time step is:

[0181]

[0182] wherein, is the optimized pursuit-evasion objective function of the evaluation network at the (l + 1)-th iteration at the k-th time step, is the first learning rate of the evaluation network;

[0183] Step S2206: Select the updated weight bias term corresponding to the minimum estimation error of the evaluation network at the (l + 1)-th iteration at the k-th time step, and make it the weight of the evaluation network at the (l + 2)-th iteration at the k-th time step, Based on the updated weight bias term corresponding to the minimum estimation error of the evaluation network at the (l + 1)-th iteration at the k-th time step, recalculate the approximate Q function corresponding to the pursuit-evasion cost function at the (l + 1)-th iteration at the k-th time step, and update the first execution weight of the execution network at the (l + 1)-th iteration at the k-th time step;

[0184] Since the purpose of updating the control strategy is to minimize the corresponding Q function, therefore, combining the definition of the evaluation network, the objective function for optimizing the execution network is expressed as:

[0185]

[0186] Based on the gradient descent method and the chain rule, the first execution weight update rule for the execution network corresponding to the $i$-th pursuer unmanned vehicle at the $(l + 1)$-th iteration at the $k$-th time step is as follows:

[0187]

[0188]

[0189] where, is the first learning rate of the execution network corresponding to the $i$-th pursuer unmanned vehicle.

[0190] Step S2207: Determine whether $l = l max holds. If it holds, set as the zero vector and go to Step S2203. Otherwise, set $l = l + 1$ and go to Step S2213;

[0191] Step S2208: Calculate the approximate Q-function corresponding to the obstacle avoidance cost function for the $i$-th pursuer unmanned vehicle at the $(l + 1)$-th iteration at the $k$-th time step based on the second evaluation weight of the evaluation network corresponding to the $i$-th pursuer unmanned vehicle at the $(l + 1)$-th iteration at the $k$-th time step; and calculate the obstacle avoidance strategy for the $i$-th pursuer unmanned vehicle at the $l$-th iteration at the $k$-th time step according to the second execution weight of the execution network corresponding to the $i$-th pursuer unmanned vehicle at the $(l + 1)$-th iteration at the $k$-th time step

[0192] The formula for calculating the approximate Q-function corresponding to the obstacle avoidance cost function for the $i$-th pursuer unmanned vehicle at the $(l + 1)$-th iteration at the $k$-th time step based on the second evaluation weight of the evaluation network corresponding to the $i$-th pursuer unmanned vehicle at the $(l + 1)$-th iteration at the $k$-th time step is:

[0193]

[0194] where, represents the approximate Q-function corresponding to the obstacle avoidance cost function for the $i$-th pursuer unmanned vehicle at the $(l + 1)$-th iteration at the $k$-th time step, is the second evaluation weight of the evaluation network corresponding to the $i$-th pursuer unmanned vehicle at the $(l + 1)$-th iteration at the $k$-th time step, is the obstacle avoidance activation function of the evaluation network corresponding to the $i$-th pursuer unmanned vehicle.

[0195] Calculate the obstacle avoidance strategy for the $i$-th pursuer unmanned vehicle at the $l$-th iteration at the $k$-th time step according to the second execution weight of the execution network corresponding to the $i$-th pursuer unmanned vehicle at the $(l + 1)$-th iteration at the $k$-th time step The formula is:

[0196]

[0197] Wherein, is the obstacle avoidance strategy of the i-th pursuer unmanned vehicle at the l-th iteration in the k-th time step, is the second execution weight of the execution network corresponding to the i-th pursuer unmanned vehicle at the l-th iteration in the k-th time step, is the obstacle avoidance activation function of the execution network corresponding to the i-th pursuer unmanned vehicle.

[0198] Step S2209: Based on the approximate Q function corresponding to the obstacle avoidance cost function of the i-th pursuer unmanned vehicle at the (l + 1)-th iteration in the k-th time step, calculate the obstacle avoidance estimation error of the evaluation network corresponding to the i-th pursuer unmanned vehicle at the (l + 1)-th iteration in the k-th time step. The formula is:

[0199]

[0200] Wherein, is the obstacle avoidance estimation error of the evaluation network corresponding to the i-th pursuer unmanned vehicle at the (l + 1)-th iteration in the k-th time step;

[0201] Step S2210: Use the gradient descent method to update the second execution weight of the execution network corresponding to the i-th pursuer unmanned vehicle at the (l + 1)-th iteration in the k-th time step. Let

[0202] Wherein, the formula for updating the second execution weight of the execution network corresponding to the i-th pursuer unmanned vehicle at the (l + 1)-th iteration in the k-th time step is:

[0203]

[0204] Wherein, is the optimized obstacle avoidance objective function of the evaluation network corresponding to the i-th pursuer unmanned vehicle at the (l + 1)-th iteration in the k-th time step, is the second learning rate of the evaluation network corresponding to the i-th pursuer unmanned vehicle;

[0205] Step S2211: Based on the updated second execution weight of the execution network corresponding to the i-th pursuer unmanned vehicle at the (l + 1)-th iteration in the k-th time step, recalculate the approximate Q function corresponding to the obstacle avoidance cost function of the i-th pursuer unmanned vehicle at the (l + 1)-th iteration in the k-th time step, and update the second execution weight of the execution network corresponding to the i-th pursuer unmanned vehicle at the (l + 1)-th iteration in the k-th time step;

[0206] Since the purpose of the update control strategy is to minimize the corresponding Q function, combined with the definition of the evaluation network, the objective function for optimizing the execution network is expressed as:

[0207]

[0208] Based on the gradient descent method and the chain rule, the weight update rule is:

[0209]

[0210] where, is the second learning rate of the execution network corresponding to the i-th pursuer unmanned vehicle.

[0211] Step S2212: Determine whether l = l max holds. If it holds, let be the zero vector and go to step S2213. Otherwise, let l = l + 1 and go to step S2208;

[0212] Step S2213: Determine whether k = k max holds. If it holds, stop the iteration and output the target pursuit and evasion strategy of the i-th pursuer unmanned vehicle at the k max -th time step. Otherwise, let k = k + 1 and go to step 3202.

[0213] In this embodiment, optionally, the condition for determining weight reset in step S2207 is:

[0214]

[0215] The condition for determining weight reset in step S2212 is:

[0216]

[0217] The condition for stopping the iteration in step S2213 is:

[0218]

[0219] For the evader unmanned vehicle, the approximate Q function corresponding to the pursuit and evasion cost function of the evader unmanned vehicle at the (l + 1)-th iteration at the k-th time step, and the approximate Q function corresponding to the obstacle avoidance cost function of the evader unmanned vehicle at the (l + 1)-th iteration at the k-th time step, the formula is:

[0220]

[0221] where, represents the approximate Q function corresponding to the pursuit and evasion cost function of the evader unmanned vehicle at the (l + 1)-th iteration at the k-th time step, The first evaluation weight of the evaluation network corresponding to the evader unmanned vehicle at the (l + 1)-th iteration in the k-th time step is ψ e (·) is the pursuit-evasion activation function of the evaluation network denotes the approximate Q-function corresponding to the obstacle avoidance cost function of the evader unmanned vehicle at the (l + 1)-th iteration in the k-th time step The second evaluation weight of the evaluation network corresponding to the evader unmanned vehicle at the (l + 1)-th iteration in the k-th time step is is the obstacle avoidance activation function of the evaluation network

[0222] The iterative formulas for updating the first evaluation weight and the second evaluation weight of the evaluation network corresponding to the evader unmanned vehicle at the (l + 1)-th iteration in the k-th time step based on gradient descent are as follows:

[0223]

[0224] where is the updated first evaluation weight of the evaluation network corresponding to the evader unmanned vehicle at the (l + 1)-th iteration in the k-th time step is the updated second evaluation weight of the evaluation network corresponding to the evader unmanned vehicle at the (l + 1)-th iteration in the k-th time step, γ ce is the first learning rate of the evaluation network corresponding to the evader unmanned vehicle is the second learning rate of the evaluation network corresponding to the evader unmanned vehicle is the pursuit-evasion estimation error of the evaluation network corresponding to the evader unmanned vehicle at the (l + 1)-th iteration in the k-th time step is the obstacle avoidance estimation error of the evaluation network corresponding to the evader unmanned vehicle at the (l + 1)-th iteration in the k-th time step, ψ e (·) is the pursuit-evasion activation function of the evaluation network corresponding to the evader unmanned vehicle is the obstacle avoidance activation function of the evaluation network corresponding to the evader unmanned vehicle is the pursuit strategy of the evader unmanned vehicle at the l-th iteration in the k-th time step is the obstacle avoidance strategy of the evader unmanned vehicle at the l-th iteration in the k-th time step

[0225] The formulas for the approximate Q-function corresponding to the pursuit-evasion cost function of the evader unmanned vehicle at the (l + 1)-th iteration in the k-th time step and the approximate Q-function corresponding to the obstacle avoidance cost function of the evader unmanned vehicle at the (l + 1)-th iteration in the k-th time step are as follows:

[0226]

[0227] where Denote the approximate Q - function corresponding to the pursuit - evasion cost function of the evader unmanned vehicle at the \(k\) - th time step and the \((l + 1)\) - th iteration. is the first evaluation weight of the evaluation network corresponding to the evader unmanned vehicle at the \(k\) - th time step and the \((l + 1)\) - th iteration, \(\psi\) e (·) is the pursuit - evasion activation function of the evaluation network corresponding to the evader unmanned vehicle. Denote the approximate Q - function corresponding to the obstacle - avoidance cost function of the evader unmanned vehicle at the \(k\) - th time step and the \((l + 1)\) - th iteration. is the second evaluation weight of the evaluation network corresponding to the evader unmanned vehicle at the \(k\) - th time step and the \((l + 1)\) - th iteration. is the obstacle - avoidance activation function of the evaluation network corresponding to the evader unmanned vehicle.

[0228] Calculate the pursuit strategy of the evader unmanned vehicle at the \(k\) - th time step and the \(l\) - th iteration according to the first execution weight of the execution network corresponding to the evader unmanned vehicle at the \(k\) - th time step and the \(l\) - th iteration. The formula is:

[0229]

[0230] The update iteration formula for the first execution weight of the execution network corresponding to the evader unmanned vehicle at the \(k\) - th time step and the \((l + 1)\) - th iteration is:

[0231]

[0232] where \(\gamma\) ae is the first learning rate of the execution network corresponding to the evader unmanned vehicle.

[0233] Calculate the obstacle - avoidance strategy of the evader unmanned vehicle at the \(k\) - th time step and the \(l\) - th iteration according to the second execution weight of the execution network corresponding to the evader unmanned vehicle at the \(k\) - th time step and the \(l\) - th iteration. The formula is:

[0234]

[0235] where is the obstacle - avoidance strategy of the evader unmanned vehicle at the \(k\) - th time step and the \(l\) - th iteration. is the second execution weight of the execution network corresponding to the evader unmanned vehicle at the \(k\) - th time step and the \(l\) - th iteration. is the obstacle - avoidance activation function of the execution network corresponding to the evader unmanned vehicle.

[0236] The update iteration formula for the second execution weight of the execution network corresponding to the evader unmanned vehicle at the \(k\) - th time step and the \((l + 1)\) - th iteration is:

[0237]

[0238] wherein, is the second learning rate of the execution network corresponding to the escapee unmanned vehicle.

[0239] The present invention introduces multiple evaluation networks to approximate the Q function, which can enable the local error of the unmanned vehicle to converge faster, so that the pursuit-evasion game can reach the Nash equilibrium faster. The pursuit-evasion game involves the strategic confrontation between the pursuer and the escapee. Both sides are constantly adjusting their own strategies according to the actions of the other side and environmental changes to maximize their own interests. Under the traditional method, due to the inaccurate and inefficient approximation of the Q function, there is a large blindness in the process of both sides adjusting their strategies, and it is difficult to quickly find the optimal confrontation strategy, resulting in a long pursuit-evasion game process and it is difficult to stably reach the Nash equilibrium state. The multiple evaluation networks, through multi-dimensional information processing and more accurate approximation of the Q function, enable both the pursuer and the escapee to more quickly and accurately understand the strategies of the other side and the impact of environmental changes on the strategies, and at the same time the pursuit-evasion game reaches the Nash equilibrium faster.

[0240] Step S3: Based on the target pursuit-evasion strategy and obstacle avoidance strategy of each unmanned vehicle at each time step, generate the control input of the target pursuit-evasion strategy and obstacle avoidance strategy of each unmanned vehicle at each time step.

[0241] Based on Embodiment 1, this Embodiment 2 is provided with four pursuer unmanned vehicles and one escapee unmanned vehicle, constituting a multi-unmanned vehicle obstacle avoidance pursuit-evasion scenario. As Figure 4 shown, Figure 4 is the communication topology diagram of the multi-unmanned vehicle. Through this diagram, the information interaction relationship between each vehicle can be clearly understood. The pursuer unmanned vehicle can obtain information such as the position of the escapee through communication to implement the pursuit strategy, and at the same time share information with each other to optimize the formation and encirclement operations, which has an important impact on the cooperative operation and strategy formulation between vehicles.

[0242] As Figure 5 shown, Figure 5 is the coordinate change trajectory diagram of the multi-unmanned vehicle. From the overall trend of the curve, the curves of the pursuer unmanned vehicles gradually approach the curve of the escapee unmanned vehicle, indicating that during the execution of the pursuit-evasion strategy by the pursuers, the position gap with the escapee is continuously reduced, verifying the effective pursuit action of the pursuers against the escapee under the pursuit-evasion game method.

[0243] As Figure 6 shown, Figure 6It is the local error trajectory diagram of multiple unmanned vehicles. As the time step increases, the error curves of each vehicle generally show a downward trend, gradually becoming stable and converging to a smaller value. The fact that the errors gradually converge indicates that during the operation of multiple unmanned vehicles, they can continuously adjust and optimize their own states, reducing the state deviations caused by various factors (such as environmental interference, model errors, etc.), which reflects that the system has good stability and self - adaptability and can maintain relatively stable operation during the pursuit - evasion game process.

[0244] As Figure 7 shown, Figure 7 It is the coordinate change trajectory diagram of multiple unmanned vehicles in a two - dimensional coordinate system. It can be seen from the simulation diagram that when an obstacle (red solid - line circle) enters the detection range (red dashed - line circle) of the unmanned vehicle, the unmanned vehicle will adopt an obstacle - avoidance strategy and thus move away from the obstacle. When a safe distance is maintained from the obstacle, all unmanned vehicles will adopt a pursuit - evasion strategy and finally reach the Nash equilibrium of the game. At this time, the pursuer unmanned vehicles can surround the evader unmanned vehicle in a preset diamond formation, which verifies the effectiveness of the multi - unmanned vehicle obstacle - avoidance and pursuit - evasion game method based on value - iteration Q - learning.

[0245] Embodiment 3 of the present invention provides a multi - unmanned vehicle obstacle - avoidance and pursuit - evasion game system, including:

[0246] A cost - function construction module, used to construct the pursuit - evasion cost function and the obstacle - avoidance cost function of the evader unmanned vehicle and each pursuer unmanned vehicle at each time step, including:

[0247] Based on the local error function of the current unmanned vehicle at each time step and the control input generated by the pursuit - evasion strategy of the current unmanned vehicle at each time step, construct the pursuit - evasion cost function of the current unmanned vehicle at each time step;

[0248] Based on the safety radius of the current unmanned vehicle, the safety radius of the obstacle, and the distance between the current unmanned vehicle and the obstacle at each time step, construct the relative safety function of the current unmanned vehicle at each time step; set a safety - distance constraint: the relative safety function of the current unmanned vehicle at each time step is not greater than the safety threshold;

[0249] Through the gradient - re - centered self - consistent barrier function, convert the safety - distance constraint into the obstacle penalty term of the current unmanned vehicle at each time step;

[0250] Based on the relative safety function of the current unmanned vehicle at each time step, the control input generated by the obstacle - avoidance strategy of the current unmanned vehicle at each time step, and the obstacle penalty term of the current unmanned vehicle at each time step, construct the obstacle - avoidance cost function of the current unmanned vehicle at each time step;

[0251] A strategy solving module, which is used to adopt a value iteration algorithm based on the pursuit-evasion cost function and the Q function corresponding to the obstacle avoidance cost function of the evader unmanned vehicle and each pursuer unmanned vehicle at each time step, so as to obtain the target pursuit-evasion strategy and obstacle avoidance strategy of each unmanned vehicle at each time step;

[0252] A control generation module, which is used to generate the control input of the target pursuit-evasion strategy and obstacle avoidance strategy of each unmanned vehicle at each time step based on the target pursuit-evasion strategy and obstacle avoidance strategy of each unmanned vehicle at each time step.

[0253] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0254] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0255] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0256] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1Steps of functions specified in one or more boxes.

[0257] Obviously, the above embodiments are merely examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or alterations can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. And the obvious changes or alterations derived therefrom still fall within the protection scope of the present invention.

Claims

1. A multi-unmanned vehicle obstacle avoidance pursuit-evasion game method, characterized in that, Including: Constructing the pursuit-evasion cost function and obstacle avoidance cost function for the evader unmanned vehicle and each pursuer unmanned vehicle at each time step, including: Based on the local error function of the current unmanned vehicle at each time step and the control input generated by the pursuit-evasion strategy of the current unmanned vehicle at each time step, constructing the pursuit-evasion cost function of the current unmanned vehicle at each time step; Based on the safety radius of the current unmanned vehicle, the safety radius of the obstacle, and the distance between the current unmanned vehicle and the obstacle at each time step, constructing the relative safety function of the current unmanned vehicle at each time step; setting a safety distance constraint: the relative safety function of the current unmanned vehicle at each time step is not greater than the safety threshold; Converting the safety distance constraint into an obstacle penalty term of the current unmanned vehicle at each time step through the gradient re-centered self-coordinated barrier function; Based on the relative safety function of the current unmanned vehicle at each time step, the control input generated by the obstacle avoidance strategy of the current unmanned vehicle at each time step, and the obstacle penalty term of the current unmanned vehicle at each time step, constructing the obstacle avoidance cost function of the current unmanned vehicle at each time step; Based on the Q functions corresponding to the pursuit-evasion cost functions and obstacle avoidance cost functions of the evader unmanned vehicle and each pursuer unmanned vehicle at each time step, using the value iteration algorithm to obtain the target pursuit-evasion strategies and obstacle avoidance strategies of each unmanned vehicle at each time step; Based on the target pursuit-evasion strategies and obstacle avoidance strategies of each unmanned vehicle at each time step, generating the control inputs of the target pursuit-evasion strategies and obstacle avoidance strategies of each unmanned vehicle at each time step.

2. The multi-unmanned vehicle obstacle avoidance pursuit and escape game method according to claim 1, wherein The relative safety function of the evader unmanned vehicle at each time step is: η e (k) = (r e + r o ) / ξ e (k), The relative safety function of each pursuer unmanned vehicle at each time step is: Among them, η e (k) is the relative safety function of the escaping unmanned vehicle at the k-th time step, r e is the safety radius of the escaping unmanned vehicle, r o is the safety radius of the obstacle, ξ e (k) is the distance between the escaping unmanned vehicle and the obstacle at the k-th time step, is the dynamic model of the escaping unmanned vehicle at the k-th time step, x e (k) is the horizontal coordinate of the escaping unmanned vehicle at the k-th time step, y e (k) is the vertical coordinate of the escaping unmanned vehicle at the k-th time step, is the two-dimensional coordinate of the obstacle, is the relative safety function of the i-th pursuing unmanned vehicle at the k-th time step, is the safety radius of the i-th pursuing unmanned vehicle, r o is the safety radius of the obstacle, is the distance between the i-th pursuing unmanned vehicle and the obstacle at the k-th time step, is the dynamic model of the i-th pursuing unmanned vehicle at the k-th time step, is the horizontal coordinate of the i-th pursuing unmanned vehicle at the k-th time step, is the vertical coordinate of the i-th pursuing unmanned vehicle at the k-th time step.

3. A multi-unmanned vehicle obstacle avoidance pursuit-evasion game method according to claim 1, characterized in that, The obstacle penalty term of the evader unmanned vehicle at each time step is as follows: where is the gradient re-centered self-concordant barrier function, and the gradient re-centered self-concordant barrier function converts the safety distance constraint η e (k + n) ≤ v e into the obstacle penalty term of the evader unmanned vehicle at the k-th time step v e is the safety threshold of the evader unmanned vehicle, and η e (k + n) is the relative safety function of the evader unmanned vehicle at the (k + n)-th time step; The obstacle penalty term of the $i$-th pursuer unmanned vehicle at each time step is: Among them, the gradient re-centered self-coordinated obstacle function converts the safety distance constraint of the $i$-th pursuer unmanned vehicle into the obstacle penalty term of the $i$-th pursuer unmanned vehicle at the $k$-th time step is the safety threshold of the $i$-th pursuer unmanned vehicle, is the relative safety function of the $i$-th pursuer unmanned vehicle at the $(k + n)$-th time step.

4. A multi-unmanned vehicle obstacle avoidance pursuit-evasion game method according to claim 1, characterized in that, The pursuit-evasion cost function of the evader unmanned vehicle at each time step is: The pursuit-evasion cost function of each pursuer unmanned vehicle at each time step is: Among them, J e (δ e (k), u e (k)) is the pursuit and evasion cost function of the evader unmanned vehicle at the k-th time step, δ e (k) is the local error function of the evader unmanned vehicle at the k-th time step, v e (k + n) is the local error function of the evader unmanned vehicle at the (k + n)-th time step, u e (k) = [V e (k), θ e (k)] T , (·) T is the transpose, u e (k) is the control input generated by the pursuit and evasion strategy of the evader unmanned vehicle at the k-th time step, u e (k + n) is the control input generated by the pursuit and evasion strategy of the evader unmanned vehicle at the (k + n)-th time step, V e (k) is the speed of the evader unmanned vehicle at the k-th time step, δ e (k) is the heading angle of the evader unmanned vehicle at the k-th time step, U e (·) is the single-step cost of the evader unmanned vehicle regarding pursuit and evasion, P e is the first symmetric positive definite matrix of the evader unmanned vehicle, R e is the second symmetric positive definite matrix of the evader unmanned vehicle, P e , is the pursuit and evasion cost function of the i-th pursuer unmanned vehicle at the k-th time step, is the local error function of the i-th pursuer unmanned vehicle at the k-th time step, the local error function of the i-th pursuer unmanned vehicle at the (k + n)-th time step, is the control input generated by the pursuit and evasion strategy of the i-th pursuer unmanned vehicle at the k-th time step, is the control input generated by the pursuit and evasion strategy of the i-th pursuer unmanned vehicle at the (k + n)-th time step, is the speed of the i-th pursuer unmanned vehicle at the k-th time step, is the heading angle of the i-th pursuer unmanned vehicle at the k-th time step, is the single-step cost of the i-th pursuer unmanned vehicle regarding pursuit and evasion, is the first symmetric positive definite matrix of the i-th pursuer unmanned vehicle, is the second symmetric positive definite matrix of the i-th pursuer unmanned vehicle, k is the time step, and n is the time step offset index.

5. A multi-unmanned vehicle obstacle avoidance pursuit and escape game method according to claim 1, characterized in that, The obstacle avoidance cost function of the evader unmanned vehicle at each time step is: The obstacle avoidance cost function of each pursuer unmanned vehicle at each time step is: Among them, is the obstacle avoidance cost function of the escaping unmanned vehicle at the k-th time step, and η e (k) is the relative safety function of the escaping unmanned vehicle at the k-th time step, and η e (k + n) is the relative safety function of the escaping unmanned vehicle at the (k + n)-th time step, (·) T is the transpose, is the control input generated by the obstacle avoidance strategy of the escaping unmanned vehicle at the k-th time step, is the control input generated by the pursuit and escape strategy of the escaping unmanned vehicle at the (k + n)-th time step, V e (k) is the speed of the escaping unmanned vehicle at the k-th time step, θ e (k) is the heading angle of the escaping unmanned vehicle at the k-th time step, β e is the obstacle avoidance parameter of the escaping unmanned vehicle, β e > 0, L e (·) is the single-step cost of the escaping unmanned vehicle regarding obstacle avoidance, is the third symmetric positive definite matrix of the escaping unmanned vehicle, is the fourth symmetric positive definite matrix of the escaping unmanned vehicle, is the obstacle avoidance cost function of the i-th pursuer unmanned vehicle at the k-th time step, is the relative safety function of the i-th unmanned vehicle at the k-th time step, is the relative safety function of the i-th unmanned vehicle at the (k + n)-th time step, is the control input generated by the obstacle avoidance strategy of the i-th pursuer unmanned vehicle at the k-th time step, is the control input generated by the pursuit and escape strategy of the i-th pursuer unmanned vehicle at the (k + n)-th time step, is the speed of the i-th pursuer unmanned vehicle at the k-th time step, is the heading angle of the i-th pursuer unmanned vehicle at the k-th time step, is the obstacle avoidance parameter of the i-th pursuer unmanned vehicle, is the single-step cost of the i-th pursuer unmanned vehicle regarding obstacle avoidance, is the third symmetric positive definite matrix of the i-th pursuer unmanned vehicle, is the fourth symmetric positive definite matrix of the i-th pursuer unmanned vehicle, k is the time step, and n is the time step offset index.

6. A multi-unmanned vehicle obstacle avoidance pursuit-evasion game method according to claim 1, characterized in that The value range of the safety threshold of the escape unmanned vehicle is The value range of the safety threshold of the i-th pursuer unmanned vehicle is where r e is the safety radius of the escape unmanned vehicle, and τ e is the detection radius of the escape unmanned vehicle. is the safety radius of the i-th pursuer unmanned vehicle, r o is the safety radius of the obstacle, is the detection radius of the i-th pursuer unmanned vehicle.

7. A multi-unmanned vehicle obstacle avoidance pursuit and escape game method according to claim 1, characterized in that The local error function of the evader unmanned vehicle at each time step is: The local error function of each pursuer unmanned vehicle at each time step is: where, δ e (k) is the local error function of the evader unmanned vehicle at the k-th time step, is the local error function of the i-th pursuer unmanned vehicle at the k-th time step, is the set of neighbor pursuer unmanned vehicles of the evader unmanned vehicle, a ei is the connection gain coefficient between the evader unmanned vehicle and the i-th pursuer unmanned vehicle, is the set of neighbor pursuer unmanned vehicles of the i-th pursuer unmanned vehicle, is the dynamic model of the evader unmanned vehicle at the k-th time step, is the dynamic model of the i-th pursuer unmanned vehicle at the k-th time step, is the expected displacement between the i-th pursuer unmanned vehicle and the j-th neighbor pursuer unmanned vehicle, is the expected displacement between the i-th pursuer unmanned vehicle and the evader unmanned vehicle, where i is the index of the pursuer unmanned vehicle and j is the index of the neighbor pursuer unmanned vehicle.

8. A multi-unmanned vehicle obstacle avoidance pursuit and escape game method according to claim 7, characterized in that, The dynamic model of the escape unmanned vehicle at the k-th time step is x e (k), and the update iteration formula for y e (k) is as follows: The dynamic model of each pursuer unmanned vehicle at the k-th time step is The update iteration formula of is: where x e (k + 1) is the horizontal coordinate of the escaping unmanned vehicle at the (k + 1)-th time step, y e (k + 1) is the vertical coordinate of the escaping unmanned vehicle at the (k + 1)-th time step, x e (k) is the horizontal coordinate of the escaping unmanned vehicle at the k-th time step, y e (k) is the vertical coordinate of the escaping unmanned vehicle at the k-th time step, T is the sampling interval, V e (k) is the speed of the escaping unmanned vehicle at the k-th time step, θ e (k) is the heading angle of the escaping unmanned vehicle at the k-th time step, is the horizontal coordinate of the i-th pursuing unmanned vehicle at the (k + 1)-th time step, y e (k + 1) is the vertical coordinate of the i-th pursuing unmanned vehicle at the (k + 1)-th time step, is the horizontal coordinate of the i-th pursuing unmanned vehicle at the k-th time step, is the vertical coordinate of the i-th pursuing unmanned vehicle at the k-th time step, is the speed of the i-th pursuing unmanned vehicle at the k-th time step, is the heading angle of the i-th pursuing unmanned vehicle at the k-th time step, k represents the time step, [·] T is the transpose.

9. A multi-unmanned vehicle obstacle avoidance pursuit and escape game method according to claim 1, characterized in that, The above-mentioned method of using the value iteration algorithm to obtain the target pursuit-evasion strategies and obstacle avoidance strategies of each unmanned vehicle at each time step based on the Q functions corresponding to the pursuit-evasion cost functions and obstacle avoidance cost functions of the evader unmanned vehicle and each pursuer unmanned vehicle at each time step includes: For each time step, judging whether the distance between each unmanned vehicle and the obstacle at the current time step is less than or equal to the detection radius of the unmanned vehicle. If it is less than or equal to, obtaining the target obstacle avoidance strategy of the unmanned vehicle at the current time step through the Q function corresponding to the obstacle avoidance cost function of the unmanned vehicle at the current time step. If it is greater than, obtaining the target pursuit-evasion strategy of the unmanned vehicle at the current time step through the Q function corresponding to the pursuit-evasion cost function of the unmanned vehicle at the current time step, so as to obtain the target pursuit-evasion strategies and obstacle avoidance strategies of each unmanned vehicle at each time step.

10. A multi-unmanned vehicle obstacle avoidance pursuit and escape game system, characterized in that, Including: A cost function construction module for constructing the pursuit-evasion cost function and obstacle avoidance cost function for the evader unmanned vehicle and each pursuer unmanned vehicle at each time step, including: Based on the local error function of the current unmanned vehicle at each time step and the control input generated by the pursuit-evasion strategy for the current unmanned vehicle at each time step, construct the pursuit-evasion cost function of the current unmanned vehicle at each time step; Based on the safety radius of the current unmanned vehicle, the safety radius of the obstacle, and the distance between the current unmanned vehicle and the obstacle at each time step, construct the relative safety function of the current unmanned vehicle at each time step; Set the safety distance constraint: the relative safety function of the current unmanned vehicle at each time step is not greater than the safety threshold; Through the gradient re-centered self-coordinated barrier function, convert the safety distance constraint into the barrier penalty term of the current unmanned vehicle at each time step; Based on the relative safety function of the current unmanned vehicle at each time step, the control input generated by the obstacle avoidance strategy for the current unmanned vehicle at each time step, and the barrier penalty term of the current unmanned vehicle at each time step, construct the obstacle avoidance cost function of the current unmanned vehicle at each time step; The policy solving module is used to obtain the target pursuit-evasion strategy and obstacle avoidance strategy of each unmanned vehicle at each time step by using the value iteration algorithm based on the Q functions corresponding to the pursuit-evasion cost function and the obstacle avoidance cost function of the evader unmanned vehicle and each pursuer unmanned vehicle at each time step; The control generation module is used to generate the control input of the target pursuit-evasion strategy and obstacle avoidance strategy of each unmanned vehicle at each time step based on the target pursuit-evasion strategy and obstacle avoidance strategy of each unmanned vehicle at each time step.