Multi-unmanned aerial vehicle hunting method and system
By combining dynamic modeling, extreme value search and adaptive entropy in the multi-UAV roundup system, the Actor network and Critic network are optimized, and the problem of multi-UAV roundup in the existing technology is easily trapped in local optimality in a three-dimensional environment, improving the roundup efficiency.
Patent Information
- Application Number
- CN202510178134.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-27
AI Technical Summary
The existing multi-UAV roundup method is prone to falling into local optimization in a three-dimensional environment, limiting the global exploration ability of the strategy, resulting in insufficiency of roundup.
By dynamically modeling the drone, a multi-UAV seven-tuple in a three-dimensional dynamic environment is constructed, and combined with extreme search and adaptive entropy, the Actor network and the Critic network are optimized to generate optimal actions to complete the roundup task.
The optimization training of the network is realized, and does not rely on gradient information, and can more efficiently approach the global optimal solution, avoid the algorithm from falling into local optimality, and improve the efficiency of drone rounding.
Smart Images

Figure CN120044983A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of UAV cooperative rescue, agriculture, military patrol and control systems, and particularly to a multi-UAV encirclement method and system. Background Art
[0002] In recent years, multi-UAV cooperation technology has developed rapidly and shown broad application prospects in many fields such as military, logistics, agriculture, and disaster relief. In the military field, multi-UAV systems have become one of the key means to achieve complex combat tasks due to their high flexibility, high mobility, and intelligent characteristics. As an important application scenario of multi-UAV cooperative operation, the encirclement task can be used for reconnaissance, intercepting enemy targets or performing defense tasks, and its importance is self-evident. However, the encirclement task poses extremely high requirements for the cooperative planning, real-time decision-making, and intelligent control of multi-UAVs. Especially in a complex three-dimensional environment, how to efficiently complete the encirclement task has become a difficult and hot issue in current research.
[0003] The solution of multi-UAV encirclement strategies is mainly divided into two paths: based on traditional control methods and based on reinforcement learning methods. Traditional control methods, such as rule-based behavior planning or optimal control methods, can provide theoretical guidance for multi-UAVs to a certain extent and ensure the stability of the system. However, when facing high-dimensional dynamic environments and multi-objective cooperative tasks, the computational complexity of these methods is often too high, and it is difficult to meet the real-time requirements. In contrast, reinforcement learning-based methods, especially multi-agent deep reinforcement learning (such as MADDPG, MAPPO, MASAC), can obtain better encirclement strategies in dynamic and uncertain environments through the self-learning ability of agents. However, when the current reinforcement learning methods are applied in a three-dimensional environment, they are prone to fall into local optima due to the too large policy search space, resulting in low encirclement efficiency. For example, a multi-UAV encirclement method combining state prediction and DDPG disclosed in Chinese Patent Publication No. CN113625775A uses a traditional DDPG network for training, and there are problems such as gradient descent optimization and possible local optima. In addition, the way that reinforcement learning methods rely too much on gradient descent to solve the loss function may limit the global exploration ability of the strategy. Therefore, existing methods still face many challenges in three-dimensional multi-UAV encirclement scenarios. Summary of the Invention
[0004] The technical problem to be solved by the present invention is that the existing multi-UAV encirclement methods are prone to fall into local optima, limit the global exploration ability of the strategy, and have low encirclement efficiency.
[0005] The present invention solves the above technical problems by the following technical means: A multi-UAV encirclement method, comprising:
[0006] S1. Perform dynamic modeling on the unmanned aerial vehicle (UAV).
[0007] S2. Construct a seven-tuple of multiple UAVs in a three-dimensional dynamic environment to describe the scenario of multiple UAVs surrounding and capturing a single UAV.
[0008] S3. Each UAV includes an Actor network and a Critic network. The Actor network generates the actions of the UAV, and the Critic network generates the evaluation of the corresponding actions. Combine extreme value search and adaptive entropy and continuously iterate to update the Actor network and the Critic network. Use the optimized Actor network to generate the optimal actions, and each UAV executes the optimal actions to complete the surrounding and capture of a single UAV by multiple UAVs.
[0009] The present invention combines extreme value search and adaptive entropy to update the network, realizes the optimized training of the network, does not rely on gradient information, can more efficiently approach the global optimal solution, avoids the algorithm falling into the local optimum, and improves the surrounding and capture efficiency of the UAV.
[0010] Further, S1 includes:
[0011] There are N UAVs, where N > 1. The discrete dynamic equation of the multiple UAVs from time t to time t + 1 is:
[0012]
[0013] where i is the index of the UAV, UAV is the abbreviation of the unmanned aerial vehicle, x, y, z are the x-axis, y-axis, and z-axis coordinates of the UAV, v is the UAV speed, ψ is the UAV heading angle, θ is the UAV pitch angle, φ is the UAV roll angle, ω ψ,i is the heading angular velocity control input, ω θ,i is the pitch angular velocity control input, ω φ,i is the roll angular velocity control input, and Δt is the time step from time t to time t + 1.
[0014] Furthermore, the constraint conditions for the flight states of the multiple UAVs are:
[0015]
[0016] where x i,min , x i,max , y i,min , y i,max , z i,min , z i,max are the minimum and maximum values of the i-th UAV on the x-axis, y-axis, and z-axis respectively, v i,min , v i,max are the minimum and maximum values of the speed of the i-th UAV, ψ i,min , ψ i,maxare the minimum and maximum values of the heading angle of the i-th UAV, θ i,min , θ i,max are the minimum and maximum values of the pitch angle of the i-th UAV, φ i,min , φ i,max are the minimum and maximum values of the roll angle of the i-th UAV, u is the acceleration control input of the UAV, u i,min , u i,max are the minimum and maximum values of the acceleration control input of the i-th UAV, ω ψ,i,min , ω ψ,i,max are the minimum and maximum values of the heading angular velocity input of the i-th UAV, ω θ,i,min , ω θ,i,max are the minimum and maximum values of the pitch angular velocity input of the i-th UAV, ω φ,i,min , ω φ,i,max are the minimum and maximum values of the roll angular velocity input of the i-th UAV; the obstacle is set as a circular area, (x O , y O , z O ) are the x-axis, y-axis and z-axis coordinates of the center of the sphere of the obstacle, R O is the radius of the obstacle.
[0017] Furthermore, S2 includes:
[0018] Construct a seven-tuple <S, A, R, P, Z, O, γ> of multiple UAVs in a three-dimensional dynamic environment, where S is the state set, S = {s 1 , s 2 ,..., s i ,..., s N}, s i is the state of the i-th UAV and s i = {x i , y i , z i , v i , ψ i , θ i , φ i}; A is the action set and A = {a 1 , a 2 ,..., a i ,..., a N}, a i is the action of the i-th UAV, a i = [u i , ω ψ,i , ω θ,i , ω φ,i T , where T represents the transpose operation; R represents the reward obtained by the UAV taking corresponding actions in a certain state; P is the probability that the UAV transfers from one state to another by taking an action; Z is the set of observations, Z = {z 1 , z 2 ,..., z i ,..., z N}, where z i is the observation information of the i-th UAV, and z i = {x i , y i , z i , v i , ψ i , θ i , φ i}; O is the observation function; γ is the discount factor, representing the preference and proportion of the current reward and future rewards.
[0019] Furthermore, S3 includes:
[0020] S31. The reinforcement learning method in this design is divided into a main network and a target network. The main network is responsible for interacting with the environment. The target network is a copy of the main network, with the same structure as the main network, but a lower parameter update frequency. The parameters of the target network are obtained by soft-updating from the main network. The target network can be understood as helping the main network calculate the target value, and then realizing the update of the actor and critic networks in the main network. Therefore, the present invention includes an Actor network and a Critic network, which are the main networks, and also includes an Actor target network and a Critic target network. The Actor target network has exactly the same structure as the Actor network, and the Critic target network has exactly the same structure as the Critic network. Initialize the parameters θ a of the Actor network and the parameters θ c of the Critic network, initialize the parameters of the Actor target network and the parameters
[0021] S32. Observe the next state and reward for each UAV's action, form samples and put them into the experience replay pool. When the number of samples reaches the preset value, execute S33 to update the Critic network and execute S34 to update the Actor network;
[0022] S33. Update the Critic network;
[0023] S34. Update the Actor network;
[0024] S35. Update the parameters of the Actor target network and the Critic target network, and return to execute S31 to S35 until the objective function of the Critic network reaches the minimum value and the objective function of the Actor network reaches the maximum value, or when the Actor network and the Critic network reach the preset number of training rounds, stop training to obtain the optimized Critic network and the optimized Actor network;
[0025] S36. Use the optimized Actor network to generate the actions of the corresponding drones, and use the optimized Critic network to generate the rewards for the actions of the corresponding drones. Each drone executes the actions of its corresponding optimized Actor network to complete the multi-drone pursuit of a single drone.
[0026] Furthermore, S33 includes:
[0027] S331. Randomly sample a batch of data (x j , a j , r j , x′ j ) from the experience replay pool and calculate the target value: Among them, x j represents the initial state of the j-th drone, a j represents the action of the j-th drone, r j represents the reward value of the j-th drone, x′ j represents the next state of the j-th drone; Q is the Q function, represents the Q value obtained by the Critic target network of the j-th drone for the next state using the action a′ j ;
[0028] S332. Add a perturbation signal S(t) to the parameters θ c of the Critic network: Among them, is the estimated value of θ c ;
[0029] S333. Construct the objective function of the Critic network: represents the Q value obtained by the Critic network of the j-th drone for the initial state using the action a j ; l represents the number of samples;
[0030] S334. Use feedback control to update Among them, K represents the gain; M(t) is the noise signal, and the updated is used as the current Return to S332, and execute S35 after looping a preset number of times.
[0031] Furthermore, S34 includes:
[0032] S341. Calculate the entropy value of the current action policy: where is the policy distribution based on the parameter θ a and represents the expected value of the action a under the distribution of the policy i (which can also be interpreted as the expected value of the function value corresponding to the action a randomly selected under the current policy distribution i by calculating the corresponding expected value).
[0033] S342. Dynamically adjust the entropy regularization coefficient to obtain the regularization coefficient at the current moment where α start is the initial regularization coefficient, α end is the final regularization coefficient, train_step is the current training time step, and t decay is the decay frequency;
[0034] S343. Add a perturbation signal S(t) to the parameter θ a of the Actor network: is the estimated value of θ a ;
[0035] S344. Construct the objective function of the Actor network: is to output the corresponding action a according to the state s of the drone, is the policy function, a mapping implemented by a neural network. The input of the neural network is the state s of the drone, and the output is the action a of the drone;
[0036] S345. Use feedback control to update Take the updated as the current Return to S343, and execute S35 after looping a preset number of times.
[0037] Furthermore, S35 includes:
[0038] Perform a soft update on the Actor target network and the Critic target network. The soft update process is as follows
[0039]
[0040] Among them, are the parameters after the update of the Actor target network, are the parameters after the update of the Critic target network, τ is the soft update coefficient and τ ∈ (0, 1], are the parameters of the Critic target network before the update, are the parameters of the Actor target network before the update.
[0041] The present invention also provides a multi-UAV encirclement and capture system, including:
[0042] A single-UAV modeling module, used for performing dynamic modeling on the UAV;
[0043] A seven-tuple construction module, used for constructing a seven-tuple of multiple UAVs in a three-dimensional dynamic environment to describe multiple UAV encirclement and capture scenarios;
[0044] A multi-UAV encirclement and capture module, where each UAV includes an Actor network and a Critic network. The Actor network generates the actions of the UAV, and the Critic network generates evaluations corresponding to the actions. By combining extreme value search and adaptive entropy and continuously iteratively updating the Actor network and the Critic network, the optimized Actor network is used to generate the optimal actions, and each UAV executes the optimal actions to complete the encirclement and capture of a single UAV by multiple UAVs.
[0045] Furthermore, the single-UAV modeling module is also used for:
[0046] There are N UAVs, where N > 1. The discrete dynamic equation of multiple UAVs from time t to time t + 1 is:
[0047]
[0048] Among them, i is the index of the UAV, UAV is the abbreviation of the unmanned aerial vehicle, x, y, z are the x-axis, y-axis and z-axis coordinates of the UAV, v is the UAV speed, ψ is the UAV heading angle, θ is the UAV pitch angle, φ is the UAV roll angle, ω ψ,i is the heading angular velocity control input, ω θ,i is the pitch angular velocity control input, ω φ,i is the roll angular velocity control input, and Δt is the time step from time t to time t + 1.
[0049] Even further, the constraint conditions for the flight states of multiple UAVs are:
[0050]
[0051] Among them, x i,min ,x i,max ,y i,min, y i,max , z i,min , z i,max is the minimum and maximum values of the \(i\)-th UAV on the \(x\)-axis, \(y\)-axis, and \(z\)-axis, \(v\) i,min , v i,max is the minimum and maximum values of the speed of the \(i\)-th UAV, \(\psi\) i,min , \(\psi\) i,max is the minimum and maximum values of the heading angle of the \(i\)-th UAV, \(\theta\) i,min , \(\theta\) i,max is the minimum and maximum values of the pitch angle of the \(i\)-th UAV, \(\varphi\) i,min , \(\varphi\) i,max is the minimum and maximum values of the roll angle of the \(i\)-th UAV, \(u\) is the acceleration control input of the UAV, \(u\) i,min , u i,max is the minimum and maximum values of the acceleration control input of the \(i\)-th UAV, \(\omega\) ψ,i,min , \(\omega\) ψ,i,max is the minimum and maximum values of the heading angular velocity input of the \(i\)-th UAV, \(\omega\) θ,i,Min , \(\omega\) θ,i,max is the minimum and maximum values of the pitch angular velocity input of the \(i\)-th UAV, \(\omega\) φ,i,min , \(\omega\) φ,i,max is the minimum and maximum values of the roll angular velocity input of the \(i\)-th UAV; The obstacle is set as a circular area, \((x\) O , y O , z O ) are the \(x\)-axis, \(y\)-axis, and \(z\)-axis coordinates of the center of the sphere of the obstacle, \(R\) O is the radius of the obstacle.
[0052] Furthermore, the seven-tuple construction module is also used for:
[0053] Construct a seven-tuple \(\langle S, A, R, P, Z, O, \gamma\rangle\) of multiple UAVs in a three-dimensional dynamic environment, where \(S\) is the state set, \(S = \{s\) 1 , s 2 ,..., s i ,..., s N \}, s i is the state of the \(i\)-th UAV and \(s\) i = \{x\) i , y i , z i , v i , \(\psi\) i , \(\theta\) i , \(\varphi\) i \}; \(A\) is the action set and \(A = \{a\) 1 , a 2 ,..., a i ,..., a N \}, ai is the action of the i-th UAV, a i = [u i , ω ψ,i , ω θ,i , ω φ,i T , where T is the transpose operation; R represents the reward obtained by the UAV taking the corresponding action in a certain state; P is the probability that the UAV transfers from one state to another by taking an action; Z is the observation set, Z = {z 1 , z 2 ,..., z i ,..., z N}, where z i is the observation information of the i-th UAV, z i = {x i , y i , z i , v i , ψ i , θ i , φ i}; O is the observation function; γ is the discount factor, representing the preference and proportion of the current reward and future rewards.
[0054] Furthermore, the multi-UAV encirclement and capture module includes:
[0055] Initialization unit. In reinforcement learning, it is divided into a main network and a target network. The main network is responsible for interacting with the environment. The target network is a copy of the main network, with the same structure as the main network, but a lower parameter update frequency. The parameters of the target network are obtained by soft-updating from the main network. The target network can be understood as helping the main network calculate the target value, and then realizing the update of the actor and critic networks in the main network. Therefore, the present invention includes an Actor network and a Critic network, which are the main networks, and also includes an Actor target network and a Critic target network. The Actor target network has exactly the same structure as the Actor network, and the Critic target network has exactly the same structure as the Critic network. Initialize the parameters θ a of the Actor network and the parameters θ c of the Critic network, initialize the parameters of the Actor target network and the parameters
[0056] Sample generation unit, which is used to execute actions on each UAV to observe the next state and reward, form samples and put them into the experience replay pool. When the number of samples reaches the preset value, execute the first update unit to update the Critic network and execute the second update unit to update the Actor network;
[0057] The first update unit is used to update the Critic network;
[0058] The second update unit is used to update the Actor network;
[0059] The third update unit is used to update the parameters of the Actor target network and the Critic target network, and return to execute the initialization unit to the third update unit until the objective function of the Critic network reaches the minimum value and the objective function of the Actor network reaches the maximum value, or when the Actor network and the Critic network reach the preset number of training rounds, stop training to obtain the optimized Critic network and the optimized Actor network;
[0060] The action execution unit is used to generate the actions of the corresponding drones by using the optimized Actor network, generate the rewards for the actions of the corresponding drones by using the optimized Critic network, and each drone executes the actions of its corresponding optimized Actor network to complete the multi-drone encirclement of a single drone.
[0061] Furthermore, the first update unit is further used for:
[0062] S331. Randomly sample a batch of data (x h , a j , r j , x′ j ) from the experience replay pool, and calculate the target value: Where x j represents the initial state of the j-th drone, a j represents the action of the j-th drone, r j represents the reward value of the j-th drone, x′ j represents the next state of the j-th drone; Q is the Q function, represents the Q value obtained by the next state x′ j of the Critic target network of the j-th drone by adopting the action a′ j ;
[0063] S332. Add a perturbation signal S(t) to the parameters θ c of the Critic network: Where is the estimated value of θ c ;
[0064] S333. Construct the objective function of the Critic network: represents the Q value obtained by the initial state of the Critic network of the j-th drone by adopting the action a j ; l represents the number of samples;
[0065] S334. Update using feedback control where K represents the gain; M(t) is the noise signal, and the updated is used as the current Return to S332, and execute S35 after looping a preset number of times.
[0066] Furthermore, the second update unit is also used for:
[0067] S341. Calculate the entropy value of the current action policy: where is the policy distribution based on the parameter θ a and represents the expected value of the action a under the distribution of the policy (which can also be interpreted as the expected value of the function value corresponding to the action a randomly selected based on the current policy distribution i ). Under the current policy distribution i , randomly select an action a
[0068] S342. Dynamically adjust the entropy regularization coefficient to obtain the regularization coefficient at the current moment where α start is the initial regularization coefficient, α end is the final regularization coefficient, train_step is the current training time step, and t decay is the attenuation frequency;
[0069] S343. Add a perturbation signal S(t) to the parameter θ a of the Actor network: is the estimated value of θ a ;
[0070] S344. Construct the objective function of the Actor network: is to output the corresponding action a according to the state s of the UAV, is the policy function, a mapping implemented by a neural network. The input of the neural network is the state s of the UAV, and the output is the action a of the UAV;
[0071] S345. Update using feedback control The updated is used as the current Return to S343, and execute S35 after looping a preset number of times.
[0072] Furthermore, the third update unit is also used for:
[0073] performing a soft update on the Actor target network and the Critic target network, and the soft update process is as follows
[0074]
[0075] wherein, are the parameters after the update of the Actor target network, are the parameters after the update of the Critic target network, τ is the soft update coefficient and τ ∈ (0, 1], are the parameters of the Critic target network before the update, are the parameters of the Actor target network before the update.
[0076] The advantages of the present invention are as follows:
[0077] (1) The present invention combines extreme value search and adaptive entropy to update the network, realizes the optimized training of the network, does not rely on gradient information, can approximate the global optimal solution more efficiently, avoids the algorithm falling into local optimum, and improves the UAV encirclement efficiency.
[0078] (2) Traditional MADDPG optimizes the Actor and Critic networks through gradient descent, but gradient descent highly depends on the local gradient information of the loss function, is prone to falling into local optimum, and may have slow convergence speed due to gradient disappearance or oscillation problems in high-dimensional dynamic environments. After using the extreme value search algorithm in the present invention, it can approximate the global optimal solution more efficiently through global search and dynamic perturbation without relying on gradient information. At the same time, an adaptive entropy mechanism is added to the Actor network, which can dynamically adjust the balance between exploration and exploitation, and enhance the strategy learning ability of multiple UAVs. When the intelligent agents (i.e., UAVs) conduct initial exploration, by setting the entropy value to a relatively large value, it encourages multiple UAVs to execute more exploratory actions and explore new strategy spaces. After the intelligent agent strategy selection converges in the later stage, the entropy value is gradually decreased to let the intelligent agent focus on using the learned strategies and improve the overall task efficiency.
[0079] (3) In the present invention, in order to expand the applicable environment of multi-UAV encirclement, POMDP modeling is adopted, and the multi-UAV encirclement task is formalized into a seven-tuple structure (state set, action set, reward set, observation set, state transition probability, observation function, discount factor) to adapt to the partially observable dynamic environment.
[0080] (4) In the present invention, the extreme value search algorithm is adopted to replace the gradient descent algorithm to update the Actor and Critic functions in the MADDPG algorithm, which can reduce the computational complexity and improve the success rate of the multi-UAV pursuit mission. By combining the extreme value search and the adaptive entropy, the UAVs can fully explore and find the global optimal solution, thereby enhancing the robustness and adaptability of the strategy and adapting to complex stochastic environments. Description of the Drawings
[0081] Figure 1 It is a flowchart of a multi-UAV pursuit method disclosed in an embodiment of the present invention;
[0082] Figure 2 It is a principle block diagram of the extreme value search algorithm in a multi-UAV pursuit method disclosed in an embodiment of the present invention;
[0083] Figure 3 It is a schematic diagram of the extreme value search algorithm process in a multi-UAV pursuit method disclosed in an embodiment of the present invention;
[0084] Figure 4 It is a flowchart of the Monte Carlo test process in a multi-UAV pursuit method disclosed in an embodiment of the present invention;
[0085] Figure 5 It is a schematic diagram of the structural advantages of a multi-UAV pursuit method disclosed in an embodiment of the present invention;
[0086] Figure 6 It is a schematic diagram of the functional advantages of a multi-UAV pursuit method disclosed in an embodiment of the present invention. Detailed Embodiments
[0087] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0088] Embodiment 1
[0089] Embodiment 1 of the present invention provides a multi-UAV pursuit method, including the following steps:
[0090] S1. Perform dynamic modeling on the UAVs; the specific process is as follows:
[0091] Now consider N UAVs, where N>1. The discrete dynamic equation of the multi-UAV from time t to time t+1 is:
[0092]
[0093] Among them, i is the index of the UAV, where UAV is the abbreviation of unmanned aerial vehicle, (x, y, z) is the position of the UAV, v is the UAV speed, ψ is the UAV heading angle, θ is the UAV pitch angle, φ is the UAV roll angle, u is the UAV acceleration control input, ω ψ,i is the heading angular velocity control input, ω θ,i is the pitch angular velocity control input, ω φ,i is the roll angular velocity control input, and Δt is the time step from time t to time t + 1. Constraints on the flight state: x i,min , x i,max , y i,min , y i,max , z i,min , z i,max are the minimum and maximum values at the position of the i-th UAV, v i,min , v i,max are the minimum and maximum values of the speed of the i-th UAV, ψ i,min , ψ i,max are the minimum and maximum values of the heading angle of the i-th UAV, θ i,min , θ i,max are the minimum and maximum values of the pitch angle of the i-th UAV, φ i,min , φ i,max are the minimum and maximum values of the roll angle of the i-th UAV, u i,min , u i,max are the minimum and maximum values of the acceleration control input of the i-th UAV, ω ψ,i,min , ω ψ,i,max are the minimum and maximum values of the heading angular velocity input of the i-th UAV, ω θ,i,min , ω θ,i,max are the minimum and maximum values of the pitch angular velocity input of the i-th UAV, ω φ,i,min , ω φ,i,max are the minimum and maximum values of the roll angular velocity input of the i-th UAV. Flight path constraints: For the convenience of processing, the obstacle is set as a standard circular area, (x O , y O , z O ) is the center position of the obstacle sphere, and R O is the radius of the obstacle.
[0094] S2. Construct a decision-making model for multiple UAVs to surround a single UAV, that is, construct a seven-tuple of multiple UAVs in a three-dimensional dynamic environment to describe multiple UAV surrounding scenarios; the specific process is as follows:
[0095] Multiple unmanned aerial vehicles (UAVs) perform a cooperative encirclement mission on a single UAV in a bounded three-dimensional space. The mission objective is for the multiple UAVs to start from random positions, avoid dynamic and static obstacles in the space (such as hills, interference areas, etc.), and encircle the target UAV within a specified time. The condition for a successful encirclement is that all encircling UAVs must enter a spherical area centered at the center position (x T , y T , z T ) of the target UAV with a radius of R T within the specified time, and during the mission execution, none of the encircling UAVs collide with any obstacles. The corresponding condition for a failed encirclement is that during the specified time, any encircling UAV collides with a static or dynamic obstacle, resulting in the damage of that UAV, or any UAV flies out of the boundary of the three-dimensional space (i.e., [x min , x max × [y min , y max × [z min , z max ) or the encircling UAVs fail to enter the encirclement area centered on the target escaping UAV.
[0096] The obstacles in the three-dimensional space include dynamic obstacles, such as flight interference sources or moving friendly UAVs, and at the same time, there are also the influences of static obstacles. The shape of the obstacles may be spherical or cylindrical structures, and their heights also need to be taken into account. During the flight of the encircling UAVs, the energy consumption requirements of the UAVs and the speed, path length, and vertical movement in the three-dimensional space that affect energy consumption should also be considered. At the same time, the control of the attitude and angle of the encircling UAVs should also be considered. Because the stability of the encircling UAVs should be ensured when exploring the position of the target UAV, the angle range and angular velocity limits of the pitch angle and roll angle of the encircling UAVs are particularly important.
[0097] The dynamic uncertainty described in the present invention means that when the environment is reset at the end of each round, the initial states of all UAVs are randomly generated. That is, in the three-dimensional simulation space x max × y max × z max , the initial state of each UAV is randomly generated by the position (x, y, z), speed v, heading angle ψ, pitch angle θ, and roll angle φ, and is updated according to the discrete dynamics model in 1. The initial position of the target UAV and the positions of the obstacles are randomly generated in the space. For the convenience of calculation and implementation, the obstacles are regarded as spherical areas, represented by their center coordinates (x O , y O , z O ) and radius R O , and the size of the obstacles remains fixed in each round.
[0098] The POMDP is a partially observable Markov decision process, which is an extension of the Markov decision process and is used to describe how an agent makes optimal decisions when the environmental state cannot be fully observed by the agent. The POMDP is used to describe the decision-making model for multiple UAVs to surround a single UAV, and this model can be defined as a seven-tuple <S, A, R, P, Z, O, γ>, where S is the set of states, S = {s 1 , s 2 ,..., s i ,..., s N}, N is the number of agents, s i = {x i , y i , z i , v i , ψ i , θ i , φ i}, where x i , y i , z i are the positions of the UAV in three-dimensional space, v i is the speed of the UAV, ψ i , θ i , φ i are the heading angle, pitch angle and roll angle of the UAV respectively, and the state set S is unknown to the agent.
[0099] A is the set of actions, A = {a i , a 2 ,..., a i ,..., a N}, a i is the action of the i-th UAV, i = {1, 2,..., N}, a i = [u i , ω ψ,i , ω θ,i , ω φ,i , T , where u i is the acceleration control input of the UAV, ω ψ,i , ω θ,i , ω φ,i are the angular velocity control inputs of the heading angle, pitch angle and roll angle, and the actions are known to the agent.
[0100] R(s, a) represents the reward obtained by the UAV when taking action a in state s. To simplify the design of the reward function, we combine multiple piecewise rewards into a comprehensive function and use weights to balance different objectives. The expression is:
[0101] R(s, a) = ω 1·r obs + ω 2 ·r goal
[0102] where the obstacle reward (r obs ):
[0103]
[0104] is the distance between the UAV and the obstacle.
[0105] The reward (r goak ) for successfully capturing the target UAV is:[[]]
[0106]
[0107] is the distance between the UAV and the target UAV. The weights ω 1 , ω 2 are used to balance the importance of obstacle avoidance and capturing the target UAV.
[0108] P is the state transition probability. P(s′|s,a) represents the probability that the UAV transfers from state s to state s′ by taking action a. The state transition probability is based on the 3D UAV dynamics model in 1 and is unknown to the agent.
[0109] Z is the observation set, Z = {z 1 , z 2 ,..., z i ,..., z N}, where z i is the observation information of the i-th UAV, z i = {x i , y i , z i , v i , ψ i , θ i , φ i}, which is known to the agent; O is the observation function. O(z|s,a) represents the probability that after executing an action a ∈ A, the state s ∈ S generates an observation z ∈ Z, which can be simplified to a Gaussian distribution based on the sensor accuracy, and is expressed by the formula:[[]] where σ is the standard deviation of the sensor noise, and the observation function is unknown to the agent; the discount factor γ ∈ [0,1] represents the preference and proportion of the current reward and future rewards.
[0110] S3. Solving the multi - UAV surrounding single - UAV strategy of the MADDPG reinforcement learning method optimized based on extremum search. That is, each UAV includes an Actor network and a Critic network. The Actor network generates the actions of the UAV, and the Critic network generates the evaluation of the corresponding actions. By combining extremum search and adaptive entropy and continuously iteratively updating the Actor network and the Critic network, the optimized Actor network is used to generate the optimal actions, and each UAV executes the optimal actions to complete the multi - UAV surrounding single - UAV. The flow chart of the whole method is as Figure 1 shown; the specific process is as follows:
[0111] First, introduce the extremum search algorithm. In Figure 2 the block diagram of the extremum search algorithm, the objective function is J(θ), which represents the performance index of the system with respect to the parameter θ. The optimization goal is to make the value of this function reach the extremum. θ is the decision variable, and by dynamically adjusting the value of θ, the system output reaches the target extremum. is the estimated variable, representing the current estimated value of the parameter, which is an intermediate variable that is continuously updated. S(t) is a noise signal, usually in the form of a periodic signal, such as S(t) = [a 1 sin(ω 1 t),..., a n sin(ω n t)], which is often used to drive the adjustment of parameters to explore the extremum point. M(t) is a noise signal used for the search result, and in form it is a signal used to extract the adjustment information near the extremum point and used to guide the update of parameters. K / S is an integral link, and by adjusting the value of the gain K, the parameter estimated value can be dynamically adjusted according to the search result
[0112] The core principle of the extremum search algorithm is: using a periodic signal S(t) to drive the change of the parameter θ, and continuously adjusting the estimated value of the parameter through the feedback loop Finally, making the system output y = J(θ) approach the target extremum. The main step flow chart is as Figure 3 shown. The flow chart of Figure 2 can be simply summarized as: performing iterative search of parameters through the periodic perturbation signal S(t), and finally finding the optimal parameter θ that makes the objective function reach the extremum. In this design, this method is used to optimize the parameters in the Actor and Critic networks, θ c and θ a .
[0113] The following introduces the extremum search optimization process of the Critic network parameters:
[0114] The Critic network is responsible for fitting the Q - value function Its goal is to minimize the TD error:
[0115] where r j is the immediate reward and γ is the discount factor.
[0116] Goal:
[0117] Considering that the essence of using the extremum search algorithm is to find the maximum value, the objective function here is selected as the reverse TD error, which is
[0118]
[0119] Using the extremum search algorithm to achieve the minimization of the TD error is shown in Table 1 below:
[0120] Table 1 Logical process of the extremum search algorithm to achieve the minimization of the TD error
[0121]
[0122] The extremum search optimization process of the Actor network parameters is as follows:
[0123] The Actor network is responsible for generating the action policy Its goal is to maximize the long-term return of the policy:
[0124] Goal:
[0125] Using the extremum search algorithm to achieve the maximization of the long-term return of the policy is shown in Table 2 below:
[0126] Table 2 Logical process of the extremum search algorithm to achieve the maximization of the long-term return of the policy
[0127]
[0128] The present invention combines the above extremum search optimization method and the reinforcement learning algorithm (MADDPG algorithm) for the policy selection of multi-capture drones. The following introduces the Actor and Critic networks in the MADDPG algorithm based on extremum search optimization.
[0129] An improved MADDPG algorithm is adopted for the policy selection of multi-capture drones. The MADDPG algorithm is a multi-agent algorithm based on deep reinforcement learning, mainly used to solve the cooperation and confrontation tasks in a multi-agent environment, and is an extension of the deep deterministic policy gradient algorithm combined with centralized training and decentralized execution.
[0130] Within each training cycle, the extreme value search and optimization processes of the Critic and Actor are alternated to ensure the collaborative optimization of the policy network (Actor) and the value function network (Critic). Reinforcement learning consists of a main network and a target network. The main network is responsible for interacting with the environment. The target network is a copy of the main network, with the same structure as the main network, but a lower parameter update frequency. The parameters of the target network are obtained through soft updates from the main network. The target network can be understood as helping the main network calculate the target value and then realizing the update of the actor and critic networks in the main network. Therefore, the present invention includes an Actor network and a Critic network, which are the main networks, and also includes an Actor target network and a Critic target network. The Actor target network has exactly the same structure as the Actor network, and the Critic target network has exactly the same structure as the Critic network. The complete process of the improved MADDPG algorithm based on extreme value search of the present invention is as shown in Table 3 below:
[0131] Table 3 Logical process of the MADDPG algorithm based on extreme value search
[0132]
[0133]
[0134]
[0135] The above Q-value function is prior art. Q(s,a) represents the expected cumulative reward that the agent (i.e., the drone) can obtain when performing action a in state s and following the optimal policy. The mathematical definition is
[0136]
[0137] where γ t is the discount factor at time t (0 ≤ γ ≤ 1), r t+1 is the reward obtained at time t + 1, π is the policy. Ε is the expectation. s 0 is the initial state, a 0 is the initial action.
[0138] In order to comprehensively evaluate the effectiveness of the multi-drone three-dimensional encirclement of a single drone strategy of the MADDPG reinforcement learning method optimized based on extreme value search, a series of simulation experiments were designed. The goal of the experiment is to verify the robustness and performance of the method in a complex dynamic environment, especially the task completion rate, flight efficiency, and stability under randomly initialized multi-scenarios. The following is the simulation training and performance index analysis of multi-drone encirclement of a single drone.
[0139] 1) Training environment setup
[0140] Create a three-dimensional virtual flight scenario that includes multi-capture drones, obstacles, and a single target drone. The scenario range is set to [x min , x max × [y miN , y maX × [z miN , z max . The static randomly distributed spherical obstacles in the three-dimensional environment are of the same size, but their positions are randomly initialized at the beginning of each round.
[0141] Set the initial states of multiple drones in the three-dimensional environment, including position (x, y, z), velocity v, heading angle ψ, pitch angle θ, roll angle φ, acceleration u, etc. Initialize the parameters of the Actor and Critic networks (θ a and θ C ) in the MADDPG algorithm, as well as the perturbation amplitude and frequency S(t) and M(t) in the extremum search algorithm. Define the time step Δt and the maximum flight time t max .
[0142] Initialize the parameters of the Actor and Critic networks in the MADDPG algorithm optimized based on extremum search, which are θ c and θ a respectively. Initialize the target network parameters θ’ c and θ‘ a , and initialize the experience replay pool D.
[0143] 2) Training process
[0144] Execute for the multi-capture drones: Initialize the parameters, generate the actions of the multi-capture drones, interact with the environment, obtain the feedback of the reward value, update the drone states, and achieve the update of the network parameters and the objective function until the execution of this round is completed. The end conditions for each round of training are: successfully achieving the capture of the target drone, reaching the maximum flight time T_max, colliding with obstacles or other drones. When these three situations occur, this round will end and enter the next round until convergence or the set number of rounds reaches the maximum value.
[0145] 3) Simulation result output and evaluation of training performance
[0146] During the training phase, the positions of the obstacles, as well as the flight trajectories of the pursuing drones and the target drone, will be recorded to visualize the paths of both drones, showing the flight trajectories of the drones from the starting point to the target point and whether the target drone has been successfully captured. In addition, performance metrics such as the cumulative reward curve, the convergence of the policy, and the average number of collisions can be used to compare the curve of the reward value against the number of training episodes and analyze the effectiveness of policy improvement.
[0147] As Figure 4 shown, to comprehensively evaluate the proposed method for three-dimensional multi-drone pursuit of a single drone based on MADDPG and extremum search algorithms in a complex real environment, the present invention uses Monte Carlo testing for verification. Through multiple rounds of simulation experiments, the effectiveness and robustness of the method in different uncertain environments are verified under various randomly generated three-dimensional scenarios.
[0148] 1) Setup of the test environment
[0149] Create a three-dimensional virtual flight scene containing multiple pursuing drones, obstacles, and a single target drone. The scene range is set to [x min , x max × [y min , y max × [z min , z max . The static randomly distributed spherical obstacles in the three-dimensional environment are of the same size, but their positions are randomly initialized at the start of each episode.
[0150] Set the initial states of multiple drones in the three-dimensional environment, including position (x, y, z), velocity v, heading angle ψ, pitch angle θ, roll angle φ, acceleration u, etc. Initialize the parameters (θ a and θ c ) of the Actor and Critic networks in the MADDPG algorithm, as well as the perturbation amplitude and frequency S(t) and M(t) in the extremum search algorithm. Define the time step Δt and the maximum flight time T max of the flight mission.
[0151] 2) Monte Carlo test procedure
[0152] Set the total number of rounds of Monte Carlo simulation to N, and each round of simulation represents a complete multi-UAV flight mission test. At the beginning of each round of experiment, it is necessary to initialize the position, speed, heading angle, pitch angle, roll angle, etc. of the multi-UAV for encirclement. At the same time, it is necessary to dynamically process some obstacles and add environmental interference factors that do not exist in the simulation training, such as random wind fields, signal interference, etc. By introducing uncertainty through randomization, the complexity of the real field is simulated, so that the effectiveness and practicality of the trained multi-UAV encirclement strategy can be tested from the perspective of the real environment.
[0153] After setting the environment and initializing the state of the UAVs, each UAV for encirclement generates an action [u i , ω ψ,i , ω θ,i , ω φ,i through the Actor policy network in MADDPG, and gives the action for the next time step based on the current environmental state (including the position of obstacles, the distance to the target UAV, the positions of other UAVs, etc.). The role of the extremum search algorithm is to optimize the actions generated by MADDPG, introduce perturbations (such as S(t)) into the loss functions of the Actor and Critic networks, fine-tune the policy parameters, and find a better path selection.
[0154] The UAVs fly in the simulated environment according to the adjusted actions and update their own states based on the three-dimensional dynamics model. During the flight, they interact with obstacles, other UAVs, etc., and the environment feeds back new state information, including the current position (x, y, z) of the UAV; whether a collision occurs; the distance d G to the target UAV; the feedback of the overall reward value, and the reward value is consistent with the previously set reward function.
[0155] After obtaining the environmental feedback data, update the parameters of the Actor and Critic networks, and use the extremum search optimization to generate the action policy and Q-value function. According to the change of the objective function (such as the reward function or the Q-value function), adjust the perturbation amplitude S(t) and feedback intensity M(t) in the extremum search algorithm.
[0156] In each round of experiment, the following key performance indicators will be recorded: the mission completion rate (MCR):
[0157]
[0158] where t success represents the number of rounds in which the mission is successfully completed, and N is the total number of Monte Carlo tests.
[0159] Average flight time: Statistically calculate the average flight time of the UAVs from the starting point to the encirclement of the target UAV.
[0160] Collision times: Count the number of collisions between the UAV and obstacles or other UAVs in each round of experiments.
[0161] Energy consumption:
[0162]
[0163] where \(P(v i ,u i ) is the energy consumption function of the UAV flight.
[0164] After the N rounds of tests are completed, the collected data will be used to calculate statistics such as the mean, standard deviation, minimum and maximum values of various performance metrics. The average flight time can reflect the efficiency of the overall mission implementation of the cooperative UAVs; the standard deviation can reveal the stability of the algorithm in a random scenario; the minimum and maximum values can show the best performance and extreme disadvantages of the multi-cooperative UAVs; and by comparing the average reward values of multi-cooperative UAVs using different algorithms in the same environment, it can be evaluated whether the improved method significantly improves the performance and whether the flight time, collision probability, etc. have been correspondingly improved.
[0165] The technical principle of the present invention is as follows: The dynamic modeling of the cooperative UAVs in the present invention is the basis of the entire algorithm and simulation environment, providing physical constraints for the subsequent definition of states, actions, rewards, etc. The subsequent state modeling and policy optimization all rely on accurate dynamic descriptions. At the same time, the dynamic modeling is directly connected to the POMDP modeling, providing a physical basis for state transitions in the POMDP modeling. The POMDP modeling establishes a connection between the dynamic modeling and policy optimization parts of the multi-UAVs, abstracting the complexity problem of multi-UAV cooperation into an optimization problem, facilitating the subsequent processing of applying extreme value search optimization to the MADDPG algorithm.
[0166] When optimizing the Actor network and the Critic network, the traditional MADDPG algorithm updates the parameters through the gradient descent method. However, since the gradient descent highly depends on the local gradient information of the loss function, problems such as gradient vanishing, oscillation, or stagnation may occur in high-dimensional environments, making it difficult for the network parameters to find the global optimal solution. In the present invention, the extremum search algorithm is used to replace the gradient descent for optimizing the mean squared error (MSE) loss, making the approximation of the Q-value function more accurate and optimizing the performance of the MADDPG algorithm. Its core idea is to utilize the characteristics of the extremum search, which does not rely on gradient information and has the ability of global search, to make up for the defect that the gradient descent algorithm is prone to fall into local optima, so as to achieve better policy optimization in complex non-linear environments. In addition, the extremum search algorithm combines the mechanism of adaptive entropy. By dynamically adjusting the balance between exploration and exploitation, it enhances the exploration ability of multiple unmanned aerial vehicles in policy learning and avoids falling into a single policy or inefficient area. The combination of the global optimization ability of the extremum search algorithm and the dynamic adjustment mechanism of adaptive entropy greatly improves the success rate and robustness of the multi-UAV cooperative pursuit mission.
[0167] Through the simulation training part, the improved MADDPG algorithm based on extremum search optimization is used for multiple unmanned aerial vehicles to interact with the three-dimensional obstacle environment in simulation to obtain the results of whether the policy converges, the average reward curve, and the number of collisions to measure the effectiveness of the proposed method. On this basis, the successfully converged multi-UAV pursuit strategy is subjected to Monte Carlo tests, and a large number of randomized simulations are carried out to simulate the uncertainties in the real scenario to further verify and detect the robustness and universality of the strategy.
[0168] Through the collaborative cooperation among the above parts, a universal strategy for multiple unmanned aerial vehicles to pursue a single target unmanned aerial vehicle in a three-dimensional environment can be obtained, thus providing a sample for multiple unmanned aerial vehicles to pursue a single target unmanned aerial vehicle in the real environment.
[0169] As Figure 5 shown, in terms of structure, the present invention adds an adaptive entropy mechanism to the Actor network, which can dynamically adjust the balance between exploration and exploitation and enhance the policy learning ability of multiple unmanned aerial vehicles. When the agent conducts initial exploration, by setting the entropy value to a relatively large value, it encourages multiple unmanned aerial vehicles to execute more exploratory actions and explore new policy spaces. After the agent's policy selection converges in the later stage, the entropy value is gradually reduced, allowing the agent to focus on utilizing the learned policies and improving the overall task efficiency. In the present invention, in order to expand the applicable environment for the pursuit of multiple unmanned aerial vehicles, POMDP modeling is adopted, and the pursuit task of multiple unmanned aerial vehicles is formalized into a seven-tuple structure (state set, action set, reward set, observation set, state transition probability, observation function, discount factor) to adapt to the partially observable dynamic environment.
[0170] As Figure 6As shown, in terms of functionality, replacing the gradient descent algorithm with the extremum search algorithm to update the Actor and Critic functions in the MADDPG algorithm can reduce the computational complexity and improve the success rate of the multi-UAV pursuit mission; combining the extremum search and adaptive entropy enables the UAVs to fully explore and find the global optimal solution, thereby enhancing the robustness and adaptability of the strategy and adapting to complex stochastic environments; through POMDP modeling and Monte Carlo testing, the robustness of this strategy under various random initial conditions can be verified. Whether it is the randomization of the target UAV's action selection or the dynamic change of the environment, the multi-UAVs can maintain a high success rate and low energy consumption.
[0171] The following illustrates the optimal usage state and its advantages of the method of the present invention from specific task scenarios:
[0172] Scenario 1 is the multi-UAV pursuit in a complex obstacle environment. In this scenario, the target UAV will escape the pursuit of the pursuing UAVs by flying quickly between complex buildings. The multi-UAV pursuit strategy designed in the present invention is modeled based on the partially observable Markov. The influence of complex obstacles on the state change of the UAVs can be captured by designing the observation function and state transition probability, enabling the pursuing UAVs to make decisions and successfully avoid obstacles in this environment and achieve the pursuit of the target UAV.
[0173] Scenario 2 is that the target UAV adopts a high-dynamic maneuvering escape strategy, such as random rapid charging, deceleration, changing the heading angle, pitch angle, roll angle, etc. In this situation, the traditional gradient descent algorithm is likely to fail. The extremum search algorithm used in the present invention can skip the local optimal points and directly search for the global optimal strategy, enabling the multi-UAVs to quickly adjust the strategy when dealing with the behavior of the target UAV with high-dynamic maneuvering ability, maintaining effective tracking and pursuit. At the same time, the adaptive entropy adopted in the present invention can increase its own exploration intensity when the target UAV quickly switches actions, encouraging the pursuing UAVs to try different action sequences to achieve the follow-up and pursuit of the target UAV as soon as possible.
[0174] Scenario 3 is the pursuit mission with time and energy consumption damaged. The mission requirement of the pursuit is to complete the pursuit of the target UAV within a limited time and minimize the energy consumption of the multi-UAVs as much as possible. In this case, the time and energy consumption of the pursuing UAVs also need to be used as optimization objectives, and these two objectives are directly optimized through the extremum search algorithm. At the same time, the adaptive entropy can be adjusted to reduce ineffective exploration in the later stage to reduce energy consumption and concentrate resources to achieve the pursuit of the target UAV.
[0175] The solution in the present invention is applicable to scenarios such as border patrol, anti-drone defense, and dynamic collaboration in complex three-dimensional environments, and can demonstrate the advantages of high robustness, adaptability, and efficiency, providing more exploration space for multi-drone pursuit.
[0176] Embodiment 2
[0177] Based on Embodiment 1, Embodiment 2 of the present invention further provides a multi-drone pursuit system, including:
[0178] A single-drone modeling module for performing dynamic modeling on the drone;
[0179] A seven-tuple construction module for constructing a seven-tuple of multiple drones in a three-dimensional dynamic environment to describe multiple drone pursuit scenarios;
[0180] A multi-drone pursuit module, where each drone includes an Actor network and a Critic network. The Actor network generates the actions of the drone, and the Critic network generates the evaluation of the corresponding actions. The extreme value search and adaptive entropy are combined to continuously iterate and update the Actor network and the Critic network. The optimized Actor network is used to generate the optimal actions, and each drone executes the optimal actions to complete the multi-drone pursuit of a single drone.
[0181] Specifically, the single-drone modeling module is further used for:
[0182] There are N drones, where N>1. The discrete dynamic equation of the multi-drones from time t to time t+1 is:
[0183]
[0184] where i is the index of the UAV (UAV is the abbreviation of unmanned aerial vehicle), x, y, z are the x-axis, y-axis, and z-axis coordinates of the UAV, v is the UAV speed, ψ is the UAV heading angle, θ is the UAV pitch angle, φ is the UAV roll angle, ω ψ,i is the heading angular velocity control input, ω θ,i is the pitch angular velocity control input, ω φ,i is the roll angular velocity control input, and Δt is the time step from time t to time t+1.
[0185] More specifically, the constraint conditions for the flight states of the multi-drones are:
[0186]
[0187] where x i,min , x i,max , y i,min , y i,max , z i,min , zi,max are the minimum and maximum values of the \(i\)-th UAV on the \(x\)-axis, \(y\)-axis, and \(z\)-axis, \(v\) i,min , \(v\) i,max are the minimum and maximum values of the speed of the \(i\)-th UAV, \(\psi\) i,min , \(\psi\) i,max are the minimum and maximum values of the heading angle of the \(i\)-th UAV, \(\theta\) i,min , \(\theta\) i,max are the minimum and maximum values of the pitch angle of the \(i\)-th UAV, \(\varphi\) i,min , \(\varphi\) i,max are the minimum and maximum values of the roll angle of the \(i\)-th UAV, \(u\) is the acceleration control input of the UAV, \(u\) i,min , \(u\) i,max are the minimum and maximum values of the acceleration control input of the \(i\)-th UAV, \(\omega\) ψ,i,min , \(\omega\) ψ,i,max are the minimum and maximum values of the heading angular velocity input of the \(i\)-th UAV, \(\omega\) θ,i,min , \(\omega\) θ,i,max are the minimum and maximum values of the pitch angular velocity input of the \(i\)-th UAV, \(\omega\) φ,i,min , \(\omega\) φ,i,max are the minimum and maximum values of the roll angular velocity input of the \(i\)-th UAV; the obstacle is set as a circular area, \((x\) O , \(y\) O , \(z\) O ) are the \(x\)-axis, \(y\)-axis, and \(z\)-axis coordinates of the center of the sphere of the obstacle, \(R\) O is the radius of the obstacle.
[0188] More specifically, the seven-tuple construction module is also used for:
[0189] Construct a seven-tuple \(\langle S, A, R, P, Z, O, \gamma\rangle\) of multiple UAVs in a three-dimensional dynamic environment, where \(S\) is the state set, \(S=\{s\) 1 , \(s\) 2 ,\(\cdots\), \(s\) i ,\(\cdots\), \(s\) N \}, \(s\) i is the state of the \(i\)-th UAV and \(s\) i =\{x\) i , \(y\) i , \(z\) i , \(v\) i , \(\psi\) i , \(\theta\) i , \(\varphi\) i \}; \(A\) is the action set and \(A = \{a\) 1 , \(a\) 2 ,\(\cdots\), \(a\) i ,\(\cdots\), \(a\) N \}, \(a\) i is the action of the \(i\)-th UAV, \(a\) i= [u i , ω ψ,i , ω θ,i , ω φ,i T , where T is the transpose operation; R represents the reward obtained by the UAV taking corresponding actions in a certain state; P is the probability that the UAV transfers from one state to another by taking an action; Z is the observation set, Z = {z 1 , z 2 ,..., z i ,..., z N}, where z i is the observation information of the i-th UAV, and z i = {x i , y i , z i , v i , ψ i , θ i , φ i}; O is the observation function; γ is the discount factor, representing the preference and proportion of the current reward and future rewards.
[0190] More specifically, the multi-UAV pursuit module includes:
[0191] An initialization unit for initializing the parameter θ a of the Actor network and initializing the parameter θ c of the Critic network; initializing the parameter of the Actor target network and initializing the parameter
[0192] A sample generation unit for observing the next state and reward for each UAV to execute an action, forming samples and putting them into the experience replay pool. When the number of samples reaches a preset value, the first update unit is executed to update the Critic network, and the second update unit is executed to update the Actor network;
[0193] A first update unit for updating the Critic network;
[0194] A second update unit for updating the Actor network;
[0195] A third update unit for updating the parameters of the Actor target network and the Critic target network, and returning to execute the initialization unit to the third update unit until the objective function of the Critic network reaches the minimum value and the objective function of the Actor network reaches the maximum value, or when the Actor network and the Critic network reach the preset number of training rounds, stop training to obtain the optimized Critic network and the optimized Actor network;
[0196] The action execution unit is used to generate actions corresponding to the drones by using the optimized Actor network, and generate rewards for the actions corresponding to the drones by using the optimized Critic network. Each drone executes the actions of its corresponding optimized Actor network to complete the multi-drone encirclement of a single drone.
[0197] More specifically, the first update unit is further used for:
[0198] S331. Randomly sample a batch of data (x j , a j , r j , x′ j ) from the experience replay pool and calculate the target value: y j = r j + γ· where x j represents the initial state of the j-th drone, a j represents the action of the j-th drone, r j represents the reward value of the j-th drone, x′ j represents the next state of the j-th drone; Q is the Q function, represents the Q value obtained by adopting the action a′ j for the next state of the Critic target network of the j-th drone;
[0199] S332. Add a perturbation signal S(t) to the parameters θ c of the Critic network: where is the estimated value of θ c ;
[0200] S333. Construct the objective function of the Critic network: represents the Q value obtained by adopting the action a j for the initial state of the Critic network of the j-th drone; l represents the number of samples;
[0201] S334. Use feedback control to update where K represents the gain; M(t) is the noise signal, and the updated is used as the current Return to S332, and execute S35 after looping a preset number of times.
[0202] More specifically, the second update unit is further used for:
[0203] S341. Calculate the entropy value of the current action policy: where is based on the parameter θ aThe policy distribution, denotes the expected value of the action a i under the distribution of the policy (which can also be interpreted as randomly selecting an action a i under the current policy distribution and calculating the corresponding expected value of the function value corresponding to this action).
[0204] S342. Dynamically adjust the entropy regularization coefficient to obtain the regularization coefficient at the current moment where α start is the initial regularization coefficient, α end is the final regularization coefficient, train_step is the current training time step, and t decay is the decay frequency;
[0205] S343. Add a perturbation signal S(t) to the parameters θ a of the Actor network: is the estimated value of θ a ;
[0206] S344. Construct the objective function of the Actor network: is to output the corresponding action a according to the state s of the drone, is the policy function, a mapping implemented by a neural network. The input of the neural network is the state s of the drone, and the output is the action a of the drone;
[0207] S345. Use feedback control to update Take the updated as the current Return to S343 and execute S35 after looping a preset number of times.
[0208] More specifically, the third update unit is also used for:
[0209] Soft-update the parameters of the Actor target network and the Critic target network. The soft-update process is as follows
[0210]
[0211]
[0212] where, are the updated parameters of the Actor target network, are the updated parameters of the Critic target network, τ is the soft-update coefficient and τ ∈ (0, 1], are the parameters before the update of the Critic target network, are the parameters before the update of the Actor target network.
[0213] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-UAV round-up method, characterized in that: include: S1. Dynamic modeling of the UAV; S2, constructing a seven-tuple of multiple drones in a three-dimensional dynamic environment to describe the multiple drone capture scenario; S3. Each drone includes an Actor network and a Critic network. The Actor network generates the drone's actions, and the Critic network generates the evaluation of the corresponding actions. The extreme value search and adaptive entropy are combined to continuously iterate and update the Actor network and the Critic network. The optimized Actor network is used to generate the optimal action. Each drone executes the optimal action to complete the multi-drone encirclement of a single drone.
2. A multi-UAV capture method according to claim 1, characterized in that: S1 includes: Suppose there are N drones, where N>1, and the discrete dynamic equation of multiple drones from time t to time t+1 is: Where i is the index of the UAV, UAV is the abbreviation of UAV, x, y, z are the x-axis, y-axis and z-axis coordinates of the UAV, v is the UAV speed, ψ is the UAV heading angle, θ is the UAV pitch angle, φ is the UAV roll angle, ω ψ,i is the heading angular velocity control input, ω θ,i is the pitch angular velocity control input, ω φ,i is the roll angular velocity control input, and Δt is the time step from time t to time t+1.
3. A multi-UAV capture method according to claim 2, characterized in that: The constraints of the flight status of multiple UAVs are: Among them, x i,min , x i,max ,y i,min ,y i,max , z i,min , z i,max are the minimum and maximum values of the i-th UAV on the x-axis, y-axis, and z-axis, respectively, v i,min , v i,max is the minimum and maximum value of the speed of the i-th UAV, ψ i,min , ψ i,max is the minimum and maximum heading angle of the i-th UAV, θ i,min ,θ i,max is the minimum and maximum pitch angle of the i-th UAV, φ i,min ,φ i,max is the minimum and maximum rolling angle of the i-th UAV, u is the acceleration control input of the UAV, and u i,min ,u i,max is the minimum and maximum value of the acceleration control input of the i-th UAV, ω ψ,i,min ,ω ψ,i,max is the minimum and maximum value of the angular velocity input of the i-th UAV, ω θ,i,min ,ω θ,i,max is the minimum and maximum value of the pitch angular velocity input of the i-th UAV, ω φ,i,min ,ω φ,i,max is the minimum and maximum value of the rolling angular velocity input of the i-th UAV; the obstacle is set as a circular area, (x O ,y O ,z O ) are the x-axis, y-axis and z-axis coordinates of the center of the obstacle, R O is the radius of the obstacle.
4. A method for capturing multiple drones according to claim 2, characterized in that S2 include: Constructing a seven-tuple of multiple drones in a three-dimensional dynamic environment<S,A,R,P,Z,O,γ> , where S is the state set, S = {s1, s2, ..., s i ,...,s N },s i is the state of the ith drone and s i ={x i ,y i ,z i ,v i ,ψ i ,θ i ,φ i }; A is a set of actions and A={a1,a2,...,a i ,...,a N }, a i is the action of the ith UAV, a i =[u i ,ω ψ,i ,ω θ,i ,ω φ,i ] T , T is the transposition operation; R represents the reward obtained by the UAV when taking the corresponding action in a certain state; P is the probability that the UAV takes an action in one state to transfer to another state; Z is the observation set, Z={z1,z2,...,z i ,...,z N }, z i is the observation information of the i-th UAV, z i ={x i ,y i ,z i ,v i ,ψ i ,θ i ,φ i }; O is the observation function; γ is the discount factor, which represents the preference and proportion of current rewards and future rewards.
5. The method for capturing multiple drones according to claim 2, wherein S3 include: S31. Initialize the parameters θ of the Actor network a And initialize the parameters θ of the Critic network c ; Initialize the parameters of the Actor target network And initialize the parameters of the Critic target network S32, perform actions on each drone to observe the next state and reward, form samples and put them into the experience playback pool. When the number of samples reaches the preset value, execute S33 to update the Critic network and execute S34 to update the Actor network. S33, update the Critic network; S34. Update the Actor network; S35, update the parameters of the Actor target network and the Critic target network, return to execute S31 to S35, until the objective function of the Critic network reaches the minimum value and the objective function of the Actor network reaches the maximum value, or when the Actor network and the Critic network reach the preset training rounds, stop training to obtain the optimized Critic network and the optimized Actor network; S36. Use the optimized Actor network to generate the corresponding drone's action, use the optimized Critic network to generate the reward for the corresponding drone's action, and each drone executes the action of its corresponding optimized Actor network to complete the multi-drone capture of a single drone.
6. A multi-UAV capture method according to claim 5, characterized in that: S33 includes: S331, randomly sample a batch of data from the experience replay pool (x j ,a j ,r j ,x ′ j ), calculate the target value: Among them, x j represents the initial state of the jth UAV, a j represents the action of the jth drone, r j represents the reward value of the jth drone, x ′ j represents the next state of the j-th UAV; Q is the Q function, Indicates that the next state of the Critic target network of the j-th drone takes action a ′ j The Q value obtained; S332, parameters θ of the Critic network c Add disturbance signal S(t): in, is θ c An estimated value of S333. Construct the objective function of the Critic network: Indicates that the initial state of the Critic network of the j-th UAV takes action a j The Q value obtained; l represents the number of samples; S334, using feedback control update Among them, K represents the gain; M(t) is the noise signal, and the updated As the current Return to S332, and execute S35 after looping for a preset number of times.
7. A multi-UAV capture method according to claim 6, characterized in that: S34 includes: S341. Calculate the entropy value of the current action strategy: in is based on the parameter θ a The strategy distribution of Indicated in strategy For action a under the distribution of i Expected value; S342, dynamically adjust the entropy regularization coefficient to obtain the regularization coefficient at the current moment Among them, α start is the initial regularization coefficient, α end is the final regularization coefficient, train_step is the current training time step, t decay is the attenuation frequency; S343, parameters θ for the Actor network a Add disturbance signal S(t): is θ a An estimated value of S344. Objective function of constructing Actor network: To output the corresponding action a according to the state s of the drone; S345, Update using feedback control The updated As the current Return to S343, and execute S35 after looping for a preset number of times.
8. A multi-UAV capture method according to claim 5, characterized in that: S35 includes: The parameters of the Actor target network and the Critic target network are soft updated. The soft update process is as follows in, is the updated parameter of the Actor target network, is the updated parameter of the Critic target network, τ is the soft update coefficient and τ∈(0,1], is the parameter of the Critic target network before updating, These are the parameters of the Actor's target network before updating.
9. A multi-UAV capture system, characterized in that: include: Single UAV modeling module, used for dynamic modeling of UAV; The seven-tuple construction module is used to construct the seven-tuple of multiple drones in a three-dimensional dynamic environment to describe the multiple drone capture scene; The multi-UAV capture module includes an Actor network and a Critic network for each UAV. The Actor network generates the UAV's actions, and the Critic network generates the evaluation of the corresponding actions. The extreme value search and adaptive entropy are combined to continuously iterate and update the Actor network and the Critic network. The optimized Actor network is used to generate the optimal action. Each UAV executes the optimal action to complete the capture of a single UAV by multiple UAVs.
10. A multi-UAV capture system according to claim 9, characterized in that: The Single Drone Modeling Module is also used to: Suppose there are N drones, where N>1, and the discrete dynamic equation of multiple drones from time t to time t+1 is: Where i is the index of the UAV, UAV is the abbreviation of UAV, x, y, z are the x-axis, y-axis and z-axis coordinates of the UAV, v is the UAV speed, ψ is the UAV heading angle, θ is the UAV pitch angle, φ is the UAV roll angle, ω ψ,i is the heading angular velocity control input, ω θ,i is the pitch angular velocity control input, ω φ,i is the roll angular velocity control input, and Δt is the time step from time t to time t+1.
Citation Information
Patent Citations
State prediction and DDPG combined multi-unmanned aerial vehicle hunting method
CN113625775A
Cited By
Decision method for patrol path of unmanned aerial vehicle in complex environment
CN120521607A
Unmanned aerial vehicle obstacle avoidance method based on multi-agent graph reinforcement learning
CN121635458A