A Distributed Cooperative Guidance Law for UAV Swarms Based on Multi-Agent Reinforcement Learning Method
The method addresses parameter tuning issues in traditional cooperative guidance laws by integrating multi-agent deep reinforcement learning with proportional guidance to enhance exploration and convergence, improving the collaborative guidance of unmanned vehicle swarms.
Patent Information
- Application Number
- CN202310415715.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-04-18
AI Technical Summary
The coordinated guidance law parameter design of traditional drone clusters is cumbersome and the generalization performance is poor, making it difficult to achieve intelligent and efficient coordinated strikes.
Combining the deep reinforcement learning and proportional guidance rate of multiple agents, a distributed collaborative guidance law for drone clusters is constructed, and the initial exploration of reinforcement learning agents is optimized through the FACMAC network optimization, and intelligent and clustered guidance law parameters are designed.
The distributed collaborative guidance of the drone cluster has been realized, the collaborative guidance capability and convergence speed have been improved, and the network dimension explosion problem caused by excessive number of agents has been avoided.
Smart Images

Figure CN116610139B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a distributed cooperative guidance law for UAV swarms based on a multi-agent reinforcement learning method, belonging to the technical field of UAV swarm guidance. Background Art
[0002] With the progress of technology and the improvement of UAV technology level, UAV swarms play an increasingly important role in the military combat field. They can not only conduct reconnaissance and target search, but also undertake important tasks such as air-ground strikes and air-to-air defenses, and gradually become the main attack agents in modern warfare. Among them, suicide UAVs, compared with traditional missiles, do not require preset targets and can achieve flexible guidance based on real-time reconnaissance information. When no target is found, they can return for recovery, with the advantages of accurate strikes, flexible tactics, and low cost; compared with traditional UAVs, they have the advantages of strong concealment, compact structure, and flexible penetration tactics. With the development of suicide UAV swarms towards collectivization and intelligence, constructing an intelligent distributed cooperative guidance law that can effectively enhance the strike effectiveness and achieve intelligent decision-making has become the current research direction.
[0003] Traditional cooperative guidance laws often need to adjust the controller parameters through optimization algorithms or manual parameter tuning, without fully realizing the intelligence of parameter design, and often need to assume certain characteristics for the scenario, with poor generalization performance. In recent years, deep reinforcement learning technology has developed rapidly. Compared with traditional supervised learning or unsupervised learning, deep reinforcement learning uses a trial-and-error method to interact with the environment to obtain the maximum cumulative reward, and can be effectively applied to agent decision-making. It has made significant breakthroughs in game fields such as Atari games, Go, and StarCraft, as well as in fields such as autonomous driving and recommendation systems.
[0004] In multi-agent reinforcement learning algorithms, the MADDPG algorithm extends the DDPG algorithm to the multi-agent field, adopts the Actor-Critic framework, and follows the idea of "centralized training - distributed execution". During the training process, global information is input into the Critic network, and during the execution process, each Actor makes decisions only based on its own local observations, providing an effective idea for solving the problem of non-stationarity of the multi-agent system environment. FACMAC introduces the value decomposition idea into the MADDPG algorithm. The Critic network of each agent no longer uses global observations, but uses local observations to obtain the corresponding Q value, and then obtains the overall Q value through a hybrid network for value decomposition of all Q values, avoiding the problem that the input dimension of the Critic network becomes too large and difficult to converge due to the excessive number of agents. However, the FACMAC algorithm uses the method of introducing random Gaussian noise to ensure the lowest exploration in exploration, with low sample quality and low exploration efficiency in the initial stage of training, and is prone to falling into local optimal solutions. Summary of the Invention
[0005] The present invention aims to solve the problem of cumbersome parameter design in traditional cooperative guidance laws. By combining multi-agent deep reinforcement learning with the proportional guidance law, a distributed cooperative guidance law for UAV swarms is constructed to enhance the cooperative strike effectiveness of UAV swarms. The proportional guidance law is used to optimize the exploration of the reinforcement learning agent in the initial stage, and the FACMAC network is used to intelligently output all the parameters required for the design of the guidance law, so as to achieve intelligent and clustered distributed cooperation of UAV swarms.
[0006] To achieve the object of the present invention, the technical solution adopted by the present invention is: a distributed cooperative guidance law for UAV swarms based on the multi-agent reinforcement learning method. The training of the reinforcement learning network in this guidance law includes the following steps:
[0007] Step 1, initialize the parameters of the reinforcement learning network, including the neural network parameters θ of the Actor network μ, the Critic network Q, and the hybrid network Q tot 、θ Q 、θ μ and and their corresponding target network parameters θ Q′ 、θ μ′ and At the same time, initialize the experience replay pool and the cooperative guidance simulation environment.
[0008] Step 2, for the UAV swarm system, all UAVs interact with the environment to obtain their current own states s i . All agents communicate according to the current communication topology network, where the communication topology network is expressed by means of graph theory to construct an undirected graph G S =(V S , E S ), where V S ={1, 2,..., n} is the vertex set representing each UAV, is the edge set of the links for direct communication between UAVs. For any two UAVs m i and m j , if each UAV obtains its own current local observation o i .
[0009] Step 3, input the corresponding local observation o i of each agent into the corresponding Actor network μ to obtain the effective navigation ratio N i and the cooperative control item a c,i of the current agent.
[0010] Step 4, input the effective navigation ratio N i and the cooperative control item a c,i of each agent into the cooperative guidance law based on reinforcement learning, where the effective navigation ratio Ni Output the corresponding part of the action through the traditional proportional navigation law, and the calculation formula is as follows:
[0011]
[0012] In the formula, N i is the current effective navigation ratio, R represents the relative distance between the i-th UAV and the target, represents the line-of-sight angular velocity between the i-th UAV and the target.
[0013] Among them, and are solved through the non-linear relative motion equation of the UAV and the target, and the form is as follows:
[0014]
[0015] In the formula, m i , m j represents the UAV, T represents the target, V T represents the target velocity, a T represents the target acceleration, v m,i represents the velocity of the i-th UAV, a m,i represents the acceleration of the i-th UAV, θ i represents the ballistic inclination angle, λ i represents the line-of-sight angle between the missile and the target, σ i is the heading angle error of the i-th UAV, θ T,i represents the ballistic inclination angle of the target relative to the i-th UAV, λ T,i represents the line-of-sight angle between the target and the i-th UAV, σ T,i represents the heading angle error of the target relative to the i-th UAV, R i is the relative distance between the i-th UAV and the target.
[0016] Finally, for agent i, its corresponding action combines the acceleration a p,i obtained by inputting the effective navigation ratio into the traditional proportional guidance law and the cooperative control term a c,i output by the reinforcement learning network to obtain the action a i :
[0017] a i = a p,i + a c,i
[0018] Step 5, execute the action a i of each agent, and interact with the environment. Each agent obtains the current reward r i and the next moment state s i′, and store the previous state, previous action, current reward, and next state [s i , a i , r i , s i ′] into the experience replay pool.
[0019] Step 6, take out I historical experiences [s j , a j , s j+1 , r j from the experience replay pool, and update the parameters of the Critic network through the following loss function:
[0020]
[0021]
[0022] where π is the Actor network in the AC framework, also known as the policy network, and its parameters are represented as θ π ; μ is the Critic network, also known as the evaluation network, and its parameters are represented as θ μ ; π′ is the target network of the Actor network, and μ′ is the target network of the Critic network; γ is the attenuation degree of the reward.
[0023] Update the gradient of the Actor network in the following way:
[0024]
[0025] where E refers to taking the expectation, π is the Actor network in the AC framework, also known as the policy network, which is used to output actions, and its parameters are represented as θ π ; μ is the Critic network, also known as the evaluation network, which is used to output a single Q value, and its parameters are represented as θ μ ; Q tot is the Mixer network, also known as the mixing network, which is used to evaluate all Q values; π′ is the target network of the Actor network, μ′ is the target network of the Critic network, and Q tot ′ is the target network of the Mixer network; γ is the attenuation degree of the reward.
[0026] Considering that in the cooperative guidance scenario of this patent, all agents are homogeneous agents and in a fully cooperative environment, there is no competition among them. Therefore, the network parameters are shared, that is, different agents use the same set of networks for learning and training. Correspondingly, when calculating the loss function and gradient, it is necessary to change to the calculation of the global loss function and global gradient, and the corresponding formulas are:
[0027]
[0028] Among them, J i represents the gradient calculated by the i-th agent sample, and ω i represents the influence weight of agent i on the global gradient calculation, and n represents the total number of samples. By this strategy, the number of networks to be trained is reduced to achieve the effect of accelerating training.
[0029] In step 7, the parameters of the Critic network and the Actor network are updated by the soft update method:
[0030]
[0031] Among them, τ is the soft update coefficient; θ μ ′ is the target network parameter of the Actor network; θ π ′ is the target network parameter of the Critic network, is the target network parameter of the Mixer network.
[0032] In step 8, the multi-agent reinforcement learning network is trained in units of episodes until the gradient of the reinforcement learning network and the overall obtained reward converge to the target requirements, and finally a policy network that can be executed distributively is obtained to achieve the distributed cooperative guidance of the UAV swarm. In this scheme, by combining the traditional proportional guidance law with the multi-agent reinforcement learning method, the initial exploration of the reinforcement learning agent is optimized by the proportional guidance law, and the key parameters of the guidance law are output by the reinforcement learning network, realizing the fully intelligent design of the guidance law parameters. Finally, through training, a cooperative guidance law for the UAV swarm that can be executed distributively is obtained.
[0033] The present invention proposes a distributed cooperative guidance law for UAV swarms based on multi-agent reinforcement learning, realizing the distributed cooperative guidance of UAV swarms. By introducing a reinforcement learning network, the intelligent design of the cooperative guidance law parameters is realized. By introducing the proportional guidance law, the exploration of the reinforcement learning agent is optimized. The FACMAC deep reinforcement learning model is adopted to avoid the non-convergence caused by the explosion of too many dimensions of the agent. Experiments show that the algorithm of the present invention can effectively improve the convergence speed of the multi-agent reinforcement learning network in the cooperative guidance environment and the cooperative guidance ability of the UAV swarm. Description of the Drawings
[0034] Figure 1 is the flow chart of the method of the present invention,
[0035] Figure 2 is the overall framework diagram of the present invention,
[0036] Figure 3 is the pseudo code of the algorithm flow. Detailed Embodiment
[0037] The present invention will be further described below in conjunction with specific implementation cases. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, those skilled in the art's various equivalent modifications of the present invention all fall within the scope defined by the appended claims of this application.
[0038] Embodiment: A cooperative guidance law for UAV swarms based on multi-agent reinforcement learning. During the guidance process of the UAV swarm, each UAV obtains the corresponding effective guidance ratio and cooperative control term based on a distributed policy network. The proportional guidance law with variable effective navigation ratio is used to guide the intelligent body through the reinforcement learning network, and the cooperative performance between intelligent bodies is ensured through the cooperative control term. The training of the reinforcement learning neural network follows the idea of distributed execution and centralized training, including the following training steps:
[0039] Step 1: Initialize the parameters of the reinforcement learning network, copy the current network parameters to the target network, and initialize the experience replay pool;
[0040] Step 2: Interact with the environment. For each agent, obtain the initial state and the corresponding local observation;
[0041] Step 3: Input the local observations obtained by each agent into the corresponding policy network to obtain the current effective navigation ratio and cooperative control term;
[0042] Step 4: Obtain the corresponding action for each agent through the cooperative guidance law for the output items of the policy network;
[0043] Step 5: Each agent executes the action, interacts with the environment, obtains the reward for the current action, the state and observation at the next moment, and stores them in the experience replay pool;
[0044] Step 6: Retrieve historical experiences from the experience replay pool, obtain the Q value of the historical action through the evaluation network and the hybrid network, and update the network parameters through the loss function;
[0045] Step 7: Update the parameters of the online network and the target network in a soft update manner;
[0046] Step 8: Repeat the above steps until the reinforcement learning network converges.
[0047] The specific implementation process is as follows:
[0048] Step 1, initialize the parameters of the reinforcement learning network, including the neural network parameters θ of the Actor network μ, the Critic network Q, and the hybrid network Q tot of Q 、θ μ and and their corresponding target network parameters θ Q′, θ μ′ and Initialize the experience replay pool and the cooperative guidance simulation environment simultaneously.
[0049] Step 2, for the UAV swarm system, all UAVs interact with the environment to obtain their current self-states s i . All agents communicate according to the current communication topology network, which is expressed by graph theory, and an undirected graph G S =(V S , E S ) is constructed, where V S ={1, 2,..., n} is the vertex set representing each UAV, is the edge set of the links for direct communication between UAVs. For any two UAVs m i and m j , if each UAV obtains its current local observation o i .
[0050] Step 3, input the corresponding local observation o i of each agent into the corresponding Actor network μ to obtain the effective navigation ratio N i and the cooperative control term a c,i of the current agent.
[0051] Step 4, input the effective navigation ratio N i and the cooperative control term a c,i of each agent into the cooperative guidance law based on reinforcement learning, where the effective navigation ratio N i outputs the corresponding partial action through the proportional navigation law and finally executes the action a i = a p,i + a c,i
[0052] Step 5, execute the action a i of each agent and interact with the environment. Each agent obtains the current reward r i and the next moment state s i ′, and stores the previous moment state, the previous moment action, the current reward, and the next moment state [s i , a i , r i , s i ′] in the experience replay pool.
[0053] Step 6, take out I historical experiences [s j , a j , s j+1 , r j, the parameters of the Critic network are updated through the following loss function:
[0054]
[0055]
[0056] Among them, π is the Actor network in the AC framework, also known as the policy network, and its parameters are represented as θ π ; μ is the Critic network, also known as the evaluation network, and its parameters are represented as θ μ ; μ′ is the target network of the Actor network, and μ′ is the target network of the Critic network; γ is the attenuation degree of the reward.
[0057] The Actor network is updated by gradient in the following way:
[0058]
[0059] Among them, represents taking the gradient, represents the policy update target, and θ i is the parameter to be optimized in the policy network, E refers to taking the expectation, means that this action is directly obtained from historical experience in the experience replay pool; a i means that this action is obtained through the policy network using historical observations in the experience replay pool.
[0060] Considering that in the cooperative guidance scenario of this patent, the agents are all homogeneous agents and in a fully cooperative environment, there is no competition among them. Therefore, the network parameters are shared, that is, different agents use the same set of networks for learning and training. Correspondingly, when calculating the loss function and gradient, it is necessary to change to the calculation of the global loss function and global gradient, and the corresponding formula is:
[0061]
[0062] Among them, J i represents the gradient calculated from the i-th agent sample, and ω i represents the influence weight of agent i on the calculation of the global gradient, and n represents the total number of samples. By this strategy, the number of networks to be trained is reduced to achieve the effect of accelerating training.
[0063] In step 7, the parameters of the Critic network and the Actor network are updated by soft update:
[0064]
[0065] Among them, τ is the soft update coefficient; θ μ′is the target network parameter of the Actor network; θ π′ is the target network parameter of the Critic network.
[0066] In step 8, the multi-agent reinforcement learning network is trained in units of episodes until the gradients of the reinforcement learning network and the overall obtained rewards converge to the target requirements, and finally a policy network capable of distributed execution is obtained.
[0067] After the training is completed, each drone will use the trained and converged policy network to obtain its own actions according to its local observation state and execute them distributively, and finally realize the distributed cooperative guidance of the drone swarm.
[0068] It should be noted that the above embodiments are not used to limit the protection scope of the present invention, and equivalent transformations or substitutions made on the basis of the above technical solutions all fall within the protection scope of the claims of the present invention.
Claims
1. A distributed cooperative guidance law for UAV swarms based on reinforcement learning, characterized in that, The training of the reinforcement learning network in this guidance law includes the following steps: Step 1: Initialize the parameters of the reinforcement learning network, copy the current network parameters to the target network, initialize the experience replay pool and the environment. Step 2: Interact with the environment. For each agent, obtain the initial state and the corresponding local observation. Step 3: Input the local observations obtained by each agent into the corresponding policy network to obtain the current effective navigation ratio and the cooperative control term. Step 4: Obtain the actions corresponding to each agent through the cooperative guidance law for the output items of the policy network. Step 5: Each agent executes the action, interacts with the environment, obtains the reward of the current action, the state and observation at the next moment, and stores them in the experience replay pool. Step 6: Retrieve historical experiences from the experience replay pool, and obtain the values of the historical actions through the evaluation network and the hybrid network, and update the gradients of the network parameters through the loss function, and update the gradients of the network parameters through the loss function, Step 7: Update the parameters of the online network and the target network in a soft update manner. Step 8: Repeat the above steps until the reinforcement learning network converges. Among them, in step 4, the effective navigation ratio of each agent and the cooperative control item are input into the cooperative guidance law based on reinforcement learning, where the effective navigation ratio is used as the output action of reinforcement learning to output the corresponding part of the action through the proportional guidance law, and the calculation formula is as follows: Wherein, is the current effective navigation ratio, represents the relative distance between the th drone and the target, represents the line-of-sight angular velocity between the th drone and the target. Among them, and is solved from the non-linear relative motion equation of the drone and the target, and has the following form: Wherein, , represents an unmanned aerial vehicle, represents a target, represents the target speed, represents the th unmanned aerial vehicle speed, represents the th unmanned aerial vehicle acceleration, represents the ballistic inclination angle, represents the line-of-sight angle between the missile and the target, is the th unmanned aerial vehicle course angle error, represents the ballistic inclination angle of the target relative to the th unmanned aerial vehicle, represents the line-of-sight angle between the target and the th unmanned aerial vehicle, represents the course angle error of the target relative to the th unmanned aerial vehicle, is the th unmanned aerial vehicle and the relative distance between the target Finally, for the agent , its corresponding action is the acceleration obtained by the effective navigation ratio compared with the traditional proportional guidance law combined with the cooperative control term output by the reinforcement learning network : 。 2. The distributed cooperative guidance law for an unmanned aerial vehicle (UAV) swarm based on reinforcement learning according to claim 1, wherein In step 1, initialize the parameters of the reinforcement learning network, including the neural network parameters of the Actor network , the Critic network and the hybrid network , as well as their corresponding target network parameters , and , and , and . At the same time, initialize the experience replay pool and the initial environment.
3. The distributed cooperative guidance law for an unmanned aerial vehicle (UAV) swarm based on reinforcement learning according to claim 2, wherein In step 2, the UAV swarm is regarded as a multi-agent system. The th agent interacts with the environment to obtain the current corresponding state , and all agents communicate according to the communication topology network to obtain the current local observation .
4. The distributed cooperative guidance law for UAV swarms based on reinforcement learning according to claim 3, characterized in that, In step 3, the corresponding local observation of each agent is input into the corresponding Actor network , and the effective navigation ratio of the current corresponding UAV and the cooperative control term are obtained.
5. A distributed cooperative guidance law for UAV swarms based on reinforcement learning according to claim 2, characterized in that, In step 5, each agent executes its corresponding action , and after interacting with the environment, each agent obtains the current reward and the next moment state , and takes the previous moment state , the previous moment action , the current reward and the next moment state , and stores them jointly in the experience replay pool in the form of for offline training.
6. The distributed cooperative guidance law for an unmanned aerial vehicle (UAV) swarm based on reinforcement learning according to claim 2, wherein, In step 6, take out from the experience replay pool historical experiences , and update the parameters of the Critic network through the following loss function: Among them, is the Actor network in the AC framework, also known as the policy network, which is used to output actions, and the parameters are represented as ; is the Critic network, also known as the evaluation network, which is used to output a single Q value, and the parameters are represented as ; is the Mixer network, also known as the mixing network, which is used to evaluate all Q values; is the target network of the Actor network, is the target network of the Critic network, is the target network of the Mixer network; is the attenuation degree of the reward, The gradient of the Actor network is updated in the following way: Among them, represents calculating the gradient, represents the policy update target, are the parameters to be optimized of the policy network, refers to calculating the expectation, indicates that this action is obtained from the corresponding observation in the historical experience, Share the network parameters, that is, different agents use the same set of networks for learning and training. Correspondingly, when calculating the loss function and the gradient, change to the calculation of the global loss function and the global gradient. The corresponding formula is: Among them, represents the gradient calculated by the th agent sample, represents the influence weight of the agent on the global gradient calculation, represents the total number of samples. By this strategy, the number of networks to be trained is reduced to achieve the effect of accelerating training.
7. A distributed cooperative guidance law for UAV swarms based on reinforcement learning according to claim 2, characterized in that, In Step 7, whenever the number of training steps reaches the specified value, update the parameters of the Critic network and the Actor network in a soft update manner. Among them, is the soft update coefficient; is the target network parameter of the Actor network; is the target network parameter of the Critic network is the target network parameter of the Mixer network.
8. The distributed cooperative guidance law for UAV swarms based on reinforcement learning according to claim 2, wherein In Step 8, train the multi-agent reinforcement learning network in units of episodes until the gradient of the reinforcement learning network and the reward obtained as a whole converge to the target requirements, and finally obtain a policy network that can be executed distributively to achieve the distributed cooperative guidance of the UAV swarm.