A Reinforcement Learning-Based Method for UAV Swarm Formation Decision-Making and Encirclement

By establishing a consensus negotiation update equation and a multi-agent dynamics model based on reinforcement learning, the problems of formation manipulation and autonomous decision-making in UAV swarm formation decision-making and control are solved, enabling flexible formation manipulation and improving the encirclement efficiency and interpretability of UAV swarms.

CN119396171BActive Publication Date: 2026-01-30BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411471199.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-22
Publication Date
2026-01-30
Estimated Expiration
2044-10-22

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively address the formation decision-making and control issues of drone swarms, especially in dynamic network topologies and complex mission environments, where the interpretability and security of formation manipulation and autonomous decision-making are insufficient.

Method used

We employ a reinforcement learning-based approach to establish a consensus negotiation update equation and a multi-agent dynamics model. We implement formation decision-making in a distributed manner and combine it with model predictive control to learn formation parameters, thereby enhancing the interpretability and security of the decision-making process.

Benefits of technology

It enables flexible manipulation of multi-agent formations, enhances the application potential of multi-agent systems, improves encirclement efficiency and interpretability, and adapts to dynamic task changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119396171B_ABST
    Figure CN119396171B_ABST
Patent Text Reader

Abstract

This invention relates to a reinforcement learning-based method for UAV swarm formation decision-making and encirclement, belonging to the field of multi-agent formation control and decision-making technology. This invention establishes a consensus negotiation update equation and a multi-agent dynamic model, obtaining the desired positions of each agent and control commands for the multi-agent dynamic model. It achieves the desired formation configuration given by the decision system in a distributed manner without a central node, solving the multi-agent formation control problem. Secondly, it can intuitively manipulate the multi-agent formation to adapt to dynamically changing tasks, solving the multi-agent formation manipulation problem. Based on actual adversarial tasks, it realizes online autonomous decision-making for the learnable and evolutionary multi-agent system, solving the multi-agent autonomous formation decision-making problem. Finally, it utilizes control model mechanisms to enhance the interpretability and security of the learned decisions, considering the interpretability and security issues in the formation strategy learning process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-agent formation control and decision-making technology, specifically to a method for UAV swarm formation decision-making and encirclement based on reinforcement learning. Background Technology

[0002] Clustering and intelligence are important trends in the technological development of future system confrontation and multi-platform warfare. The core technology involved is multi-agent control and decision-making. Multi-agent control and decision-making technology can organize a group of individuals with simple structures and limited capabilities, and through reasonable information interaction, evolve and emerge complex organizational structures and agent behaviors, thereby accomplishing tasks that traditional single platforms cannot achieve.

[0003] The realization of these complex tasks undoubtedly relies on formation decision-making and control methods for multi-agent systems. Formation decision-making mainly addresses the question of what formation configuration the swarm should have to accomplish a specific task; while formation control studies how to safely (collision avoidance within the formation, environmental threat avoidance), controllably (maintaining formation shape, maintaining communication connections), and efficiently (fastest convergence speed, optimal position allocation) achieve the desired formation configuration given by the formation decision-making system. Currently, there is relatively more research on distributed control problems of UAV swarms, but less research on distributed formation decision-making problems, which is also more challenging. Considering the dynamic network topology of UAV swarms, applying intelligent learning algorithms for formation decision-making requires solving the "curse of dimensionality" and the problem of dynamic local information dimensionality. Furthermore, there is relatively little research on combining supervised learning, reinforcement learning, and other intelligent learning algorithms with cooperative control mechanism models to improve the guidance and interpretability of swarm decision-making results. Considering the dynamic expansion requirements and the complexity and variability of tasks in networked multi-agent systems, formation control and decision-making methods should possess typical characteristics such as distributed self-organization, scalable number of platforms, dynamic reconfigurability of network topology, and learnable evolution of policies. This undoubtedly makes distributed formation control and decision-making for multi-agent systems a challenging technical problem. Summary of the Invention

[0004] In view of the above problems, this invention provides a method for UAV swarm formation decision-making and encirclement based on reinforcement learning. This invention establishes a consensus negotiation update equation and a multi-agent dynamic model, obtaining the desired positions of each agent and control commands for the multi-agent dynamic model. It achieves the desired formation configuration given by the decision system in a distributed manner without a central node, thus solving the multi-agent formation control problem. Secondly, it can intuitively manipulate the formation of multiple agents to adapt to dynamically changing tasks, solving the multi-agent formation manipulation problem. Based on actual adversarial tasks, such as encirclement problems, it realizes online autonomous decision-making for the multi-agent system, enabling learning and evolution, thus solving the multi-agent autonomous formation decision-making problem. Finally, it utilizes control model mechanisms to enhance the interpretability and security of the learned decisions, considering the interpretability and security issues in the formation strategy learning process.

[0005] This invention provides a reinforcement learning-based method for drone swarm formation decision-making and encirclement, comprising:

[0006] Step 1: Determine the formation parameters of the multi-agent formation, and obtain the reference position of the multi-agent formation based on the formation parameters;

[0007] Preferably, the formation parameter expression of the multi-agent formation is:

[0008]

[0009] Where θ represents the formation parameters of the multi-intelligent formation, and p c ∈ 2 Indicates the center position of the formation; represents a real number. ζ represents the formation orientation, β represents the neighbor spacing, and T represents the transformation angle;

[0010] Preferably, the multi-agent formation includes translation, rotation, scaling, and deformation actions; the multi-agent formation has a better ability to cope with changes in the environment or task.

[0011] In one embodiment of the present invention, for the encirclement task, the multi-agent formation is an open encirclement formation and / or a closed encirclement formation;

[0012] Preferably, the expression for the reference position of the multi-agent formation in step 1 is:

[0013]

[0014] Where, p i Let q be the reference position of the i-th agent. i Let be the displacement of the reference position of the i-th agent relative to the reference position of the first agent.

[0015] When i = 1, qi Let be the displacement of the reference position of the i-th agent relative to the reference position of the first agent;

[0016] When i ≠ 1, i = 2, 3, ..., I, where I is the total number of agents. ψ i The direction of the reference position of the i-th agent relative to the reference position of the (i-1)-th agent.

[0017] Step 2: Establish a multi-agent formation parameter decision model; set constraints for the multi-agent formation parameter decision model;

[0018] Acquire local observation information of multi-agent formations;

[0019] A consensus negotiation update equation is established based on the multi-agent formation parameter decision model, constraints, and local observation information of the multi-agent formation.

[0020] Preferably, the expression for the consensus negotiation update equation in step 2 is:

[0021]

[0022] in, For updating the formation parameters of the i-th agent, π i Let represent the decision model for the formation parameters of the i-th agent. Based on the local observation information of the i-th agent, it outputs the update rate of the formation parameters, s. a Let o(·) be the attacker's state; o(·) be the agent's local observation information function; c θ The weighting coefficients for consensus negotiation, Let be the estimated formation parameters for the i-th agent. For the estimated formation parameters of the j-th agent adjacent to the i-th agent, Let i be the set of neighbors of the i-th agent who is a defender.

[0023] Preferably, the expression for the constraint condition in step 2 is:

[0024]

[0025] Where, θ * The optimal value of the formation parameters is if and only if When the equality holds, A collection of defensive intelligent agents.

[0026] Step 3: Update the formation parameters of each agent based on the consensus negotiation update equation to obtain the updated formation parameters of each agent;

[0027] The desired position is obtained by updating the formation parameters of each agent;

[0028] Step 4: Establish a dynamic model of the multi-agent system;

[0029] Control commands for the dynamic model of the multi-agent system are set based on the desired positions of each agent.

[0030] Preferably, the dynamic model of the multi-agent in step 4 includes a dynamic model of a holistic constrained multi-agent and a dynamic model of a non-holistic constrained multi-agent;

[0031] Furthermore, the control commands for the dynamic model of the nonholonomic constrained multi-agent system include linear velocity commands and angular velocity commands;

[0032] Furthermore, the expression for the complete constrained multi-agent dynamic model is as follows:

[0033]

[0034] in, Let v be the velocity vector of the i-th agent's position change. i Let C be the linear velocity of the i-th agent. d u is the drag coefficient of the intelligent agent. i For the i-th agent control command. This indicates that the control instruction for the i-th agent has a maximum value constraint. This represents the maximum value of the control input for the intelligent agent.

[0035] Furthermore, the expression for the dynamic model of the nonholonomically constrained multi-agent system is as follows:

[0036]

[0037] Among them, v i Let ω be the linear velocity of the i-th agent. i Let be the angular velocity of the i-th agent. Let x be the projection of the position change velocity of the i-th agent in the x-direction. Let be the projection of the position change velocity of the i-th agent in the y-direction. Let be the velocity of the i-th agent's orientation angle change. The orientation angle of the i-th agent.

[0038] Preferably, the expression for the control command of the multi-agent dynamics model in step 4 is:

[0039]

[0040] Where, k p These are the weighting coefficients corresponding to the position error. Let k be the desired position of the i-th agent as the defender. v These are the weighting coefficients corresponding to the speed error. It is the movement speed of the formation center.

[0041] Step 5: Establish a framework and learning environment model for reinforcement learning; embed the learning environment model into the reinforcement learning framework to obtain the reinforcement learning method;

[0042] Determine the learning samples for reinforcement learning methods;

[0043] Determine the loss function for the multi-agent formation parameter decision model;

[0044] The multi-agent formation parameter decision model is trained using reinforcement learning methods, learning samples, and the loss function of the multi-agent formation parameter decision model to obtain the final multi-agent formation parameter decision model.

[0045] Preferably, the learning environment model includes an action and state space, a reward system, and a learning objective;

[0046] Furthermore, the action space described in step 5 is the update rate θ of each component of the formation parameter θ. i ;

[0047] The state space is selected based on local observations from multiple agents, including the estimation of formation parameters. and attacker status s a ;

[0048] The reward is the reward for each state transition; the reward includes end-game rewards and immediate rewards.

[0049] The instant rewards include encirclement angle reward, alignment angle reward, average distance reward, and control energy reward;

[0050] Furthermore, the learning objectives include model learning objectives, evaluation learning objectives, and action learning objectives;

[0051] Furthermore, the expression for the learning objective is:

[0052]

[0053] Where, ω * For the model learning objective, ψ * To evaluate the learning objectives, φ * The learning objective is t, where t represents the learning time.

[0054] Furthermore, the learning samples include: model learning samples, evaluation learning samples, and action learning samples;

[0055] Furthermore, the model learns the sample expression as follows:

[0056]

[0057] in, This is the model replay sample corresponding to the k-th reference position of the attacking agent. For the a-th j The state of an attacking agent at the k-th reference position. For the a-th j The displacement of an attacking agent from the k-th reference position to the nearest attacking agent. For the a-th j The minimum displacement of an attacking agent from the k-th reference position to the encirclement formation. For the a-th j The state of an attacking agent at the (k+1)th reference position.

[0058] The expression for evaluating the learning sample is:

[0059]

[0060] in, Let q be the evaluation replay sample corresponding to the k-th reference position of the attacker agent, and s be the evaluation replay. k Let a be the k-th reference position of the attacking agent. k For the action of the attacker agent at the k-th reference position, s k+1 Let R(s) be the (k+1)th reference position of the attacking agent. k ,a k ) is the reward for the state transition at the k-th reference position of the attacking agent.

[0061] The expression for the action learning sample is:

[0062]

[0063] in, This is the action replay sample at the k-th reference position of the attacking agent. This is the first action of the attacking agent at the k-th reference position.

[0064] Furthermore, the loss function of the formation parameter decision model includes a model learning loss function, an evaluation learning loss function, and an action learning loss function.

[0065] Furthermore, the model learning loss function is expressed as follows:

[0066]

[0067] Among them, J m The loss function for model learning, The sample is any one of the mini-batch samples used by the model for learning, and the sample corresponds to the reference position of the attacker agent. For small batch samples, As a sample, For the a-th collected from the real environment j The state of an attacking agent at the (k+1)th reference position;

[0068] Furthermore, the loss function expression for the evaluation learning is:

[0069]

[0070] Among them, J q To evaluate the loss function of learning, To evaluate the learning of mini-batch samples, the samples correspond to reference positions, q ψ The objective evaluation network, where γ is the discount rate used to balance short-term and long-term rewards, As an independent objective evaluation network, It is an independent target policy network.

[0071] Furthermore, the expression for the action learning loss function is:

[0072]

[0073] in, Let be the policy gradient loss function. These are mini-batch samples representing the policy gradient, and the samples correspond to reference positions. π φ For the target policy network.

[0074] Step 6: When executing the encirclement mission, the multi-agents act as defenders, using the final multi-agent formation parameter decision model to encircle and capture the attacking cluster, and bring the attacking cluster into the safe zone.

[0075] Preferably, the capture task described in step 6 includes a capture phase and a transportation phase;

[0076] Preferably, the specific steps for encircling and capturing the attacking group based on the final formation parameter decision model include:

[0077] Step 61: When the attacking cluster is far away, multiple agents enter the encirclement phase; in the encirclement phase, each agent inputs the final formation parameter decision model and outputs the formation parameters of each agent.

[0078] Step 62: Adjacent agents negotiate formation parameters according to the consensus negotiation update equation to obtain the updated formation parameters of each agent;

[0079] Step 63: Input the updated formation parameters of each agent into the multi-agent dynamics model, and output the control commands of each agent; obtain the state of the corresponding agent based on the control commands of each agent;

[0080] Step 64: The attack cluster makes corresponding state transitions based on the states of each agent to obtain the updated state of the attack cluster.

[0081] Step 65: Determine whether the updated state of the attacking cluster is enclosed in the environment model. If yes, proceed to the next step. If no, repeat steps 62-64 until the updated state of the attacking cluster is enclosed in the environment model.

[0082] Step 66: After the updated state of the attacking cluster is surrounded by the environment model, it enters the transportation phase; the vector of the center velocity of the multi-agent formation is optimized based on the velocity obstacle method to maintain the closed and surrounded formation while avoiding environmental obstacles.

[0083] Furthermore, the transportation phase described in step 66 includes multiple transportation tasks; these multiple transportation tasks include avoiding environmental obstacles, preventing the attacking cluster from escaping, and transporting the attacking cluster to a designated safe area.

[0084] Compared with the prior art, the present invention has at least the following beneficial effects:

[0085] (1) This invention proposes a parameterized formation control method that can flexibly manipulate the formation of multi-agent systems, further enhancing the application potential of multi-agent systems;

[0086] (2) This invention establishes a multi-agent formation decision framework based on parameterized formation, which can achieve more efficient and generalizable multi-agent encirclement and capture.

[0087] (3) This invention designs a model-based reinforcement learning method that combines the idea of ​​model predictive control for trapping strategy learning, which has higher sample efficiency and better interpretability and guidance. Attached Figure Description

[0088] The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of the invention.

[0089] Figure 1 (a)-(b) are schematic diagrams illustrating the problem scenarios of the capture and transportation stages in the embodiments of the present invention;

[0090] Figure 2This is a schematic diagram of the overall framework of the multi-agent formation decision-making and encirclement method in an embodiment of the present invention;

[0091] Figure 3 (a)-(b) are schematic diagrams of parameterized encirclement formation design for open and closed encirclement formations in the embodiments of the present invention;

[0092] Figure 4 This is a schematic diagram of multi-agent encirclement based on formation parameter decision-making in an embodiment of the present invention;

[0093] Figure 5 This is a schematic diagram of the formation parameter decision algorithm architecture based on model reinforcement learning in an embodiment of the present invention;

[0094] Figure 6 This is a schematic diagram illustrating the non-escape angle in an embodiment of the present invention;

[0095] Figure 7 This is a schematic diagram of the alignment angle in an embodiment of the present invention;

[0096] Figure 8 This is a schematic diagram illustrating the algorithm execution and sample collection process of the learning method proposed in this embodiment of the invention;

[0097] Figure 9 (a)-(b) are schematic diagrams of the policy network loss curve and the evaluation network loss curve during the training process in the embodiments of the present invention;

[0098] Figure 10 (a)-(b) are schematic diagrams of the cumulative discount reward and win rate curves during the training process in the embodiments of the present invention;

[0099] Figure 11 (a)-(b) are schematic diagrams showing the simulation comparison and verification results of the encirclement stage and the traditional model predictive control in the embodiments of the present invention;

[0100] Figure 12 (a)-(b) are schematic diagrams of the formation parameter decision command and formation parameter change curve during the simulation process in the embodiments of the present invention;

[0101] Figure 13 (a)-(b) are schematic diagrams of the velocity curves of the defender and the attacker during the simulation process in the embodiments of the present invention;

[0102] Figure 14 (a)-(d) are schematic diagrams of the quantitative generalization simulation results in the embodiments of the present invention;

[0103] Figure 15 This is a schematic diagram of the experimental platform composition in an embodiment of the present invention;

[0104] Figure 16 This is a schematic diagram of the results of a multi-robot encirclement experiment in an embodiment of the present invention;

[0105] Figure 17 (a)-(c) are schematic diagrams of the robot's many-to-many driving experiment data in the embodiments of the present invention. Detailed Implementation

[0106] To better understand the above-described objectives, features, and advantages of the present invention, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other. Furthermore, the present invention can be implemented in other ways different from those described herein; therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.

[0107] A specific embodiment of the present invention, such as Figure 1-17 A method for drone swarm formation decision-making and encirclement based on reinforcement learning is disclosed.

[0108] To illustrate the effectiveness of the method proposed in this invention, the following detailed description of the above technical solution is provided through a specific embodiment. The specific implementation steps are as follows:

[0109] This invention provides a reinforcement learning-based method for drone swarm formation decision-making and encirclement, comprising:

[0110] Step 1: Determine the formation parameters of the multi-agent formation, and obtain the reference position of the multi-agent formation based on the formation parameters;

[0111] Preferably, the formation parameter expression of the multi-agent formation is:

[0112]

[0113] Where θ represents the formation parameters of the multi-intelligent formation, and p c ∈ 2 Indicates the center position of the formation; represents a real number. ζ represents the formation orientation, β represents the neighbor spacing, and T represents the transformation angle;

[0114] Preferably, the multi-agent formation includes translation, rotation, scaling, and deformation actions; the multi-agent formation has a better ability to cope with changes in the environment or task.

[0115] In one embodiment of the present invention, for the encirclement task, the multi-agent formation is an open encirclement formation and / or a closed encirclement formation;

[0116] Preferably, the expression for the reference position of the multi-agent formation in step 1 is:

[0117]

[0118] Where, p i Let q be the reference position of the i-th agent. i Let be the displacement of the reference position of the i-th agent relative to the reference position of the first agent.

[0119] When i = 1, q i Let be the displacement of the reference position of the i-th agent relative to the reference position of the first agent;

[0120] When i ≠ 1, i = 2, 3, ..., I, where I is the total number of agents. ψ i The direction of the reference position of the i-th agent relative to the reference position of the (i-1)-th agent.

[0121] Step 2: Establish a multi-agent formation parameter decision model; set constraints for the multi-agent formation parameter decision model;

[0122] Acquire local observation information of multi-agent formations;

[0123] A consensus negotiation update equation is established based on the multi-agent formation parameter decision model, constraints, and local observation information of the multi-agent formation.

[0124] Preferably, the expression for the consensus negotiation update equation in step 2 is:

[0125]

[0126] in, For updating the formation parameters of the i-th agent, π i Let represent the decision model for the formation parameters of the i-th agent. Based on the local observation information of the i-th agent, it outputs the update rate of the formation parameters, s. a Let o(·) be the attacker's state; o(·) be the agent's local observation information function; c θ The weighting coefficients for consensus negotiation, Let be the estimated formation parameters for the i-th agent. For the estimated formation parameters of the j-th agent adjacent to the i-th agent, Let i be the set of neighbors of the i-th agent who is a defender.

[0127] Preferably, the expression for the constraint condition in step 2 is:

[0128]

[0129] Where, θ *The optimal value of the formation parameters is if and only if When the equality holds, A collection of defensive intelligent agents.

[0130] Step 3: Update the formation parameters of each agent based on the consensus negotiation update equation to obtain the updated formation parameters of each agent;

[0131] The desired position is obtained by updating the formation parameters of each agent;

[0132] Step 4: Establish a dynamic model of the multi-agent system;

[0133] Control commands for the dynamic model of the multi-agent system are set based on the desired positions of each agent.

[0134] Preferably, the dynamic model of the multi-agent in step 4 includes a dynamic model of a holistic constrained multi-agent and a dynamic model of a non-holistic constrained multi-agent;

[0135] Furthermore, the control commands for the dynamic model of the nonholonomic constrained multi-agent system include linear velocity commands and angular velocity commands;

[0136] Furthermore, the expression for the complete constrained multi-agent dynamic model is as follows:

[0137]

[0138] in, Let v be the velocity vector of the i-th agent's position change. i Let C be the linear velocity of the i-th agent. d u is the drag coefficient of the intelligent agent. i For the i-th agent control command. This indicates that the control instruction for the i-th agent has a maximum value constraint. This represents the maximum value of the control input for the intelligent agent.

[0139] Furthermore, the expression for the dynamic model of the nonholonomically constrained multi-agent system is as follows:

[0140]

[0141] Among them, v i Let ω be the linear velocity of the i-th agent. i Let be the angular velocity of the i-th agent. Let x be the projection of the position change velocity of the i-th agent in the x-direction. Let be the projection of the position change velocity of the i-th agent in the y-direction. Let be the velocity of the i-th agent's orientation angle change. The orientation angle of the i-th agent.

[0142] Preferably, the expression for the control command of the multi-agent dynamics model in step 4 is:

[0143]

[0144] Where, k p These are the weighting coefficients corresponding to the position error. Let k be the desired position of the i-th agent as the defender. v These are the weighting coefficients corresponding to the speed error. It is the movement speed of the formation center.

[0145] Step 5: Establish a framework and learning environment model for reinforcement learning; embed the learning environment model into the reinforcement learning framework to obtain the reinforcement learning method;

[0146] Determine the learning samples for reinforcement learning methods;

[0147] Determine the loss function for the multi-agent formation parameter decision model;

[0148] The multi-agent formation parameter decision model is trained using reinforcement learning methods, learning samples, and the loss function of the multi-agent formation parameter decision model to obtain the final multi-agent formation parameter decision model.

[0149] Preferably, the learning environment model includes an action and state space, a reward system, and a learning objective;

[0150] Furthermore, the action space described in step 5 is the update rate of each component of the formation parameter θ.

[0151] The state space is selected based on local observations from multiple agents, including the estimation of formation parameters. and attacker status s a ;

[0152] The reward is the reward for each state transition; the reward includes end-game rewards and immediate rewards.

[0153] The instant rewards include encirclement angle reward, alignment angle reward, average distance reward, and control energy reward;

[0154] Furthermore, the learning objectives include model learning objectives, evaluation learning objectives, and action learning objectives;

[0155] Furthermore, the expression for the learning objective is:

[0156]

[0157] Where, ω * For the model learning objective, ψ * To evaluate the learning objectives, φ * The learning objective is t, where t represents the learning time.

[0158] Furthermore, the learning samples include: model learning samples, evaluation learning samples, and action learning samples;

[0159] Furthermore, the model learns the sample expression as follows:

[0160]

[0161] in, This is the model replay sample corresponding to the k-th reference position of the attacking agent. For the a-th j The state of an attacking agent at the k-th reference position. For the a-th j The displacement of an attacking agent from the k-th reference position to the nearest attacking agent. For the a-th j The minimum displacement of an attacking agent from the k-th reference position to the encirclement formation. For the a-th j The state of an attacking agent at the (k+1)th reference position.

[0162] The expression for evaluating the learning sample is:

[0163]

[0164] in, Let q be the evaluation replay sample corresponding to the k-th reference position of the attacker agent, and s be the evaluation replay. k Let a be the k-th reference position of the attacking agent. k For the action of the attacker agent at the k-th reference position, s k+1 Let R(s) be the (k+1)th reference position of the attacking agent. k ,a k ) is the reward for the state transition at the k-th reference position of the attacking agent.

[0165] The expression for the action learning sample is:

[0166]

[0167] in, This is the action replay sample at the k-th reference position of the attacking agent. This is the first action of the attacking agent at the k-th reference position.

[0168] Furthermore, the loss function of the formation parameter decision model includes a model learning loss function, an evaluation learning loss function, and an action learning loss function.

[0169] Furthermore, the model learning loss function is expressed as follows:

[0170]

[0171] Among them, J m The loss function for model learning, The sample is any one of the mini-batch samples used by the model for learning, and the sample corresponds to the reference position of the attacker agent. For small batch samples, As a sample, For the a-th collected from the real environment j The state of an attacking agent at the (k+1)th reference position;

[0172] Furthermore, the loss function expression for the evaluation learning is:

[0173]

[0174] Among them, J q To evaluate the loss function of learning, To evaluate the learning of mini-batch samples, the samples correspond to reference positions, q ψ The objective evaluation network, where γ is the discount rate used to balance short-term and long-term rewards, As an independent objective evaluation network, It is an independent target policy network.

[0175] Furthermore, the expression for the action learning loss function is:

[0176]

[0177] in, Let be the policy gradient loss function. These are mini-batch samples representing the policy gradient, and the samples correspond to reference positions. π φ For the target policy network.

[0178] Step 6: When executing the encirclement mission, the multi-agents act as defenders, using the final multi-agent formation parameter decision model to encircle and capture the attacking cluster, and bring the attacking cluster into the safe zone.

[0179] Preferably, the capture task described in step 6 includes a capture phase and a transportation phase;

[0180] Preferably, the specific steps for encircling and capturing the attacking group based on the final formation parameter decision model include:

[0181] Step 61: When the attacking cluster is far away, multiple agents enter the encirclement phase; in the encirclement phase, each agent inputs the final formation parameter decision model and outputs the formation parameters of each agent.

[0182] Step 62: Adjacent agents negotiate formation parameters according to the consensus negotiation update equation to obtain the updated formation parameters of each agent;

[0183] Step 63: Input the updated formation parameters of each agent into the multi-agent dynamics model, and output the control commands of each agent; obtain the state of the corresponding agent based on the control commands of each agent;

[0184] Step 64: The attack cluster makes corresponding state transitions based on the states of each agent to obtain the updated state of the attack cluster.

[0185] Step 65: Determine whether the updated state of the attacking cluster is enclosed in the environment model. If yes, proceed to the next step. If no, repeat steps 62-64 until the updated state of the attacking cluster is enclosed in the environment model.

[0186] Step 66: After the updated state of the attacking cluster is surrounded by the environment model, it enters the transportation phase; the vector of the center velocity of the multi-agent formation is optimized based on the velocity obstacle method to maintain the closed and surrounded formation while avoiding environmental obstacles.

[0187] Furthermore, the transportation phase described in step 66 includes multiple transportation tasks; these multiple transportation tasks include avoiding environmental obstacles, preventing the attacking cluster from escaping, and transporting the attacking cluster to a designated safe area.

[0188] The present invention also includes step 7, a simulation platform, and simulation results;

[0189] 71. Simulation Scene Setting

[0190] The simulation parameters are set as shown in the table below; the parameters of the proposed reinforcement learning method are set as shown in Table 1; simulation training is conducted using a baseline scenario of 8 defenders and 3 attackers.

[0191] Table 1 Simulation Scene Parameter Settings

[0192]

[0193] The parameter settings for the learning algorithm are shown in Table 2.

[0194]

[0195] 72. Simulation Training Results

[0196] The policy network is a traditional fully connected neural network, consisting of a normalization layer, two hidden layers, and an anti-normalization layer, which finally outputs updated formation parameters.

[0197] During simulation training, the loss function curve of the formation parameter decision model is as follows: Figure 9 As shown, the loss function value of the formation parameter decision model gradually decreases during training, dropping from a maximum of approximately 0.4 to below 0.05. This indicates that the direct output of the formation parameter decision model gradually converges to the action output optimized by the PSO algorithm. In subsequent simulations after training, if there are high real-time requirements for control command decisions, the action optimization process under the model predictive control framework during the execution phase is not necessary, and the formation parameter decision model can be directly used for formation parameter update decisions. Subsequent simulations and experiments will verify this.

[0198] Figure 10 The cumulative discounted reward and win rate curves for each round during training are presented. The cumulative discounted reward is a crucial indicator of policy performance, and maximizing it is the starting point for reinforcement learning algorithm design. During training, a -1 penalty is added to the immediate reward at each step to guide the agent to capture the target as quickly as possible. Therefore, when the capture task fails or takes a long time, the cumulative discounted reward will exhibit a large negative value; as the win rate increases and the capture time decreases during training, the cumulative discounted reward will gradually increase. Figure 10 (a) The cumulative discount rewards shown are still processed by moving index averaging with a smoothing factor of 0.9.

[0199] Another metric that can more intuitively reflect the performance of a strategy is the win rate of a task. Figure 10 (b) The win rates shown are obtained using a moving window averaging method with a window size of 20. This means that each point on the curve represents the win rate calculated from the wins and losses of the next 20 rounds. Initially, since no data has been collected for training, the policy at this stage is equivalent to a traditional model predictive control method without an evolutionary learning mechanism. As can be seen, the win rate is approximately 60% at this point, indicating that the learning method proposed in this chapter provides acceptable policy performance in the initial training stage rather than relying entirely on random trial and error. As the simulation training progresses, the model network, evaluation network, and policy network are gradually optimized, and the win rate can be stably maintained at 100% within 40 rounds.

[0200] 73. Simulation Verification of Comparison Methods

[0201] Using the baseline formation parameter decision model as a comparison method, the proposed reinforcement learning-based formation decision method is demonstrated in the application of a many-to-many encirclement problem. The simulation scenario uses 8 defenders and 3 attackers. The environment includes two polygonal obstacle areas and one circular obstacle area (impassable to defenders but permissible to attackers). The defenders must encircle all attackers within a closed formation and transport them all to a safe area for the mission to succeed.

[0202] Figure 11 Simulation results for two methods are presented, with the large left image showing the entire trajectory and the smaller right image showing scene snapshots at different times. Figure 11 (a) shows the simulation results of formation parameter decision-making using the traditional model predictive control method. Due to significant errors in the predictive model of the attacker's state transition and insufficient time for the PSO algorithm to fully optimize the decision-making actions, in this simulation, although the defenders were able to exhibit a strategy of coordinating to trap the attackers in an encirclement formation, the defenders formed an encirclement trend too early, causing two attackers to successfully escape the encirclement formation and eventually enter the protected area, resulting in the failure of our mission. Figure 11 (b) Simulation results of formation parameter decisions using the policy network trained by the proposed method. In this simulation, the defender can use the attacker's proximity to the protected area to lure them further into the encirclement formation before quickly closing it off, ultimately successfully trapping all attackers inside the encirclement formation.

[0203] Figure 12 The decision-making instructions and formation parameter variation curves of the formation parameter decision-making model during the simulation are presented. All defenders estimate the formation parameters θ based on their respective local observations. i Decisions are made using independent formation parameter decision models, depending on the attacker's state. Each defender's initial formation parameters differ randomly, so their initial decision commands are not entirely identical. However, this difference is quickly eliminated as the formation parameters converge through negotiation. During the simulation, the translation command fluctuates significantly when the attacking group is close enough to complete the encirclement; after the encirclement is complete, the translation command guides the encircling formation to move at a constant linear speed towards a safe area. The deformation command increases slowly at first and then rapidly during the initial defensive phase, thus completing the encirclement of the attacking group. Thereafter, the deformation parameter β remains at its maximum value.

[0204] Figure 13The simulation demonstrates the velocity changes of all defenders and attackers during a coordinated expulsion. The defenders' speeds consistently do not exceed their maximum speed of 2 m / s, while the attackers' speeds can reach 2.5 m / s. During the encirclement process, the velocity changes of the different defenders remain continuous and stable, and in the expulsion phase after the encirclement is complete, the speeds of all defenders are essentially the same. After being trapped in the encirclement formation, the attackers experience drastic fluctuations in both the amplitude and direction of their speeds. This is because the attacking group, while being forced to move to a safe area, continues to attempt to break free of the encirclement and reach the protected area.

[0205] 74. Quantitative Generalization Simulation Verification

[0206] This method makes decisions based on formation parameters, eliminating the need to directly use the global state. Therefore, it is suitable for distributed systems, and the learned formation parameter decisions can be flexibly applied to various encirclement and drive-away tasks with different numbers of defenders and attackers. This reuse of formation parameter decisions not only eliminates the need for redesigning the strategy, but also allows the direct use of the policy network's parameters without retraining. The figure below illustrates the performance of a strategy trained with 8 defenders and 3 attackers, directly applied to encirclement and drive-away scenarios with varying numbers of enemy and friendly forces.

[0207] Figure 14 (a)-(d) correspond to task scenarios with reduced attackers (8 vs. 2), increased attackers (8 vs. 4), simultaneous reduction of both defenders and attackers (6 vs. 2), and simultaneous increase of both defenders and attackers (10 vs. 5), respectively. In the simulation, each defender uses an independent pre-trained network to make formation parameter update decisions based on their own observations. The instructions after the policy network decision are directly used for formation parameter updates without going through the PSO algorithm optimization step. Simulation results show that the pre-trained network can successfully complete the task for task scenarios with different numbers of enemy and friendly forces. During the encirclement phase, it can confine the entire attacking group within the encirclement formation; during the transport phase, it can maintain the encirclement formation and successfully transport all attackers to the designated safe area.

[0208] Fixed-dimensional decision instructions can achieve formation control of different numbers of agents, which verifies that the method proposed in this invention does not depend on the number of agents and has good scalability and application prospects.

[0209] This invention also includes step 8, robot experimental verification; designing a robot physical experiment to verify the effectiveness of the method of this invention;

[0210] 81. Introduction to the composition of the experimental system

[0211] The experiment utilized a larger number of mobile robots, including six defensive robots and two offensive robots. A motion capture system was used for positioning, and commands were distributed based on the ROS2 robot operating system. To make the encirclement formation more visually clear, adjacent defensive robots were connected by elastic ropes, as detailed below. Figure 15 The experimental program was written in Python, using TensorFlow 2 as the reinforcement learning library, Ubuntu 18.04 as the computer system, and ROS2foxy as the robot's operating system version.

[0212] 82. Multi-robot encirclement experiment

[0213] This part of the work designs a physical experiment to verify the application of the proposed distributed formation decision-making method in a practical multi-robot encirclement and drive problem. The mobile robot is a wheeled robot with nonholonomic constraints, and its control commands are linear velocity and angular velocity. The maximum linear velocity and maximum angular velocity of the defender are set as v0 and v0, respectively. d =0.15m / s and ω d =1.2 rad / s; the attacker's is v a =0.2m / s and ω a = 3.6 rad / s.

[0214] like Figure 16 As shown, the six red robots at the bottom are defenders; the two robots with blue lights at the top are attackers; the cyan area at the bottom represents the protected zone; the green area at the top represents the safe zone; the defenders are connected by a net. Two attackers attempt to invade the protected zone below but are blocked by the defender formation. The defenders update their formation parameters through distributed decision-making, forming an encirclement around the attackers. Then, through formation deformation, they completely confine all attackers within the closed encirclement. The transport phase then begins, where the defenders move the encirclement formation to transport all attackers to the safe zone.

[0215] Figure 17 The data analysis results for the entire experiment are presented. Subgraph (a) shows the formation parameter update rate decision instructions for all defenders. It can be seen that, except for slight differences in the decisions made by different defenders regarding formation parameter updates at the very beginning when the formation parameters had not converged, the subsequent decision results were basically consistent. Subgraph (b) shows the formation parameters understood by each defender during the experiment. In subgraph (c), from top to bottom, the linear velocities v of the defenders are... d Defender's angular velocity ω d Attacker's linear velocity v a attacker's angular velocity ω a and the convergence error of the encirclement formation e fThe defender's speed never exceeds 0.15 m / s and converges to zero after reaching the safe region. Experiments show that the method of this invention can successfully complete the task.

[0216] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for formation decision and hunting of a UAV swarm based on reinforcement learning, characterized in that, Comprise: Step 1, determine the formation parameters of the multi-agent formation, obtain the reference position of the multi-agent formation based on the formation parameters of the multi-agent formation; Step 2, establish a multi-agent formation parameter decision model; Set the constraint conditions of the multi-agent formation parameter decision model; Obtain the local observation information of the multi-agent formation; Based on the multi-agent formation parameter decision model, the constraint condition and the local observation information of the multi-agent formation, a consistent negotiation update equation is established, The expression is: in, For the first i Update the formation parameters of each agent. Indicates the first i The decision model for the formation parameters of the first agent, based on the first... i The local observation information of each agent is used to output the update rate of the formation parameters. The attacker's state; For the local observation information function of the intelligent agent; The weighting coefficients for consensus negotiation, For the first i Estimated formation parameters for each agent In order to be with the first i The neighboring agents of the first agent j Estimated formation parameters for each agent For the first i Each agent is a set of neighbors of the defender; Step 3, update the formation parameters of each agent based on the consistent negotiation update equation, obtain the updated formation parameters of each agent; based on the updated formation parameters of each agent, the corresponding expected position is obtained; Step 4, establish a multi-agent dynamics model; based on the corresponding expected position of each agent, set the control instruction of the multi-agent dynamics model; Step 5, design a reinforcement learning method and determine the learning sample; determine the loss function of the multi-agent formation parameter decision model; use the reinforcement learning method, the learning sample and the loss function of the multi-agent formation parameter decision model to learn and train the multi-agent formation parameter decision model, and obtain the final multi-agent formation parameter decision model; Step 6, when performing the encirclement task, the attacking swarm is encircled based on the final multi-agent formation parameter decision model, and the attacking swarm is made to enter the safe area.

2. The UAV swarm formation decision and encirclement method according to claim 1, wherein the expression of the reference position of the multi-agent formation in step 1 is: I, I wherein, is the reference position of the i th agent, denotes the formation center position, is the reference position of the i th agent relative to the reference position of the first agent, i = 1, 2, 3… 3. The UAV swarm formation decision and encirclement method according to claim 1, wherein the multi-agent dynamics model in step 4 comprises a complete constraint multi-agent dynamics model and a non-complete constraint multi-agent dynamics model. denotes the total number of agents.

4. The UAV swarm formation decision and encirclement method according to claim 2, wherein the expression of the control instruction of the multi-agent dynamics model in step 4 is:

5. The UAV swarm formation decision and encirclement method according to claim 1, wherein the specific steps of designing the reinforcement learning method and determining the learning sample in step 5 comprise: Establish a reinforcement learning framework and a learning environment model; embed the learning environment model into the reinforcement learning framework to obtain a reinforcement learning method.

6. The UAV swarm formation decision and encirclement method according to claim 1, wherein the loss function of the formation parameter decision model comprises a model learning loss function, an evaluation learning loss function and an action learning loss function. wherein, is the control instruction of the i-th agent, i is the weight coefficient corresponding to the position error, is the expected position of the i-th agent as a defender, i is the reference position of the i-th agent, i is the weight coefficient corresponding to the speed error, is the moving speed of the formation center, is the linear speed of the i-th agent, i is the drag coefficient of the agent.​​​​ 7. The UAV swarm formation decision and encirclement method according to claim 1, wherein the specific steps of encircling the attacking swarm based on the final multi-agent formation parameter decision model in step 6 comprise: Step 61, when the attacking swarm is far away, multiple agents enter an encirclement stage; in the encirclement stage, each agent inputs the final formation parameter decision model and outputs the formation parameters of each agent. ​ ​ ​ ​ ​ ​ Step 62, the formation parameters of the adjacent intelligent agents are negotiated according to the consistency negotiation update equation to obtain the updated formation parameters of each intelligent agent; Step 63, the updated formation parameters of each intelligent agent are input into the multi-agent dynamics model to output the control instructions of each intelligent agent; and the state of the corresponding intelligent agent is obtained based on the control instructions of each intelligent agent; Step 64, the attack cluster makes corresponding state transition according to the state of each intelligent agent to obtain the updated state of the attack cluster; Step 65, it is judged whether the updated state of the attack cluster is surrounded into the environment model, if yes, the next step is entered, if not, steps 62-64 are repeated until the updated state of the attack cluster is surrounded into the environment model; Step 66, after the updated state of the attack cluster is surrounded into the environment model, the transportation stage is entered; and the vector of the center speed of the multi-agent formation is optimized based on the speed obstacle method to keep the closed surrounding formation and avoid the environmental obstacles.

8. The UAV cluster formation decision and hunting method according to claim 7, characterized in that, the transportation stage of step 66 comprises a plurality of transportation tasks; and the plurality of transportation tasks comprise avoiding environmental obstacles, preventing the attack cluster from escaping, and transporting the attack cluster to a designated safe area.

Citation Information

Patent Citations

  • Unmanned aerial vehicle cluster adaptive self-reconstruction method and system based on reinforcement learning

    CN115047912A

  • Multi-unmanned aerial vehicle air combat decision-making method based on multi-agent layered reinforcement learning

    CN115291625A