Heterogeneous UAV layered formation control method and system based on Shapley value and coupled SAC
Through the Shapley value and coupled SAC heterogeneous UAV layered formation control method, combined with the artificial potential field method and multi-agent reinforcement learning, the stability and adaptability problems of multi-UAV formations in complex environments are solved, and efficient formation tracking and anti-collision control are achieved.
Patent Information
- Application Number
- CN202411373011.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-29
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-09-29
AI Technical Summary
Existing multi-UAV formation tracking strategies experience performance degradation when the communication range is insufficient or a node fails. Homogeneous groups have limited fault tolerance and poor adaptability in complex environments. Dynamic nonlinearity leads to unstable control strategies, making it difficult for traditional methods to handle formation chaos and collision problems in complex environments.
A heterogeneous UAV hierarchical formation control method based on Shapley value and coupled SAC is adopted. By constructing an undirected graph to describe the relationship between UAVs, combining the artificial potential field method and multi-agent reinforcement learning, a neural network adaptive sliding mode control algorithm is used to design trajectory and attitude control signals to achieve collaborative tracking control of UAVs. The network weight update law is designed through the Lyapunov method to reduce dependence on the system dynamics model.
It improves the stability and adaptability of multi-UAV swarms in complex environments, enhances the robustness and tracking efficiency of the formation, solves the formation maintenance problem under insufficient communication distance and node failure, and improves the overall performance and decision-making level of the system.
Smart Images

Figure CN119396201B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of unmanned aerial vehicle (UAV) technology, and in particular to a heterogeneous UAV layered formation control method and system based on Shapley value and coupled SAC. Background Art
[0002] With the widespread use of drones in both military and civilian applications, the difficulty and complexity of the missions they can perform are constantly increasing. For these challenging and complex missions, multiple drones often need to work together. Forming multiple drones in formation can expand the field of view, increasing reconnaissance and search ranges. For this reason, multi-drone formation flight control technology is crucial.
[0003] Existing multi-UAV formation tracking strategies rely heavily on the overall system's network topology. Consequently, their performance often degrades significantly when the communication range between UAVs fails to meet control requirements. Furthermore, existing tracking strategies place high demands on the perception capabilities of network nodes. If some nodes fail, resulting in a loss of detection functionality or a lack of target detection capability, this inevitably negatively impacts the coordinated tracking performance of the entire UAV formation.
[0004] At the same time, most current drone swarm formation tracking control systems only consider a single type of drone. This makes them poorly adaptable to complex and changing mission environments, with limited fault tolerance across the entire homogeneous swarm and a single advantage in collaborative scenarios. As application scenarios become increasingly complex, a single type of drone can no longer meet mission requirements.
[0005] Furthermore, dynamic nonlinearity is prevalent in multi-UAV cooperative control systems. Due to the inherent complexity of the system or the external environment, these nonlinearities often prevent mathematical models from accurately describing the actual system dynamics. Consequently, the control system cannot utilize accurate state information, leading to unstable control strategies and potentially degraded system performance, disrupted stability, and even formation failure.
[0006] While some existing multi-UAV collaborative control methods based on traditional multi-agent reinforcement learning and artificial potential field methods exist, these approaches suffer from difficulties in convergence and chaotic formation topology. When the system environment becomes complex, multi-agent reinforcement learning algorithms and artificial potential field methods struggle to cope with the conflicts between different factors, resulting in poor formation tracking and even collisions, loss of control, and formation failure. Summary of the Invention
[0007] In order to address some or all of the problems in the prior art, the present invention provides, in a first aspect, a method for controlling a layered formation of heterogeneous UAVs based on Shapley values and coupled SAC, comprising:
[0008] Confirm the position and posture of the tracking drone and the target drone, as well as their relative position information;
[0009] Based on the position and attitude information, a position control module determines a trajectory control signal to drive the tracking UAV to track the target UAV and maintain a desired formation;
[0010] Determining an attitude control signal through an attitude control module to drive the tracking drone to track a desired yaw trajectory;
[0011] Performing collaborative tracking control on the UAV based on the trajectory control signal and the attitude control signal, and calculating a tracking error variable; and
[0012] Determine whether the tracking error variable meets the preset conditions. If so, use the current trajectory control signal and attitude control signal as the optimal control strategy. Otherwise, update the trajectory control signal and attitude control signal until the preset conditions are met.
[0013] Furthermore, the heterogeneous UAV layered formation control method further includes:
[0014] Constructing a dynamic model of a UAV, wherein the UAV is a quadrotor UAV; and
[0015] An undirected graph is established to describe the interaction relationship of multiple UAV systems.
[0016] Furthermore, the desired formation includes an outer first drone group, an inner second drone group, and a target drone in the center, wherein the first drone group does not have a target detection function, and the second drone group has a target detection function.
[0017] Furthermore, based on the position and posture information, determining the trajectory control signal by the position control module includes:
[0018] A multi-agent reinforcement learning method that combines the artificial potential field method and Shapley value is used to determine the trajectory control signal of the UAV.
[0019] Furthermore, for each UAV, the determination of its trajectory control signal includes:
[0020] Determine the first, second, and third action signals through the Actor submodule according to the position information of the homogeneous neighbor drones, heterogeneous neighbor drones, and the target drone respectively;
[0021] determining a fourth action signal based on the position information and the first, second, and third action signals;
[0022] Based on the first, second, and third action signals, determining a first auxiliary control signal through a homogeneous subgroup interaction potential field, a target interaction potential field, and a heterogeneous subgroup interaction potential field; and
[0023] Based on the first auxiliary control signal, a trajectory control signal is determined by a neural network adaptive sliding mode position control algorithm.
[0024] Furthermore, the determination of the trajectory control signal also includes:
[0025] The marginal contribution Shapley Q value function is used to evaluate the Actor submodule based on the fourth action signal and position information, and perform iterative updates.
[0026] Furthermore, for each UAV, the determination of its attitude control signal includes:
[0027] Determining a roll angle and a pitch angle of the virtual drone based on the trajectory control signal;
[0028] generating a second auxiliary control signal using a PD control algorithm according to the roll angle and the pitch angle; and
[0029] Based on the second auxiliary control signal, a posture control signal is determined by a neural network adaptive sliding mode position control algorithm.
[0030] Based on the control method described above, the second aspect of the present invention provides a heterogeneous UAV hierarchical formation control system based on Shapley value and coupled SAC, comprising:
[0031] A position control module, which is used to determine the trajectory control signal of the UAV based on the position and posture of the tracking UAV and the target UAV;
[0032] an attitude control module, configured to determine an attitude control signal of the UAV according to the trajectory control signal and the attitude; and
[0033] The feedback module is used to calculate the tracking error variable to perform feedback iteration of the control signal.
[0034] Furthermore, the position control module includes:
[0035] An actor submodule, comprising first, second, third, and fourth sub-actors, wherein the first sub-actor is configured to generate a first action signal based on the position information of a homogeneous neighboring drone, the second sub-actor is configured to generate a second action signal based on the position information of a heterogeneous neighboring drone, the third sub-actor is configured to generate a third action signal based on the position information of a target drone, and the fourth sub-actor is configured to determine a fourth action signal based on the first, second, and third action signals and the position information;
[0036] a critic submodule, configured to generate a marginal contribution Shapley Q value based on the fourth action signal and position information;
[0037] an interaction potential field, comprising a homogeneous subgroup interaction potential field, a heterogeneous subgroup interaction potential field, and a target interaction potential field, wherein the interaction potential field is used to generate a first auxiliary control signal according to the first, second, and third action signals; and
[0038] A calculation submodule adopts a neural network adaptive sliding mode position control algorithm to determine a trajectory control signal based on the first auxiliary control signal.
[0039] Furthermore, the first, second, and third sub-actors all adopt LSTM networks, and the fourth sub-actor and the critic sub-module all adopt FC networks.
[0040] The present invention provides a hierarchical formation control method and system for heterogeneous UAVs based on Shapley values and coupled SAC. First, graph theory is used to represent the relationships between UAVs and their neighbors, establishing an undirected graph to represent the communication relationships between UAVs. The artificial potential field method is then combined with multi-agent reinforcement learning technology that incorporates Shapley values from game theory. A behaviorally coupled multi-agent soft actor critic (SAC) hierarchical heterogeneous multi-UAV collaborative tracking control method based on Shapley values is proposed. This method addresses the formation maintenance and collision avoidance issues when some UAVs in the swarm lack target detection capabilities, as well as the difficulty of maintaining communication when heterogeneous subgroups have different communication distances. It enhances the tracking efficiency of the heterogeneous multi-UAV collaborative formation, enabling stable operation in complex and uncertain environments. Secondly, a neural network adaptive sliding mode control method is employed, using a radial basis function neural network to fit the unknown system dynamics function to compensate for auxiliary control signals, thereby designing a controller. The adaptive update law for the network weights is then designed based on the Lyapunov method. Consequently, the controller design reduces the reliance on an accurate system dynamics model, eliminating the need for an accurate system dynamics model. This helps the system adapt to dynamic environmental changes and effectively improves its robustness. Based on the Shapley value method, the control method and system utilize a centralized training and distributed execution architecture. The overall joint action value is distributed to the action value of each drone according to the agent member's marginal contribution to the swarm, determining the average marginal benefit created by each drone member within the swarm. This addresses the problem of lazy agents, promotes the development of cooperative behavior and improved decision-making within the swarm, enhances overall system performance, and strengthens the adaptability and robustness of multi-drone swarms in complex dynamic environments. The improved multi-agent soft actor critic algorithm utilizes a decomposition-then-coupling actor network architecture. The output actions of the decomposed sub-actor networks serve as the strength of the layered interaction potential field in the artificial potential field method to generate more refined and accurate auxiliary control signals. The actions of the sub-actor networks are then coupled to form the final action for evaluation. This helps the drone swarm generate higher-quality decisions, enhances the dynamic path planning capabilities of the artificial potential field method, and improves the multi-agent reinforcement learning's ability to handle conflicts between different reward functions, leading to algorithm convergence. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] To further illustrate the above and other advantages and features of various embodiments of the present invention, a more detailed description of various embodiments of the present invention will be presented with reference to the accompanying drawings. It will be understood that these drawings depict only typical embodiments of the present invention and are not to be considered as limiting the scope thereof. In the drawings, for clarity, identical or corresponding components will be represented by the same or similar reference numerals.
[0042] Figure 1A schematic flow chart illustrating a method for controlling a layered formation of heterogeneous UAVs based on Shapley values and coupled SAC according to an embodiment of the present invention is shown;
[0043] Figure 2 A schematic structural diagram of a quad-rotor drone according to an embodiment of the present invention is shown;
[0044] Figure 3 Schematic diagram of an instantiated smooth potential function with finite cutoff value showing one embodiment of the present invention;
[0045] Figure 4 A schematic diagram illustrating the structure of a heterogeneous UAV hierarchical formation control system based on Shapley value and coupled SAC according to one embodiment of the present invention; and
[0046] Figure 5a and 5b A schematic diagram showing the initial position and ideal formation of a target formation tracked by a heterogeneous UAV layered formation control method based on Shapley value and coupled SAC according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0047] In the following description, the present invention is described with reference to various embodiments. However, those skilled in the art will recognize that the various embodiments can be implemented without one or more of the specific details or with other alternative and / or additional methods, materials, or components. In other cases, well-known structures, materials, or operations are not shown or described in detail to avoid obscuring the inventive aspects of the present invention. Similarly, for the purpose of explanation, specific quantities, materials, and configurations are described to provide a comprehensive understanding of the embodiments of the present invention. However, the present invention is not limited to these specific details. In addition, it should be understood that the various embodiments shown in the drawings are illustrative representations and are not necessarily drawn to scale.
[0048] In this specification, reference to "one embodiment" or "the embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. The appearances of the phrase "in one embodiment" in various places in this specification are not necessarily all referring to the same embodiment.
[0049] It should be noted that the embodiments of the present invention describe the process steps in a specific order. However, this is only for the purpose of illustrating the specific embodiment and does not limit the order of the steps. On the contrary, in different embodiments of the present invention, the order of the steps can be adjusted according to the process.
[0050] In the present invention, drones include those with target detection capabilities and those without. A homogeneous neighbor drone for any drone refers to a drone that is adjacent to the drone and has the same expected target detection capabilities, while a heterogeneous neighbor drone refers to a drone that is adjacent to the drone and has a different expected target detection capabilities. For example, for a drone with target detection capabilities, its homogeneous neighbor drone refers to a drone that is adjacent to the drone and has the expected target detection capabilities, while its heterogeneous neighbor drone refers to a drone that is adjacent to the drone and does not have the expected target detection capabilities.
[0051] To achieve heterogeneous multi-UAV formation tracking of moving targets, this paper uses graph theory to construct a multi-UAV formation tracking scenario model. Then, considering the internal and external environments of the UAVs, the UAVs are modeled based on the position subsystem and attitude subsystem.
[0052] For the position subsystem, the desired relative position information of each individual drone, its neighboring drones, and the tracking target is used to construct a potential function. Using the Shapley value-based behaviorally coupled multi-agent soft actor critic hierarchical heterogeneous multi-UAV collaborative tracking control method proposed in this paper, a Shapley value-based behaviorally coupled multi-agent soft actor critic algorithm is used to adjust the potential field strength of the interaction potential field to construct auxiliary control signals for formation tracking. A neural network adaptive sliding mode control algorithm is then used to address errors in position control caused by system and environmental disturbances, and a position controller is designed.
[0053] For the attitude subsystem, the error variable is constructed according to the set expected yaw angle and the yaw angle of the virtual drone, and the PD control algorithm is used to design the attitude auxiliary control signal. Then, the neural network adaptive sliding mode control algorithm is used to solve the error caused by system and environmental disturbances on the attitude control, and the attitude controller is designed.
[0054] In summary, the position control of the quadrotor drone affects the effect and dynamic response of attitude control. At the same time, attitude control is the basis for achieving precise position control. The two enable the control system to work closely together through coordinated feedback control to achieve flight missions.
[0055] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings of the embodiments.
[0056] Figure 1 A flow chart illustrating a method for controlling a layered formation of heterogeneous UAVs based on Shapley values and coupled SAC according to an embodiment of the present invention is shown.
[0057] like Figure 1 As shown in FIG, a hierarchical formation control method for heterogeneous UAVs based on Shapley value and coupled SAC includes:
[0058] First, in step 101, the initial position and attitude information is confirmed. The position and attitude of the tracking drone and the target drone, as well as their relative position information, are confirmed. In one embodiment of the present invention, in order to achieve formation control of a multi-quadrotor heterogeneous drone group, dynamic modeling is performed and the information exchange relationship is described. In the embodiment of the present invention, the drone is a quadrotor drone, and its structure is as follows: Figure 2 The rotor UAV is a highly nonlinear underactuated system with less control quantity than the system state variables. The change of motion state is achieved through the control input generated by the four rotors.
[0059] Based on this, according to the Lagrange equation, the dynamic model of the i-th quadrotor drone is as follows:
[0060]
[0061] in:
[0062] i∈N={N R ,N F}={v1,v2,...,v N} is the total set of drones, where N R N is a collection of drones with target detection capabilities. F A collection of drones that do not have the ability to detect targets;
[0063] are pitch angle, roll angle and yaw angle respectively;
[0064] [x i ,y i ,z i ] is the position of the center of mass of the quadrotor drone in the inertial coordinate system;
[0065] m i is the total load mass;
[0066] [I i,x ,I i,y ,I i,z ] is the inertial rotation matrix;
[0067] g is the acceleration due to gravity;
[0068] [U i,1 ,U i,2 ,U i,3 ,U i,4 ] is the control input of the four rotors;
[0069] External interference on the three position channels and the three attitude channels, respectively, is caused by environmental uncertainty such as wind speed and system uncertainty; and
[0070] [τ x,i ,τ y,i ,τ z,i ] are the control components on the three position channels of the quadrotor drone, where:
[0071]
[0072] In one embodiment of the present invention, in order to describe the communication relationship between N quadrotor drones, the interaction relationship of the multi-drone system is described by establishing an undirected graph as G = (V, E, A), where each drone node is described as V = N = {v1, v2, ..., v N}, the interaction relationship between two drone nodes is described by the edge set E(G)∈{(v i ,v j )|v i ,v j ∈V(G),i≠j}, and define a ij is the i-th row and j-th column element of the cluster system spatial topology adjacency matrix A(G), a ij ≥0 means edge (v i ,v j )’s weight:
[0073]
[0074] where r c is the communication distance between the two machines, p i =[x i ,y i ,z i ] is the three-dimensional coordinate position of any UAV i.
[0075] It should be understood that for any UAV i that does not have the target detection function, there is always a connection path to a UAV with the target detection function. Otherwise, the UAV i cannot obtain the target UAV information and the formation fails.
[0076] Next, in step 102, a trajectory control signal is determined. Based on the position and attitude information, a position control module determines a trajectory control signal to drive the tracking drone to track the target drone and maintain a desired formation. In one embodiment of the present invention, the desired formation includes a first drone swarm in the outer layer, a second drone swarm in the inner layer, and the target drone in the center, wherein the first drone swarm does not have a target detection function, while the second drone swarm does.
[0077] In one embodiment of the present invention, a potential function is constructed based on the system states and expected parameters of drone i and its neighboring drones. To ensure that multiple drones form a regular polygonal double-layer formation to track target g, the expected distance between individual drone i and its neighboring drones is defined as:
[0078]
[0079] in represents the expected distance between individual i of the drone group with target detection function and its neighboring individual drone j in the same group, represents the expected distance between individual i of the UAV swarm and its neighboring UAV j of the same type with the non-local detection function; r R represents the expected distance between the UAV swarm with target detection function and the target, r F Indicates the expected distance between the drone swarm and the target without the target detection function.
[0080] In one embodiment of the present invention, from the perspective of heterogeneous drone nodes, due to mission requirements or environmental interference, a certain type of drone group does not have the ability to detect the target and needs to rely on neighboring nodes to form a stable encirclement. To avoid obstruction of target detection, a drone group that can effectively detect the target is formed in the inner layer, and a drone group that cannot detect the target is formed in the outer layer. The distance between the two types of drone groups and the target satisfies r R <r F Double-layer encirclement tracking target.
[0081] In one embodiment of the present invention, a smooth attraction / repulsion function ψ is introduced α (z), used to generate control signals to enable the drone swarm to form a desired formation, is defined as follows:
[0082]
[0083] where z is the distance between the two particles, d α is the minimum point of the potential function. Figure 3 A schematic diagram of a smooth pair potential function with a finite cutoff value in one embodiment of the present invention is shown. Based on this, the interaction potential field of the homogeneous cluster is defined as and Defining the interaction potential field of heterogeneous clusters and The interaction potential field with the target is and Then instantiate the potential function:
[0084] make The instantiation potential function parameters are i,j∈N R , The instantiation potential function parameters are i,j∈N F , r R With r F is the communication distance within the same group of two types of UAVs;
[0085] make The instantiation potential function parameters are i∈N R ,j∈N F ,make The instantiation potential function parameters are i∈N F ,j∈N R ,in is the communication distance between two types of heterogeneous drone groups, Expressed as the desired spacing between the two types of drone swarms; and
[0086] make The instantiation potential function parameters are (r,d)=(r det ,r R ),make To instantiate the potential function parameters are (r, d) = (r det ,r F ), where r det is the distance at which a drone normally detects a target. In one embodiment of the present invention, r / d=1.2.
[0087] The construction of the potential function is used to describe the interaction force between two intelligent agents. The potential function can be used to calculate the attraction and repulsion between the intelligent agents. By adjusting the positions of the intelligent agents, the total potential energy of the system can be minimized, so that the multi-agent system can achieve a stable formation structure and realize anti-collision function.
[0088] In one embodiment of the present invention, a multi-agent reinforcement learning method that integrates an artificial potential field method and Shapley values is used to determine the trajectory control signals of drones. Specifically, for each drone, the first, second, and third action signals are first determined using the Actor submodule based on the position information of the drone's homogeneous neighbor drones, heterogeneous neighbor drones, and target drone. Then, based on the first, second, and third action signals, a first auxiliary control signal is determined using the homogeneous subgroup interaction potential field, the target interaction potential field, and the heterogeneous subgroup interaction potential field. Based on the first auxiliary control signal, a trajectory control signal is determined using a neural network adaptive sliding mode position control algorithm. Furthermore, a fourth action signal is determined based on the position information and the first, second, and third action signals. The marginal contribution Shapley Q-value function is then used to evaluate and iteratively update the Actor submodule based on the fourth action signal and the position information.
[0089] In one embodiment of the present invention, the position subsystem is defined as:
[0090]
[0091] in,
[0092] Based on this, the virtual drone position model is constructed as follows:
[0093]
[0094] where x i,s ,y i,s ,z i,s is the position signal of the virtual drone of the i-th drone, and is the auxiliary control signal to be designed.
[0095] Affected by the state constraints of position, speed, acceleration, heading angle, heading angular velocity, and heading angular acceleration in the UAV model, there is a positive constant B x 、B y 、B z 、B ψ , making In order to avoid collisions between drones and their neighbors and targets, the safe collision avoidance distance between individual drones in the tracking drone group is set to ρ1, and the safe collision avoidance distance with the target is set to ρ2.
[0096] Based on this, a Shapley value-based behavior-coupled multi-agent soft actor critic hierarchical heterogeneous multi-UAV collaborative tracking control method is constructed for heterogeneous multi-UAV collaborative formation tracking control to efficiently achieve tracking tasks. In one embodiment of the present invention, the design of the position control module includes:
[0097] First, we construct the marginal contribution Shapley Q-value function. Let N be a collective alliance of multiple drones, and let the individual drones in the multi-drone system form an arbitrary alliance C. We introduce the Shapley value in the cooperative game field and distribute the benefits according to the marginal contribution rate of each member to the alliance. The marginal contribution value of drone i is:
[0098]
[0099] where a C =[a j |j∈C] is the joint action of any UAV in the coalition C, μ C is the strategy set of alliance C.
[0100] Then the marginal contribution Shapley Q value function of UAV i is:
[0101]
[0102] Since the computational complexity of traversing alliance C is huge, the sub-alliances are processed to obtain the approximate Shapley Q value:
[0103]
[0104] Where M represents the number of sampling times; C k ∈{C1,C2,…,C M} is the sampled sub-coalition.
[0105] The behavior-coupled multi-agent soft actor critic algorithm based on Shapley value adopts the actor-critic architecture to construct a global value function Q(s,a) for evaluating the state s=[o 1 ,…,o ∣N∣ ], execute the action set a=[a 1 ,…,a ∣N∣ ], where o i is the incomplete observation of agent i, individual action a i is the deterministic action function π(·|o i ). Using a deep neural network to fit the relevant functions, each agent i has a deterministic policy network and a value network Q(s,a;ω i ). Each agent i in the behavior-coupled multi-agent soft actor critic algorithm based on Shapley value uses two deep neural Shapley soft action value networks (with parameters ω1 and ω2 respectively), and takes the network that outputs the minimum action value as the target network to solve the overestimation problem. Then the value network Q(s,a;ω i ) is:
[0106]
[0107] By introducing the concept of Shapley value in game theory, the search for the optimal solution is transformed into the search for an equilibrium solution of a strategy combination that satisfies all agents. Then, the strategy learning of all agents has a common objective function:
[0108]
[0109] in Represents the randomness of the policy π in the state s, then the policy learning can be described as the following optimization problem:
[0110]
[0111] In one embodiment of the present invention, different entropies are required at different training stages. In order to automatically adjust the entropy regularization term, a loss function of α is obtained:
[0112]
[0113] When the entropy of the policy is lower than When the entropy of the strategy is higher than the target value, minimizing the loss function L(α) will increase the value of α and enhance the exploration. When , minimizing the loss function L(α) will reduce the value of α, making the strategy training more focused on improving the target value.
[0114] The introduction of Shapley value can solve the problem of lazy agents. At the same time, the introduction of entropy regularization term will further enhance the exploration of agent actions during training, strengthen the development of cooperative behavior and improve the decision-making level of the system.
[0115] Next, construct the state and action space. A Markov decision process usually consists of a tuple In the formation tracking scenario limited by communication, each UAV cannot perceive all the information of the environment, and each agent has a local observation o i .
[0116] In an embodiment of the present invention, the observation consists of three parts at the current time step t: defining the two-dimensional plane coordinate ζ = {x, y}, the first part is the distance of the drone i relative to the target g The second part is the state information of UAV i relative to its isomorphic neighbors where i,j∈N R / N F ; The third part is the status information of UAV i relative to its heterogeneous neighbors where i∈N F / N R , j∈N R / N F ;
[0117] To achieve continuous action and precise control, in one embodiment of the present invention, a continuous action space is adopted. For the continuous action space, the policy of the behavior-coupled multi-agent soft actor critic algorithm based on Shapley value outputs the mean and standard deviation of the Gaussian distribution. However, it is not differentiable for the actions sampled from the parameterized probability distribution. The reparameterization technique is used to transform the sampling process into a deterministic function. The mean μ and standard deviation σ of the action distribution output by the policy network are then sampled from the unit Gaussian distribution. The reparameterized action is:
[0118] a t =fθ (∈ t ;s t )=a=μ+σ⊙∈;
[0119] Next, we design a reward function. This reward function guides the agent's strategy learning, using environmental feedback as a real-time evaluation of its behavior to adjust its strategy and drive the agent toward its goals. In one embodiment of the present invention, the reward function achieves three goals: maintaining tracking of the target g, avoiding collisions with neighbors and the target, and maintaining a desired distance from neighbors and the target to form a desired formation. Based on this, the reward function is as follows:
[0120] Approaching the target: The reward function in this part drives UAV i to approach the dynamic target UAV and maintain the desired distance, so that the relative position error between the UAV and the target is Then the negative reward of drone i approaching the target at time step t is set as:
[0121]
[0122] Avoid collision: This part of the reward function drives UAV i to maintain a safe distance from its neighbors and targets to avoid collisions. The negative reward for avoiding collisions for UAV i at time step t is set as:
[0123] r2=-r g -r n ,
[0124] where r n The negative reward function to avoid collision between drone i and its neighbors:
[0125]
[0126] r g The negative reward function to avoid collision between drone i and the target:
[0127]
[0128] Formation maintenance: This part of the reward function drives UAV i to maintain the desired distance with the homogeneous and heterogeneous swarms, thereby forming the desired formation. The negative reward for UAV i’s formation maintenance at time step t is set as:
[0129] r3=-r hm -r ht ,
[0130] where r hm is the negative reward function maintained between drone i and the homogeneous group formation, and the relative position error between drone i and its homogeneous neighbors is but:
[0131]
[0132] where r ht is the negative reward function maintained between drone i and the heterogeneous fleet formation, and the relative position error between drone i and its homogeneous neighbors is but:
[0133] as well as
[0134] Quick completion: Let ρ3 be the threshold of the reward for the formation to be designed. When the value of r3 is greater than or equal to ρ3, this part of the reward function drives the UAV system to quickly form the target formation. Design the negative reward function:
[0135]
[0136] In summary, the entire reward function of drone i at the current time step t can be derived as:
[0137] r i,t =r1+r2+r3+r4;
[0138] Next, the behavior-coupled soft actor critic algorithm is used. The drone state information s is divided into the position information of homogeneous neighbor drones, the position information of heterogeneous neighbor drones, and the position information of the target drone. The position information of homogeneous neighbor drones is used to adjust the interaction potential field strength between homogeneous drones, which is set as Ω∈{R,F}, the location information of heterogeneous neighboring UAVs is used to adjust the interaction potential field strength between heterogeneous UAVs, which is set as The target UAV position information is used to adjust the interaction potential field strength between the tracking UAV and the target UAV, which is set as s g .
[0139] In one embodiment of the present invention, the behavior-coupled multi-agent soft actor critic uses three different sub-actor networks to process these three types of information respectively and output a g Three actions, where Ω∈{R,F}, use the above three actions as the interaction potential field strength between the current drone and its homogeneous neighbors, heterogeneous neighbors and targets to achieve more refined and accurate drone position control. Then, these three actions and state information s are combined into a new vector as the input of the fourth sub-actor network to generate a final action a, which is used as the input of the critic network for value evaluation and guidance of the actor network.
[0140] Compared to fully connected layers that only focus on information in the current time step, the introduction of LSTM networks helps drones better understand information about neighboring drones and target drones. Furthermore, LSTM networks maintain stable gradients through a special gate structure, mitigating the impact of vanishing and exploding gradients in RNNs. To reduce network computational complexity, the RNN sequence length is set to 4, processing only the state information of the first four time steps. The first three sub-actors use LSTM networks, while the fourth sub-actor uses a fully connected layer. Using different network structures for different training objectives can reduce model complexity and enable the actor network to make more accurate decisions.
[0141] Next, the formation tracking auxiliary control signal is designed. According to the behavior-coupled multi-agent soft actor critic hierarchical heterogeneous multi-UAV cooperative tracking control method based on Shapley value, the auxiliary control signal is designed for the UAV i∈N with the ability to detect the target. R Design auxiliary control signals:
[0142]
[0143] For the undetectable target drone i∈N F Design auxiliary control signals:
[0144]
[0145] Design height tracking auxiliary control signal: where i∈N; and
[0146] Finally, the trajectory control signal is determined by designing a neural network adaptive sliding mode control strategy. Let η = {x, y, z} and the sliding mode function is designed as:
[0147]
[0148] in c η is a positive constant to be designed.
[0149] Derivative of the sliding mode function:
[0150]
[0151] From the analysis, we can see that is an unknown nonlinear function that includes system uncertainty and environmental interference. Using radial basis function neural network to approximate the unknown nonlinear function, we can get:
[0152]
[0153] in, is the ideal weight of the neural network, is the basis function vector,
[0154] is the approximation error.
[0155] Using the current radial basis neural network prediction The predicted value is:
[0156]
[0157] Then the prediction error of the unknown function is
[0158]
[0159]
[0160] The purpose of using radial basis neural network is to make e η,i Approaching 0, the formation tracking task is completed, and the weight update law is designed according to the Lyapunov stability theory. The designed Lyapunov function is
[0161]
[0162] Where γ>0. Then
[0163]
[0164] Set the sliding mode reaching law to:
[0165]
[0166] Among them, α η,i >0,β η,i >0, sgn(·) is the sign function.
[0167] The designed position control signal is:
[0168]
[0169] According to sin(·)+cos(·)=1, the trajectory control signal can be obtained:
[0170]
[0171] but but
[0172] Design the adaptive update law as:
[0173]
[0174] Where, ρη,i >0.
[0175] In summary, the position control module is used to drive drones to track targets and maintain the desired formation. On the one hand, the multi-agent soft actor critic algorithm incorporating Shapley values is used to solve the lazy agent problem during strategy training and achieve a Nash equilibrium for cooperative games in complex dynamic environments. On the other hand, by constructing a layered artificial potential field combined with reinforcement learning and designing a neural adaptive sliding mode control algorithm to address system uncertainty and environmental interference, a two-layer formation structure is formed, with the target drone at the center and weaker drones relying on more capable drones. This greatly enhances the robustness and adaptability of the entire system.
[0176] Next, in step 103, a posture control signal is determined. The posture control module determines the posture control signal to drive the tracking drone to follow a desired yaw trajectory. Specifically, the virtual drone's roll and pitch angles are determined based on the trajectory control signal. A PD control algorithm is then used to generate a second auxiliary control signal based on the roll and pitch angles. Finally, a neural network adaptive sliding mode position control algorithm is used to determine the posture control signal based on the second auxiliary control signal.
[0177] The position control and attitude control of a quadrotor drone are interdependent. Position control affects the effectiveness and dynamic response of attitude control. At the same time, attitude control is the basis for achieving precise position control. The two need to work closely together through a coordinated feedback control system to achieve flight missions. In one embodiment of the present invention, the design of the attitude control module includes:
[0178] First, the position control module obtained through the above steps can be used to calculate the roll angle and pitch angle of the virtual drone:
[0179]
[0180] The yaw angle of the i-th virtual drone The desired yaw angle ψ is designed i,d The attitude subsystem is:
[0181]
[0182] in:
[0183] Design the desired yaw angle ψ i,d , then the yaw angle of the virtual drone is With ψ i,d The tracking error between is:
[0184]
[0185] The virtual UAV yaw angle model is constructed as follows:
[0186]
[0187] Design attitude tracking auxiliary control signal:
[0188]
[0189] Among them, k3>0, k4>0, then:
[0190] e ψ =B(ψ s -ψ d 1 N ),
[0191] Among them, e ψ =[e ψ,1 ,…,e ψ,N ], but
[0192] According to Lyapunov stability theory, the Lyapunov function is constructed as follows:
[0193]
[0194] Where λ1 is the largest eigenvalue of B, a is a positive constant, and λ1 is the largest characteristic root of B. Select appropriate parameter values k3, k4, λ1, and a so that
[0195] Next, we design a neural network adaptive sliding mode control strategy to design the attitude controller. The sliding mode function is designed as:
[0196]
[0197] in c ξ is a positive constant to be designed.
[0198] Derivative of the sliding mode function:
[0199]
[0200] Among them, κ={2,3,4}.
[0201] From the analysis, we can see that is an unknown nonlinear function that includes system uncertainty and environmental interference. Using radial basis function neural network to approximate the unknown nonlinear function, we can get:
[0202]
[0203] in, is the ideal weight of the neural network, is the basis function vector,
[0204] is the approximation error.
[0205] Using the current radial basis neural network prediction The predicted value is:
[0206]
[0207] Then the prediction error of the unknown function is
[0208]
[0209] The purpose of using radial basis neural network is to make e ξ,i Approaching 0, the formation tracking task is completed, and the attitude control module U is designed by designing the Lyapunov function. i,κ for
[0210]
[0211] Among them, α ξ,i >0,β ξ,i >0.
[0212] According to Lyapunov stability theory, we have:
[0213]
[0214] but The design weight update law is as follows
[0215]
[0216] Where, ρ ξ,i >0.
[0217] In summary, by designing an attitude control module to drive the UAV to track the target and maintain the desired yaw angle, on the one hand, PD control is used to quickly reduce the attitude error, making the system response faster; on the other hand, by designing a neural adaptive sliding mode control algorithm to address system uncertainty and environmental interference, the system can form more precise maneuvers, allowing the system to achieve the encirclement target;
[0218] Next, in step 104, the tracking error variable is calculated. The UAV is tracked and controlled collaboratively based on the trajectory control signal and the attitude control signal, and the tracking error variable is calculated. as well as
[0219] Finally, in step 105, iterative update is performed. Is it satisfied where c p >0 is a custom smaller constant. If the above conditions are met, the current control strategy for tracking UAV i is the optimal control strategy. Otherwise, repeat the above steps to update the control strategy until the above conditions are met.
[0220] Figure 4 FIG2 shows a schematic diagram of a structure of a heterogeneous UAV layered formation control system based on Shapley value and coupled SAC according to an embodiment of the present invention. Figure 4 As shown, a heterogeneous UAV hierarchical formation control system based on Shapley values and coupled SAC includes a position control module, an attitude control module, and a feedback module. The position control module is used to determine the trajectory control signal of the tracking UAV and the target UAV based on the position and attitude of the tracking UAV and the target UAV. The attitude control module is used to determine the attitude control signal of the UAV based on the trajectory control signal and attitude. The feedback module is used to calculate the tracking error variable to perform feedback iteration of the control signal.
[0221] As shown in the figure, in one embodiment of the present invention, the position control module includes:
[0222] An actor submodule includes first, second, third, and fourth sub-actors, wherein the first sub-actor is configured to generate a first action signal based on the location information of a homogeneous neighboring drone, the second sub-actor is configured to generate a second action signal based on the location information of a heterogeneous neighboring drone, the third sub-actor is configured to generate a third action signal based on the location information of a target drone, and the fourth sub-actor is configured to determine a fourth action signal based on the first, second, and third action signals and the location information. In one embodiment of the present invention, the first, second, and third sub-actors each include an LSTM network, and the fourth sub-actor and the critic submodule include an FC network.
[0223] a Critic submodule, configured to generate a marginal contribution Shapley Q value based on the fourth action signal and position information;
[0224] an interaction potential field, comprising a homogeneous subgroup interaction potential field, a heterogeneous subgroup interaction potential field, and a target interaction potential field, wherein the interaction potential field is used to generate a first auxiliary control signal according to the first, second, and third action signals; and
[0225] A calculation submodule adopts a neural network adaptive sliding mode position control algorithm to determine a trajectory control signal based on the first auxiliary control signal.
[0226] Based on the control system, according to the Shapley value-based behavior-coupled multi-agent soft actor critic hierarchical heterogeneous multi-UAV cooperative tracking control method described in the previous steps, first initialize the tracking UAV position and velocity and the tracking target position p i,t and p g,t , initialize the relative position information and instantiate the potential function; next, initialize the actor network and divide the overall state information into three types of position information, including homogeneous neighbor drone position information, heterogeneous neighbor drone position information and target drone position information, as the input of the three sub-actor networks constructed with LSTM networks, obtain three actions as the potential field strength of the hierarchical interactive potential field, and finally use a new sub-actor network to couple the three actions to generate the final action; then construct the individual action value function Individual dual soft action value function j∈{1,2}, scores the final action, and updates the individual action value function through the global Shapley Q-value loss function back gradient propagation Promote the training of the optimal strategy; repeat the above steps to obtain the formation tracking auxiliary control signal Then, the position control module U is designed through the neural network adaptive sliding mode control algorithm. i,1 At the same time, the desired yaw angle is set, and the yaw angle attitude tracking auxiliary control signal is designed according to the PD algorithm. Then the virtual drone roll angle calculated by the designed horizontal position controller is and pitch angle and obtained from the virtual UAV yaw angle model The attitude control module U is designed by using the neural network adaptive sliding mode control algorithm. i,κ , k={2,3,4}. Based on this, the tracking error variable is calculated And determine whether it satisfies c p >0 is a custom smaller constant. If the above conditions are met, the current control strategy for tracking UAV i is the optimal control strategy. Otherwise, repeat the above steps to update the control strategy until the above conditions are met.
[0227] Figure 5a and 5bThe initial position and ideal formation diagram of target formation tracking using the Shapley value and coupled SAC-based heterogeneous UAV formation control method according to one embodiment of the present invention are shown. In this task, the total set of tracking UAVs is represented as N = {N R ,N F}, where N R is the set of drones that can effectively detect the target, N F A collection of drones that cannot detect targets; it contains at least 3 quadrotor drones with detection capabilities, 3 quadrotor drones without target detection capabilities, and 1 target drone; the communication distance between the two types of drones is The communication relationship between tracked UAVs is represented by an undirected graph.
[0228] If we use 5 quad-rotor drones with detection capabilities and 5 quad-rotor drones without target detection capabilities, the communication distance between the two types of drones is They need to form a double-layer regular pentagonal formation with the target drone as the center, drones without target detection capability in the outer layer, and drones with target detection capability in the inner layer to track the target drone. Five quadrotors with detection capability and five quadrotors without target detection capability start from any initial position and form a formation to track a moving target with an unknown motion trajectory. The initial drone position state is as follows: Figure 5a As shown, the height of each drone can be different. After adopting the cooperative tracking control method of the present invention, the effect of the ideal multi-drone cooperative hunting game confrontation is as follows: Figure 5b shown.
[0229] The heterogeneous UAV layered formation control method and system of the present invention solves the problems of formation maintenance and collision avoidance when some UAV groups do not have the ability to detect targets, and the difficulty of maintaining communication when the communication distances between heterogeneous subgroups are different.
[0230] According to the Lagrange equation, the dynamic model of the i-th quadrotor UAV is considered, and the virtual UAV position model is constructed for the position subsystem. Through the proposed behavior-coupled multi-agent soft actorcritic hierarchical heterogeneous multi-UAV cooperative tracking control method based on Shapley value, auxiliary control signals are designed for two types of UAV swarms, among which the auxiliary control signals are designed for the UAV i∈N with the ability to detect the target. R Design auxiliary control signals:
[0231]
[0232] For the undetectable target drone i∈N F Design auxiliary control signals:
[0233]
[0234] Output of the behavior-coupled multi-agent soft actor critic algorithm based on Shapley value a g , used to adjust the potential field strength of three types of interactive potential fields, and to enhance the artificial potential field method for formation maintenance and collision avoidance control of multiple quadrotor drones, thereby solving the problem of communication maintenance difficulties and enabling drones to make more intelligent decisions.
[0235] At the same time, the heterogeneous UAV layered formation control method and system also solves the problem of dynamic nonlinearity in multi-UAV systems. The design of the UAV i controller is divided into position subsystem and attitude subsystem:
[0236]
[0237] For the unknown nonlinear function and where η={x,y,z}, Through the neural network adaptive sliding mode control method, the sliding mode function is designed as:
[0238]
[0239] By using radial basis neural network to approximate unknown nonlinear functions,
[0240]
[0241] in and is the predicted value of the ideal weight of the neural network, and is the basis function vector.
[0242] According to Lyapunov stability theory, the adaptive update law is designed as:
[0243]
[0244] By introducing the neural network adaptive sliding mode control algorithm, the multi-UAV system has strong robustness in the face of system uncertainty and external interference, and can quickly adjust the control law according to system changes, solving the dynamic nonlinearity problem of the multi-UAV system, thereby improving the tracking accuracy of the control system.
[0245] The heterogeneous UAV layered formation control method and system improves the training stability and convergence speed of multi-agent reinforcement learning in complex environments. The multi-agent soft actor critic algorithm is a powerful machine learning method, but it faces the problem of difficult convergence of policy training in highly complex environments. The present invention adopts a Shapley value-based method and a centralized training distributed execution implementation architecture. The overall joint action value is allocated to the action value of each UAV according to the marginal contribution rate of the agent member to the group. The average value of the marginal benefit created by each UAV member in the group is determined. The introduced Shapley Q value function is:
[0246]
[0247] Since the computational complexity of traversing alliance C is huge, the sub-alliances are processed to obtain the approximate Shapley Q value:
[0248]
[0249] By introducing the Shapley Q value method, the lazy agent problem is solved.
[0250] This approach leverages a decomposition-then-coupling actor network architecture, using the output actions of the decomposed sub-actor networks as the strength of the hierarchical interaction potential field. The sub-actor network actions are then coupled to form the final action for evaluation. This approach helps drone swarms generate higher-quality decisions, enhances the dynamic path planning capabilities of the artificial potential field method, and sets appropriate sub-actor networks for different training objectives, enhancing multi-agent reinforcement learning's ability to handle conflicts between different reward functions.
[0251] By incorporating Shapley Q-value into the multi-agent soft actor critic algorithm to solve the credit allocation problem and the lazy agent problem, and using a decomposition-then-coupling actor network architecture to address the difficulty of reinforcement learning in converging when training in complex environments, the network architecture of the combined methods improves the training stability and convergence speed of multi-agent reinforcement learning in complex environments.
[0252] Although various embodiments of the present invention have been described above, it should be understood that they are presented by way of example only and not limitation. It will be apparent to those skilled in the relevant art that various combinations, modifications, and variations may be made thereto without departing from the spirit and scope of the present invention. Therefore, the breadth and scope of the present invention disclosed herein should not be limited by the exemplary embodiments disclosed above, but should be defined solely in accordance with the appended claims and their equivalents.
Claims
1. A heterogeneous UAV layered formation control method based on Shapley value and coupled SAC, characterized by: Including steps: Confirm the location and posture information of the tracking drone and the target drone; Based on the position and attitude information, a position control module is used to integrate an artificial potential field method and a multi-agent reinforcement learning method of Shapley value to determine the trajectory control signal of the UAV, so as to drive the tracking UAV to track the target UAV and maintain the desired formation. For each UAV, the Actor submodule is used to determine the first, second, and third action signals according to the position information of the UAV's homogeneous neighbor UAVs, heterogeneous neighbor UAVs, and target UAV respectively. A fourth action signal is determined based on the position information and the first, second, and third action signals. Based on the first, second, and third action signals, a first auxiliary control signal is determined through the homogeneous subgroup interaction potential field, the target interaction potential field, and the heterogeneous subgroup interaction potential field. Finally, based on the first auxiliary control signal, a trajectory control signal is determined through a neural network adaptive sliding mode position control algorithm. The marginal contribution Shapley Q value function is used to evaluate the Actor submodule based on the fourth action signal and position information, and iteratively update the Actor submodule. Determining an attitude control signal through an attitude control module to drive the tracking drone to track a desired yaw trajectory; Performing collaborative tracking control on the tracking UAV based on the trajectory control signal and the attitude control signal, and calculating a tracking error variable; as well as Determine whether the tracking error variable meets the preset conditions. If so, use the current trajectory control signal and attitude control signal as the optimal control strategy. Otherwise, update the trajectory control signal and attitude control signal until the preset conditions are met.
2. The heterogeneous UAV layered formation control method according to claim 1, characterized in that: Also includes the steps: Constructing a dynamic model of a UAV, wherein the UAV is a quadrotor UAV; and An undirected graph is established to describe the interaction relationship of multiple UAV systems.
3. The heterogeneous UAV layered formation control method according to claim 1, characterized in that: The desired formation includes a first drone group in an outer layer, a second drone group in an inner layer, and a target drone in the center, wherein the first drone group does not have a target detection function, and the second drone group has a target detection function.
4. The heterogeneous UAV layered formation control method according to claim 1, characterized in that: For each UAV, the determination of its attitude control signal includes the following steps: Determining a roll angle and a pitch angle of the virtual drone based on the trajectory control signal; generating a second auxiliary control signal using a PD control algorithm according to the roll angle and the pitch angle; and Based on the second auxiliary control signal, a posture control signal is determined by a neural network adaptive sliding mode position control algorithm.
5. A heterogeneous UAV hierarchical formation control system based on Shapley value and coupled SAC, characterized by: include: The position control module is configured to determine the trajectory control signal of the UAV according to the position and posture of the tracking UAV and the target UAV, so as to drive the tracking UAV to track the target UAV and maintain the desired formation. The position control module includes an Actor submodule, a Critic submodule, an interactive potential field, and a calculation submodule, wherein the Actor submodule includes a first, a second, a third and a fourth sub-Actor, wherein the first sub-Actor is configured to generate a first action signal according to the position information of the homogeneous neighbor UAV, the second sub-Actor is configured to generate a second action signal according to the position information of the heterogeneous neighbor UAV, and the third sub-Actor is configured to generate a second action signal according to the position information of the heterogeneous neighbor UAV. A third action signal is generated according to the position information of the target UAV, and the fourth sub-Actor is configured to determine a fourth action signal according to the first, second, and third action signals and the position information, the Critic sub-module is configured to generate a marginal contribution Shapley Q value based on the fourth action signal and the position information, the interaction potential field includes a homogeneous subgroup interaction potential field, a heterogeneous subgroup interaction potential field, and a target interaction potential field, the interaction potential field is configured to generate a first auxiliary control signal according to the first, second, and third action signals, and the calculation sub-module is configured to adopt a neural network adaptive sliding mode position control algorithm to determine a trajectory control signal based on the first auxiliary control signal; an attitude control module, configured to determine an attitude control signal of the UAV according to the trajectory control signal and the attitude, so as to drive the tracking UAV to track a desired yaw trajectory; as well as A feedback module is configured to calculate a tracking error variable for performing feedback iteration of the control signal.
6. The heterogeneous UAV layered formation control system according to claim 5, characterized in that: The first, second, and third sub-actors all use LSTM networks, and the fourth sub-actor and critic sub-module all use FC networks.
Citation Information
Patent Citations
Multi-time-varying formation tracking control method and system for network heterogeneous robot system
CN111522341A
Multi-robot layered formation control method and system based on deep reinforcement learning
CN118377304A