Distributed security formation control method for cross-domain cluster system in urban interference environment
By constructing a distributed elastic observer and an online dual-integrator reinforcement learning method, the network attack problem of cross-domain air-to-ground unmanned cluster systems in urban interference environments was solved, the optimal safe formation control was achieved under unknown dynamic models, and the formation tracking performance and collaborative operation efficiency were improved.
Patent Information
- Application Number
- CN202510731851.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies for the formation control of cross-domain air-to-ground unmanned swarm systems in urban interference environments are vulnerable to external network attacks, resulting in degradation or even failure of formation tracking performance. Furthermore, control strategies that rely on precise dynamic models make it difficult to achieve global optimal control and reduce fuel consumption.
A distributed elastic observer is constructed, combined with the dynamic model and augmented system of the air-ground unmanned swarm system, and the optimal control gain matrix is determined using the online double integrator reinforcement learning method. The optimal formation controller is constructed to eliminate the continuous excitation assumption, reduce storage requirements, and achieve safe formation control.
Under distributed denial of service network attacks, it ensures formation tracking performance, reduces memory consumption, achieves optimal formation control, and improves collaborative work effectiveness and efficiency.
Smart Images

Figure CN120686895A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of unmanned cluster systems, and in particular to a distributed security formation control method for a cross-domain cluster system in an urban interference environment. Background Art
[0002] As the demands for collaborative missions in unmanned swarm systems become increasingly complex, cross-domain air-ground unmanned swarm systems, comprised of multiple drones and multiple unmanned vehicles, can overcome the functional limitations of single-configuration unmanned systems and provide effective solutions for complex, cross-regional, and cross-dimensional collaborative missions. Capable of carrying multiple payloads, collecting data, and allocating tasks, these systems demonstrate significant application potential in areas such as tracking and pursuit, search and rescue, resource exploration, and coordinated material transportation. For example, in urban environments, drones can quickly transport small items in the air and provide good visual guidance for ground-based unmanned vehicles, which in turn carry large or heavy items. Through real-time communication and dynamic path planning, both drones and ground-based unmanned vehicles can avoid traffic jams and ensure the rapid and precise delivery of materials to their destinations.
[0003] Collaborative formation control is a key issue in the coordinated control of cross-domain air-ground unmanned swarm systems. Its primary goal is to enable different unmanned systems in a swarm to achieve a mission-specific formation configuration through local information exchange between individuals. However, the effective execution of collaborative formation control for air-ground unmanned swarm systems relies on open, shared communication networks. The open nature of these networks makes cross-domain air-ground swarm formation systems extremely vulnerable to external network attacks, which can degrade formation tracking performance and even lead to mission failure.
[0004] Therefore existing technology still needs to be improved and improved. Summary of the Invention
[0005] The technical problem to be solved by this application is to provide a distributed security formation control method for a cross-domain cluster system in an urban interference environment in response to the shortcomings of the existing technology.
[0006] In order to solve the above technical problems, the first aspect of the present application provides a distributed safety formation control method for a cross-domain cluster system in an urban interference environment, which is applied to an air-ground unmanned cluster system, wherein the air-ground unmanned cluster system includes a virtual leader and multiple unmanned systems following the virtual leader, and the multiple unmanned systems include drones and unmanned vehicles; the distributed safety formation control method for a cross-domain cluster system in an urban interference environment specifically includes:
[0007] Constructing a distributed elastic observer, wherein the distributed elastic observer is used to estimate the virtual machine leader state to provide a formation output reference for safe formation control when the air-ground unmanned cluster system is attacked by a distributed denial of service network;
[0008] Constructing an augmented system of the air-ground unmanned cluster system according to the distributed elastic observer and the dynamic model of the air-ground unmanned cluster system, and constructing a formation controller of the air-ground unmanned cluster system according to the augmented system;
[0009] Collecting system data of the air-to-ground unmanned swarm system subject to control inputs including initial excitations under a preset initial excitation condition, and determining an optimal control gain matrix based on the system data using an online double-integrator reinforcement learning method to obtain an optimal formation controller;
[0010] The optimal formation controller is used to perform formation control on the air-to-ground unmanned cluster system.
[0011] The distributed safety formation control method for a cross-domain cluster system in an urban interference environment, wherein the parameter matrix of the dynamic model of the air-ground unmanned cluster system includes system parameters that are easily affected during the movement of the air-ground unmanned cluster system, and the system parameters include the aerodynamic drag coefficient of the UAV, the mass of the UAV, the moment of inertia of the UAV, the control gain of the UAV autopilot and the mass of the unmanned vehicle.
[0012] The distributed safety formation control method for a cross-domain cluster system in an urban interference environment, wherein the construction process of the distributed elastic observer specifically includes:
[0013] Determining an observation control gain constraint value based on the attack intensity of the denial of service attack on the communication link;
[0014] Constructing a linear matrix inequality about the observation control gain based on the observation control gain constraint value, and solving the linear matrix inequality to obtain the observation control gain;
[0015] A distributed elastic observer is constructed based on the communication topology neighbor weights between the unmanned systems, the state data of the virtual leader and the observed state data of the unmanned systems and the observation control gain.
[0016] The distributed safety formation control method for a cross-domain cluster system in an urban interference environment, wherein the distributed elastic observer is specifically:
[0017]
[0018] in, represents the derivative of the observed state data, represents the first unmanned system in the air-ground unmanned cluster system The observed state data of each subsystem, represents the first unmanned system in the air-ground unmanned cluster system The observed output data of each subsystem, represents the positive gain constant, represents the control gain of the distributed elastic observer, represents the neighborhood leader-follower error.
[0019] The distributed safety formation control method for a cross-domain cluster system in an urban interference environment, wherein the initial excitation condition is:
[0020]
[0021]
[0022] Among them, vec(·) represents the vectorized operation of the matrix, vec -1 (·) represents the inverse operation of vec(·), vech(·) represents the semi-vectorized operation of the matrix without folding, ΔT represents the time window length, δ i,κ Indicates the degree of motivation, z i,κ (t) represents the augmented vector, u i,κ (t) represents the control input, m i,κ (t), h i,κ (t), and η i,κ (t) represents intermediate variables.
[0023] The distributed safety formation control method for a cross-domain cluster system in an urban interference environment, wherein determining the optimal control gain matrix based on the system data using an online double integrator reinforcement learning method to obtain an optimal formation controller specifically includes:
[0024] Solving a preset data equation according to the system data to obtain an iterative learning matrix;
[0025] Using the iterative learning matrix to update the initial gain matrix, and determining whether the updated initial gain matrix meets the preset conditions;
[0026] When the updated initial gain matrix meets the preset conditions, the updated initial gain matrix is used as the optimal control gain matrix to obtain the optimal formation controller;
[0027] When the updated initial gain matrix does not meet the preset conditions, the step of solving the preset data equation according to the system data to obtain the iterative learning matrix is re-executed.
[0028] The distributed security formation control method for a cross-domain cluster system in an urban interference environment, wherein the optimal formation controller is specifically:
[0029]
[0030] Among them, R i,κ Both represent weight matrices, represents the optimal formation controller, represents the optimal control gain matrix, represents the symmetric matrix obtained by solving the discounted factor algebraic Riccati equation, Represents model information.
[0031] A second aspect of the present application provides a distributed safety formation control device for a cross-domain cluster system in an urban interference environment, which is applied to an air-ground unmanned cluster system, wherein the air-ground unmanned cluster system includes a virtual leader and multiple unmanned systems, wherein the multiple unmanned systems include drones and / or unmanned vehicles; the distributed safety formation control device for a cross-domain cluster system in an urban interference environment includes:
[0032] A construction module is used to construct a distributed elastic observer, construct an augmented system of the air-ground unmanned cluster system based on the distributed elastic observer and the dynamic model of the air-ground unmanned cluster system, and construct a formation controller of the air-ground unmanned cluster system based on the augmented system; collect system data of the air-ground unmanned cluster system subjected to control inputs including initial excitations with a preset initial excitation condition as a constraint, and determine the optimal control gain matrix based on the system data using an online double integrator reinforcement learning method to obtain an optimal formation controller, wherein the distributed elastic observer is used to estimate the virtual machine leader state when the air-ground unmanned cluster system is attacked by a distributed denial of service network to provide a formation output reference for safe formation control, and the initial excitation satisfies the preset initial excitation condition;
[0033] A control module is used to use the optimal formation controller to perform formation control on the air-to-ground unmanned cluster system.
[0034] A third aspect of the present application provides a computer-readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the distributed security formation control method of a cross-domain cluster system in an urban interference environment as described above.
[0035] A fourth aspect of the present application provides a terminal device, comprising: a processor and a memory;
[0036] The memory stores a computer-readable program executable by the processor;
[0037] When the processor executes the computer-readable program, the steps in the distributed security formation control method of the cross-domain cluster system in an urban interference environment as described above are implemented.
[0038] Beneficial effects:
[0039] (1) This application adopts a data-driven online policy reinforcement learning algorithm. Without the need to obtain an accurate system dynamics model, it only uses the system data during the operation of the air-ground cluster to learn the optimal elastic formation control strategy to ensure the safe formation of the air-ground unmanned cluster system.
[0040] (2) This application constructs a distributed elastic observer. When the cross-domain air-ground unmanned cluster system is attacked by a distributed denial of service network, the distributed elastic observer can estimate the state of the virtual leader to provide a trajectory reference for the formation tracking control of the air-ground unmanned cluster system, thereby ensuring the formation tracking performance.
[0041] (3) This application constructs the initial incentive conditions during online reinforcement learning, eliminates the restrictive assumption of continuous incentives in reinforcement learning, and constructs the data equation through the double integration technology. By solving the data equation, the iterative learning matrix is obtained, bypassing the calculation of the state derivatives of the air-ground cluster system, eliminating the algorithm's need for finite window integration, and reducing memory consumption; it avoids the need for reinforcement learning to use full rank conditions to intelligently collect sufficiently rich data. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0043] Figure 1 This is a flowchart of a distributed security formation control method for a cross-domain cluster system in an urban interference environment provided by an embodiment of the present application.
[0044] Figure 2 This is a principle block diagram of a distributed security formation control device for a cross-domain cluster system in an urban interference environment provided by an embodiment of the present application.
[0045] Figure 3 This is a block diagram of the principles of the terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0046] The embodiments of this application provide a distributed security formation control method for a cross-domain cluster system in an urban interference environment. To make the objectives, technical solutions, and effects of this application more clear and explicit, the application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only intended to explain this application and are not intended to limit this application.
[0047] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.
[0048] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0049] It should be understood that the sequence numbers and sizes of the steps in this embodiment do not imply the order of execution. The order of execution of each process is determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of this application.
[0050] Research has found that as the complexity of collaborative missions required by unmanned swarm systems continues to increase, cross-domain air-ground unmanned swarm systems, comprised of multiple drones and multiple unmanned vehicles, can overcome the functional limitations of single-configuration unmanned systems and provide an effective solution for complex cross-regional and cross-dimensional collaborative missions. Capable of carrying multiple payloads, collecting data, and allocating tasks, cross-domain air-ground unmanned swarm systems demonstrate significant application potential in areas such as tracking and pursuit, search and rescue, resource exploration, and coordinated material transportation. For example, in urban environments, drones can quickly transport small items through the air and provide good visual guidance for ground-based unmanned vehicles, which in turn carry large or heavy items. Through real-time communication and dynamic path planning, the two can avoid traffic jams and ensure that materials reach their destination quickly and accurately.
[0051] Collaborative formation control is a key topic in the coordinated control of cross-domain air-ground unmanned swarm systems. Its primary goal is to enable the different unmanned systems in a swarm to achieve a mission-specific formation configuration through local information exchange between individuals. However, the effective execution of collaborative formation control for air-ground unmanned swarm systems relies on open, shared communication networks. This inherent openness makes cross-domain air-ground swarm formation systems extremely vulnerable to external network attacks, which can degrade formation tracking performance and even lead to mission failure.
[0052] Furthermore, existing air-ground swarm formation control methods rely on accurate dynamic model information of the swarm system and only consider the steady-state performance of the control system while ignoring the transient performance of the system. These control strategies based on prior knowledge of air-ground swarm models are not only difficult to implement in practical applications, but also fail to achieve global optimal control to reduce fuel consumption and improve collaborative work effectiveness and efficiency. To address this issue, current research is collecting system operation data and using data-driven algorithms such as reinforcement learning or adaptive dynamic programming to learn the optimal formation controller using the collected system data, thereby achieving data-driven optimal formation tracking control for cross-domain air-ground unmanned swarms.
[0053] However, these data-driven reinforcement learning-based optimal formation control schemes for air-ground swarms fail to account for the vulnerability of air-ground swarm systems to distributed denial-of-service network attacks. Furthermore, these reinforcement learning methods require the assumption of continuous excitation conditions and intelligent data storage to achieve closed-loop data collection and optimal algorithm convergence. However, continuous excitation conditions are difficult to monitor online and lack practicality for online implementation. Intelligent data storage also results in significant memory consumption and delays in meeting the required full-rank condition.
[0054] In order to solve the above problems, in an embodiment of the present application, a distributed elastic observer is constructed, an augmented system of the air-ground unmanned cluster system is constructed based on the distributed elastic observer and the dynamic model of the air-ground unmanned cluster system, and a formation controller of the air-ground unmanned cluster system is constructed based on the augmented system. The system data of the air-ground unmanned cluster system subjected to the control input including the initial excitation is collected with the preset initial excitation condition as a constraint, and the optimal control gain matrix is determined based on the system data using the online double integrator reinforcement learning method to obtain the optimal formation controller, and the air-ground unmanned cluster system is formation controlled using the optimal formation controller. In an embodiment of the present application, when the cross-domain air-ground unmanned cluster system is attacked by a distributed denial of service network, the state of the virtual leader is estimated by the distributed elastic observer to provide a trajectory reference for the formation tracking control of the air-ground unmanned cluster system, thereby ensuring the formation tracking performance. At the same time, the initial incentive conditions are constructed during online reinforcement learning, eliminating the restrictive assumptions of continuous incentives in reinforcement learning. The data equation is constructed through the double integration technique, and the iterative learning matrix is obtained by solving the data equation, bypassing the calculation of the state derivatives of the air-ground cluster system, eliminating the algorithm's need for finite window integration, and reducing memory consumption; it avoids the need for reinforcement learning to use full rank conditions to intelligently collect sufficiently rich data.
[0055] The application content will be further explained below through description of embodiments in conjunction with the accompanying drawings.
[0056] This embodiment provides a distributed safety formation control method for a cross-domain cluster system in an urban interference environment. The method applies an air-ground unmanned cluster system, wherein the air-ground unmanned cluster system includes a virtual leader and multiple unmanned systems following the virtual leader, each unmanned system is a follower, and the multiple unmanned systems include drones and unmanned vehicles. For example, the air-ground unmanned cluster system includes a virtual leader and N unmanned systems (i.e., N followers), and the N unmanned systems include M drones and NM unmanned vehicles, where N and M are both positive integers and N is greater than M. The unmanned systems communicate with each other through a topological graph, and the communication between any two unmanned systems is directional, that is, the communication topology between the air-ground unmanned cluster systems is a directed graph.
[0057] The system model of the unmanned air-ground swarm system is a globally unified description model that includes the dynamic models of drones, unmanned vehicles, a dynamic model of a virtual leader, and the desired formation. The parameter matrix in this system model includes system parameters that are susceptible to change during the swarm's movement. This allows reinforcement learning to construct a data equation that excludes the system parameter matrix, thereby learning an optimal formation controller based on the system data. This allows the reinforcement learning-driven approach to utilize the system data generated during the swarm's movement to learn a new optimal controller that is appropriate for the current parameter information online, even when the susceptible system parameters change.
[0058] The system parameters in the parameter matrix can include the drone's aerodynamic drag coefficient (affected by wind speed), the drone's mass (affected by load release and fuel consumption), the drone's moment of inertia (affected by aerodynamic forces and dynamic changes in mass), and the unmanned vehicle's mass (affected by load release and fuel consumption). Based on the analysis of these susceptible parameters, their values may change during the movement of the air-to-ground drone swarm system. For model-based controllers, when these parameters change due to the system's movement, they cannot adaptively learn a new controller that is suitable for the current parameters, thereby reducing the controller's control performance.
[0059] The following is a detailed description of the system model of the air-to-ground unmanned cluster system.
[0060] (1) The dynamic model of the UAV is:
[0061]
[0062] in, is the position of drone i, is the posture of UAV i, φ i is the roll angle, θ i is the pitch angle, ψ i is the yaw angle; For p i The second derivative of is Θ i The second derivative of k i,x ,k i,y and k i,z is the aerodynamic drag coefficient of UAV i, m qi is the mass of UAV i, g is the gravitational acceleration of UAV i; b qi,φ =(l / J qi,φ ),b qi,θ =(l / J qi,θ ) and b qi,ψ =(1 / J qi,ψ) are intermediate quantities, l is the arm length of UAV i, J qi,φ ,J qi,θ and J qi,ψ is the moment of inertia of UAV i; and is the control gain of the UAV i autopilot, and are the desired translational velocity and yaw rate.
[0063] make:
[0064]
[0065] Among them, u i,1 ,u i,2 ,u i,3 and u i,4 is the control input of UAV i, x ai is the status data of drone i.
[0066] Then, the dynamic model of the UAV can be expressed as a linear model, which is:
[0067]
[0068] A ai =diag{A i,1 ,A i,2 ,A i,3 ,A i,4}
[0069] B ai =diag{B i,1 ,B i,2 ,B i,3 ,B i,4}
[0070] u ai =[u i,1 ,u i,2 ,u i,3 ,u i,4 ] T
[0071] Among them, A i,1 ,B i,1 ,A i,2 ,B i,2 ,A i,3 ,B i,3 ,A i,4 ,B i,4 Corresponding to drone i and p respectively i,x ,p i,y ,p i,z 3D position and ψ iThe subsystem parameter matrix related to the yaw angle contains the system parameters that are easily affected by the movement of UAV i.
[0072] Among them, the parameter matrices of each subsystem are:
[0073]
[0074] (2) The dynamic model of the unmanned vehicle is:
[0075]
[0076] in, is the position of the unmanned vehicle i in the two-dimensional plane, is the speed of the unmanned vehicle i in the two-dimensional plane, ψ gi is the yaw angle of unmanned vehicle i, ω gi is the yaw rate of the unmanned vehicle i, F gi,x and F gi,y are the total forces acting on the unmanned vehicle i in the two-dimensional position direction, C A,i is the aerodynamic drag coefficient of unmanned vehicle i, C f,i is the rolling friction coefficient of unmanned vehicle i, m gi The mass of the unmanned vehicle i.
[0077] Assume that the unmanned vehicle i moves at a small angle and ω gi ≈0, and using feedback linearization technology, the control input and state data of the unmanned vehicle i can be expressed as:
[0078]
[0079] u i,6 =F gi,y ,
[0080] x gi =[p i,x ,v i,x ,p i,y ,v i,y ] T ,
[0081] Among them, u i,5 ,u i,6 represents the control input of autonomous vehicle i, x gi Represents the status data of the unmanned vehicle i.
[0082] Then, the dynamic model of the unmanned vehicle can be expressed as a linear model, which is:
[0083]
[0084] A gi=diag{A i,5 ,A i,6},
[0085] B gi =diag{B i,5 ,B i,6},
[0086] Among them, A gi , B gi Both represent parameter matrices, u gi =[u i,5 ,u i,6 ] T is the control input, A i,5 ,B i,5 and A i,6 ,B i,6 Corresponding to the p of unmanned vehicle i i,x and p i,y Two-dimensional position-dependent subsystem parameter matrix.
[0087] The parameter matrix of each subsystem is expressed as:
[0088]
[0089] B i,5 =B i,6 =[0,1 / m gi ] T .
[0090] (3) The dynamic model and expected formation of the virtual leader in the air-ground unmanned swarm system are:
[0091] y0=C0x0,
[0092]
[0093] Among them, A0 and C0 are given constant matrices, x0=[x 0,x ,x 0,y ,x 0,z ,x 0,ψ ] T is the position (angle) of the virtual leader, y0=[y 0,x ,y 0,y ,y 0,z ,y 0,ψ ] T is the speed (angular velocity) of the virtual leader, and are the position deviations of the i-th follower and virtual leader, is the position deviation between the jth follower and the virtual leader, is the expected 3D position deviation of the ith follower and the jth follower based on the virtual leader.
[0094] (4) Based on the dynamic model of the UAV, the dynamic model of the unmanned vehicle, the dynamic model of the virtual leader, and the expected formation, a global unified description model (i.e., system model) is constructed. The global unified description model is:
[0095] Followers: y i,κ =C i,κ χ i,κ ,
[0096] Leader:
[0097] Formation:
[0098] in, and They represent the state data, control input and output data of the κth subsystem of the i-th unmanned system in the air-to-ground UAV cluster system, n, f and p represent the dimensions of the state data, control input and output data of the subsystem respectively, i=1,2,...,N,κ=1,2,...,6,A i,κ ,B i,κ and C i,κ is the dynamic matrix matching the UAV and the UGV. If κ∈{1,2}, C i,κ =[1 0 0 0], otherwise C i,κ =[1 0]. and Represents the virtual leader Status data and output data of each subsystem, and Represents the virtual leader The system matrix and output matrix of each subsystem, The i-th and j-th followers and the virtual leader's The position deviation of the dimension,
[0099] like Figure 1 As shown, the distributed security formation control method for a cross-domain cluster system in an urban interference environment provided by an embodiment of the present application specifically includes:
[0100] S10. Build a distributed elastic observer.
[0101] Specifically, a distributed elastic observer is used to observe followers. This allows the observer to estimate the virtual machine leader state and formation information to determine a formation tracking trajectory reference when the air-ground unmanned swarm system is attacked by a distributed denial of service (DDoS) network. This enables the observer to safely control distributed multi-channel asynchronous DDoS network attacks. The observer is constructed based on the attack intensity of the air-ground unmanned swarm system, the neighbor weights of the communication topology between the unmanned systems, and the neighborhood leader-follower error.
[0102] Exemplarily, the construction process of the distributed elastic observer specifically includes:
[0103] Determining an observation control gain constraint value based on the attack intensity of the denial of service attack on the communication link;
[0104] Constructing a linear matrix inequality about the observation control gain based on the observation control gain constraint value, and solving the linear matrix inequality to obtain the observation control gain;
[0105] A distributed elastic observer is constructed based on the communication topology neighbor weights between the unmanned systems, the state data of the virtual leader and the observed state data of the unmanned systems and the observation control gain.
[0106] Specifically, when the air-ground unmanned cluster system is attacked by a distributed denial of service network, the communication links between the unmanned systems will be attacked. The attack intensity of each communication link can be obtained, and the observation control gain constraint value can be determined based on the attack intensity. The observation control gain constraint value is obtained by solving an inequality group related to the attack intensity, wherein the inequality group can be expressed as:
[0107]
[0108] in, represents the equivalent decay rate of the active state of (i, j), represents the equivalent decay rate of the inactive state of (i, j), (i, j) is a communication link in the communication network topology established by the air-ground unmanned cluster system, E is the set of all communication links, that is, the directed graph communication network topology between the air-ground unmanned cluster systems, Γ is the distributed denial of service attack model, α Γ Mean observed control gain constraint value, ν i,j is the attack intensity of the communication link (i, j)∈E under the denial of service attack.
[0109] In addition, in the embodiment of the present application, in order to more accurately construct a distributed elastic observer, the distributed denial of service network attack suffered by the air-ground unmanned cluster system is constrained, that is, the attack intensity of the communication link (i, j)∈E suffered by the denial of service attack is constrained. Among them, the constraint conditions corresponding to the distributed denial of service network attack suffered by the air-ground unmanned cluster system can be: Among them, D i,j (t0,t) is the union of the time intervals during which the communication link (i,j)∈E is attacked by the denial of service within the time interval (t0,t), t0 is the initial time of the system operation, is the fixed attack effect of the communication link (i, j)∈E under the distributed denial of service attack. Of course, in practical applications, the constraints corresponding to the distributed denial of service network attack on the air-ground unmanned cluster system can also adopt other forms, which will not be explained here one by one.
[0110] Furthermore, after determining the observed control gain constraint value, a linear matrix inequality about the observed control gain can be constructed based on the observed control gain constraint value, and the observed control gain can be obtained by solving the linear matrix inequality, wherein the inequality can be expressed as:
[0111]
[0112] in, is the identity matrix, is the observed control gain, To adjust the parameters, is the positive gain constant, is the system matrix, for The transposed matrix of .
[0113] Furthermore, after obtaining the observation control gain, a distributed elastic observer is constructed based on the communication topology neighbor weights between the unmanned systems, the state data of the virtual leader, the observation state data of the unmanned system, and the observation control gain. Specifically, the distributed elastic observer is:
[0114]
[0115]
[0116] in, represents the derivative of the observed state data, Represents the i-th unmanned system in the air-ground unmanned cluster system The observed state data of each subsystem, represents the number of the i-th unmanned system in the air-ground unmanned cluster system The observed output data of each subsystem, represents the control gain of the distributed elastic observer, represents the neighborhood leader-follower error, w ij is the communication topology adjacency weight between the i-th and j-th unmanned systems, b i is the connection weight between the virtual leader and the unmanned system i, and All are intermediate amounts, is the impact of the distributed denial of service attack model on the communication topology adjacency weight, is the impact of the distributed denial of service attack model on the communication topology adjacency weight.
[0117] Furthermore, if there is communication between (i, j), then w ij Greater than zero, otherwise w ij = zero; if there is communication between the virtual leader and the unmanned system i, then b i Greater than zero, otherwise b i is equal to zero; for any distributed denial of service attack model Γ∈E, if (i,j) is attacked by the network, then Greater than zero, otherwise =zero; Since the virtual leader is only used to provide the desired reference tracking trajectory, the communication link between the virtual leader and the unmanned system is vulnerable to cyber attacks. Equal to zero.
[0118] S20. Construct an augmented system of the air-ground unmanned cluster system based on the distributed elastic observer and the dynamic model of the air-ground unmanned cluster system, and construct a formation controller of the air-ground unmanned cluster system based on the augmented system.
[0119] Specifically, after obtaining the distributed elastic observer and the dynamic model of the air-ground unmanned cluster system, the state data χ of the κth subsystem of the i-th unmanned system in the air-ground unmanned cluster system can be combined i,κ and the state data of the distributed elastic observer that matches it Construct augmented vector, augmented vector z i,κ It can be expressed as
[0120] After the augmented vector is constructed, the augmented system of the augmented subsystem of the air-to-ground unmanned cluster system is constructed based on the augmented vector and the formation tracking error. The dynamic model of the augmented system can be expressed as:
[0121]
[0122]
[0123] Among them, e i,κ is the formation tracking error, is the system matrix of the augmented system, is the input matrix of the augmented system, is the output matrix of the augmented system, n z is the dimension of the state data of the augmented subsystem, m z is the dimension of the control input of the augmented subsystem, u i,κ is the control input of the augmented subsystem, and are all control variables. If κ∈{1,2}, otherwise When the distributed elastic observer converges, will converge asymptotically to zero; is the control gain matrix to be learned, y i,κ Output data for the κth subsystem of the i-th unmanned system in the air-to-ground UAV cluster system. is the number of the i-th unmanned system in the air-ground unmanned cluster system Observation output data of each subsystem.
[0124] S30. Collect system data of the air-to-ground unmanned cluster system subject to control inputs including initial excitations with a preset initial excitation condition as a constraint, and determine the optimal control gain matrix based on the system data using an online dual-integrator reinforcement learning method to obtain an optimal formation controller.
[0125] Specifically, the initial excitation condition is used to converge the optimal controller parameters. The initial excitation condition is constructed using the state data and control input of the air-ground unmanned swarm system during its movement. It is used as the initial excitation constraint contained in the control input of the air-ground unmanned swarm system to eliminate the restrictive assumption of continuous excitation in reinforcement learning. The initial excitation condition can be:
[0126]
[0127] Among them, vec(·) represents the vectorized operation of the matrix, vec -1 (·) represents the inverse operation of vec(·), vech(·) represents the semi-vectorized operation of the matrix without folding, ΔT represents the time window length, δ i,κ Indicates the degree of motivation, z i,κ (t) represents the augmented vector, u i,κ (t) represents input data, m i,κ (t), h i,κ (t), and η i,κ (t) represents intermediate variables.
[0128] Furthermore, when collecting system data, a preset initial excitation determination condition is used as the termination condition. That is, when applying control inputs, including initial excitations, to the unmanned air-ground swarm system and collecting system data during its movement, the initial excitation determination condition is used to determine whether sufficient system data has been received. Specifically, if the initial excitation determination condition is met, sufficient system data has been collected; if the initial excitation determination condition is not met, further system data collection is necessary. System data includes control inputs, state data, and observed state data from the performance observer.
[0129] Exemplarily, the initial excitation determination condition may be:
[0130] det(Π i,κ (t0+ΔT))>0,
[0131]
[0132] Among them, Π i,κ (t0+ΔT)) is an intermediate variable.
[0133] Based on this, the process of collecting system data of the air-ground unmanned cluster system after the control input including the initial excitation is applied can be: select the initial gain matrix And set the number of iterations k = 0, and then input the control containing the initial excitation Acting on air-to-ground unmanned cluster system, where u noise is the detection noise. When the initial excitation judgment condition det(Π i,κ When (t0+ΔT)>0 is satisfied, it indicates that sufficient data information has been collected. In addition, it should be noted that the initial excitation determination condition may also adopt other conditions, such as the amount of system data collected reaching a preset amount.
[0134] Furthermore, after collecting enough system data, the initial excitation is removed and the data equation constructed by the quadratic integration technique is used to learn the optimal control gain matrix online through iterative calculation. To obtain the optimal formation controller. That is, after obtaining the system data, the optimal control gain matrix is determined based on the system data using the online dual integrator reinforcement learning method to obtain the optimal formation controller. The method of determining the optimal control gain matrix based on the system data using the online dual integrator reinforcement learning method to obtain the optimal formation controller specifically includes:
[0135] Solving a preset data equation according to the system data to obtain an iterative learning matrix;
[0136] Using the iterative learning matrix to update the initial gain matrix, and determining whether the updated initial gain matrix meets the preset conditions;
[0137] When the updated initial gain matrix meets the preset conditions, the updated initial gain matrix is used as the optimal control gain matrix to obtain the optimal formation controller;
[0138] When the updated initial gain matrix does not meet the preset conditions, the step of solving the preset data equation according to the system data to obtain the iterative learning matrix is re-executed.
[0139] Specifically, the preset data equation is constructed by removing the initial excitation and using the quadratic integration technique, wherein the preset data equation can be expressed as:
[0140]
[0141] in, is the iterative learning matrix, and are P at the kth iteration respectively i,κ Matrix and K i,κ matrix, and are the data matrices constructed by the quadratic integration technique, Z i,κ (t), Z i,κ (t0), r i,κ 、 as well as are intermediate variables. is the identity matrix, R i,κ is the weight matrix, is the weight matrix.
[0142] Iteratively solve the preset data equation according to the system data to obtain the iterative learning matrix Then use Update the control gain matrix And judge whether the updated initial gain matrix meets the preset conditions, where the preset conditions are ∈ is a small constant value greater than zero. If If it holds, the iteration stops, and the control gain matrix is The optimal formation controller is That is, to obtain the optimal control gain matrix To obtain the optimal formation controller On the contrary, if If it does not hold, then k=k+1, and continue to iteratively solve the preset data equation according to the system data to obtain the iterative learning matrix Until The condition is established or the number of iterative solutions reaches the preset threshold.
[0143] Based on this, the optimal formation controller can be expressed as:
[0144]
[0145] in, and R i,κ is the weight matrix, is the intermediate variable, is the optimal formation controller, represents the optimal control gain matrix, is the symmetric matrix obtained by solving the discounted factor algebraic Riccati equation, which is α i,κ is a discount factor used to ensure the convergence of the optimal formation controller.
[0146] S40: Utilize the optimal formation controller to perform formation control on the air-to-ground unmanned cluster system.
[0147] Specifically, after obtaining the optimal formation controller, the optimal formation controller is used to perform formation control on the air-ground unmanned cluster system to improve the ability of the air-ground unmanned cluster system to cope with distributed denial of service network attacks. It realizes the optimal safe formation tracking control of the cross-domain air-ground unmanned cluster system in the case of unknown system dynamic model and distributed denial of service attacks, avoiding the problem of poor formation tracking performance of the air-ground unmanned cluster system due to attacks on the service network, and even the failure of the formation mission.
[0148] In summary, this embodiment provides a distributed safe formation control method for a cross-domain cluster system in an urban interference environment. The method establishes a dynamic model for a cross-domain air-ground unmanned cluster system; constructs a distributed elastic observer to estimate the state of a virtual leader, providing a formation output reference for collaborative formation tracking control; combines the state data of the distributed elastic observer with the state data of the air-ground unmanned cluster system to construct an augmented system for the air-ground unmanned cluster system, and uses a feedforward-feedback control method to construct a control gain matrix. The initial excitation is applied to the augmented system of the air-ground unmanned cluster system. When the running time meets the initial excitation condition, the state data of the air-ground unmanned cluster system is collected, and the online dual-integrator reinforcement learning method is used to learn the optimal formation tracking control to achieve the optimal formation tracking control of the air-ground unmanned cluster system. The present invention can achieve the optimal safe formation tracking control of a cross-domain air-ground unmanned cluster system when the system dynamic model is unknown and a distributed denial of service attack occurs.
[0149] Based on the above-mentioned distributed safety formation control method for a cross-domain cluster system in an urban interference environment, this embodiment provides a distributed safety formation control device for a cross-domain cluster system in an urban interference environment, which is applied to an air-ground unmanned cluster system, wherein the air-ground unmanned cluster system includes a virtual leader and multiple unmanned systems, and the multiple unmanned systems include drones and / or unmanned vehicles; Figure 2 As shown, the distributed security formation control device of the cross-domain cluster system in the urban interference environment includes:
[0150] A construction module 100 is used to construct a distributed elastic observer, construct an augmented system of the air-ground unmanned cluster system based on the distributed elastic observer and the dynamic model of the air-ground unmanned cluster system, and construct a formation controller of the air-ground unmanned cluster system based on the augmented system; collect system data of the air-ground unmanned cluster system subjected to control inputs including initial excitations with a preset initial excitation condition as a constraint, and determine the optimal control gain matrix based on the system data using an online double integrator reinforcement learning method to obtain an optimal formation controller, wherein the distributed elastic observer is used to estimate the virtual machine leader state when the air-ground unmanned cluster system is attacked by a distributed denial of service network to provide a formation output reference for safe formation control, and the initial excitation satisfies the preset initial excitation condition;
[0151] The control module 200 is used to perform formation control on the air-to-ground unmanned cluster system using the optimal formation controller.
[0152] Based on the above-mentioned distributed security formation control method for a cross-domain cluster system under an urban interference environment, this embodiment provides a computer-readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the distributed security formation control method for a cross-domain cluster system under an urban interference environment as described in the above-mentioned embodiment.
[0153] Based on the above-mentioned distributed security formation control method for cross-domain cluster system in urban interference environment, this application also provides a terminal device, such as Figure 3 As shown, it includes at least one processor 20; a display screen 21; and a memory 22. It may also include a communications interface 23 and a bus 24. The processor 20, display screen 21, memory 22, and communications interface 23 can communicate with each other via bus 24. The display screen 21 is configured to display a preset user guidance interface in the initial setup mode. The communications interface 23 can transmit information. The processor 20 can call the logic instructions in the memory 22 to execute the method in the above embodiment.
[0154] In addition, the logic instructions in the memory 22 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.
[0155] The memory 22, as a computer-readable storage medium, can be configured to store software programs or computer-executable programs, such as program instructions or modules corresponding to the methods in the embodiments of the present disclosure. The processor 20 executes the software programs, instructions, or modules stored in the memory 22 to perform functional applications and data processing, thereby implementing the methods in the above embodiments.
[0156] The memory 22 may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data created based on the use of the terminal device. In addition, the memory 22 may include high-speed random access memory and non-volatile memory. For example, various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, may also be transient storage media.
[0157] In addition, the specific process of loading and executing the multiple instructions in the storage medium and the processor in the terminal device has been described in detail in the above method and will not be described here one by one.
[0158] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A distributed security formation control method for a cross-domain cluster system in an urban interference environment, characterized in that: Applicable to an air-to-ground unmanned swarm system, the air-to-ground unmanned swarm system comprising a virtual leader and multiple unmanned systems following the virtual leader, the multiple unmanned systems comprising drones and unmanned vehicles; The distributed security formation control method for a cross-domain cluster system in an urban interference environment specifically includes: Constructing a distributed elastic observer, wherein the distributed elastic observer is used to estimate the virtual machine leader state to provide a formation output reference for safe formation control when the air-ground unmanned cluster system is attacked by a distributed denial of service network; Constructing an augmented system of the air-ground unmanned cluster system according to the distributed elastic observer and the dynamic model of the air-ground unmanned cluster system, and constructing a formation controller of the air-ground unmanned cluster system according to the augmented system; Collecting system data of the air-to-ground unmanned swarm system subject to control inputs including initial excitations under a preset initial excitation condition, and determining an optimal control gain matrix based on the system data using an online double-integrator reinforcement learning method to obtain an optimal formation controller; The optimal formation controller is used to perform formation control on the air-to-ground unmanned cluster system.
2. The distributed security formation control method for a cross-domain cluster system in an urban interference environment according to claim 1 is characterized in that: The parameter matrix of the dynamic model of the air-to-ground unmanned cluster system includes system parameters that are easily affected during the movement of the air-to-ground unmanned cluster system. The system parameters include the aerodynamic drag coefficient of the UAV, the mass of the UAV, the moment of inertia of the UAV, the control gain of the UAV autopilot and the mass of the unmanned vehicle.
3. The distributed security formation control method for a cross-domain cluster system in an urban interference environment according to claim 1 is characterized in that: The construction process of the distributed elastic observer specifically includes: Determining an observation control gain constraint value based on the attack intensity of the denial of service attack on the communication link; Constructing a linear matrix inequality about the observation control gain based on the observation control gain constraint value, and solving the linear matrix inequality to obtain the observation control gain; A distributed elastic observer is constructed based on the communication topology neighbor weights between the unmanned systems, the state data of the virtual leader and the observed state data of the unmanned systems and the observation control gain.
4. The distributed security formation control method for a cross-domain cluster system in an urban interference environment according to claim 1 or 3, characterized in that: The distributed elastic observer is specifically: in, represents the derivative of the observed state data, represents the first unmanned system in the air-ground unmanned cluster system The observed state data of each subsystem, represents the first unmanned system in the air-ground unmanned cluster system The observed output data of each subsystem, represents the positive gain constant, represents the control gain of the distributed elastic observer, represents the neighborhood leader-follower error.
5. The distributed security formation control method for a cross-domain cluster system in an urban interference environment according to claim 1 is characterized in that: The initial incentive conditions are: Among them, vec(·) represents the vectorization operation of the matrix, vec -1 (·) represents the inverse operation of vec(·), vech(·) represents the semi-vectorized operation of the matrix without folding, ΔT represents the time window length, δ i,κ Indicates the degree of motivation, z i,κ (t) represents the augmented vector, u i,κ (t) represents the control input, m i,κ (t), h i,κ (t), and η i,κ (t) represents intermediate variables.
6. The distributed security formation control method for a cross-domain cluster system in an urban interference environment according to claim 1 is characterized in that: Determining the optimal control gain matrix based on the system data using an online double integrator reinforcement learning method to obtain an optimal formation controller specifically includes: Solving a preset data equation according to the system data to obtain an iterative learning matrix; Using the iterative learning matrix to update the initial gain matrix, and determining whether the updated initial gain matrix meets the preset conditions; When the updated initial gain matrix meets the preset conditions, the updated initial gain matrix is used as the optimal control gain matrix to obtain the optimal formation controller; When the updated initial gain matrix does not meet the preset conditions, the step of solving the preset data equation according to the system data to obtain the iterative learning matrix is re-executed.
7. The distributed security formation control method for a cross-domain cluster system in an urban interference environment according to claim 1 or 6, characterized in that: The optimal formation controller is specifically: Among them, R i,κ Both represent weight matrices, represents the optimal formation controller, represents the optimal control gain matrix, represents the symmetric matrix obtained by solving the discounted factor algebraic Riccati equation, Represents model information.
8. A distributed security formation control device for a cross-domain cluster system in an urban interference environment, characterized in that: Applicable to an air-to-ground unmanned swarm system, the air-to-ground unmanned swarm system comprising a virtual leader and multiple unmanned systems, the multiple unmanned systems comprising drones and / or unmanned vehicles; The distributed security formation control device for the cross-domain cluster system in the urban interference environment includes: A construction module is used to construct a distributed elastic observer, construct an augmented system of the air-ground unmanned cluster system based on the distributed elastic observer and the dynamic model of the air-ground unmanned cluster system, and construct a formation controller of the air-ground unmanned cluster system based on the augmented system; collect system data of the air-ground unmanned cluster system subjected to control inputs including initial excitations with a preset initial excitation condition as a constraint, and determine the optimal control gain matrix based on the system data using an online double integrator reinforcement learning method to obtain an optimal formation controller, wherein the distributed elastic observer is used to estimate the virtual machine leader state when the air-ground unmanned cluster system is attacked by a distributed denial of service network to provide a formation output reference for safe formation control, and the initial excitation satisfies the preset initial excitation condition; A control module is used to use the optimal formation controller to perform formation control on the air-to-ground unmanned cluster system.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the distributed security formation control method of the cross-domain cluster system in an urban interference environment as described in any one of claims 1-7.
10. A terminal device, characterized in that: include: processor and memory; The memory stores a computer-readable program executable by the processor; When the processor executes the computer-readable program, the steps of the distributed security formation control method of the cross-domain cluster system in an urban interference environment are implemented as described in any one of claims 1-7.
Citation Information
Patent Citations
Unmanned aerial vehicle and unmanned vehicle heterogeneous cluster formation tracking control method under topology switching
CN111665848A
Multi-agent adaptive cooperative fault-tolerant tracking control method
CN117687434A
Distributed formation control method of heterogeneous cluster unmanned system based on reinforcement learning
CN117873122A
Finite time tracking method, system and device for time-varying heterogeneous formation with unknown leader system matrix, and medium
CN119440098A
Self-adaptive elastic formation control method for quad-rotor unmanned aerial vehicles under DoS attack
CN119596969A
Cited By
Safety reinforcement learning cooperation method for unmanned cluster system in urban confrontation environment
CN121386439A
Urban counter-environment collaborative method for security reinforcement learning of unmanned swarm system
CN121386439B
Unmanned aerial vehicle cluster hierarchical disturbance control compliance consistency control method based on data driving
CN121857736A