A reinforcement learning driven cooperative formation control method for air-ground unmanned cluster system

By using a reinforcement learning-driven approach, the dynamic models of UAVs and unmanned vehicles are decoupled, and a distributed predefined time observer and augmented dynamic model are constructed. This solves the problem of optimal formation tracking control of air-to-ground unmanned swarm systems under unknown model conditions, achieving optimal control with low computational resource consumption.

CN120044982BActive Publication Date: 2026-03-31BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing collaborative control methods for air-to-ground unmanned swarm systems struggle to achieve optimal formation tracking control when the system model is unknown. Furthermore, they consume excessive computational resources, cannot adapt to changes in system parameters, and existing reinforcement learning schemes rely on initial stable control strategies.

Method used

By employing a reinforcement learning-driven approach, the dynamic models of UAVs and unmanned vehicles are decoupled through linear feedback. A distributed predefined time observer and an augmented dynamic model are constructed, and the optimal formation controller is learned through a data-driven method, reducing computational burden and avoiding dependence on the system model.

Benefits of technology

Without relying on a system model, optimal formation control of an air-to-ground unmanned swarm system was achieved, reducing computational resource consumption, adapting to changes in system parameters, and quickly providing trajectory references, thereby improving the learning efficiency and accuracy of the controller.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120044982B_ABST
    Figure CN120044982B_ABST
Patent Text Reader

Abstract

This disclosure provides a reinforcement learning-driven cooperative formation control method for an air-to-ground unmanned swarm system. A dynamic model of the air-to-ground unmanned swarm system is established; a distributed predefined time observer is constructed to estimate the virtual leader state, providing a state reference for cooperative formation tracking control; the dynamic model of the augmented air-to-ground unmanned swarm system is reconstructed, and a control gain matrix is ​​then constructed. Piecewise constant initial stimuli are applied to the augmented system, and its early operational state data is collected and stored. Then, based on the early operational data of the air-to-ground unmanned system, an initial stable control strategy is obtained using a data-driven method. Finally, based on an offline policy reinforcement learning algorithm and a data-based initial stable control strategy, the optimal formation tracking controller is learned, achieving optimal formation tracking control of the air-to-ground unmanned system. This invention can achieve optimal time-varying formation tracking control of an air-to-ground unmanned swarm system according to complex task requirements, even when the system model is unknown.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data-driven and air-to-ground unmanned swarm systems technology, specifically to a reinforcement learning-driven collaborative formation control method for air-to-ground unmanned swarm systems. Background Technology

[0002] In recent years, due to its characteristics of distribution, robustness, fault tolerance, and strong adaptability, distributed cooperative control of multiple unmanned systems has been widely applied in many fields such as intelligent transportation, multi-UAV and multi-unmanned vehicle cooperative systems, and microgrids. Among them, formation control is an important topic in the cooperative control of unmanned systems. Its main goal is to design a distributed control strategy that utilizes local communication between the various unmanned systems to ensure that the entire unmanned system maintains the corresponding relative positions according to the set formation.

[0003] However, existing research on formation control largely focuses on homogeneous multi-UAV or multi-RV systems. When performing increasingly complex collaborative tasks such as autonomous surveillance, cooperative target search, and cargo transport, a single homogeneous UAV system cannot effectively accomplish these tasks. Heterogeneous air-to-ground UAV swarm systems combine the advantages of UAVs' rapid detection of vast areas and RVs' precise target localization and high payload capacity. Therefore, air-to-ground UAV swarm systems can provide an effective solution for these complex cross-domain and cross-dimensional collaborative tasks.

[0004] Furthermore, in the design of most existing collaborative control schemes for air-to-ground unmanned systems, the realization of collaborative control requires a complete understanding of the unmanned system's model information, which is difficult to achieve in many practical applications. At the same time, existing controller design techniques only consider the system's steady-state performance and not its transient performance, failing to guarantee the optimality of collaborative control.

[0005] Therefore, how to design a controller to ensure optimal coordination of an air-to-ground unmanned swarm system when the system model is unknown is a current hot and difficult issue.

[0006] To address this issue, current research has focused on collecting system data and employing reinforcement learning control strategies to learn the optimal controller from the collected system data, thereby ensuring optimal cooperative tracking control even when the system model is unknown.

[0007] However, this optimal cooperative tracking control scheme requires significant computational resources. When the number of unmanned systems and the state dimension are large, learning the optimal controller incurs expensive computational costs. Furthermore, this reinforcement learning scheme relies on the initial stable control strategy of the system model for learning.

[0008] Therefore, how to propose an optimal formation tracking and control scheme that does not rely on system model information and can guarantee low computational resource consumption when the model of the air-to-ground unmanned swarm system is unknown is an urgent problem to be solved. Summary of the Invention

[0009] This invention provides a reinforcement learning-driven cooperative formation control method for air-to-ground unmanned swarm systems, which is applicable to optimal formation cooperative control under unknown system model conditions and reduces the computational burden on the controller during the learning process.

[0010] To solve the above-mentioned technical problems, the present invention is implemented as follows.

[0011] A reinforcement learning-driven cooperative formation control method for an air-to-ground unmanned swarm system, comprising a virtual leader and followers consisting of M unmanned aerial vehicles (UAVs) and NM unmanned vehicles (UAVs); the method includes the following steps:

[0012] Step 1: Establish the dynamic model of the air-to-ground unmanned swarm system, including the dynamic models of UAVs, unmanned vehicles, the virtual leader, and the desired formation, and convert it into a globally unified description model; wherein, the dynamic models of UAVs and unmanned vehicles are decoupled into linear models including parameter matrices and system states through a linear feedback method; the system parameters that are easily affected during the movement of the air-to-ground unmanned swarm system are designed into the parameter matrix;

[0013] Step 2: Construct a distributed predefined time observer to estimate the virtual leader state and formation information, providing trajectory reference for formation tracking in the air-to-ground unmanned swarm system; the convergence time of the distributed predefined time observer is set to a precise predefined time T independent of the observer parameters. o ;

[0014] Step 3: Using the aforementioned globally unified description model and distributed predefined time observers, reconstruct the augmented dynamics model of the air-to-ground unmanned swarm system, design an optimal formation controller based on the augmented dynamics model, and construct the control gain matrix K. i,κ ;

[0015] Step four: Apply piecewise constant initial excitation to each augmented subsystem of the augmented dynamics model and collect system operation data from the previous operation; based on the system operation data from the previous operation, obtain the initial stable control gain matrix using a data-driven method.

[0016] Step 5: Perform offline policy iteration for reinforcement learning: Construct a data storage matrix using only the integral of the system running data, and further construct a non-singular data matrix according to the data index; use the system running data in the non-singular data matrix to solve the generalized Hillwest-transpose data equation that does not contain a parameter matrix, and learn the optimal control gain matrix through iterative calculation. This leads to the optimal formation controller for the air-to-ground unmanned swarm system.

[0017] Preferably, in step one, the system parameters that are easily affected during the movement of the air-to-ground unmanned swarm system include: the aerodynamic drag coefficient of the UAV, the mass of the UAV, the moment of inertia of the UAV, the control gain of the UAV autopilot, and the mass of the unmanned vehicle.

[0018] Preferably, the characteristic is that, in step one, the dynamic model for establishing the air-to-ground unmanned cluster system is:

[0019] Assuming drones and unmanned vehicles are uniformly ordered as followers, the dynamic model of the i-th follower in the air-to-ground unmanned swarm system is as follows:

[0020] (1) The dynamic model of the UAV is:

[0021]

[0022] Where, p i =[p i,x ,p i,y ,p i,z ] T and Θ i =[φ i ,θ i ,ψ i ] T Let φ represent the three-dimensional position and attitude of the i-th UAV, respectively; i θ is the roll angle. i Let ψ be the pitch angle. i Yaw angle; and p i and Θ i The second derivative; k i,x ,k i,y and k i,z Let m be the aerodynamic drag coefficient of UAV i. qi For drones i The mass of b, g is the acceleration due to gravity; qi,φ =(l / J qi,φ ),b qi,θ =(l / J qi,θ ) and b qi,ψ =(1 / J) qi,ψ); where l is the arm length of the drone; J qi,φ J qi,θ and J qi,ψ It is the moment of inertia; and For the control gain of the drone autopilot; and Given the desired translational velocity and yaw rate; define and As the control input for UAV i, and the state vector of the UAV is defined as... Therefore, the dynamic model of UAV i can be re-expressed as a linear model:

[0023]

[0024] Wherein, the parameter matrix is ​​A ai =diag(A i,1 A i,2 A i,3 A i,4 ),B ai =diag(B i,1 B i,2 B i,3 B i,4 The control input is u. ai =[u i,1 ,u i,2 ,u i,3 ,u i,4 ] T ; Parameter matrix A i,1 B i,1 A i,2 B i,2 A i,3 B i,3 and A i,4 B i,4 p, corresponding to drone i respectively i,x ,p i,y ,p i,z 3D position and ψ i The yaw angle-related subsystem contains system parameters in its parameter matrix that are susceptible to impact during the movement of the air-to-ground unmanned swarm system.

[0025] The parameter matrix is:

[0026]

[0027] (2) Dynamics model of unmanned vehicles

[0028]

[0029] J gi ωgi =τ gi ,

[0030] Where, p i =[p i,x ,p i,y ] T and v i =[v i,x ,v i,y ] T Let ψ represent the position and velocity of the unmanned vehicle i in the two-dimensional plane, respectively. gi and ω gi These are the yaw angle and yaw rate of the unmanned vehicle i, respectively; F gi,x and F gi,y Let C represent the total forces acting on the unmanned vehicle i in terms of position and orientation, respectively. A,i and C f,i These are the aerodynamic drag coefficient and rolling friction coefficient of the unmanned vehicle i, respectively; m gi For the quality of the driverless car i;

[0031] Assume the driverless car moves at a small angle ψ gi ≈0 and ω gi ≈0, and using feedback linearization technology, the control input of the autonomous vehicle is defined as and u i,6 =F gi,y The state vector of the autonomous vehicle is defined as x. gi =[p i,x ,v i,x ,p i,y ,v i,y ] T Then the dynamic model of the autonomous vehicle can be re-expressed as a linear model as follows:

[0032]

[0033] Wherein, the parameter matrix is ​​A gi =diag(A i,5 A i,6 ),B gi =diag(B i,5 B i,6 );

[0034] The control input is u gi =[u i,5 ,u i,6 ] T ;

[0035] (3) Dynamics model of virtual leaders:

[0036]

[0037] Where, p0 = [x 0,x ,x 0,y ,x 0,z ,x 0,ψ ] T v0 = [v 0,x ,v 0,y ,v 0,z ,v 0,ψ ] T and a0 = [a 0,x ,a 0,y ,a 0,z ,a 0,ψ ] T These represent the virtual leader's position, velocity, and acceleration, respectively.

[0038] (4) Desired formation

[0039] H ij =H i0 -H j0 ,

[0040] Among them, H i,0 =[h i0,x ,h i0,y ,h i0,z ] T H i0 and H j0 denoted as the expected state deviations of the i-th and j-th followers and the virtual leader, respectively;

[0041] Next, based on the dynamic models of drones and unmanned vehicles, the dynamic model of the virtual leader, and the desired formation, a globally unified description model is constructed:

[0042]

[0043] Among them, i=1, 2,...,N; κ=1, 2,..., 6; χ i,κ ,u i,κ and y i,κ These represent the status, control inputs, and outputs of each subsystem in the air-to-ground UAV swarm system; A i,κ B i,κ and C i,κ The dynamic matrix is ​​used to describe the model globally and uniformly.

[0044] Preferably, in step two, the constructed distributed predefined time observer is:

[0045]

[0046] in, and These represent the position and velocity states of observer i, respectively. Where m∈(0,1), λ min (·) represents the smallest eigenvalue of the matrix. For the information transfer matrix, T o >0 represents the preset convergence time;

[0047] By constructing the Lyapunov function of the observer and taking its derivative, we can obtain... It satisfies the predefined time lemma form:

[0048]

[0049] Therefore, the constructed observer can converge within a predefined time; k3 and k4 are positive gain constants; w ij b is the adjacency weight between the i-th and j-th followers. It is greater than zero if there is communication between i and j, and equal to zero otherwise. i The connection weight between the virtual leader and followers.

[0050] Preferably, in step three, the augmented dynamic model is constructed as follows:

[0051] The distributed predefined time observer is rewritten in the following form:

[0052]

[0053] in,

[0054]

[0055] in, and Each drone interacts with its neighboring drones for position and speed information. 0i = [1,0]; I2 is a two-dimensional identity matrix;

[0056] definition As the formation tracking error, and combined with the state χ in equations (I) and (II) i,κ and Construct augmented vectors Thus, an augmented subsystem dynamic model of the air-to-ground UAV swarm system is constructed:

[0057]

[0058] in, r = p + 2, where r is the augmented vector X i,κ The dimension of p is the state vector χ after the global unified description of the unmanned system. i,κ Dimensions and ui,κ The control input of the augmented subsystem is s, where s is the control input u of the augmented subsystem. i,κ The dimension; if κ∈{1,2}, otherwise

[0059] Preferably, the optimal formation controller is obtained by solving the discount factor algebraic Riccati equation; the optimal formation controller is expressed as:

[0060]

[0061] in, and For optimal formation controller, Represents the optimal control gain matrix; By solving the algebraic Riccati equation of the discount factor We obtain, where α i,κ This is a discount factor used to ensure the convergence of the controller.

[0062] Preferably, in step four, the initial stable control gain matrix Obtained through a data-driven approach:

[0063]

[0064] The data storage matrix is ​​constructed using only the integral of the system's initial operating data. and and These are data storage matrices The pseudoinverse and a set of basis vectors of the null space, for any Its satisfaction in, This is the matrix selected using the pole placement method.

[0065] Preferably, step five is as follows:

[0066] S501. Define the data storage matrix and in, It stores the information about the augmented system during its operation and its state X. i,κ Related system operation data, It stores the data related to the input u during the operation of the augmented system. i,κ Related system operation data, It stores the state differential during the operation of the augmented system. Related system operation data;

[0067]

[0068] Wherein, the time interval [t0, t l ] represents the sampling time of data collection, where lΔT=t l -t0, l≥(r+1)s+r represents the length of the collected data, ΔT>0 represents the sampling step size, t0 is the initial time of data collection, t1 is the time after the system has passed the sampling step size at the initial data collection time t0, and so on, t l The termination time for system data collection;

[0069] These represent the κ-th augmenting subsystem of the i-th follower in the corresponding air-to-ground unmanned swarm system within the subinterval [t0,t1], [t1,t2], ..., [t], respectively. l-1 ,t l System status data collected within the system; These represent the κ-th augmenting subsystem of the i-th follower in the corresponding air-to-ground unmanned swarm system within the subinterval [t0,t1], [t1,t2], ..., [t], respectively. l-1 ,t l The system input data collected within the system; These represent the κ-th augmenting subsystem of the i-th follower in the corresponding air-to-ground unmanned swarm system within the subinterval [t0,t1], [t1,t2], ..., [t], respectively. l-1 ,t l The system state collected internally The integral data, among which express In [t] l-1 ,t l Integrals within ];

[0070] S502. From respectively and Select a set of data with indices σ = {1, 2, ..., r + s} to construct a data storage matrix. Make It is a non-singular matrix;

[0071] S503. Strategy Evaluation: Relying on system operation data in the data storage unit, the following data-driven generalized Hill West transpose matrix equation is solved iteratively:

[0072]

[0073] in, The superscript σ indicates a data matrix reconstructed by index; Each iteration updates the data, where the superscripts k and k+1 represent the data corresponding to the k-th and k+1-th iterations, respectively. For matrix The block submatrix is ​​a positive definite matrix. For matrix The block submatrix, This represents the controller gain matrix of the κ augmented subsystems of the i-th follower in the k-th iteration of the spatiotemporal unmanned swarm system.

[0074] S504. Strategy Iteration: Using Update control gain matrix

[0075] S505. Judgment Whether it is true or not, ∈ is a small constant value greater than zero;

[0076] ①If If the condition is met, the iteration stops, and the control gain matrix is ​​then... Formation controller is Proceed to step S506;

[0077] ②If If the condition is not met, increment k by one, and then proceed to S503 and S504 to continue solving.

[0078] S506. Obtain the optimal control gain matrix Simultaneously obtain the optimal formation controller

[0079] Beneficial effects:

[0080] (1) This invention provides a reinforcement learning-driven collaborative formation control method for air-to-ground unmanned swarm systems. This method employs a data-driven, efficient offline policy reinforcement learning algorithm to learn the initial stable control policy and the optimal control policy without needing to acquire the system model. This reduces the computational burden during the optimal controller learning process while ensuring the collaborative formation control of the air-to-ground unmanned swarm system.

[0081] Secondly, the dynamic models of UAVs and unmanned vehicles are decoupled into linear models containing parameter matrices and system states through a linear feedback method, so as to build data-based generalized Sylvester-transpose equations in the later reinforcement learning process.

[0082] Meanwhile, by applying piecewise constant initial excitation to the air-to-ground unmanned system to obtain state data for learning, and by solving the data-based generalized Hillwest-transpose equation to obtain the optimal control strategy, the dependence on full-rank conditions and Kronecker product properties is avoided, reducing the computational burden in the optimal controller learning process.

[0083] (2) This invention utilizes the collected state data and adopts a data-driven approach to obtain the initial stable control strategy required in the strategy iteration algorithm, thereby avoiding the dependence of the initial stable strategy on the dynamic input of the system model.

[0084] (3) The present invention designs a predefined time distributed observer, which can estimate the virtual leader state under complex time-varying formation tasks within a predefined time, thereby quickly providing necessary and accurate trajectory references for the distributed control of air-to-ground unmanned swarm systems.

[0085] (4) In the offline iterative learning process, the present invention constructs a data storage matrix that contains only the system data integral and converts the data storage matrix into a non-singular data matrix, thereby reducing the amount of data and thus reducing the computational burden in the learning process of the optimal controller. Attached Figure Description

[0086] Figure 1 The flowchart illustrates the design of a reinforcement learning-driven collaborative formation control method for an air-to-ground unmanned swarm system, as provided by this invention. Detailed Implementation

[0087] To enable those skilled in the art to better understand the present invention, the technical solution of the present invention will be described in detail below with reference to the accompanying drawings and examples.

[0088] like Figure 1 As shown, the air-to-ground unmanned swarm system described in this invention consists of one virtual leader and N followers; the followers include M drones and NM unmanned vehicles. They are connected through a network communication topology.

[0089] The specific steps of the cooperative formation control method for air-to-ground unmanned swarm systems based on reinforcement learning provided by this invention are as follows:

[0090] Step 1: Establish an air-to-ground unmanned swarm system model, including the dynamics models of drones and unmanned vehicles, the dynamics model of the virtual leader, and the desired formation, and construct a globally unified description model.

[0091] In this step, the dynamic model of the UAV and unmanned vehicle is decoupled into a linear model containing a parameter matrix and system state through a linear feedback method; the system parameters that are susceptible to change during the movement of the air-to-ground unmanned swarm system are incorporated into the parameter matrix. This design allows the construction of the Guangxi Hillwest-transpose data equation without the parameter matrix during reinforcement learning optimal controller learning in step 5. This enables the controller to be learned and solved solely based on system operating data. The advantage of this approach is that even when the susceptible system parameters change, the reinforcement learning-driven method can still utilize the system data generated during the movement of the air-to-ground unmanned swarm system to learn a new optimal controller suitable for the current parameter information.

[0092] The parameters of unmanned vehicles and drones that are susceptible to the influence of the system's movement include: the drone's aerodynamic drag coefficient (affected by wind speed), the drone's mass (affected by payload release and fuel consumption), the drone's moment of inertia (affected by aerodynamic forces and dynamic changes in mass), and the unmanned vehicle's mass (affected by payload release and fuel consumption).

[0093] Based on the analysis of these susceptible parameters, their values ​​may change during the movement of an air-to-ground unmanned aerial vehicle (UAV) swarm system. For model-based controllers, when these parameters change due to system movement, they cannot adaptively learn a new controller suitable for the current parameters, thus reducing controller performance. The dynamic model of the i-th follower in the air-to-ground UAV swarm system constructed in this invention is as follows:

[0094] (1) Dynamics model of unmanned aerial vehicles

[0095]

[0096] in, and Let φ represent the three-dimensional position and attitude of the i-th UAV, respectively. i θ is the roll angle. i Let ψ be the pitch angle. i (yaw angle) and p i and Θ i The second derivative of k. i,x ,k i,y and k i,z Let m be the aerodynamic drag coefficient of UAV i. qi Let b be the mass of the drone i, and g be the acceleration due to gravity. qi,φ =(l / J qi,φ ),b qi,θ =(l / J qi,θ ) and b qi,ψ =(1 / J)qi,ψ ), where l represents unmanned

[0097] The arm length of the machine, J qi,φ J qi,θ and J qi,ψ Let be the moment of inertia. and This refers to the control gain of the drone's autopilot. and Let these be the desired translational velocity and yaw rate. Defined. and As the control input for UAV i, and the state vector of the UAV is defined as... Therefore, the dynamic model of UAV i can be re-expressed as a linear model:

[0098]

[0099] Among them, A ai =diag(A i,1 A i,2 A i,3 A i,4 B ai =diag(B i,1 B i,2 B i,3 B i,4 ), and u ai =[u i,1 ,u i,2 ,u i,3 ,u i,4 ] T Parameter matrix A i,1 B i,1 A i,2 B i,2 A i,3 B i,3 and A i,4 B i,4 Corresponding to drones i and p respectively i,x ,p i,y ,p i,z 3D position and ψ i The yaw angle-related subsystems contain system parameters that are susceptible to impact during the movement of the air-to-ground unmanned swarm system.

[0100] The parameter matrices are as follows:

[0101]

[0102] (2) Dynamics model of unmanned vehicles

[0103] in, and Let ψ represent the position and velocity of the unmanned vehicle i in the two-dimensional plane, respectively. gi and ω gi F represents the yaw angle and yaw rate of the unmanned vehicle i, respectively. gi,x and F gi,y Let C represent the total forces acting on the unmanned vehicle i in terms of position and orientation, respectively. A,i and C f,i The aerodynamic drag coefficient and rolling friction coefficient of the unmanned vehicle i, respectively, m gi For the quality of the unmanned vehicle i.

[0104] Assume the driverless car moves at a small angle ψ gi ≈0 and ω gi If the input is approximately 0, and feedback linearization is used, then we can define the control input as... and u i,6 =F gi,y Then the state vector of the autonomous vehicle is defined as x. gi =[p i,x ,v i,x ,p i,y ,v i,y ] T Then the dynamic model of the autonomous vehicle can be re-expressed as a linear model as follows:

[0105]

[0106] Wherein, the parameter matrix is ​​A gi =diag(A i,5 A i,6 ),B gi =diag(B i,5 B i,6 The control input is u. gi =[u i,5 ,u i,6 ] T A i,5 B i,5 and A i,6 B i,6 These correspond to the unmanned vehicle i and the location-related subsystem, respectively. and B i,5 =B i,6 =[0,1 / m gi ] T .

[0107] (3) The dynamic model and expected formation of the virtual leader in the constructed air-to-ground unmanned swarm system are as follows:

[0108]

[0109] Where, p0 = [x 0,x ,x 0,y ,x 0,z ,x 0,ψ ] T v0 = [v 0,x ,v 0,y ,v 0,z ,v 0,ψ ] T and a0 = [a 0,x ,a 0,y ,a 0,z ,a 0,ψ ] T These represent the virtual leader's position (angle), velocity (angular velocity), and acceleration (angular acceleration), respectively. H i,0 =[h i0,x ,h i0,y ,h i0,z ] T H i0 and H j0 denoted as the expected state deviations of the i-th and j-th followers and the virtual leader, respectively.

[0110] Next, based on the aforementioned dynamic models of drones, unmanned vehicles, the virtual leader, and the desired formation, a globally unified description model is constructed:

[0111]

[0112] Where i = 1, 2, ..., N; k = 1, 2, ..., 6; and Let A represent the state, control input, and output of each subsystem in the air-to-ground UAV swarm system, respectively. Let n, f, and p represent the dimensions of the subsystem state, control input, and output, respectively. i,κ B i,κ and C i,κ Let C be the dynamical matrix of appropriate dimensions. If κ∈{1,2}, then C i,κ =[e 4,1 ] T =

[1000] , otherwise C i,κ =[e 2,1 ] T =

[10] .

[0113] Step 2: Construct a distributed predefined time observer to estimate the virtual leader state and formation information, providing necessary trajectory references for formation tracking of the air-to-ground unmanned swarm. The convergence time of the distributed predefined time observer is set to a precise predefined time T, independent of the observer parameters. o .

[0114] In this step, a distributed formation observer is established using a symbolic function and a predefined time lemma, as follows:

[0115]

[0116] in, and These represent the position and velocity states of observer i, respectively. Where m∈(0,1), λ min (·) represents the smallest eigenvalue of the matrix, H is the information transfer matrix, and T is the smallest eigenvalue of the matrix. o >0 represents the preset convergence time. By constructing the Lyapunov function of the observer and taking its derivative, we can obtain... It satisfies the predefined time lemma form: Therefore, the constructed observer can converge within a predefined time. k3 and k4 are positive gain constants. ij b is the adjacency weight between the i-th and j-th followers. It is greater than zero if there is communication between i and j, and equal to zero otherwise. i The connection weight between virtual leaders and followers.

[0117] Step 3: Using the aforementioned globally unified description model and distributed predefined time observers, reconstruct the augmented dynamics model of the air-to-ground unmanned swarm system, design an optimal formation controller based on the augmented dynamics model, and construct the control gain matrix K. i,κ .

[0118] According to equation (7), the distributed predefined time observer is rewritten in the following form:

[0119]

[0120] in, Z 0i =I2, where I2 is a two-dimensional identity matrix.

[0121]

[0122] in,

[0123] and Each drone interacts with its neighboring drones for position and speed information. 0i =[1,0].

[0124] definition As the formation tracking error, and combined with the state χ in equations (6) and (8) i,κ and Construct augmented vectors Thus, an augmented subsystem dynamic model of the air-to-ground UAV swarm system is constructed:

[0125]

[0126] in, r = m + 2, where r is the augmented vector X i,κ The dimension of p is the state vector χ after the global unified description of the unmanned system. i,κ Dimensions and u i,κ The control input of the augmented subsystem is s, where s is the control input u of the augmented subsystem. i,κ The dimension. If κ∈{1,2}, otherwise:

[0127]

[0128] Design a model-based optimal formation controller.

[0129]

[0130] in, and For optimal formation controller, This represents the optimal feedback gain matrix. This can be achieved by solving the discount factor algebraic Riccati equation. We obtain, where α i,κ This is a discount factor used to ensure the convergence of the controller.

[0131] Step 4: Apply piecewise constant initial excitation to each augmented subsystem of the augmented dynamics model of the air-to-ground unmanned swarm system, collect and store the system operation data from its early operation, and then obtain the initial stable control strategy based on the early operation data of the air-to-ground unmanned system using a data-driven method.

[0132]

[0133] The data storage matrix is ​​constructed using only the integral of the system's initial operating data. and and These are data storage matrices The pseudo-inverse and a set of basis vectors of the null space satisfy the following: For any in, This refers to the appropriate matrix selected using the pole placement method.

[0134] Step 5: Perform offline policy iteration for reinforcement learning: Construct a data storage matrix using only the integral of the system running data, avoiding the use of the Kronecker product, and further construct a small non-singular data matrix according to the data index; use the system running data in the non-singular data matrix to solve the generalized Hillwest-transpose data equation that does not contain a parameter matrix, and learn the optimal control gain matrix through iterative calculation. This ensures the formation controller of the air-to-ground unmanned swarm system This optimizes the learning process of the optimal formation controller and reduces the computational burden.

[0135] In this embodiment of the invention, the specific process of the offline policy iterative learning algorithm based on efficient reinforcement learning is as follows:

[0136] S501. Define the data storage matrix and in, It stores the information about the augmented system during its operation and its state X. i,κ Related system operation data, It stores the data related to the input u during the operation of the augmented system. i,κ Related system operation data, It stores the state differential during the operation of the augmented system. Related system operation data.

[0137]

[0138] Wherein, the time interval [t0, t l ] represents the sampling time of data collection, where lΔT=t l -t0, l≥(r+1)s+r represents the length of the collected data, ΔT>0 represents the sampling step size, t0 is the initial time of data collection, t1 is the time after the system has passed the sampling step size at the initial data collection time t0, and so on, t l The point at which the system stops collecting data. These represent the κ-th augmenting subsystem of the i-th follower in the corresponding air-to-ground unmanned swarm system within the subinterval [t0,t1], [t1,t2], ..., [t], respectively. l-1 ,t l The system status data collected internally, These represent the κ-th augmenting subsystem of the i-th follower in the corresponding air-to-ground unmanned swarm system within the subinterval [t0,t1], [t1,t2], ..., [t], respectively. l-1 ,t l The system input data collected within the system, These represent the κ-th augmenting subsystem of the i-th follower in the corresponding air-to-ground unmanned swarm system within the subinterval [t0,t1], [t1,t2], ..., [t], respectively. l-1 ,t l The system state collected internally The integral data, among which

[0139] S502. From respectively and Select a set of data with indices σ = {1, 2, ..., r + s} to construct a data storage matrix. Make It is a non-singular matrix.

[0140] S503. Strategy Evaluation: Relying on system operation data in the data storage unit, iteratively solve the following data-driven generalized Hillwest transpose matrix equation.

[0141]

[0142] in, The superscript σ indicates a data matrix reconstructed by index. In each iteration update, the superscripts k and k+1 represent the data corresponding to the k-th and k+1-th iterations, respectively. matrix The block submatrix is ​​a positive definite matrix. For matrix The block submatrix, This represents the controller gain matrix of the κ augmented subsystems of the i-th follower in the (k+1)-th iteration of the spatiotemporal unmanned swarm system.

[0143] S504. Strategy Iteration: Using Update control gain matrix

[0144] S505. Judgment Whether it is true or not, ∈ is a small constant value greater than zero;

[0145] ③If If the condition is met, the iteration stops, and the control gain is... The controller is Proceed to step S506;

[0146] ④If If this is not true, then k = k + 1, and then proceed to S503 and S504 to continue solving;

[0147] S506. Obtain the optimal control gain matrix Simultaneously obtain the optimal formation controller

[0148] The above specific embodiments only describe the design principles of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, those skilled in the art can modify or make equivalent substitutions to the technical solutions described in the foregoing embodiments; and such modifications and substitutions do not depart from the inventive spirit and technical solutions of the present invention, and should all fall within the scope of protection of the present invention.

Claims

1.A method for cooperative formation control of air-ground unmanned cluster system driven by reinforcement learning, characterized in that, The air-ground unmanned cluster system comprises a virtual leader and followers composed of M unmanned aerial vehicles and N-M unmanned ground vehicles; the method comprises the following steps: Step one: a dynamic model of the air-ground unmanned cluster system is established, comprising a dynamic model of the unmanned aerial vehicles, a dynamic model of the unmanned ground vehicles, a dynamic model of the virtual leader and a desired formation, and is converted into a globally unified description model; wherein the dynamic model of the unmanned aerial vehicles and the dynamic model of the unmanned ground vehicles are decoupled into a linear model comprising a parameter matrix and a system state through a linear feedback method; system parameters susceptible to influence during the air-ground unmanned cluster system advancing are designed into the parameter matrix; Step two: Constructing a distributed predefined time observer, using the observer to estimate the virtual leader state and formation information, providing trajectory reference for formation tracking of the unmanned aerial swarm system; the convergence time of the distributed predefined time observer is set to be an accurate predefined time independent of the observer parameters ; Step three: reconstruct the augmented dynamics model of the air-ground unmanned cluster system by using the global unified description model and the distributed predefined time observer, design the optimal formation controller based on the augmented dynamics model and construct the control gain matrix ; Step four: apply the piecewise constant initial excitation to each augmented subsystem of the augmented dynamic model and collect the system running data of the previous run; obtain the initial stable control gain matrix by using the data-driven method according to the system running data of the previous run ; Step five: offline policy iteration of reinforcement learning: only using the integral of the system running data to build a data storage matrix, and further constructing a non-singular data matrix according to the data index; using the system running data in the non-singular data matrix to solve the generalized Hilbert-transposed data equation without parameter matrix, and learning the optimal control gain matrix through iterative calculation , and then obtaining the optimal formation controller of the air-ground unmanned cluster system ; the optimal platoon controller obtained by solving a discounted factor algebraic Riccati equation; the optimal platoon controller is represented as: wherein ; denotes the optimal control gain matrix; is the state vector of each augmented subsystem of the swarm of unmanned aerial vehicles system; is obtained by solving the discounted factor algebraic Riccati equation wherein is the discounted factor to ensure the convergence of the controller; wherein , ; , , ; , ; and is the kinetic matrix of the global unified description model; The step five is: S501. Define data storage matrix and wherein, is stored in the data storage matrix associated system operation data, is stored in the data storage matrix associated system operation data, is stored in the data storage matrix associated system operation data; is the state vector of each augmented subsystem of the swarm of unmanned aerial vehicles; is the derivative of the state vector ; is the control input of each augmented subsystem of the swarm of unmanned aerial vehicles; ; ; wherein the time interval represents the sampling time of data collection, wherein, , represents the length of collecting data, represents the sampling step, t 0 is the initial time of system collecting data, t 1 is the time of system collecting data at the initial time t 0 through the sampling step, and so on, t l is the termination time of system collecting data; r is the dimension of state vector , s is the dimension of augmented subsystem control input . These respectively represent the corresponding air-to-ground unmanned swarm system number 1 i The first follower An augmented subsystem in a sub-interval Internally collected system status data; These respectively represent the corresponding air-to-ground unmanned swarm system number 1 i The first follower An augmented subsystem in a sub-interval Internally collected system input data; These respectively represent the corresponding air-to-ground unmanned swarm system number 1 i The first follower An augmented subsystem in a sub-interval Internally collected system state The integral data, among which ,express exist Integrals within; S502. Selecting a set of indices and from to construct a data storage matrix such that is a non-singular matrix; S503. Strategy evaluation: depending on system running data in the data storage unit, iterative solving is performed according to the following data-driven generalized Sylvester transpose matrix equation: wherein the superscript denotes the data matrix reconstructed by the index; the update is performed at each iteration, the superscript k and k +1 denotes the data corresponding to the k and the k +1th iteration, wherein is a block submatrix of the matrix and is a positive definite matrix, is a block submatrix of the matrix , denotes the controller gain matrix of the k th follower's i th augmented subsystem in the airspace unmanned swarm system at the th iteration, ; S504. Policy iteration: use updating the control gain matrix ; S505. judging whether the following is true, a small constant value greater than zero; If is true, the iteration is stopped, the control gain matrix is , the formation controller is , and the process goes to step S506. If is not true, let k be incremented by one, and then go to S503 and S504 to continue solving. S506. Obtain optimal control gain matrix while obtaining optimal platoon controller . 2.The method of claim 1, wherein, In step one, the system parameters susceptible to influence during the air-ground unmanned cluster system advancing comprise an unmanned aerial vehicle aerodynamic drag coefficient, an unmanned aerial vehicle mass, an unmanned aerial vehicle moment of inertia, an unmanned aerial vehicle autopilot control gain and an unmanned ground vehicle mass. 3.The method of claim 1 or 2, wherein, In step one, the dynamic model of the air-ground unmanned cluster system is established as follows: The UAV and the unmanned vehicle are uniformly sequenced as followers, and a dynamic model of a first follower in an air-ground unmanned cluster system is as follows: i ​ (1) the dynamic model of the unmanned aerial vehicles is: where and denote the three-dimensional position and attitude of the quadrotor, respectively; i is the roll angle, is the pitch angle, are the second derivatives of are the aerodynamic drag coefficients of the quadrotor i is the mass of the quadrotor i g is the gravitational acceleration; ; where is the arm length of the quadrotor; are the moments of inertia; are the control gains of the autopilot of the quadrotor; are the desired translational velocity and yaw rate; define as the control inputs of the quadrotor i and define the state vector of the quadrotor as ; then the dynamics model of the quadrotor i can be re-expressed as a linear model:​​​​​​​​​​​​​ Wherein, the parameter matrix is The control input is ; parameter matrix and Corresponding to drones i of Three-dimensional position and The yaw angle-related subsystem contains system parameters in its parameter matrix that are susceptible to impact during the movement of the air-to-ground unmanned swarm system. The parameter matrix is: (2) the dynamic model of the unmanned ground vehicles wherein, and respectively represent the position and velocity of the unmanned vehicle i in a two-dimensional plane, and respectively represent the yaw angle and yaw rate of the unmanned vehicle i ; and respectively represent the total force in position and direction experienced by the unmanned vehicle i , and respectively represent the aerodynamic drag coefficient and rolling friction coefficient of the unmanned vehicle i ; is the mass of the unmanned vehicle i . Assume the car moves with a small angle and and using feedback linearization techniques, the control input of the car is defined as and ; the state vector of the car is defined as The dynamic model of the car can be re-expressed as a linear model: wherein the parameter matrix is , and the control input is ; (3) the dynamic model of the virtual leader: wherein, , and represent the position, velocity and acceleration of the virtual leader, respectively; (4) the desired formation wherein, , and are the desired state deviations of the i and j followers and virtual leaders, respectively; Then, a globally unified description model is constructed according to the dynamic model of the unmanned aerial vehicles, the dynamic model of the unmanned ground vehicles, the dynamic model of the virtual leader and the desired formation: (I) wherein, ; and represent the state, control input and output of each subsystem of the unmanned cluster system respectively; and are the dynamics matrices of the globally unified description model. 4.The method of claim 3, wherein, In step two, the constructed distributed pre-defined time observer is: wherein and are the position and velocity states of the observer i respectively; wherein , denotes the smallest eigenvalue of the matrix is the information transfer matrix, is a preset convergence time; Taking the derivative of the Lyapunov function constructed for the observer, we have which satisfies the pre-defined time lemma form: Thus, the constructed observer is able to converge within the pre-defined time; and are positive gain constants; is the i and j adjacency weight between the ij is greater than zero if there is a communication between the is the connection weight between the virtual leader and the follower. 5.The method of claim 4, wherein, In step three, an augmented dynamic model is constructed as follows: The distributed pre-defined time observer is rewritten in the following form: (I) wherein, ; ; ; wherein and interacting position and velocity information of each UAV with its neighbor UAVs respectively, is a two-dimensional identity matrix;​ Definitions As platoon tracking error, in combination with the states in Equations (I) and (II) and Constructing the augmented vector thereby constructing an augmented subsystem dynamics model for the air-ground unmanned swarm system: wherein wherein r is the dimension of the augmented vector p is the dimension of the state vector of the unified global description of the unmanned system and , is the control input of the augmented subsystem, wherein s is the dimension of the augmented subsystem control input ; if , , otherwise .​​ 6.The method of claim 1, wherein In step four, the initial stable control gain matrix Obtained by data-driven approach: wherein the data storage matrix is constructed only using integrals of the previous system operation data and , and are a set of basis vectors of the pseudo-inverse and null space of the data storage matrix respectively, for any which satisfies wherein is a matrix selected by the pole placement method.

Citation Information

Patent Citations

  • Unmanned system air-ground collaborative navigation and obstacle avoidance method based on vision

    CN116540784A

  • Event triggering optimal cooperative control method for heterogeneous unmanned cluster system

    CN118707844A