Reinforced learning-driven air-ground unmanned cluster system cooperative formation control method

Through the reinforcement learning-driven method, the global unified description model and distributed predefined time observer are used to learn the optimal control gain matrix, which solves the optimal formation tracking control problem of the open-ground unmanned cluster system in the unknown system model, and realizes efficient control of low computing resource consumption.

CN120044982AActive Publication Date: 2025-05-27BEIJING INST OF TECH

Patent Information

Application Number
CN202510169736.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-27
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

The existing collaborative control scheme for unmanned cluster systems requires system model information and consumes a lot of computing resources, making it difficult to achieve optimal formation tracking control when the system model is unknown.

Method used

Using reinforcement learning-driven methods, the augmented dynamics model is reconstructed by establishing a global unified description model and distributed predefined time observer, and using the data-driven generalized Hillwest-transpose data equation to learn the optimal control gain matrix to achieve optimal formation control.

Benefits of technology

In the case of unknown system model, optimal formation tracking control with low computing resource consumption is realized, which avoids dependence on the system model and improves the controller's learning efficiency and control performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120044982A_ABST
    Figure CN120044982A_ABST
Patent Text Reader

Abstract

The invention provides an air-ground unmanned cluster system cooperative formation control method driven by reinforcement learning. Establishing a dynamic model for the air-ground unmanned cluster system; a distributed predefined time observer is constructed to estimate the state of a virtual leader, and a state reference is provided for cooperative formation tracking control; and reconstructing a dynamic model of the augmented air-ground unmanned cluster system, and further constructing a control gain matrix. Segmented constant initial excitation is applied to the augmentation system, state data of early-stage operation of the augmentation system is collected and stored, then an initial stability control strategy is obtained by using a data driving method according to the early-stage operation data of the air-ground unmanned system, and finally, the stability of the air-ground unmanned system is improved based on an offline strategy reinforcement learning algorithm and the initial stability control strategy based on the data. And learning an optimal formation tracking controller to realize optimal formation tracking control of the air-ground unmanned system. According to the invention, under the condition that a system model is unknown, the optimal time-varying formation tracking control of the air-ground unmanned cluster system can be realized according to complex task requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of data-driven and air-ground unmanned cluster systems, and particularly relates to a cooperative formation control method for an air-ground unmanned cluster system driven by reinforcement learning. Background Art

[0002] In recent years, due to characteristics such as distribution, robustness, fault tolerance, and strong adaptability, the distributed cooperative control of multi-unmanned systems has been widely applied in many fields such as intelligent transportation, cooperative systems of multi-unmanned aerial vehicles and multi-unmanned ground vehicles, and microgrids. Among them, formation control is an important topic in the cooperative control of unmanned systems, and its main goal is to design a distributed control strategy that uses local communication between individual unmanned systems to ensure that the entire unmanned system maintains the corresponding relative position relationship according to the set formation shape.

[0003] However, most of the existing formation control research focuses on homogeneous multi-unmanned aerial vehicle systems or multi-unmanned ground vehicle systems. When performing increasingly complex cooperative tasks such as autonomous surveillance, cooperative target search, and cargo transportation, a single form of homogeneous unmanned system cannot effectively complete these tasks. The heterogeneous air-ground unmanned cluster system combines the advantages of the rapid detection of a wide area by unmanned aerial vehicles, the precise positioning of ground targets by unmanned ground vehicles, and high payloads. Therefore, the air-ground unmanned cluster system can provide effective solutions for these complex cross-domain and cross-dimensional cooperative tasks.

[0004] In addition, in the design of the vast majority of existing cooperative control schemes for air-ground unmanned systems, the realization of cooperative control requires a complete understanding of the model information of the unmanned system, which is difficult to achieve in many practical applications. At the same time, the existing controller design techniques only consider the steady-state performance of the system and do not consider the transient performance of the system, and cannot guarantee the optimality of the system cooperative control.

[0005] Therefore, how to design a controller under unknown system models to ensure the optimal cooperation of the air-ground unmanned cluster system is a current hot and difficult problem.

[0006] To solve this problem, current research collects system data and uses a reinforcement learning control strategy to learn the optimal controller through the collected system data, thereby ensuring optimal cooperative tracking control under unknown system models.

[0007] However, this optimal cooperative tracking control scheme requires a large amount of computing resources. When the number and state dimensions of the unmanned system are large, learning the optimal controller requires expensive computing resources. In addition, this reinforcement learning scheme needs to rely on the initial stable control strategy of the system model for learning.

[0008] Therefore, in the case where the model of the air-ground unmanned cluster system is unknown, how to propose an optimal formation tracking control scheme that is completely independent of the system model information and can ensure low computational resource consumption is an urgent problem to be solved at present. Summary of the Invention

[0009] The present invention provides a cooperative formation control method for an air-ground unmanned cluster system driven by reinforcement learning, which can be applied to optimal formation cooperation control in the case of an unknown system model and reduce the computational burden in the learning process of the controller.

[0010] To solve the above technical problems, the present invention is implemented as follows.

[0011] A cooperative formation control method for an air-ground unmanned cluster system driven by reinforcement learning, where the air-ground unmanned cluster system includes 1 virtual leader and followers composed of M unmanned aerial vehicles and N - M unmanned ground vehicles; the method includes the following steps:

[0012] Step 1: Establish the dynamic model of the air-ground unmanned cluster system, including the dynamic models of unmanned aerial vehicles, unmanned ground vehicles, the dynamic model of the virtual leader, and the desired formation shape, and convert it into a globally unified description model; among them, the dynamic models of unmanned aerial vehicles and unmanned ground vehicles are decoupled into linear models including parameter matrices and system states through the linear feedback linearization method; the system parameters that are vulnerable to influence during the movement of the air-ground unmanned cluster system are designed into the parameter matrices;

[0013] Step 2: Construct a distributed predefined-time observer, use the observer to estimate the virtual leader state and formation information, and provide a trajectory reference for the formation tracking of the air-ground unmanned cluster system; the convergence time of the distributed predefined-time observer is set to an exact predefined time T independent of the observer parameters o ;

[0014] Step 3: Use the globally unified description model and the distributed predefined-time observer to reconstruct the augmented dynamic model of the air-ground unmanned cluster system, design an optimal formation controller based on the augmented dynamic model, and construct a control gain matrix K i,κ ;

[0015] Step 4: Apply piecewise constant initial excitation to each augmented subsystem of the augmented dynamic model, and collect the system operation data of the previous operation; according to the system operation data of the previous operation, use a data-driven method to obtain an initially stable control gain matrix

[0016] Step 5: Perform reinforcement learning off-policy iteration: Only use the integral of the system operation data to construct a data storage matrix, and further construct a non-singular data matrix according to the data index; Use the system operation data in the non-singular data matrix to solve the generalized Sylvester-transpose data equation without the parameter matrix, and through iterative calculation, learn the optimal control gain matrix Furthermore, obtain the optimal formation controller of the air-ground unmanned cluster system

[0017] Preferably, in step 1, the system parameters that are easily affected during the movement of the air-ground unmanned cluster system include: the aerodynamic drag coefficient of the unmanned aerial vehicle, the mass of the unmanned aerial vehicle, the moment of inertia of the unmanned aerial vehicle, the control gain of the unmanned aerial vehicle autopilot, and the mass of the unmanned vehicle

[0018] Preferably, in step 1, the establishment of the dynamic model of the air-ground unmanned cluster system is as follows

[0019] Sort the unmanned aerial vehicles and unmanned vehicles as followers uniformly. The dynamic model of the i-th follower in the air-ground unmanned cluster system is as follows

[0020] (1) The dynamic model of the unmanned aerial vehicle is as follows

[0021]

[0022] Among them, p i =[p i,x , p i,y , p i,z T and Θ i =[φ i , θ i , ψ i T respectively represent the three-dimensional position and attitude of the i-th unmanned aerial vehicle; φ i is the roll angle, θ i is the pitch angle, ψ i is the yaw angle and are respectively the second derivatives of p i and Θ i ; k i,x , k i,y and k i,z are the aerodynamic drag coefficients of the i-th unmanned aerial vehicle, m qi is the mass of the unmanned aerial vehicle i ; g is the acceleration due to gravity; b qi,φ =(l / J qi,φ ), b qi,θ =(l / J qi,θ ) and b qi,ψ =(1 / J qi,ψ ​​); where \(l\) is the arm length of the UAV; \(J\) qi,φ , \(J\) qi,θ and \(J\) qi,ψ are the moments of inertia; and are the control gains of the UAV autopilot; and are the desired translational velocity and yaw rate; Define and as the control input of UAV \(i\), and define the state vector of the UAV as Then the dynamic model of UAV \(i\) can be re-expressed as a linear model:

[0023]

[0024] where the parameter matrix is \(A\) ai =\(diag(A\) i,1 , \(A\) i,2 , \(A\) i,3 , \(A\) i,4 ), \(B\) ai =\(diag(B\) i,1 , \(B\) i,2 , \(B\) i,3 , \(B\) i,4 ), the control input is \(u\) ai =\([u\) i,1 , \(u\) i,2 , \(u\) i,3 , \(u\) i,4 T ; The parameter matrices \(A\) i,1 , \(B\) i,1 , \(A\) i,2 , \(B\) i,2 , \(A\) i,3 , \(B\) i,3 and \(A\) i,4 , \(B\) i,4 correspond to the subsystems related to the \(p\) i,x , \(p\) i,y , \(p\) i,z three-dimensional position and \(\psi\) i yaw angle of UAV \(i\), and the parameter matrices contain the system parameters that are vulnerable to influence during the movement of the air-ground unmanned cluster system;

[0025] The parameter matrix is:

[0026]

[0027] (2) Unmanned vehicle dynamic model

[0028]

[0029] J gi \(\omega\)​gi = τ gi ,

[0030] where p i = [p i,x , p i,y T and v i = [v i,x , v i,y T represent the position and velocity of the autonomous vehicle i in the two-dimensional plane, ψ gi and ω gi are the yaw angle and yaw rate of the autonomous vehicle i respectively; F gi,x and F gi,y are the total forces acting on the autonomous vehicle i in the position and direction respectively, C A,i and C f,i are the aerodynamic drag coefficient and rolling friction coefficient of the autonomous vehicle i respectively; m gi is the mass of the autonomous vehicle i;

[0031] Assume that the autonomous vehicle moves at a small angle ψ gi ≈ 0 and ω gi ≈ 0, and using the feedback linearization technique, the control input of the autonomous vehicle is defined as and u i,6 = F gi,y ; the state vector of the autonomous vehicle is defined as x gi = [p i,x , v i,x , p i,y , v i,y T , then the dynamic model of the autonomous vehicle can be re-expressed as a linear model:

[0032]

[0033] where the parameter matrix is A gi = diag(A i,5 , A i,6 ), B gi = diag(B i,5 , B i,6 );

[0034] The control input is u gi = [u i,5 , u i,6 T ;

[0035] (3) Dynamics model of the virtual leader:

[0036]

[0037] ​​​​where p 0 =[x 0,x ,x 0,y ,x 0,z ,x 0,ψ T , v 0 =[v 0,x , v 0,y , v 0,z , v 0,ψ T and a 0 =[a 0,x , a 0,y , a 0,z , a 0,ψ T represent the position, velocity, and acceleration of the virtual leader, respectively;

[0038] (4) Desired formation

[0039] H ij = H i0 - H j0 ,

[0040] where H i,0 =[h i0,x , h i0,y , h i0,z T , H i0 and H j0 are the desired state deviations of the i-th and j-th followers and the virtual leader, respectively;

[0041] Then, based on the dynamic models of the UAVs, UGVs, the dynamic model of the virtual leader, and the desired formation, a global unified description model is constructed:

[0042]

[0043] where i = 1, 2,..., N; κ = 1, 2,..., 6; χ i,κ , u i,κ and y i,κ represent the states, control inputs, and outputs of each subsystem of the air-ground UAV swarm system, respectively; A i,κ , B i,κ and C i,κ are the dynamic matrices of the global unified description model.

[0044] Preferably, in step two, the constructed distributed predefined-time observer is:

[0045]

[0046] where ​​​​and are the position and velocity states of observer i, respectively; where m ∈ (0, 1), λ min (·) represents the minimum eigenvalue of the matrix, is the information transfer matrix, T o > 0 is the preset convergence time;

[0047] By constructing the Lyapunov function of the observer and taking its derivative, we can obtain which satisfies the form of the predefined time lemma:

[0048]

[0049] Therefore, the constructed observer can converge within the predefined time; k 3 and k 4 are positive gain constants; w ij is the adjacency weight between the i-th and j-th followers, which is greater than zero if there is communication between ij, otherwise equal to zero, b i is the connection weight between the virtual leader and the followers.

[0050] Preferably, in step three, the augmented dynamics model is constructed as:

[0051] Rewrite the distributed predefined time observer as the following form:

[0052]

[0053] where,

[0054]

[0055] where, and are the position and velocity information interactions between each UAV and its neighbor UAVs, C 0i = [1, 0]; I 2 is the two-dimensional identity matrix;

[0056] Define as the formation tracking error, and combine the states χ in equations (I) and (II) i,κ and to construct the augmented vector Thus, construct the augmented subsystem dynamics model of the air-ground UAV cluster system:

[0057]

[0058] where, r = p + 2, where r is the augmented vector X i,κThe dimension of p is the state vector χ after the global unified description of the unmanned system. i,κ The dimension of and u i,κ is the control input of the augmented subsystem, where s is the control input u of the augmented subsystem i,κ dimension; if κ∈{1,2}, otherwise

[0059] Preferably, the optimal formation controller is obtained by solving the discounted factor algebraic Riccati equation; the optimal formation controller is expressed as:

[0060]

[0061] in, and is the optimal formation controller, represents the optimal control gain matrix; By solving the discounted factor algebraic Riccati equation We get, where α i,κ is the discount factor used to ensure the convergence of the controller.

[0062] Preferably, in step 4, the initial stable control gain matrix Obtained through a data-driven approach:

[0063]

[0064] Among them, only the integral of the previous system operation data is used to construct the data storage matrix and and The data storage matrix The pseudo-inverse and the set of basis vectors of the null space are given by Its satisfaction in, is the matrix selected by the pole placement method.

[0065] Preferably, the step five is:

[0066] S501. Define data storage matrix and in, The data stored in the augmented system are related to the state X during operation. i,κ Related system operation data, The data stored in the augmented system are related to the input u i,κ Related system operation data, What is stored in the augmented system is the state differential during operation Associated system operation data;

[0067]

[0068] Among them, the time interval [t 0 , t l represents the sampling time of data collection. Among them, lΔT = t l - t 0 , l ≥ (r + 1)s + r represents the length of the collected data, ΔT > 0 represents the sampling step, t 0 is the initial moment when the system collects data, t 1 is the moment after the sampling step at the initial data collection moment t 0 of the system. And so on, t l is the termination moment when the system collects data;

[0069] respectively represent the system state data collected by the κ-th augmented subsystem of the i-th follower in the corresponding unmanned aerial and ground cluster system in the sub-intervals [t 0 , t 1 , [t 1 , t 2 ,..., [t l-1 , t l ; respectively represent the system input data collected by the κ-th augmented subsystem of the i-th follower in the corresponding unmanned aerial and ground cluster system in the sub-intervals [t 0 , t 1 , [t 1 , t 2 ,..., [t l-1 , t l ; respectively represent the integral data of the system state collected by the κ-th augmented subsystem of the i-th follower in the corresponding unmanned aerial and ground cluster system in the sub-intervals [t 0 , t 1 , [t 1 , t 2 ,..., [t l-1 , t l , where represents the integral within [t , t l-1 , t l ;

[0070] S502. Select a set of data with indexes σ = {1, 2,..., r + s} from and respectively to construct a data storage matrix such that is a non-singular matrix;

[0071] S503. Policy evaluation: Depending on the system operation data in the data storage unit, iteratively solve according to the following data-driven generalized Sylvester transpose matrix equation:

[0072]

[0073] where The superscript σ represents the data matrix reconstructed by index; It is updated at each iteration. The superscripts k and k + 1 represent the data corresponding to the kth and (k + 1)th iterations. Among them, is the matrix is a block sub-matrix of and is a positive definite matrix, is the matrix is a block sub-matrix of, represents the controller gain matrix of the κ augmented subsystems of the ith follower in the air-ground unmanned cluster system at the kth iteration,

[0074] S504. Policy iteration: Use to update the control gain matrix

[0075] S505. Judge whether it holds, ∈ is a small constant value greater than zero;

[0076] ① If holds, stop the iteration. At this time, the control gain matrix is The formation controller is Go to step S506;

[0077] ② If does not hold, let k increment by one, and then go to S503 and S504 to continue the solution;

[0078] S506. Obtain the optimal control gain matrix At the same time, obtain the optimal formation controller

[0079] Beneficial effects:

[0080] (1) The present invention provides a cooperative formation control method for an air-ground unmanned cluster system driven by reinforcement learning. This method adopts a data-driven efficient offline policy reinforcement learning algorithm. Without the need to obtain the system model, it learns the initial stable control policy and the optimal control policy, while ensuring the formation cooperation control of the air-ground unmanned cluster system, reducing the computational burden in the learning process of the optimal controller.

[0081] Secondly, the dynamic models of the UAV and the UGV are decoupled into linear models including parameter matrices and system states through the linear feedback linearization method, so as to construct a data-based generalized Sylvester-transpose equation in the later reinforcement learning process.

[0082] Meanwhile, by applying piecewise constant initial excitation to the air-ground unmanned system, state data for learning is obtained, and an optimal control strategy is obtained by solving the data-based generalized Sylvester-transpose equation, avoiding the dependence on the full rank condition and the properties of the Kronecker product, and reducing the computational burden in the learning process of the optimal controller.

[0083] (2) The present invention uses the collected state data and adopts a data-driven method to obtain the initial stable control strategy necessary in the policy iteration algorithm, avoiding the dependence of the initial stable strategy on the input dynamics of the system model.

[0084] (3) The present invention designs a predefined time-distributed observer, which can realize the state estimation of the virtual leader under complex time-varying formation tasks within a predefined time, so as to quickly provide necessary and accurate trajectory references for the decentralized control of the air-ground unmanned cluster system.

[0085] (4) During offline iterative learning, the present invention constructs a data storage matrix only containing the integral of system data, and converts the data storage matrix into a non-singular data matrix, thereby reducing the data volume and reducing the computational burden in the learning process of the optimal controller. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] Figure 1 It is a design flow chart of a cooperative formation control method for an air-ground unmanned cluster system driven by reinforcement learning provided by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0087] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the present invention will be described in detail below with reference to the drawings in the present invention and by way of examples.

[0088] As Figure 1 shown, the air-ground unmanned cluster system targeted by the embodiment of the present invention consists of 1 virtual leader and N followers; the followers include M UAVs and N-M UGVs. They are connected through a connected network communication topology.

[0089] The specific steps of the cooperative formation control method for the air-ground unmanned cluster system based on reinforcement learning provided by the present invention are as follows:

[0090] Step 1: Establish an air-ground unmanned cluster system model, including the dynamic models of unmanned aerial vehicles (UAVs), unmanned ground vehicles (UGVs), the dynamic model of the virtual leader, and the desired formation shape, and construct a global unified description model.

[0091] In this step, the dynamic models of UAVs and UGVs are decoupled into linear models containing parameter matrices and system states through the linear feedback linearization method; the system parameters vulnerable to influence during the movement of the air-ground unmanned cluster system are designed into the parameter matrices. Such a design enables the construction of a Guangxi Sylvester-transposed data equation without parameter matrices when performing reinforcement learning optimal controller learning in Step 5, so that the controller can be learned and solved only based on the system operation data. The advantage of this is that when the vulnerable system parameters change, the reinforcement learning-driven method can still utilize the system data generated during the movement of the air-ground unmanned cluster system to learn a new optimal controller suitable for the current parameter information.

[0092] The parameters of UGVs and UAVs vulnerable to the influence of the system movement process include: the aerodynamic drag coefficient of UAVs (affected by wind speed), the mass of UAVs (affected by payload release and fuel consumption), the moment of inertia of UAVs (affected by aerodynamic forces and dynamic changes in mass), and the mass of UGVs (affected by payload release and fuel consumption).

[0093] Based on the analysis of these vulnerable parameters, during the movement of the air-ground UAV cluster system, the values of these vulnerable parameters may change. For a model-based controller, when these parameters change due to the influence of the system movement process, it cannot adaptively learn a new controller suitable for the current parameters, thus reducing the control performance of the controller. The dynamic model of the i-th follower in the air-ground unmanned cluster system constructed in the present invention is as follows:

[0094] (1) UAV dynamic model

[0095]

[0096] Where, and respectively represent the three-dimensional position and attitude of the i-th UAV (φ i is the roll angle, θ i is the pitch angle, ψ i is the yaw angle), and are the second-order derivatives of p i and Θ i respectively. k i,x , k i,y and k i,z are the aerodynamic drag coefficients of UAV i, m qiis the mass of drone i, and g is the acceleration due to gravity. qi,φ =(l / J qi,φ ),b qi,θ =(l / J qi,θ ) and b qi,ψ =(1 / J qi,ψ ), where l is no one

[0097] The arm length of the machine, J qi,φ ,J qi,θ and J qi,ψ is the moment of inertia. and is the control gain of the UAV autopilot. and is the desired translation velocity and yaw rate. Define and As the control input of UAV i, the state vector of the UAV is defined as Then the dynamic model of UAV i can be re-expressed as a linear model:

[0098]

[0099] Among them, A ai =diag(A i,1 ,A i,2 ,A i,3 ,A i,4 ), B ai =diag(B i,1 ,B i,2 ,B i,3 ,B i,4 ), and u ai =[u i,1 ,u i,2 ,u i,3 ,u i,4 ] T . Parameter matrix A i,1 ,B i,1 ,A i,2 ,B i,2 ,A i,3 ,B i,3 and A i,4 ,B i,4 Corresponding to drone i and p respectively i,x ,p i,y ,p i,z 3D position and ψ i Subsystems related to yaw angle, these parameter matrices contain system parameters that are easily affected during the movement of the air-to-ground unmanned cluster system.

[0100] Among them, the parameter matrices are:

[0101]

[0102] (2) Kinematic Model of Driverless Vehicle

[0103] where and represent the position and velocity of driverless vehicle i in the two-dimensional plane, respectively, and ψ gi and ω gi are the yaw angle and yaw rate of driverless vehicle i, respectively. F gi,x and F gi,y are the total forces acting on driverless vehicle i in the position and direction, respectively. C A,i and C f,i are the aerodynamic drag coefficient and rolling friction coefficient of driverless vehicle i, respectively, and m gi is the mass of driverless vehicle i.

[0104] Assume that the driverless vehicle moves at a small angle ψ gi ≈0 and ω gi ≈0, and using the feedback linearization technique, then we can define the control input as and u i,6 =F gi,y . Then define the state vector of the driverless vehicle as x gi =[p i,x ,v i,x ,p i,y ,v i,y T , then the dynamic model of the driverless vehicle can be re-expressed as a linear model:

[0105]

[0106] where the parameter matrix is A gi =diag(A i,5 ,A i,6 ), B gi =diag(B i,5 ,B i,6 ), the control input is u gi =[u i,5 ,u i,6 T . A i,5 , B i,5 and A i,6 , B i,6 correspond to the subsystems related to the position of driverless vehicle i, respectively, where and B i,5 =B i,6 =[0,1 / m gi T .

[0107] ​​​(3) The dynamic model and desired formation of the virtual leader in the constructed unmanned aerial and ground cluster system are as follows:

[0108]

[0109] Where, p 0 = [x 0,x , x 0,y , x 0,z , x 0,ψ T , v 0 = [v 0,x , v 0,y , v 0,z , v 0,ψ T and a 0 = [a 0,x , a 0,y , a 0,z , a 0,ψ T represent the position (angle), velocity (angular velocity), and acceleration (angular acceleration) of the virtual leader respectively. H i,0 = [h i0,x , h i0,y , h i0,z T , H i0 and H j0 are the desired state deviations of the i-th and j-th followers and the virtual leader respectively.

[0110] Next, based on the dynamic models of the above-mentioned unmanned aerial vehicles and unmanned ground vehicles, the dynamic model of the virtual leader, and the desired formation, a global unified description model is constructed:

[0111]

[0112] Where, i = 1, 2,..., N; k = 1, 2,..., 6; and represent the states, control inputs, and outputs of each subsystem of the unmanned aerial and ground cluster system respectively. n, f, and p represent the dimensions of the subsystem states, control inputs, and outputs respectively. A i,κ , B i,κ and C i,κ are dynamic matrices of appropriate dimensions. If κ ∈ {1, 2}, C i,κ = [e 4,1 T =

[1000] , otherwise C i,κ = [e 2,1 T =

[10] .

[0113] ​​​​​​Step 2: Construct a distributed predefined-time observer to estimate the virtual leader state and formation information using the observer, providing the necessary trajectory reference for the formation tracking of the air-ground unmanned cluster. The convergence time of the distributed predefined-time observer is set to an exact predefined time \(T\) independent of the observer parameters o .

[0114] In this step, a distributed formation observer is established using the sign function and the predefined-time lemma as follows:

[0115]

[0116] where and are the position and velocity states of observer \(i\) respectively where \(m\in(0,1)\), \(\lambda min (\cdot)\) represents the minimum eigenvalue of the matrix, \(H\) is the information transfer matrix, \(T o >0\) is the preset convergence time. By constructing the Lyapunov function of the observer and taking its derivative, we can obtain which satisfies the form of the predefined-time lemma: Therefore, the constructed observer can converge within the predefined time. \(k 3 and \(k 4 are positive gain constants. \(w ij is the adjacency weight between the \(i\)-th and \(j\)-th followers, which is greater than zero if there is communication between \(i\) and \(j\), otherwise equal to zero, and \(b i is the connection weight between the virtual leader and the followers

[0117] Step 3: Using the global unified description model and the distributed predefined-time observer, reconstruct the augmented dynamics model of the air-ground unmanned cluster system, design an optimal formation controller based on the augmented dynamics model and construct the control gain matrix \(K i,κ .

[0118] According to Equation (7), rewrite the distributed predefined-time observer in the following form:

[0119]

[0120] where Z 0i =I 2 , \(I 2 is the 2D identity matrix

[0121]

[0122] where

[0123] and Interact with the position and velocity information of each UAV and its neighboring UAVs, respectively, C 0i = [1, 0].

[0124] Define as the formation tracking error, and combine the state χ in Eqs. (6) and (8) i,κ and Construct the augmented vector Thus, construct the augmented subsystem dynamics model of the air-ground UAV swarm system:

[0125]

[0126] where r = m + 2, where r is the dimension of the augmented vector X i,κ of, and p is the dimension of the state vector χ after the global unified description of the unmanned system i,κ of. and u i,κ is the control input of the augmented subsystem, where s is the dimension of the control input u of the augmented subsystem i,κ of. If κ ∈ {1, 2}, Otherwise:

[0127]

[0128] Design a model-based optimal formation controller

[0129]

[0130] where and are the optimal formation controllers, represents the optimal feedback gain matrix. can be obtained by solving the discounted factor algebraic Riccati equation where α i,κ is the discount factor to ensure the convergence of the controller.

[0131] Step 4: Apply a piecewise constant initial excitation to each augmented subsystem of the augmented dynamics model of the air-ground UAV swarm system, collect and store the system operation data of its previous operation, and then obtain the initial stable control strategy using a data-driven method based on the previous operation data of the air-ground unmanned system

[0132]

[0133] where only the integral of the previous system operation data is used to construct the data storage matrices and and They are respectively the pseudo-inverse of the data storage matrix and a set of basis vectors of the null space, which satisfy For any where is an appropriate matrix selected by the pole placement method.

[0134] Step 5: Perform off-policy iteration of reinforcement learning: Only use the integral of the system operation data to construct the data storage matrix, avoid using the Kronecker product, and further construct a non-singular data matrix with a smaller amount of data according to the data index; Use the system operation data in the non-singular data matrix to solve the generalized Sylvester-transposed data equation without the parameter matrix, and through iterative calculation, learn the optimal control gain matrix Furthermore, ensure the formation controller of the air-ground unmanned cluster system is optimal and reduce the computational burden in the learning process of the optimal formation controller.

[0135] In the embodiment of the present invention, the specific process of the off-policy iteration learning algorithm based on efficient reinforcement learning is as follows:

[0136] S501. Define the data storage matrix and where stores the system operation data associated with the state X i,κ during the operation of the augmented system, stores the system operation data associated with the input u i,κ during the operation of the augmented system, stores the system operation data associated with the state differential during the operation of the augmented system.

[0137]

[0138] where the time interval [t 0 , t l represents the sampling time for data collection, where lΔT = t l - t 0 , l ≥ (r + 1)s + r represents the length of the collected data, ΔT > 0 represents the sampling step, t 0 is the initial moment when the system collects data, t 1 is the moment after the sampling step at the initial data collection moment t 0 of the system, and so on, t l is the termination moment when the system collects data. respectively represent the k-th augmented subsystem of the i-th follower in the corresponding air-ground unmanned cluster system in the sub-interval [t 0 , t 1 , [t1 ,t 2 ,...,[t l-1 ,t l The system state data collected within respectively represents the system input data collected within the sub - intervals [t 0 ,t 1 ,[t 1 ,t 2 ,...,[t l-1 ,t l of the κ - th augmented subsystem of the i - th follower in the corresponding unmanned aerial - ground cluster system, respectively represents the integral data of the system state collected within the sub - intervals [t 0 ,t 1 ,[t 1 ,t 2 ,...,[t l-1 ,t l of the κ - th augmented subsystem of the i - th follower in the corresponding unmanned aerial - ground cluster system, where

[0139] S502. Select a set of data with indices σ = {1, 2, …, r + s} from and to construct a data storage matrix such that is a non - singular matrix.

[0140] S503. Policy evaluation: Relying on the system operation data in the data storage unit, perform iterative solution according to the following data - driven generalized Sylvester transpose matrix equation

[0141]

[0142] where, The superscript σ represents the data matrix reconstructed according to the indices. For each iteration update, the superscripts k and k + 1 represent the data corresponding to the k - th and (k + 1) - th iterations, where, The matrix is a block sub - matrix and is a positive - definite matrix, is a block sub - matrix of the matrix represents the controller gain matrix of the κ - th augmented subsystem of the i - th follower in the unmanned aerial - ground cluster system at the (k + 1) - th iteration,

[0143] S504. Policy iteration: Use Update control gain matrix

[0144] S505. Judgment whether it holds, ∈ is a small constant value greater than zero;

[0145] ③ If holds, stop the iteration. At this time, the control gain is The controller is Go to step S506;

[0146] ④ If does not hold, then k = k + 1, and then go back to S503 and S504 to continue the solution;

[0147] S506. Obtain the optimal control gain matrix At the same time, obtain the optimal formation controller

[0148] The above specific embodiments only describe the design principle of the present invention and are not used to limit the protection scope of the present invention. Therefore, those skilled in the art of the present invention can modify or equivalently replace the technical solutions recorded in the foregoing embodiments; and these modifications and replacements do not deviate from the gist and technical solutions of the present invention, and shall all fall within the protection scope of the present invention.

Claims

1. A reinforcement learning driven air-to-ground unmanned cluster system collaborative formation control method, characterized in that: The air-to-ground unmanned swarm system includes a virtual leader and followers consisting of M unmanned aerial vehicles and NM unmanned vehicles; the method includes the following steps: Step 1: Establish a dynamic model of the air-to-ground unmanned swarm system, including the dynamic models of unmanned aerial vehicles, unmanned vehicles, the dynamic model of the virtual leader, and the expected formation, and convert it into a global unified description model; wherein the dynamic models of the unmanned aerial vehicles and unmanned vehicles are decoupled into a linear model including a parameter matrix and a system state through a linear feedback method; and the system parameters that are easily affected during the movement of the air-to-ground unmanned swarm system are designed into the parameter matrix; Step 2: Construct a distributed predefined time observer and use it to estimate the virtual leader state and formation information to provide a trajectory reference for the formation tracking of the air-to-ground unmanned swarm system; the convergence time of the distributed predefined time observer is set to an accurate predefined time T that is independent of the observer parameters. o ; Step 3: Using the global unified description model and the distributed predefined time observer, reconstruct the augmented dynamics model of the air-to-ground unmanned swarm system, design the optimal formation controller based on the augmented dynamics model and construct the control gain matrix K i,κ ; Step 4: Apply the piecewise constant initial excitation to each augmented subsystem of the augmented dynamics model, and collect the system operation data of the previous operation; according to the system operation data of the previous operation, use the data-driven method to obtain the initial stable control gain matrix Step 5: Perform offline reinforcement learning strategy iteration: construct a data storage matrix using only the integral of the system operation data, and further construct a non-singular data matrix according to the data index; use the system operation data in the non-singular data matrix to solve the generalized Sylvester-transposed data equation that does not contain a parameter matrix, and learn the optimal control gain matrix through iterative calculation Then the optimal formation controller of the air-to-ground unmanned swarm system is obtained 2. The reinforcement learning driven air-to-ground unmanned cluster system collaborative formation control method according to claim 1, characterized in that: In step one, the system parameters that are easily affected during the movement of the air-to-ground unmanned cluster system include: the aerodynamic drag coefficient of the UAV, the mass of the UAV, the moment of inertia of the UAV, the control gain of the UAV autopilot and the mass of the unmanned vehicle.

3. The reinforcement learning driven air-to-ground unmanned cluster system collaborative formation control method as claimed in claim 1 or 2, characterized in that: In step 1, the dynamic model of establishing the air-to-ground unmanned cluster system is: Taking drones and unmanned vehicles as followers and sorting them uniformly, the dynamic model of the i-th follower in the air-to-ground unmanned cluster system is as follows: (1) The UAV dynamics model is: Among them, p i =[p i,x ,p i,y ,p i,z ] T and θ i =[φ i ,θ i ,ψ i ] T Respectively represent the 3D position and attitude of the i-th UAV; φ i is the roll angle, θ i is the pitch angle, ψ i is the yaw angle; and They are p i and θ i The second derivative of k i,x ,k i,y and k i,z is the aerodynamic drag coefficient of UAV i, m qi For drones i The mass of g is the acceleration due to gravity; b is the mass of g qi,φ =(l / J qi,φ ),b qi,θ =(l / J qi,θ ) and b qi,ψ =(1 / J qi,ψ ), where l is the arm length of the drone; J qi,φ ,J qi,θ and J qi,ψ is the moment of inertia; and is the control gain of the UAV autopilot; and is the desired translation speed and yaw rate; define and As the control input of UAV i, the state vector of the UAV is defined as Then the dynamic model of UAV i can be re-expressed as a linear model: Among them, the parameter matrix is ​​A ai =diag(A i,1 ,A i,2 ,A i,3 ,A i,4 ),B ai =diag(B i,1 ,B i,2 ,B i,3 ,B i,4 ), the control input is u ai =[u i,1 ,u i,2 ,u i,3 ,u i,4 ] T ; Parameter matrix A i,1 ,B i,1 ,A i,2 ,B i,2 ,A i,3 ,B i,3 and A i,4 ,B i,4 Corresponding to the p of drone i i,x ,p i,y ,p i,z 3D position and ψ i The parameter matrix of the subsystem related to the yaw angle contains the system parameters that are easily affected during the movement of the air-to-ground unmanned cluster system; The parameter matrix is: (2) Unmanned vehicle dynamics model J gi oh gi =t gi , Among them, p i =[p i,x ,p i,y ] T and v i =[v i,x ,v i,y ] T They represent the position and speed of the unmanned vehicle i in the two-dimensional plane, ψ gi and ω gi are the yaw angle and yaw rate of unmanned vehicle i respectively; F gi,x and F gi,y are the total forces on the position and direction of the unmanned vehicle i, C A,i and C f,i are the aerodynamic drag coefficient and rolling friction coefficient of unmanned vehicle i respectively; m gi is the mass of the unmanned vehicle i; Assume that the unmanned vehicle moves at a small angle ψ gi ≈0 and ω gi ≈0, and using feedback linearization techniques, the control input of the unmanned vehicle is defined as and u i,6 =F gi,y ; Define the state vector of the unmanned vehicle as x gi =[p i,x ,v i,x ,p i,y ,v i,y ] T , then the dynamic model of the unmanned vehicle can be re-expressed as a linear model: Among them, the parameter matrix is ​​A gi =diag(A i,5 ,A i,6 ),B gi =diag(B i,5 ,B i,6 ), the control input is u gi =[u i,5 ,u i,6 ] T ; (3) Dynamic model of virtual leaders: Where p0 = [x 0,x ,x 0,y ,x 0,z ,x 0,ψ ] T 、v0=[v 0,x ,v 0,y ,v 0,z ,v 0,ψ ] T and a0=[a 0,x ,a 0,y ,a 0,z ,a 0,ψ ] T represent the position, velocity and acceleration of the virtual leader respectively; (4) Expected formation H ij =H i0 -H j0 , Among them, H i,0 =[h i0,x ,h i0,y ,h i0,z ] T , H i0 and H j0 are the expected state deviations of the ith and jth follower and virtual leader, respectively; Then, based on the dynamic models of drones and unmanned vehicles, the dynamic model of the virtual leader, and the expected formation, a global unified description model is constructed: Among them, i=1,2,...,N,κ=1,2,...,6, l=x,y,z;χ i,κ ,u i,κ and i,κ Respectively represent the status, control input and output of each subsystem of the air-to-ground UAV cluster system; A i,κ ,B i,κ and C i,κ It is a global unified description of the dynamical matrix of the model.

4. The reinforcement learning driven air-to-ground unmanned cluster system collaborative formation control method as claimed in claim 3, characterized in that: In step 2, the distributed predefined time observer constructed is: in, and are the position and velocity states of observer i respectively; where m∈(0,1), λ min (·) represents the minimum eigenvalue of the matrix, is the information transfer matrix, T o >0 is the preset convergence time; Constructing the Lyapunov function of the observer and taking its derivative, we can get It satisfies the predefined time lemma form: Therefore, the constructed observer can converge within a predefined time; k3 and k4 are positive gain constants; w ij is the adjacency weight between the ith and jth followers, which is greater than zero if there is communication between the ijs, otherwise it is equal to zero. i is the connection weight between the virtual leader and the follower.

5. The reinforcement learning driven air-to-ground unmanned cluster system collaborative formation control method as claimed in claim 4, characterized in that: In step 3, the augmented dynamics model is constructed as: The distributed predefined time observer is rewritten as follows: in, WITH 0i =I2; in, and The position and speed information of each UAV interacts with its neighboring UAVs, C 0i =[1,0]; I2 is the two-dimensional unit matrix; definition As the formation tracking error, combined with the state χ in equation (I) and equation (II) i,κ and Constructing augmented vectors Thus, the augmented subsystem dynamics model of the air-to-ground UAV cluster system is constructed: in, r = p + 2, where r is the augmented vector X i,κ The dimension of p is the state vector χ after the global unified description of the unmanned system. i,κ The dimension of and u i,κ is the control input of the augmented subsystem, where s is the control input u of the augmented subsystem i,κ dimension; if κ∈{1,2}, otherwise 6. The reinforcement learning driven air-to-ground unmanned cluster system collaborative formation control method according to claim 5, wherein the optimal formation controller is obtained by solving the discounted factor algebraic Riccati equation; the optimal formation controller is expressed as: in, and is the optimal formation controller, represents the optimal control gain matrix; By solving the discounted factor algebraic Riccati equation We get, where α i,κ is the discount factor used to ensure the convergence of the controller.

7. The reinforcement learning driven air-to-ground unmanned cluster system collaborative formation control method as claimed in claim 1, characterized in that: In step 4, the initial stable control gain matrix Obtained through a data-driven approach: Among them, only the integral of the previous system operation data is used to construct the data storage matrix and and The data storage matrix The pseudo-inverse and the set of basis vectors of the null space are given by Its satisfaction in, is the matrix selected by the pole placement method.

8. The reinforcement learning driven air-to-ground unmanned cluster system collaborative formation control method as claimed in claim 4, characterized in that: The step five is: S501. Define data storage matrix and in, The data stored in the augmented system is related to the state of the system during operation. Related system operation data, The data stored in the augmented system are related to the input u i,κ Related system operation data, What is stored in the augmented system is the state differential during operation associated system operation data; Among them, the time interval [t0,t l ] represents the sampling time of data collection, where lΔT = t l -t0, l≥(r+1)s+r represents the length of the collected data, ΔT>0 represents the sampling step, t0 is the initial time when the system collects data, t1 is the time when the system passes the sampling step at the initial data collection time t0, and so on. l The end time of data collection for the system; They represent the k-th augmented subsystem of the ith follower of the corresponding air-ground unmanned cluster system in the subinterval [t0, t1], [t1, t2], ..., [t l-1 ,t l ] system status data collected within; They represent the k-th augmented subsystem of the ith follower of the corresponding air-ground unmanned cluster system in the subinterval [t0, t1], [t1, t2], ..., [t l-1 ,t l ] system input data collected within; They represent the k-th augmented subsystem of the ith follower of the corresponding air-ground unmanned cluster system in the subinterval [t0, t1], [t1, t2], ..., [t l-1 ,t l ] collected system status The integral data of express In [t l-1 ,t l ]; S502. and Select a set of data with index σ={1,2,…,r+s} to construct a data storage matrix Make is a non-singular matrix; S503. Strategy evaluation: Relying on the system operation data in the data storage unit, iteratively solves the following data-driven generalized Sylvester transposed matrix equation: in, The superscript σ indicates the reconstructed data matrix according to the index; Each iteration is updated, and the superscripts k and k+1 represent the data corresponding to the kth and k+1th iterations, where For the matrix The block submatrix of is positive definite, For the matrix The block submatrix of represents the controller gain matrix of the κ augmented subsystems of the ith follower in the space-time unmanned swarm system at the kth iteration, S504. Policy iteration: Use Update the control gain matrix S505. Judgment Is it true? ∈ is a small constant value greater than zero; ①If If it holds, the iteration stops, and the control gain matrix is The formation controller is Go to step S506; ②If If it is not true, then let k increase by 1, and then go to S503 and S504 to continue solving; S506. Obtaining the optimal control gain matrix At the same time, get the optimal formation controller

Citation Information

Patent Citations

  • Multi-unmanned-aerial-vehicle and multi-unmanned-ship inspection control system based on reinforcement learning

    CN113671994A

  • Heterogeneous cluster system robust output formation tracking control method and system

    CN113900380A

  • Heterogeneous cluster unmanned system event triggering cooperative control method based on reinforcement learning

    CN116430899A

  • Unmanned system air-ground collaborative navigation and obstacle avoidance method based on vision

    CN116540784A

  • Air-ground cooperative routing method based on reinforcement learning

    CN116939761A

Cited By

  • Distributed security formation control method for cross-domain cluster system in urban interference environment

    CN120686895A

  • Data-driven distributed learning control method for multi-vehicle network system

    CN121523009A

  • Spherical formation tracking optimization control method based on dual execution-evaluation network

    CN121832593A

  • Unmanned vehicle cluster fixed time distributed cooperative positioning method under positioning fault

    CN121977533A