Output synchronization optimization control method for unmanned swarm systems with unknown internal states
By building the topological structure and Laplace connection matrix of the unmanned cluster system, using the Actor-Critic network structure of the state estimator and Q learning algorithm, the control strategy of the unmanned cluster system is optimized, and the optimization problem in the unknown internal state is solved, and the stability and convergence of the system are improved.
Patent Information
- Application Number
- CN202211488163.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-25
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-11-25
AI Technical Summary
The optimization control method of existing unmanned cluster systems is rarely considered in the case of unknown internal state, and the existing methods are prone to system instability and energy consumption problems.
By building the topological structure and Laplace connection matrix of the unmanned cluster system, the internal state is estimated using a state estimator, the leader and follower drones are divided, and the Actor-Critic network structure in the Q learning algorithm is used to approximate the control action and performance functions, and combined with experience playback and target network technology to optimize the control strategy.
It improves the stability and convergence of the system, reduces the interactive resource consumption between the drone and the environment, and realizes optimized control under unknown internal states.
Smart Images

Figure CN115903901B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of unmanned aerial vehicle (UAV) control methods, and in particular to an output synchronization optimization control method for an unmanned cluster system with unknown internal states. Background Art
[0002] In recent years, inspired by the clustering behavior of organisms in nature, experts and scholars have applied the consistency of unmanned cluster systems to the collaborative control of complex systems. The consistency problem of unmanned cluster systems has important application prospects in the fields of smart grids, formation control, and drone clusters.
[0003] Consistency is a fundamental issue for unmanned swarm systems. All drones achieve the same state through information exchange under a consistent control protocol. Existing work uses measurable system input / output data to reconstruct unknown internal states. This reconstruction requires an augmented matrix, which is not fully controllable and can easily lead to tracking errors. Furthermore, existing work has rarely considered the energy consumption of unmanned swarm systems during mission execution. Summary of the Invention
[0004] The technical problem to be solved by the present invention is: how to solve the optimization problem of unmanned cluster systems when the internal state is unknown, which is rarely considered in existing research work. However, in practice, on the one hand, the internal state of drones is not easy to measure, and on the other hand, due to the limited computing power and storage of drones and the complexity of tasks, a method for output synchronization optimization control of unmanned cluster systems with unknown internal states is considered.
[0005] The present invention solves the above technical problems through the following technical solutions, which include the following steps:
[0006] S1: Constructing the topology structure and Laplace connection matrix of the unmanned swarm system based on the connection status between each drone in the unmanned swarm system;
[0007] S2: Estimate the unknown internal state of the unmanned swarm system through a state estimator, divide the UAVs into leader UAVs and follower UAVs, reconstruct the local state error system of the UAVs, and define a performance function;
[0008] S3: The Actor-Critic network structure in the Q-learning algorithm is used to approximate the control action and performance function of the UAV respectively. The Critic network is used to approximate the performance function, and the Actor network updates the control action of the UAV according to the performance function.
[0009] S4: The Critic network evaluates the approximate control actions of the Actor network, and the Actor network adjusts the control actions based on the Critic network's evaluation. The entire process is updated using gradient descent. The experience replay strategy and target network technology are used when training the neural network parameters. When the neural network parameters of the Actor-Critic network structure are no longer updated, an approximate optimized output consistent control strategy is obtained.
[0010] Compared with the prior art, the present invention has the following advantages:
[0011] 1. This invention proposes a new method for obtaining internal states. Specifically, it designs a state estimator based on an output feedback mechanism. Compared with existing internal state reconstruction methods, this method does not require augmented matrices, which can lead to system instability. This state estimator can improve system stability.
[0012] 2. Based on the estimated internal state, the present invention proposes an equivalent system model to characterize the dynamic characteristics of the local output synchronization error, which is conducive to optimal control.
[0013] 3. The unmanned cluster system of the present invention is a system whose precise model is unknown. Compared with traditional optimization methods that require the precise model to be known, in many practical situations, the precise model of the system is unknown or difficult to obtain. The Actor-Critic framework used in the present invention can better solve the situation where the precise model of the system is unknown.
[0014] 4. The experience replay strategy and target network technology proposed in this invention enable full interaction between the drone and the environment while avoiding the use of incentive conditions. Simulation results have verified this. The proposed method of increasing the experience pool and target network technology can effectively enhance the drone's exploration of the environment, ultimately improving the system's convergence. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is the overall flow chart of the present invention;
[0016] Figure 2 It is a topological diagram that may appear during the system convergence process of the present invention;
[0017] Figure 3 This is the evolution diagram of the UAV tracking error of the comparative experiment of the present invention;
[0018] Figure 4 This is a diagram showing the evolution of the UAV tracking trajectory in the comparative experiment of the present invention;
[0019] Figure 5 This is the approximate optimization control evolution diagram of the comparative experiment of the present invention;
[0020] Figure 6 This is the evolution diagram of the UAV tracking error of the present invention;
[0021] Figure 7 This is the UAV tracking trajectory evolution diagram of the present invention;
[0022] Figure 8 This is the approximate optimization control evolution diagram of the present invention;
[0023] Figure 9 This is a relationship diagram among Algorithm 1, Algorithm 2, and Algorithm 3 in the present invention. DETAILED DESCRIPTION
[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0025] Figure 1 It is the overall flow chart of the present invention; Figure 1 As shown, this embodiment provides an output synchronization optimization control method for an unmanned cluster system with unknown internal state. This method is an anti-synchronization optimization control method for an unmanned cluster system under a cooperative and competitive relationship. The method includes but is not limited to the following steps:
[0026] S1: According to the connection status between each drone in the unmanned cluster system, the topological structure and Laplace connection matrix of the unmanned cluster system are constructed; in the embodiment of the present invention, each drone in the unmanned cluster system is connected in a certain way. Figure 2 The following is an example topology diagram of a five-node leader-follower swarm system, where nodes 1, 2, 3, and 4 represent follower drones and node 0 represents the leader drone. Note that follower drone 4 obtains information from leader drone 0.
[0027] Therefore, the relevant parameters of the Laplace connection matrix in the present invention are:
[0028] The topology diagram is represented as: G=diag{1,1,0,0};
[0029] The connectivity matrix is expressed as:
[0030] The Laplace matrix is expressed as:
[0031] S2: Estimate the unknown internal state of the unmanned swarm system through a state estimator, divide the UAVs into leader UAVs and follower UAVs, reconstruct the local state error system of the UAVs, and define a performance function;
[0032] In the embodiment of the present invention, the unmanned swarm system includes a leader and follower mode, each drone contains its own status information, x i (k) represents the status information of the i-th UAV, x j (k) represents the state information of the j-th UAV; through the above method, the leader UAV and the follower UAV are updated using the corresponding leader dynamic equation and follower dynamic equation, which are expressed as:
[0033] The dynamic equation of the leader UAV is expressed as:
[0034]
[0035] The dynamic equation of the follower drone is expressed as:
[0036]
[0037] Among them, x0(k+1) represents the state value of the leader drone at time k+1, x0(k) represents the state value of the leader drone at time k, y0(k) represents the control output of the leader drone at time k, and y i (k) represents the control output of follower drone i at time k, x i (k+1) represents the state value of follower drone i at time k+1, x i (k) represents the state value of follower drone i at time k, μ i (k) represents the control input of follower drone i at time k, A, B i , C are different unknown constant matrices. These constant matrices have certain appropriate dimensions, but their specific values are unknown.
[0038] In the embodiment of the present invention, it is also necessary to reconstruct the local state error system of the UAV according to the leader dynamic equation and the follower dynamic equation, which can be expressed as:
[0039]
[0040] in, represents the local state error system of follower UAV i at time k, b i Indicates whether the follower drone can receive the status information of the leader drone, b i =1 means the follower drone can receive the information of the leader drone, b i= 0 means that the follower UAV cannot receive the information of the leader UAV; x0(k) represents the status information of the leader UAV, a ij ≥0 indicates that follower UAV i receives the status information of follower UAV j, a ij >0 means that follower drone i can receive the status information of follower drone j, a ij = 0 means that follower drone i cannot receive the status information of follower drone j; N i represents a collection of follower drones;
[0041] Therefore, the local state error system is equivalent to:
[0042]
[0043] Where W is the observation gain matrix; represents the difference between the state observation value and the true value of follower drone i at time k+1, A, B i ,C is a matrix of different unknown constants; represents the local error system based on the reconstructed internal state value, μ i (k) represents the control input of follower drone i at time k, e iy (k) represents the reconstruction output error of follower UAV i at time k based on internal reconstruction, N represents the difference between the output error and the reconstruction error of follower UAV i at time k. i represents the set of neighbor follower drones of follower drone i; represents the state observation value of follower drone i at time k, represents the state observation value of follower UAV j at time k.
[0044] In an embodiment of the present invention, a performance function is further determined based on the local state error system of the UAV, which is expressed as:
[0045]
[0046] in, The performance function of follower drone i at time k, c i (e iy (k),μ i (k)) represents the control strategy μ that the follower UAV i makes during its interaction with the environment at time k. i (k),In this process, the follower drone obtains the specific performance consumption value through the built-in device,, Q i represents the weight matrix of follower drone i, Q i ≥0; R irepresents the symmetric matrix of follower drone i, R i >0; 0<γ<1 is the discount factor.
[0047] S3: The Actor-Critic network structure in the Q-learning algorithm is used to approximate the control action and performance function of the UAV respectively. The Critic network is used to approximate the performance function, and the Actor network updates the control action of the UAV according to the performance function.
[0048] In the embodiment of the invention, a Critic network is used to approximate the performance function, and an Actor network updates the control action according to the performance function;
[0049] Neural network approximation of performance function:
[0050] Update the control action according to the performance function:
[0051] in, represents the performance function of follower drone i at time k, W ci,now represents the current weight parameter of follower drone i in the critic network, f(·)=tanh(·) represents the activation function, z ci (k) represents the input vector in the critic network containing the action information and related position information of follower drone i and its neighbor drones; the superscript T represents transposition; W ai,now Represents the current weight parameter of follower drone i in the Actor network, represents the approximation of the control action of follower drone i at time k through the Actor network.
[0052] S4: The Critic network evaluates the approximate control actions of the Actor network, and the Actor network adjusts the control actions based on the Critic network's evaluation. The entire process is updated using gradient descent. The experience replay strategy and target network technology are used when training the neural network parameters. When the neural network parameters of the Actor-Critic network structure are no longer updated, an approximate optimized output consistent control strategy is obtained.
[0053] In an embodiment of the present invention, the Critic network evaluates the approximate control action of the Actor network, including the Critic network evaluating the quality of the drone control action through the output value of the performance function, using the network approximate structure to obtain the approximate performance function, using the Bellman equation to obtain the Bellman performance function, and using the difference function to obtain the differential performance function between the approximate performance function and the Bellman performance function; minimizing the differential performance function, and using the gradient descent method to train and adjust the neural network parameters of the Critic network.
[0054] Among them, for the Critic network:
[0055] The critic network evaluates the quality of the drone's control actions through the output value of the performance function. The performance function is approximated by the following network structure:
[0056]
[0057] in is the input vector of the Critic network containing the action information and related position information of drone i and its neighbors, μ -i (k) represents the neighbor control input of UAV i, f(·) = tanh(·) represents the activation function;
[0058] The Bellman performance function is derived using the Bellman equation, which is obtained from the Bellman equation:
[0059]
[0060] in and Use critic network and target critic network to approximate respectively, and the neural network parameters are W ci,now and W ci,now- ;
[0061] The differential performance function between the approximate performance function and the Bellman performance function is obtained by using the differential function, which is expressed as:
[0062]
[0063] The goal is to train the Critic network so that the function Minimum, here the gradient descent method is used to adjust the neural network parameters, so the weight update of the Critic network is as follows:
[0064]
[0065] Among them, W ci,new represents the updated weight parameter of follower drone i in the Critic network, W ci,now represents the current weight parameter of follower drone i in the Critic network, β ci represents the learning rate of follower drone i in the Critic network, e ci (k) represents the local state error system of follower drone i at time k in the critic network, f(·) = tanh(·) represents the activation function, z ci(k) represents the input vector in the Critic network containing the action information and related position information of follower drone i and its neighbor drones.
[0066] In an embodiment of the present invention, the Actor network uses local state error system information including the follower UAV itself and its neighboring UAVs to approximate the control action; uses an approximate performance function to approximate the control action; calculates the difference loss between the approximate performance function and the desired final consumption target, minimizes the loss difference, and uses the gradient descent method to train and adjust the neural network parameters of the Actor network.
[0067] Among them, for the Actor network:
[0068] Actor networks are used to approximate control strategies, which are represented as follows:
[0069]
[0070] in, is the input of the Actor network containing information about agent i and its neighbors, U i The desired final consumption cost target is that after the system reaches consistency, no additional control or consumption is required. Therefore, the error of the Actor network can be described as:
[0071]
[0072] Update the network parameters using gradient descent so that The function is minimal.
[0073] Therefore, the network weight update of the Actor network is expressed as:
[0074]
[0075] Among them, W ai,new represents the update weight parameter of follower drone i in the Actor network, W ai,now represents the current weight parameter of follower drone i in the Actor network, β ai represents the learning rate of follower drone i in the Actor network, represents the local state error system of follower drone i at time k in the Actor network, represents f(z ci (k)) About z ci The partial derivative of (k).
[0076] In the embodiment of the present invention, the control strategy update process in the Q learning algorithm, hereinafter referred to as Algorithm 1, is processed in the following manner. The specific update method is as follows:
[0077] Step 1) Initialize the processing of any follower drone and adjust the iteration index l and Q function value And the parameter ε is initialized;
[0078] Step 2) Update the Q function value using dynamic programming, expressed as:
[0079]
[0080] Step 3) Update the control strategy as follows:
[0081]
[0082] Step 4) If the norm of the difference between the two updated Q values is less than a given smaller parameter ε, that is, the formula If it holds, terminate the iteration. Otherwise, set l+l+1 and repeat steps 2 and 3.
[0083] In this embodiment of the present invention, an experience replay strategy is also used to train network parameters, hereinafter referred to as Algorithm 2, which specifically includes the following process:
[0084] Step 1) Initialize the capacity D of the experience pool c
[0085] Step 2) Store the quadruple in the experience pool:
[0086] Step 3) For the stored ith quadruple, if the number of stored quadruple is greater than the capacity of the experience pool, delete the first stored tuple; otherwise, directly
[0087] Randomly select a tuple from the experience pool for subsequent updates.
[0088] In the embodiment of the present invention, the present invention adopts the following method to train the improved Deep Q-learning algorithm (hereinafter referred to as Algorithm 3), which includes the following steps: training the modified Deep Q-learning algorithm
[0089] Step 1: Initialization: parameters ε, discount factor γ, learning rate β ai and β ci , hyperparameter τ. and The network weight of
[0090] Step 2: Calculate the difference of the Critic network
[0091]
[0092] Step 3: Update the parameters of the Critic network using gradient descent
[0093]
[0094] Step 4: Update the parameters of the Actor network using gradient descent
[0095]
[0096] Step 5: τ is a hyperparameter that needs to be manually adjusted to perform weighted averaging to update the parameters of the target network:
[0097]
[0098] Step 6: If the norm of the difference between two updates of the neural network parameters is less than the predetermined smaller parameter ε, the system is considered to have converged and the iteration is terminated.
[0099] Step 7: Otherwise, return to step 2
[0100] To make the update process more intuitive, use Figure 9 To verify the effectiveness of the proposed optimal output synchronization control method with unknown internal state, Matlab is used for simulation verification. Figure 2 This is the experimental topology diagram of a leader-follower drone swarm system consisting of 5 nodes, where nodes 1, 2, 3, and 4 represent follower drones and node 0 represents the leader drone. It is worth noting that agent 4 obtains information from the leader drone 0.
[0101] The relevant parameters in the present invention are:
[0102] The topological structure is represented as: G = diag{1,1,0,0};
[0103] Connection Matrix:
[0104] Laplacian matrix:
[0105] The system-related parameters are:
[0106]
[0107] Q 11 =Q 22 =Q 33 =Q 44 =I 2×2 ,R 11 =R 12 =R 14 =R 21 =R 22 =R33 =R 34 =R 41 =R 43 =R 44 =1,
[0108] R 13 =R 23 =R 24 =R 31 =R 32 =R 42 =0, learning rate β ai =β ci =0.05, discount rate γ=0.95.
[0109] From the simulation results, we can see that Figure 3 and Figure 4 The error convergence graph and trajectory evolution graph of the unmanned swarm system are shown respectively. Then, we can conclude that the unmanned swarm system finally achieves consistency. The evolution of the approximate optimization controller is as follows Figure 5 As shown, the controller oscillates strongly before reaching convergence.
[0110] In order to further verify the advantages of the present invention, the same UAV dynamic system, topology, initial value of the system state, critic weight, actor weight and other relevant parameters as those in the comparative experiment can be used.
[0111] Will Figure 3 and Figure 6 、 Figure 4 and Figure 7 Comparing under the same parameters, we found Figure 6 and Figure 7 The convergence speed of the state is faster, which means that the algorithm proposed in this invention can improve the convergence speed of the unmanned cluster system. In order to further illustrate the advantages of the proposed algorithm, Figure 5 and Figure 8 Comparative simulation results are given to describe the evolution of the controller. It is clear that Figure 8 The curve in is more stable and converges faster. This shows that the algorithm can generate a better performance controller.
[0112] Previous studies have shown that choosing inappropriate cooperation and competition intensity parameters can lead to instability in unmanned cluster systems. Therefore, this invention designs an adaptive cooperation and competition intensity function. When the unmanned cluster system with cooperative and competitive interaction finally reaches anti-synchronous consistency, the corresponding cooperation intensity and competition intensity parameters converge to the optimal value (see Figure 9 ), without the need to manually adjust the coordination parameters, ensuring the stability of the system and improving the robustness of the system.
[0113] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware through a program, and the program can be stored in a computer-readable storage medium, which may include: ROM, RAM, disk or CD, etc.
[0114] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A method for output synchronization optimization control of an unmanned cluster system with unknown internal state, characterized in that: The following steps are involved: S1: Constructing the topology structure and Laplace connection matrix of the unmanned swarm system based on the connection status between each drone in the unmanned swarm system; S2: Estimate the unknown internal state of the unmanned swarm system through a state estimator, divide the UAVs into leader UAVs and follower UAVs, reconstruct the local state error system of the UAVs, and define a performance function; In step S2, the leader UAV and the follower UAV are updated using the corresponding leader dynamic equation and follower dynamic equation; The local state error system of the UAV is reconstructed according to the leader dynamic equation and the follower dynamic equation, and a performance function is determined based on the local state error system of the UAV, where: The dynamic equation of the leader UAV is expressed as: The dynamic equation of the follower drone is expressed as: Among them, x0(k+1) represents the state value of the leader drone at time k+1, x0(k) represents the state value of the leader drone at time k, y0(k) represents the control output of the leader drone at time k, and x i (k+1) represents the state value of follower drone i at time k+1, y i (k) represents the control output of follower drone i at time k, x i (k) represents the state value of follower drone i at time k, μ i (k) represents the control input of follower drone i at time k, A, B i ,C is a matrix of different unknown constants; The local state error system of the UAV is reconstructed as: in, represents the local state error system of follower UAV i at time k, b i Indicates whether the follower drone can receive the status information of the leader drone, b i =1 means the follower drone can receive the information of the leader drone, b i =0 means that the follower drone cannot receive information from the leader drone; a ij ≥0 indicates that follower UAV i receives the status information of follower UAV j, a ij >0 means that follower drone i can receive the status information of follower drone j, a ij = 0 means that follower drone i cannot receive the status information of follower drone j; N i represents the set of neighbor follower drones of follower drone i; represents the state observation value of follower drone i at time k, represents the state observation value of follower drone j at time k; The consumption performance function is expressed as: in, The consumption performance function of follower drone i at time k, c i (e iy (k),μ i (k)) represents the control strategy μ that the follower UAV i makes during its interaction with the environment at time k. i (k), Q i represents the weight matrix of follower drone i, Q i ≥0; R i represents the symmetric matrix of follower drone i, R i >0; 0<γ<1 is the discount factor; S3: The Actor-Critic network structure in the Q-learning algorithm is used to approximate the control action and performance function of the UAV respectively. The Critic network is used to approximate the performance function, and the Actor network updates the control action of the UAV according to the performance function. S4: The Critic network evaluates the approximate control actions of the Actor network, and the Actor network adjusts the control actions based on the Critic network's evaluation. The entire process is updated using gradient descent. The experience replay strategy and target network technology are used when training the neural network parameters. When the neural network parameters of the Actor-Critic network structure are no longer updated, an approximate optimized output consistent control strategy is obtained.
2. The method for output synchronization optimization control of an unmanned cluster system with unknown internal state according to claim 1, characterized in that: The state estimator in step S2 is as follows: Where W is the observation gain matrix; represents the state observation value of follower drone i at time k+1, A, B i ,C is a matrix of different unknown constants; The state observation value of follower drone i at time k, μ i (k) represents the control input of follower UAV i at time k, y i (k) represents the control output of follower UAV i at time k, represents the control output observation value of follower UAV i at time k.
3. The method for output synchronization optimization control of an unmanned cluster system with unknown internal state according to claim 1, characterized in that: In step S4, the Critic network evaluates the control action approximated by the Actor network, including the Critic network evaluating the quality of the UAV control action through the output value of the performance function, using the network approximate structure to obtain the approximate performance function, using the Bellman equation to obtain the Bellman performance function, and using the difference function to obtain the differential performance function between the approximate performance function and the Bellman performance function; minimizing the differential performance function, and using the gradient descent method to train and adjust the neural network parameters of the Critic network.
4. The method for output synchronization optimization control of an unmanned cluster system with unknown internal state according to claim 3, characterized in that: In step S4, the neural network parameter update formula of the critic network is expressed as: Among them, W ci,new represents the updated weight parameter of follower drone i in the Critic network, W ci,now represents the current weight parameter of follower drone i in the Critic network, β ci represents the learning rate of follower drone i in the Critic network, e ci (k) represents the performance function differential error in the critic network, f(·) = tanh(·) represents the activation function, z ci (k) represents the input vector in the Critic network containing the action information and related position information of follower drone i and its neighbor drones.
5. The output synchronization optimization control method for an unmanned cluster system with unknown internal state according to claim 1, characterized in that: In step S4, the Actor network uses the local state error system information of the follower UAV itself and its neighboring UAVs to approximate the control action using an approximate performance function; calculates the difference loss between the approximate performance function and the desired final consumption target, and uses the gradient descent method to train and adjust the neural network parameters of the Actor network.
6. The method for output synchronization optimization control of an unmanned cluster system with unknown internal state according to claim 5, characterized in that: In step S4, the neural network parameter update formula of the Actor network is expressed as: Among them, W ai,new represents the update weight parameter of follower drone i in the Actor network, W ai,now represents the current weight parameter of follower drone i in the Actor network, β ai represents the learning rate of follower drone i in the Actor network, represents the local state error system of follower drone i at time k in the Actor network, represents f(z ci (k)) About z ci The partial derivative of (k); z ci (k) represents the input vector in the Critic network containing the action information and related position information of follower drone i and its neighbor drones.