Optimal two-part cooperative control method for intelligent unmanned cluster system

By building a dual Q learning network in an unmanned cluster system and using the ALR-TQDPG algorithm, the learning rate is dynamically adjusted to cope with the intensity of cooperation competition, and the impact of cooperative competition intensity in an unmanned cluster system on system performance is solved, and more efficient system collaborative control and resource utilization are achieved.

CN120065840AActive Publication Date: 2025-05-30CHONGQING UNIV OF POSTS & TELECOMM
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510202720.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-30
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the impact of the intensity of cooperative competition on system performance in unmanned cluster systems, and there is a "dimensional curse" problem in optimal control solution and an overestimation bias problem in Q learning.

Method used

A two-part collaborative control method for the optimal two-part collaborative control method of intelligent unmanned cluster systems is proposed. By building a dual Q learning network and using ALR-TQDPG algorithm to train the dual Q learning network, dynamically adjust the learning rate to cope with the intensity of cooperation and competition, and the coordinated control of the system is realized through the target network output control information.

Benefits of technology

Effectively balance the impact of cooperative competition intensity on system performance, improve the convergence speed of the system, reduce resource consumption, and reduce the deviation of Q value estimation to achieve better system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120065840A_ABST
    Figure CN120065840A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of intelligent unmanned cluster system control, and particularly relates to an optimal two-part cooperative control method for an intelligent unmanned cluster system. The method comprises the following steps: determining a topological structure of the system according to an interconnection relationship between unmanned aerial vehicles in the unmanned cluster system; based on a topological structure of the system, each unmanned aerial vehicle intelligent agent sends own state information to a neighbor intelligent agent, and calculates a local state error of each unmanned aerial vehicle; constructing a double-Q learning network, training the double-Q learning network by adopting an ALR-TQDPG algorithm according to the local state error until the intelligent unmanned cluster system is consistent, and outputting control information by the target network; the intelligent agent is controlled by adopting the control information output by the target network, so that cooperative control of the intelligent unmanned cluster system is realized; the method can cope with the influence of the cooperative competition intensity on the system performance; the problem of underestimation of the Q value caused by minimizing the Q value in the action selection process is reduced, and the method has a good application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent unmanned cluster system control, and particularly relates to an optimal bipartite cooperative control method for an intelligent unmanned cluster system. Background Art

[0002] In recent years, inspired by the collective behavior of organisms in nature, experts and scholars have applied the consensus in unmanned cluster systems to the cooperative control of complex systems. The consensus problem of unmanned cluster systems has important application prospects in fields such as swarm behavior control, smart grid, satellite formation, and unmanned aerial vehicle clusters. This problem can be regarded as a basic phenomenon in unmanned cluster systems, that is, all agents reach the same state through information exchange under the guidance of the consensus control protocol.

[0003] In the past few years, the research on the consensus control of multi-agent systems has mostly focused on the assumption that agents are completely cooperative. However, in reality, there are often both cooperative and competitive relationships among agents. Existing research has begun to pay attention to the cooperative and competitive relationships among agents, but most only consider the existence of cooperation and competition, without further exploring the specific impacts brought by their intensities. In fact, it has been found that the cooperation-competition intensity not only affects the convergence speed of the system, but may even lead to system instability.

[0004] Some research has started to address the impact of the cooperation-competition intensity on system performance, designing corresponding algorithms to optimize the convergence speed and reduce energy consumption through techniques such as target networks and experience replay. However, this method requires dynamically adjusting the intensities of cooperation and competition, and in practical applications, dynamically adjusting the intensity is usually complex and difficult to implement. Therefore, seeking an effective method to cope with the impact of the cooperation-competition intensity on system performance has become an urgent problem to be solved.

[0005] In addition, when considering the consensus problem of unmanned cluster systems, it is also necessary to consider the resources consumed during the minimization process, that is, to achieve optimal consensus. The solution of optimal control depends on the Hamilton-Jacobi-Bellman (HJB) equation. The core problem of optimal control lies in the difficulty of analytically solving the HJB equation, which will cause the "curse of dimensionality" problem. The reinforcement learning method plays an important role in solving the HJB equation of optimal control. The deterministic policy gradient algorithm (DDPG) in reinforcement learning is applicable to continuous action and state spaces, uses experience replay to alleviate sample correlation, and introduces target networks and policy delayed updates to improve stability. The stability and performance of the algorithm are relatively good. However, DDPG belongs to Q-learning, and overestimation bias is a characteristic of Q-learning. The maximization of the noise value estimation will lead to consistent overestimation. How to solve the overestimation bias problem is the second problem to be solved. Summary of the Invention

[0006] In view of the deficiencies in the prior art, the present invention proposes an optimal bipartite cooperative control method for an intelligent unmanned cluster system, and the method includes:

[0007] S1: Determine the topological structure of the system according to the interconnection relationship among the unmanned aerial vehicles in the unmanned cluster system;

[0008] S2: Based on the topological structure of the system, each unmanned aerial vehicle agent sends its own state information to its neighbor agents, and calculates the local state error of each unmanned aerial vehicle;

[0009] S3: Construct a dual Q-learning network and, according to the local state error, train the dual Q-learning network using the ALR-TQDPG algorithm until the intelligent unmanned cluster system reaches consensus, and the target network outputs control information;

[0010] S4: Control the agents using the control information output by the target network to achieve the cooperative control of the intelligent unmanned cluster system.

[0011] Preferably, the formula for calculating the local state error is expressed as:

[0012]

[0013] where e i (k) represents the consensus error of the state information x i (k) at the k-th moment on the i-th follower agent, x i (k) represents the state information of the i-th follower agent at the k-th moment, x 0 (k) represents the state information of the leader agent at the k-th moment, c i represents the connection weight between the leader and agent i, g i represents the pinning control on agent i, a ij represents the connection weight from agent j to agent i, x j (k) represents the state information of the j-th follower agent at the k-th moment, N i represents the set of neighbor nodes of agent i, represents the value summation of the neighbor agent j of agent i, and sign() represents the sign function.

[0014] Preferably, the process of training the dual Q-learning network using the ALR-TQDPG algorithm includes:

[0015] S31: Initialize the parameters of the critic network, the actor network, and the target network; the actor network is the actor network; the critic network includes the critic1 network and the critic2 network; the target network includes the target networks corresponding to the critic1 network, the critic2, and the actor network;

[0016] S32: The actor network outputs the control information at the current moment, and the local state error of the UAV at the current moment, the local state error at the next moment, and the control information at the current moment are stored in the experience pool as experience information;

[0017] S33: Select experience information from the experience pool to calculate the Q value of the critic network;

[0018] S34: The target network outputs control information and calculates the Q value of the target network;

[0019] S35: Update the weight between the historical TD error and the cooperation and competition intensity;

[0020] S36: Calculate the TD error of the two critic networks according to the Q value of the target network; calculate the historical TD error according to the TD error of the critic network;

[0021] S37: Calculate the adaptive learning rate according to the historical TD error and the weight between the historical TD error and the cooperation and competition intensity; update the weights of the two critic networks according to the adaptive learning rate;

[0022] S38: Update the weights of the actor network according to the Q value of the critic network, and update the weights of the target network corresponding to the actor network according to the weights of the actor network;

[0023] S39: Update the weights of the target network corresponding to the critic network according to the weights of the two critic networks in the critic;

[0024] S310: Determine whether the intelligent unmanned cluster system reaches consensus. If it reaches consensus, the target network outputs control information; otherwise, return to step S32.

[0025] Further, the formula for updating the weight between the historical TD error and the cooperation and competition intensity is:

[0026] κ = tanh(b·ln(l + 1))

[0027] where κ represents the weight between the historical TD error and the cooperation and competition intensity, b represents a constant, and l represents the number of training iteration rounds.

[0028] Further, the formula for calculating the TD error of the two critic networks is:

[0029]

[0030] where represents the TD error of critic 1 training at time k, represents the TD error of critic2 training at time k, denotes the performance function, denotes the control input of agent i, denotes the Q function trained by the target network corresponding to the critic 1 network, denotes the Q function trained by the target network corresponding to the critic 2 network, denotes the Q function trained by critic 1, denotes the Q function trained by the critic 2 network; e i (k) represents the state information x at time k on the i-th follower agent i (k) represents the consensus error under the state information x i (k + 1) represents the state information x at time k + 1 on the i-th follower agent i (k) represents the consensus error under the state information x.

[0031] Furthermore, the formula for calculating the historical TD error is:

[0032]

[0033] where r1 i (k) represents the historical TD error trained by critic 1 at time k, r2 i (k) represents the historical TD error trained by critic 2 at time k, Θ represents the decay factor, r1 i (k - 1) represents the historical TD error trained by critic 1 at time k - 1, r2 i (k - 1) represents the historical TD error trained by critic 2 at time k - 1, represents the TD error trained by critic 1 at time k, represents the TD error trained by critic 2 at time k.

[0034] Furthermore, the formula for calculating the adaptive learning rate is:

[0035]

[0036] where β 1 ′ represents the adaptive learning rate of the critic 1 network, β 2 ′ represents the adaptive learning rate of the critic 2 network, β c represents the fixed learning rate, m is a constant, N i represents the neighbor set of agent i, a ij represents the connection weight from agent j to agent i, s ij represents the cooperation and competition intensity between agent i and agent j, k represents the weight between the historical TD error and the cooperation and competition intensity, r1i (k) represents the historical TD error of the training of critic 1 at time k, r2 i (k) represents the historical TD error of the training of critic 2 at time k.

[0037] Furthermore, the formula for determining that the intelligent unmanned cluster system reaches consensus is:

[0038]

[0039] Among them, represents the weight of critic 1 in the l-th iteration, represents the weight of critic 1 in the (l + 1)-th iteration, represents the weight of critic 2 in the l-th iteration, represents the weight of critic 2 in the (l + 1)-th iteration, and ε represents the stability threshold.

[0040] The beneficial effects of the present invention are as follows:

[0041] 1. The present invention proposes an adaptive learning rate adjustment formula, aiming to dynamically adjust the learning rate of the critic network according to the cooperation-competition intensity and TD error. This adjustment aims to balance the impact of the cooperation-competition intensity on the system performance, accelerate the convergence speed, and reduce the resources required to achieve optimal control.

[0042] 2. Since the cooperation-competition intensity affects the Q-value training of the critic network, the accuracy of Q-value estimation becomes crucial. To reduce the problem of underestimating the Q-value caused by minimizing the Q-value during the action selection process, the present invention adopts two critic networks, and each critic network corresponds to a target critic network. In each iterative update equation of the TD error, the larger one of the two target Q-values is selected for training.

[0043] 3. The DDPG algorithm is applicable to the continuous action space and can obtain more stable and higher-quality learning results. Therefore, the present invention adopts an improved DDPG algorithm, namely the ALR-TQDPG algorithm, which uses experience replay to eliminate the correlation of time-series data and uses the target network to enhance the stability of the training process. This method allows the unmanned cluster system to explore without continuous excitation conditions. Description of the Drawings

[0044] Figure 1 is the overall flowchart of the present invention;

[0045] Figure 2 is the topological graph that may appear during the system convergence process of the present invention;

[0046] Figure 3Evolution diagram of the tracking error e1 of the UAV system in the comparative experiment of the present invention;

[0047] Among them, (a) is the adaptive cooperation and competition intensity, and (b) is the algorithm of the present invention;

[0048] Figure 4 Evolution diagram of the tracking error e2 of the UAV system in the comparative experiment of the present invention;

[0049] Among them, (a) is the adaptive cooperation and competition intensity, and (b) is the algorithm of the present invention;

[0050] Figure 5 Evolution diagram of the tracking trajectory X1 of the UAV system in the comparative experiment of the present invention;

[0051] Among them, (a) is the adaptive cooperation and competition intensity, and (b) is the algorithm of the present invention;

[0052] Figure 6 Evolution diagram of the tracking trajectory X2 of the UAV system in the comparative experiment of the present invention;

[0053] Among them, (a) is the adaptive cooperation and competition intensity, and (b) is the algorithm of the present invention;

[0054] Figure 7 Evolution diagram of the performance consumption of the UAV system in the comparative experiment of the present invention;

[0055] Among them, (a) is the adaptive cooperation and competition intensity, and (b) is the algorithm of the present invention;

[0056] Figure 8 Evolution diagram of the weights of the actor network of the present invention;

[0057] Figure 9 Evolution diagram of the weights of the critic network of the present invention;

[0058] Figure 10 Evolution diagram of the adaptive learning rate in the present invention. Detailed implementation manners

[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0060] The present invention proposes an optimal bipartite cooperative control method for an intelligent unmanned cluster system, as Figure 1 shown, and the method includes the following contents:

[0061] S1: Determine the topological structure of the system according to the interconnection relationship among the drones in the unmanned cluster system.

[0062] In the specific application scenario of the present invention, the drones are connected to each other in a specific manner. For example, Figure 2 for illustration, the system is an unmanned cluster with a leader-follower architecture consisting of 5 nodes. Among them, nodes 1, 2, 3, and 4 play the role of follower drones, while node 0 exists as the leader drone. In particular, follower drone 1 and follower drone 2 are designed to receive information from leader drone 0.

[0063] Determine the topological structure of the system according to the interconnection relationship among the drones in the unmanned cluster system. For example, Figure 2 in it, the topological structure diagram is expressed as: G = diag{1, 1, 0, 0}.

[0064] The connection matrix is expressed as:

[0065] The Laplacian matrix is expressed as:

[0066] S2: Based on the topological structure of the system, each drone agent sends its own state information to its neighbor agents and calculates the local state error of each drone.

[0067] The unmanned cluster system of the present invention adopts a leader-follower mode, in which each drone maintains its own state information. Specifically, use symbols to represent the state information of a certain specific drone, and so on, which can represent the state information of any drone. To achieve the dynamic update of the system, the leader drone and the follower drones respectively follow specific leader dynamic models and follower dynamic models for state iteration, and this update process can be accurately described by corresponding equations. The leader drone dynamic equation is expressed as:

[0068] x 0 (k + 1)=A 0 x 0 (x)

[0069] The follower drone dynamic equation is expressed as:

[0070] x i (k + 1)=A i x i (k)+B i μ i (j),i = 1,2,...,N

[0071] where, x 0 (k), x 0(k + 1) represents the state value of the leader UAV at times k and k + 1, x i (k), x i (k + 1) represents the state value of the follower UAV i at times k and k + 1, μ i (k) represents the control input of the follower UAV i at time k, A 0 、A i and B i are different unknown constant matrices with appropriate dimensions.

[0072] In the implementation process of the present invention, the present invention constructs the local state error of the UAV by using the leader and follower dynamic equations and expresses it as the following formula:

[0073]

[0074] where s ij ≠0 represents the cooperation and competition intensity between agent i and agent j, e i (k) represents the consensus error of the new state information x i (k) on the i-th agent at time k; N i represents the set of neighbor nodes of agent i; a ij represents the connection weight from agent j to agent i; g i represents the pinning control on agent i, g i =0 means that the follower cannot receive the leader information, a ij >0 means that the follower UAV i can receive the state information of the follower UAV j, a ij =0 means that it cannot receive.

[0075] Based on the state equations of the UAV leader and follower, the iterative form of the local state error is expressed as follows:

[0076]

[0077] The consumption performance index function is expressed as:

[0078]

[0079] where is the effectiveness function, is the control input of the neighbor agent i, Q ii 、R ii are the weight matrix and symmetric matrix of the follower UAV i respectively; the performance index function is used as an evaluation index to evaluate the performance of the controller μ i (k) under the current state e i (e i (k)).

[0080] S3: Construct a dual Q - learning network and, based on the local state error, train the dual Q - learning network using the ALR - TQDPG algorithm until the intelligent unmanned cluster system reaches consensus, and the target network outputs control information.

[0081] Construct a dual Q - learning network including an evaluator network, an actor network, and a target network. Both networks are of the actor - critic network structure. The actor network is the actor network, and the evaluator network includes the critic1 network and the critic2 network; the target network includes the target networks corresponding to the actor network, the critic1 network, and the critic2 network.

[0082] The present invention designs the ALR - TQDPG algorithm to train the dual Q - learning network, and the specific process is as follows:

[0083] S31: Initialize the parameters of the evaluator network, the actor network, and the target network.

[0084] Initialize appropriate parameter values. The initial learning rate β of the actor network a , the initial learning rates β of the two critic networks c , the weight update value ι of the target network, the stability threshold ε. Let l = 1, k = 1, and set the maximum number of iterations l max , initialize the weight of the actor network the network weight of critic 1 the network weight of critic 2 and the weight of the target network corresponding to the actor network the weight of the target network corresponding to the critic 1 network weight

[0085] the weight of the target network corresponding to the critic 2 network weight Initialize the experience pool The size of the experience pool is M.

[0086] S32: The actor network outputs the control information at the current moment, and stores the local state error of the UAV at the current moment, the local state error at the next moment, and the control information at the current moment as experience information in the experience pool.

[0087] At the initial moment, the actor network outputs the control information at the current moment:

[0088]

[0089] Among them, represents the control input of agent i at time k, Denotes the transpose of the actor network weights, Denotes the activation function of the actor network of agent i at time k.

[0090] Subsequently, it is iteratively updated using the policy gradient method, and the update of the control policy is defined as follows:

[0091]

[0092] where β a Denotes the learning rate, Denotes the partial derivative with respect to .

[0093] Stores the dataset e i (k), e i (k + 1), into β M .

[0094] S33: Selects experience information from the experience pool to calculate the Q value of the critic network.

[0095]

[0096] where, Denotes the Q value estimate for the training of critic 1, Denotes the Q value estimate for the training of critic 2, Denotes the transpose of the network weights of critic 1, Denotes the transpose of the network weights of critic 2, Denotes the activation function of the critic network, where, Denotes the control input of agent i at time k, where, Denotes the control input of neighbor agent i at time k.

[0097] S34: The target network outputs control information and calculates the Q value of the target network.

[0098] Through Calculate the action value of the target network, i.e., the control information; calculate the Q value of the target network:

[0099]

[0100] where, Denotes the Q value estimate for the training of the target network of critic 1, Denotes the Q value estimate for the training of the target network of critic 2, Denotes the control input of the target network under the error at time k + 1, Denote the transpose of the target network weights of critic 1, Denote the transpose of the target network weights of critic 2.

[0101] Among them, Denote the control input of agent i at time k + 1, Denote the control input of neighbor agent i at time k + 1.

[0102] S35: Update the weight between the historical TD error and the cooperation-competition intensity.

[0103] κ = tanh(b·ln(l + 1))

[0104] Among them, κ denotes the weight between the historical TD error and the cooperation-competition intensity, b denotes a constant, and l denotes the number of training iterations.

[0105] S36: Calculate the TD error of the two critic networks according to the Q value of the target network; calculate the historical TD error according to the TD error of the critic network.

[0106] The formula for calculating the TD error of the two critic networks is:

[0107]

[0108]

[0109] Among them, Denote the TD error of critic 1 training at time k, Denote the TD error of critic2 training at time k, Denote the efficacy function, Denote the control input of agent i, Denote the Q function trained by the target network corresponding to critic 1, Denote the Q function trained by the target network corresponding to critic 2, Denote the Q function trained by critic 1 network, Denote the Q function trained by critic 2 network, e i (k) denotes the consensus error under the state information x i (k) of the i-th follower agent at time k, e i (k + 1) denotes the consensus error under the state information x i (k) of the i-th follower agent at time k + 1.

[0110] The formula for calculating the historical TD error is:

[0111]

[0112] Among them, r1 i (k) represents the historical TD error of critic 1 training at time k, r2 i (k) represents the historical TD error of critic2 training at time k, Θ represents the decay factor, r1 i (k - 1) represents the historical TD error of critic 1 training at time k - 1, r2 i (k - 1) represents the historical TD error of critic 2 training at time k - 1, represents the TD error of critic 1 training at time k, represents the TD error of critic 2 training at time k.

[0113] S37: Calculate the adaptive learning rate according to the historical TD error and the weight between the historical TD error and the cooperation - competition intensity; update the weights of the two critic networks according to the adaptive learning rate.

[0114] The formula for calculating the adaptive learning rate is:

[0115]

[0116] Among them, β 1 ′ represents the adaptive learning rate of the critic1 network, β 2 ′ represents the adaptive learning rate of the critic2 network, β c represents the fixed learning rate, m represents a constant, represents the sum of of neighbor agent j of agent i, κ = tanh(b·ln(l + 1)), κ represents the weight between the historical TD error and the cooperation - competition intensity.

[0117] Considering that the change of the learning rate should be within a reasonable range, so clip operations are performed on β 1 ′, β 2 ′, and the conversion formula of the clip operation is as follows:

[0118] Among them, β are the upper and lower bounds of the appropriate adaptive learning rate.

[0119] Update the weights of the critic network according to the adaptive learning rate:

[0120]

[0121] Among them, denotes the weights of the critic 1 network at the (l + 1)-th iteration, denotes the weights of the critic 2 network at the (l + 1)-th iteration, E1 ci (k) represents the loss function of critic 1 at time k, E2 ci (k) represents the loss function of critic 2 at time k, where, denotes the control input of agent i at time k, denotes the control input of neighbor agent i at time k. denotes the TD error for training critic 1 at time k, denotes the TD error for training critic 2 at time k.

[0122] S38: Update the weights of the actor network according to the Q-values of the critic network, and update the weights of the target network corresponding to the actor network according to the weights of the actor network.

[0123] Update the weights of the actor network:

[0124]

[0125] where, denotes the partial derivative with respect to of.

[0126] Update the weights of the target network corresponding to the actor network:

[0127]

[0128] S39: Update the weights of the target network corresponding to the critic network according to the weights of the two critic networks in the critic network.

[0129]

[0130] S310: Determine whether the intelligent unmanned cluster system reaches consensus. If it reaches consensus, the target network outputs control information; otherwise, return to step S32.

[0131] The formula for determining that the intelligent unmanned cluster system reaches consensus is:

[0132]

[0133] where, denotes the weights of critic 1 at the 1st iteration, denotes the weights of critic 1 at the (l + 1)-th iteration, denotes the weights of critic 2 at the 1st iteration, denotes the weight of critic 2 in the (l + 1)-th iteration, and ε denotes the stability threshold.

[0134] If the above conditions are met, it is determined that the intelligent unmanned cluster system reaches consensus; otherwise, return to step S32 and continue the iteration.

[0135] S4: Control the agents using the control information output by the target network to achieve the cooperative control of the intelligent unmanned cluster system.

[0136] Evaluate the present invention:

[0137] Simulate the present invention. Figure 3 、 Figure 4 respectively show the state convergence diagrams of the unmanned cluster system. Figure 5 , Figure 6 shows the error convergence diagram of the UAV cluster system. It can be concluded from the simulation results that the unmanned cluster system finally achieves consensus. To further verify the advantages of the present invention, use the same UAV dynamic system, topology structure, initial values of system states, critic weights, actor weights, and other relevant parameters as in the comparative experiment, where a represents the algorithm effect of the adaptive cooperation-competition intensity, and b represents the ALR-TQDPG algorithm proposed by the present invention. It can be seen that the present invention has a faster convergence speed.

[0138] Compare from the perspective of performance consumption. Figure 7 a represents the performance consumption of the algorithm of the adaptive cooperation-competition intensity, and 7b represents the performance consumption of the ALR-TQDPG algorithm proposed by the present invention. It can be clearly seen that the algorithm proposed by the present invention has lower performance consumption.

[0139] In addition, from Figure 8 and Figure 9 it can be seen that the evolution of the actor network weights and critic network weights finally tends to be stable, the neural network has converged, and the training has achieved an ideal result.

[0140] From previous studies, it is known that selecting inappropriate cooperation-competition intensity parameters will cause the unmanned cluster system to be unstable. By designing an ALR-TQDPG algorithm, when the unmanned cluster system with cooperative-competitive interaction relationships finally reaches bipartite consensus, the corresponding adaptive learning rate parameter converges to the optimal value (see Figure 10 ), without the need to manually adjust the learning rate parameter, ensuring the stability of the system and reducing the energy consumption of the system.

[0141] In summary, the present invention dynamically adjusts the learning rate of the evaluator network according to the cooperation-competition intensity and TD error, which can cope with the influence of the cooperation-competition intensity on the system performance; reduces the problem of underestimating the Q value caused by minimizing the Q value during the action selection process; achieves superior performance compared to the comparative method, and has good application prospects.

[0142] The above-mentioned embodiments further elaborate on the purpose, technical solution, and advantages of the present invention. It should be understood that the above-mentioned embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An optimal two-part collaborative control method for an intelligent unmanned cluster system, characterized in that: include: S1: Determine the topology of the system based on the interconnection relationship between drones in the unmanned swarm system; S2: Based on the topological structure of the system, each UAV agent sends its own state information to its neighboring agents and calculates the local state error of each UAV; S3: Construct a dual-Q learning network and train it using the ALR-TQDPG algorithm based on the local state error until the intelligent unmanned cluster system reaches a consensus and the target network outputs control information; S4: Use the control information output by the target network to control the intelligent agent and realize the collaborative control of the intelligent unmanned cluster system.

2. The optimal two-part collaborative control method for an intelligent unmanned cluster system according to claim 1 is characterized in that: The formula for calculating the local state error is expressed as: Among them, e i (k) represents the state information x of the i-th follower agent at time k i (k) consistency error, x i (k) represents the state information of the i-th follower agent at time k, x0(k) represents the state information of the leader agent at time k, c i represents the connection weight between the leader and agent i, g i represents the pinning control on agent i, a ij represents the connection weight from agent j to agent i, x j (k) represents the state information of the j-th follower agent at time k, N i represents the set of neighbor nodes of agent i, and sign() represents the sign function.

3. The optimal two-part collaborative control method for an intelligent unmanned cluster system according to claim 1 is characterized in that: The process of training the dual Q learning network using the ALR-TQDPG algorithm includes: S31: Initialize the parameters of the evaluator network, the action network and the target network; the action network is the actor network; the evaluator network includes the critic1 network and the critic2 network; the target network includes the critic1 network, the critic2 network and the target network corresponding to the actor network; S32: The actor network outputs the control information at the current moment, and stores the local state error of the drone at the current moment, the local state error at the next moment, and the control information at the current moment as experience information into the experience pool; S33: Selecting experience information from the experience pool to calculate the Q value of the evaluator network; S34: the target network outputs control information and calculates the Q value of the target network; S35: Update the weight between historical TD error and cooperative competition intensity; S36: Calculate the TD error of the two critic networks according to the Q value of the target network; calculate the historical TD error according to the TD error of the critic network; S37: Calculate the adaptive learning rate according to the historical TD error and the weight between the historical TD error and the cooperative competition intensity; Update the weights of the two critic networks according to the adaptive learning rate; S38: updating the weight of the actor network according to the Q value of the evaluator network, and updating the weight of the target network corresponding to the actor network according to the weight of the actor network; S39: Update the weight of the target network corresponding to the critic network according to the weights of the two critic networks in the evaluator; S310: Determine whether the intelligent unmanned cluster system has reached a consensus. If so, the target network outputs control information; otherwise, return to step S32.

4. The optimal two-part collaborative control method for an intelligent unmanned cluster system according to claim 3 is characterized in that: The formula for updating the weight between historical TD error and cooperative competition intensity is: κ=tanh(b·ln(l+1)) Among them, κ represents the weight between the historical TD error and the intensity of cooperative competition, b represents a constant, and l represents the number of training iterations.

5. The optimal two-part collaborative control method for an intelligent unmanned cluster system according to claim 3 is characterized in that: The formula for calculating the TD error of two critic networks is: in, represents the TD error of critic 1 training at time k, represents the TD error of critic2 training at time k, represents the performance function, represents the control input of agent i, represents the Q function of the target network training corresponding to the critic 1 network, represents the Q function of the target network training corresponding to the critic 2 network, represents the Q function trained by critic 1, represents the Q function of critic 2 network training; e i (k) represents the state information x of the i-th follower agent at time k i (k) The consistency error under e i (k+1) represents the state information x of the i-th follower agent at time k+1 i (k) The consistency error under 6. The optimal two-part collaborative control method for an intelligent unmanned cluster system according to claim 3 is characterized in that: The formula for calculating historical TD error is: Among them, r1 i (k) represents the historical TD error of critic 1 training at time k, r2 i (k) represents the historical TD error of critic2 training at time k, Θ represents the decay factor, r1 i (k-1) represents the historical TD error of critic 1 training at time k-1, r2 i (k-1) represents the historical TD error of critic 2 training at time k-1, represents the TD error of critic 1 training at time k, represents the TD error of critic 2 training at time k.

7. The optimal two-part collaborative control method for an intelligent unmanned cluster system according to claim 3 is characterized in that: The formula for calculating the adaptive learning rate is: Among them, β1 ′ represents the adaptive learning rate of the critic1 network, β2 ′ represents the adaptive learning rate of the critic2 network, β c represents a fixed learning rate, m is a constant, N i represents the neighbor set of agent i, represents the neighbor agent j of agent i The sum of the values ​​of a ij represents the connection weight from agent j to agent i, s ij represents the intensity of cooperation and competition between agents i and j, κ represents the weight between historical TD error and cooperation and competition intensity, r1 i (k) represents the historical TD error of critic 1 training at time k, r2 i (k) represents the historical TD error of critic 2 training at time k.

8. The optimal two-part collaborative control method for an intelligent unmanned cluster system according to claim 3 is characterized in that: The formula for judging whether the intelligent unmanned cluster system has reached consensus is: in, represents the weight of critic 1 in the lth iteration, represents the weight of critic 1 in the l+1th iteration, represents the weight of critic 2 in the lth iteration, represents the weight of critic 2 at the l+1th iteration, and ε represents the stable threshold.

Citation Information

Patent Citations

  • Unmanned aerial vehicle autonomous formation intelligent control method based on reinforcement learning

    CN114815882A

  • Intelligent unmanned cluster system optimal consistency cooperative control method under influence of multiple time delays

    CN115793448A

  • Output synchronization optimization control method for unmanned cluster system with unknown internal state

    CN115903901A

  • Cooperative multi-agent deep Q learning method based on target Q value correction

    CN116468108A

  • Multi-agent cluster motion method based on layering and DDPG

    CN117471974A