An input-constrained optimal formation tracking control method based on adaptive dynamic programming
The optimal formation tracking control algorithm designed by the adaptive dynamic programming method solves the cluster formation tracking control problem under input-constrained conditions, realizes efficient and low-energy cluster formation control, and meets the efficiency and flexibility requirements of cluster formation tracking.
Patent Information
- Application Number
- CN202411684304.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-11-22
AI Technical Summary
Existing formation tracking control methods are difficult to ensure the control performance and energy consumption of the cluster under input-constrained conditions. In addition, traditional methods are computationally complex and have high communication requirements, making it difficult to meet the efficiency and flexibility requirements of cluster formation tracking control.
An adaptive dynamic programming method is used to design the optimal formation tracking control algorithm. By establishing a cluster communication model and performance index function, the Hamilton-Jacobi-Bellman equation is used to solve the optimal control rate, and a single evaluation network structure is constructed for approximate solution. The optimal controller design is based on the Bellman optimality principle to achieve distributed control.
Under limited input conditions, the cluster can track the specified trajectory and maintain the specified geometric formation configuration, with good control effect, low energy consumption, small computational burden, and real-time and adaptability.
Smart Images

Figure CN119472737B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of cluster collaborative control, and specifically relates to an input-constrained optimal formation tracking control method based on adaptive dynamic programming, which can complete the cluster's tracking of a specified trajectory according to a specific formation configuration under input-constrained conditions and obtain optimal control performance. Background Art
[0002] With the rapid development of drone technology, the demand for swarm formation applications is increasing. Multi-agent systems are required to complete tasks in formation, such as emergency search and rescue, formation performances, and environmental monitoring. However, swarm formation tracking control requires the agents to follow predetermined trajectories while simultaneously satisfying formation geometry constraints. This makes the task of formation tracking control more demanding and challenging. Researching control algorithms for swarm formation tracking is key to ensuring that swarms can fully leverage the advantages of multi-agent intelligence.
[0003] Currently, commonly used formation control methods include behavioral methods, leader-follower methods, virtual structure methods, and artificial potential field methods. Leader-follower and artificial potential field methods are computationally simple and easy to implement; the virtual structure method avoids the interference of the actual leader in the leader-follower method. Compared to behavioral methods, the artificial potential field method offers better real-time performance and is suitable for systems with high latency requirements, such as quadrotor drones. However, it struggles to address nonholonomic kinematic constraints. The artificial potential field method is more flexible than the pilot method in formation changes and can meet the functional requirements of coordinated flight in systems such as quadrotor drone swarms. However, these control methods struggle to guarantee swarm control performance, such as energy consumption and transition performance. Furthermore, the control quantity available in practical application scenarios is bounded, a constraint that should be fully considered when designing control algorithms. Therefore, research on optimal formation tracking control algorithms for swarms under input constraints is crucial for improving swarm efficiency and fully leveraging their capabilities.
[0004] Adaptive dynamic programming methods offer a promising approach to solving optimal control problems. Currently, control algorithms based on adaptive dynamic programming have achieved considerable research results in the areas of single-vehicle trajectory tracking control, swarm consistency control, and inclusion control. However, further research is needed on optimal control algorithms for formation tracking control with constrained swarm inputs. Summary of the Invention
[0005] In order to complete the cluster formation tracking control task under input constraints, the present invention provides an optimal formation tracking control algorithm based on an adaptive dynamic programming method. The optimal controller is designed according to the Bellman optimality principle and solved using the adaptive dynamic programming method. This algorithm can ensure cluster control performance while completing the formation tracking task.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] An input-constrained optimal formation tracking control method based on adaptive dynamic programming includes the following steps:
[0008] Step (1) Establish the tracking consistency error system of each agent formation and the cluster communication model, and design the feedforward control rate ;
[0009] Step (2), define the performance indicator function of the input restricted type for each agent;
[0010] Step (3), combined with the performance index function of each agent in step (2) , derive the corresponding Hamilton-Jacobi-Bellman equation and solve the optimal control rate Theoretical formula;
[0011] Step (4) uses the adaptive dynamic programming method to establish a single evaluation network structure and calculate the optimal control rate Perform approximate solutions and design the evaluation network weight update rate;
[0012] Furthermore, the step (1) includes:
[0013] The cluster system consists of The same intelligent agents are composed and the leader-follower formation control method is adopted. The dynamic equation of each intelligent agent is:
[0014] ;
[0015] in, , The requirement is that it is Lipschitz continuous and contains the origin. , represents the feedforward control quantity, Represents the feedback control quantity.
[0016] The leader dynamics model is expressed as:
[0017] ;
[0018] in, Indicates the virtual leader status.
[0019] Agent The formation information dynamics model is:
[0020] ;
[0021] Definition The error of an agent is:
[0022] ;
[0023] Taking the derivative of the above error we can get:
[0024] ;
[0025] To save communication resources and reduce computing pressure, some followers cannot communicate directly with the leader and can only communicate with adjacent followers. To ensure that the entire cluster can synchronize with the leader and form a designated formation, it is necessary to ensure that the leader can communicate with some followers. The cluster communication model adopts a directed graph. describe, represents a non-empty set of points, Represents a set of edges between ordered sets of points. Respectively represent intelligent agent, that is ; Indicated by arrive If the edge , indicating the The agent can receive The adjacency matrix corresponding to this graph is used to represent the adjacency relationship of each agent. Indicates that if ,but ,on the contrary The degree matrix corresponding to this graph is ,in . Define the leader adjacency matrix , Representing an agent It can receive the leader's information, otherwise it cannot be received. The following consistency error is defined for each follower:
[0026] ;
[0027] in, Representing an agent A collection of individuals capable of communicating.
[0028] The following consistency error dynamics model is obtained by derivatizing the above consistency error:
[0029] ;
[0030] in, express With the agent A collection of components.
[0031] The feedforward control solution formula is:
[0032] ;
[0033] assumed exists, then the agent The feedforward control can be obtained as:
[0034] ;
[0035] Furthermore, the performance indicator function designed in step (2) is:
[0036] ;
[0037] in, is a positive definite symmetric matrix, is a positive definite integral function, defined as:
[0038] ;
[0039] in, , , , is the maximum control input;
[0040] Furthermore, the step (3) includes:
[0041] The corresponding HJB equation is obtained by deriving the performance index function of each agent:
[0042] ;
[0043] The optimal feedback control under input constraints can be solved as:
[0044] ;
[0045] Furthermore, the step (4) includes:
[0046] For intelligent agents , its value function and its derivative can be approximated as:
[0047] ;
[0048] in, Representing an agent The ideal weight parameters of the evaluation network are Representing an agent The activation function, represents the partial derivative of the activation function with respect to the consistency error, represents the approximation error and is bounded, represents the partial derivative of the approximation error with respect to the consistency error.
[0049] The following standard assumptions are made for the approximation formula:
[0050] (1) Approximation error of each agent’s value function and its partial derivative with respect to the consistency error is bounded, satisfying the condition , , A constant greater than 0
[0051] (2) Activation function of each agent and its partial derivative with respect to the consistency error is bounded, satisfying the condition , , A constant greater than 0
[0052] (3) Ideal weight of each agent’s evaluation network is bounded, satisfying the condition , A constant greater than 0
[0053] Define the actual weight of the evaluation network as , the optimal feedback control rate can be approximated as:
[0054] ;
[0055] Substituting into the HJB equation, the approximate error is:
[0056] ;
[0057] in,
[0058] The optimization objective function is defined as:
[0059] ;
[0060] The weight update rate of the evaluation network designed using the gradient descent method is:
[0061] ;
[0062] in,
[0063] The present invention can control the cluster to track a specified trajectory and maintain a specified geometric formation configuration.
[0064] Compared with the existing technology, the present invention has the following beneficial effects:
[0065] (1) Considering the constraints of limited input of intelligent agents, the present invention designs an input-constrained optimal formation tracking control algorithm based on adaptive dynamic programming. Compared with the traditional formation tracking control method, the control algorithm of the present invention has controllable control rate amplitude, better control effect, and lower energy consumption.
[0066] (2) The control algorithm designed by the present invention only uses the information of the intelligent agent and its neighboring intelligent agents, and is a distributed controller. Compared with the centralized controller, the control algorithm of the present invention has lower requirements for cluster communication and a smaller computational burden.
[0067] (3) The present invention uses a neural network to construct a single evaluation network structure for online approximate solution of the optimal controller. Compared with the offline iterative solution of the optimal control method, the present invention has real-time and adaptive properties. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 Flowchart of an input-constrained optimal formation tracking control method based on adaptive dynamic programming of the present invention;
[0069] Figure 2 This is a schematic diagram of formation tracking control in leader-follow mode;
[0070] Figure 3 Schematic diagram of the neural network structure of a single evaluation network. DETAILED DESCRIPTION
[0071] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0072] According to one embodiment of the present invention, Figure 1 As shown, the specific implementation steps of the input-constrained optimal formation tracking control method based on adaptive dynamic programming of the present invention are:
[0073] Step (1) establishes the tracking consistency error system of each agent formation and the cluster communication model, the cluster object is composed of The same intelligent agents are composed and the leader-follower formation control method is adopted. The dynamic equation of each intelligent agent is:
[0074] ;
[0075] in, represents the agent number, Representing an agent The state quantity, express dimensional real space, Representing an agent The state derivative of represents the nonlinear term of the state quantity in the dynamic equation, represents the control quantity coefficient term in the dynamic equation, Representing an agent The total control volume, express dimensional real space, The requirement is that it is Lipschitz continuous and contains the origin. , represents the feedforward control quantity, Represents the feedback control quantity.
[0076] The leader dynamics model is expressed as:
[0077] ;
[0078] in, Indicates the virtual leader state, represents the virtual leader state derivative, Represents the nonlinear term of the state quantity in the virtual leader dynamics model.
[0079] Agent The formation information dynamics model is:
[0080] ;
[0081] in, Indicates the state of the formation, represents the derivative of the formation state quantity, Represents the nonlinear term of the formation state in the formation information dynamics model.
[0082] Definition The error of the agent for:
[0083] ;
[0084] Taking the derivative of the above error we can get:
[0085] ;
[0086] For clusters, use Figure 2In the leader-follower formation tracking control mode shown in the figure, in order to save communication resources and reduce computing pressure, some followers cannot communicate directly with the leader and can only communicate with adjacent followers. In order to ensure that the entire cluster can synchronize with the leader and form a designated formation, it is necessary to ensure that the leader can communicate with some followers. The cluster communication model adopts a directed graph describe, represents a non-empty set of points, Represents a set of edges between ordered sets of points. Respectively represent intelligent agent, that is ; Indicated by arrive If the edge , indicating the The agent can receive Information about an agent. Directed graph The corresponding adjacency relationship is represented by the adjacency matrix Represents the adjacency matrix No. Rank The column elements are ,like ,but ,on the contrary . Directed graph The corresponding degree matrix ,in , represents a diagonal matrix, Indicates the number of agents in the cluster. Define the leader adjacency matrix , Representing an agent Can receive leader information, otherwise it cannot be received.
[0087] For each follower, the following consistency error is defined: :
[0088] ;
[0089] in, Representing an agent A collection of individuals capable of communicating.
[0090] The following consistency error dynamics model is obtained by derivatizing the above consistency error:
[0091]
[0092] in, express With the agent A collection of represents the derivative of the consistency error, Representing an agent The state quantity, Representing an agent The total control volume, Represents the leader adjacency matrix No. Rank Column elements, Represents a directed graph The Laplace matrix of No. Rank Column elements, directed graph The Laplace matrix of Defined as .
[0093] The feedforward control solution formula is:
[0094] ;
[0095] assumed exists, then the agent Feedforward control It can be obtained as:
[0096] ;
[0097] Step (2): define the performance indicator function of the input restricted type for each agent; for the agent Designed performance indicator function for:
[0098] ;
[0099] in, Indicates the starting time, Representing an agent The feedback control quantity, represents the integration variable, is a positive definite symmetric matrix, the superscript represents the transpose of the matrix, is a positive definite integral function, defined as:
[0100] ;
[0101] in, represents the integration variable, represents the feedback control weight matrix, Indicates the dimensional feedback control weight, represents the weight vector of feedback control quantity, All 1 dimensional column vector, is the maximum control input;
[0102] Step (3), combined with the performance index function of each agent in step (2) , derive the corresponding Hamilton-Jacobi-Bellman equation and solve the optimal control rate Theoretical formula of
[0103] The corresponding Hamilton-Jacobi-Bellman equation is obtained by deriving the performance index function of each agent:
[0104] ;
[0105] in, represents the Hamilton-Jacobi-Bellman equation, Representing an agent The optimal feedback control quantity, represents the integration variable, Representing an agent The optimal performance indicator function The derivative of .
[0106] The optimal feedback control under input constraints can be solved as:
[0107] ;
[0108] in, Representing an agent The set of allowable control quantities, represents the inverse matrix of the feedback control weight matrix, Represents a directed graph The Laplace matrix of No. Rank Column elements, Represents the leader adjacency matrix No. Rank Column elements, is the maximum control input, Representing an agent The optimal performance indicator function The derivative of .
[0109] Step (4) uses the adaptive dynamic programming method to establish a single evaluation network structure and calculate the optimal control rate Perform an approximate solution and design the evaluation network weight update rate.
[0110] The designed single evaluation network structure of each intelligent agent is as follows Figure 3As shown, Represents the agent consistency error No. Item element. For the agent , its value function and its derivatives It can be approximated as:
[0111] ;
[0112] in, Representing an agent The ideal weight parameters of the evaluation network are Representing an agent The activation function, represents the partial derivative of the activation function with respect to the consistency error, represents the approximation error and is bounded, represents the partial derivative of the approximation error with respect to the consistency error.
[0113] The following standard assumptions are made for the approximation formula:
[0114] (1) Approximation error of each agent’s value function and its partial derivative with respect to the consistency error is bounded, satisfying the condition , , represents the modulus of the orientation quantity, is a constant greater than 0;
[0115] (2) Activation function of each agent and its partial derivative with respect to the consistency error is bounded, satisfying the condition , , is a constant greater than 0;
[0116] (3) Ideal weight of each agent’s evaluation network is bounded, satisfying the condition , is a constant greater than 0;
[0117] Define the actual weight of the evaluation network as , the approximate value of the optimal feedback control rate for:
[0118] ;
[0119] in, represents the approximate value of the optimal feedback control rate, represents the inverse matrix of the feedback control weight matrix, represents the partial derivative of the activation function with respect to the consistency error, Representing an agent The ideal weight parameters of the evaluation network.
[0120] Substituting into the Hamilton-Jacobi-Bellman equation, the approximation error is:
[0121] ;
[0122] in, Representing an agent The approximation error of the Hamilton-Jacobi-Bellman equation is, .
[0123] Define the optimization objective function for:
[0124] ;
[0125] The weight update rate of the evaluation network designed using the gradient descent method is:
[0126] ;
[0127] in, Representing an agent Evaluation network weight learning rate,
[0128] in, Representing an agent The approximate value of the optimal feedback control quantity is: Representing an agent The formation status information.
[0129] Although the above describes the illustrative specific embodiments of the present invention to facilitate understanding of the present invention by those skilled in the art, and it should be clear that the present invention is not limited to the scope of the specific embodiments, it is obvious to those skilled in the art that as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.
Claims
1. An input-constrained optimal formation tracking control method based on adaptive dynamic programming, characterized in that: The steps include: Step (1) establishes the tracking consistency error system and cluster communication model of each agent formation, and designs the feedforward control rate, including: The cluster system consists of The same intelligent agents are composed and the leader-follower formation control method is adopted. The dynamic equation of each intelligent agent is: ; in, , The requirement is that it is Lipschitz continuous and contains the origin. , represents the feedforward control quantity, represents the feedback control quantity, The leader dynamics model is expressed as: ; in, Indicates the leader state, Agent The formation information dynamics model is: ; Definition The error of an agent is: ; Taking the derivative of the above error we get: ; In order to save communication resources and reduce computing pressure, some followers cannot communicate directly with the leader and can only communicate with neighboring followers. To ensure that the entire cluster can synchronize with the leader and form a designated formation, it is necessary to ensure that the leader can communicate with some followers. The cluster communication model adopts a directed graph. describe, represents a non-empty set of points, represents the set of edges between ordered sets of points, Respectively represent intelligent agent, that is ; Indicated by arrive If the edge , indicating the The agent can receive The information of each agent, the adjacency relationship corresponding to this graph is represented by the adjacency matrix Indicates that if ,but ,on the contrary , the degree matrix corresponding to this graph is ,in , define the leader adjacency matrix , Representing an agent It can receive leader information, otherwise it cannot receive it. The following consistency error is defined for each follower: ; in, Representing an agent A collection of individuals capable of communicating, The following consistency error dynamics model is obtained by derivatizing the above consistency error: ; in, express With the agent A collection of represents the Laplace matrix diagonal elements, the directed graph Laplacian matrix is defined as , The feedforward control solution formula is: ; assumed exists, then the agent The feedforward control is obtained as: ; Step (2) defines a performance indicator function of the input restricted type for each agent; the performance indicator function is: ; in, is a positive definite symmetric matrix, is a positive definite integral function, defined as: ; in, , , , is the maximum control input; Step (3), combined with the performance index function of each agent in step (2) , derive the corresponding Hamilton-Jacobi-Bellman equation and solve the optimal control rate Theoretical formula; Step (4) uses the adaptive dynamic programming method to establish a single evaluation network structure and calculate the optimal control rate Perform an approximate solution and design the evaluation network weight update rate.
2. The input-constrained optimal formation tracking control method based on adaptive dynamic programming according to claim 1, characterized in that: The step (3) includes: The corresponding HJB equation is obtained by deriving the performance index function of each agent: The optimal feedback control under input constraints can be solved as: 。 3. The input-constrained optimal formation tracking control method based on adaptive dynamic programming according to claim 1, characterized in that: The step (4) includes: For intelligent agents , its value function and its derivative are approximately: ; in, Representing an agent The ideal weight parameters of the evaluation network are Representing an agent The activation function, represents the partial derivative of the activation function with respect to the consistency error, represents the approximation error and is bounded, represents the partial derivative of the approximation error with respect to the consistency error, The following standard assumptions are made for the approximation formula: (1) Approximation error of each agent’s value function and its partial derivative with respect to the consistency error is bounded, satisfying the condition , , is a constant greater than 0, (2) Activation function of each agent and its partial derivative with respect to the consistency error is bounded, satisfying the condition , , is a constant greater than 0, (3) Ideal weight of each agent’s evaluation network is bounded, satisfying the condition , is a constant greater than 0, Define the actual weight of the evaluation network as , the optimal feedback control rate is approximately: ; Substituting into the HJB equation, the approximate error is: ; in, ; The optimization objective function is defined as: ; The weight update rate of the evaluation network designed using the gradient descent method is: ; in, .
Citation Information
Patent Citations
Multi-agent output formation tracking control method and system
CN113485344A
Time-varying formation tracking optimization control method and system of nonlinear cluster system
CN114967677A