Optimal time-varying lag formation control method of high-order nonlinear multi-agent system

By combining sliding mode control and actor-critic neural network reinforcement learning technology, the optimal time-varying lag formation control problem of high-order nonlinear multi-agent systems is solved, and the stable and efficient control of the system is achieved.

CN120215277APending Publication Date: 2025-06-27TIANJIN POLYTECHNIC UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510430388.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the optimal time-varying lag formation control problem of higher-order nonlinear multiagent systems, especially when the system state variables have differential relationships.

Method used

Using the sliding mode control method and actor-critic neural network reinforcement learning technology, an interference observer and neural network recognizer are built, appropriate adaptive rates are designed, sliding mode vectors and performance index functions are defined, HJB equations are derived, and optimal time-varying lag formation control is achieved.

Benefits of technology

It realizes the optimal time-varying lag formation control of high-order nonlinear multi-agent system, with the characteristics of concise structure and convenient implementation, and can effectively manage state variables and compensate for unknown system dynamics and external disturbances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215277A_ABST
    Figure CN120215277A_ABST
Patent Text Reader

Abstract

The invention is suitable for the field of intelligent control, and provides an optimal time-varying lag formation control method for a high-order nonlinear multi-agent system, and the method comprises the steps: firstly, constructing a mathematical model and a virtual leader model of the high-order nonlinear multi-agent system with external interference and unknown system dynamics; secondly, designing an interference observer and a neural network recognizer to compensate the influence of external interference and unknown system dynamics on a single agent; then, an optimal time-varying lagging formation control strategy is designed by using a sliding mode control method and a reinforcement learning technology, and the controller not only can realize actual time-varying lagging formation of a multi-agent system, but also can minimize a performance index function; finally, the effectiveness of the proposed control scheme is verified through a multi-agent system composed of four unmanned ships.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent control, and particularly to an optimal time-varying lag formation control method for high-order nonlinear multi-agent systems. Background Art

[0002] With the rapid development of fields such as the energy Internet, multi-circuit systems, multi-satellite systems, and unmanned ship clusters, the cooperative control technology of multi-agent systems has become a research hotspot in the field of intelligent control. In particular, formation control, as an important research direction of multi-agent system cooperative control, has received extensive attention and obtained many meaningful results. It is worth noting that most of the research results discuss the case where the formation shape is fixed. However, in many practical applications (such as resource exploration and rescue missions), multi-agent systems need to frequently change their formation shapes. Therefore, it is very meaningful to study the time-varying formation control of multi-agent systems.

[0003] However, these studies on the formation control of multi-agent systems, whether for fixed formation shapes or time-varying formation shapes, do not consider optimal performance. In fact, optimal control has received increasing attention from researchers because it can achieve a balance between system stability and control resources by minimizing performance indices. The bottleneck in its application to multi-agent systems lies in its dependence on the solution of the HJB equation, and it is extremely difficult to obtain an analytical solution to this equation. To make up for this deficiency, reinforcement learning algorithms are usually considered in recent years to achieve optimal formation control. Unfortunately, the obtained optimal formation control results only discuss first-order or second-order nonlinear multi-agent systems.

[0004] Considering that many actual engineering systems should be described by high-order mathematical models with nonlinear dynamics (such as wheeled mobile robot systems and benchmark mechanical systems), some researchers have begun to preliminarily study the time-varying formation control problem of high-order nonlinear multi-agent systems. In particular, for a class of high-order nonlinear multi-agent systems with differential relationships among state variables, sliding mode control has always been the most popular method because it can easily manage these state variables by defining a sliding mode surface in the state space. Unfortunately, few researchers consider using sliding mode control strategies to deal with the optimal time-varying formation control problem of high-order nonlinear multi-agent systems. Moreover, considering that each agent obtains its own state information significantly later than the acquisition of the virtual leader's state information. Therefore, it is also very meaningful and challenging to study the optimal time-varying lag formation control problem of high-order nonlinear multi-agent systems using sliding mode control methods and reinforcement learning techniques. Summary of the Invention

[0005] The present invention provides an optimal time-varying lag formation control method for high-order nonlinear multi-agent systems to solve the technical problems mentioned in the background art.

[0006] Optimal time-varying lag formation control method for high-order nonlinear multi-agent systems, the method comprising:

[0007] S1. Construct a mathematical model of the high-order nonlinear multi-agent system and a virtual leader model;

[0008] S2. Construct an interference observer and a neural network identifier, define an error model, and design a suitable adaptation rate to ensure that the observer error and the identifier error are bounded;

[0009] S3. Define a sliding mode surface vector and a performance index function to obtain the HJB equation, and use the sliding mode control method and the actor-critic neural network reinforcement learning technology to design an optimal time-varying lag formation control strategy;

[0010] S4. For the closed-loop system of the sliding mode surface vector, construct a Lyapunov functional to ensure that the multi-agent system achieves optimal time-varying lag formation.

[0011] Preferably, in step S1, when constructing the mathematical model of the high-order nonlinear multi-agent system, the mathematical model of the th agent is:

[0012] ;

[0013] wherein, and respectively represent the state vector and the control input of the th agent, is an unknown nonlinear continuous function, is an unknown continuous external interference and satisfies and , and respectively represent dimensional and dimensional spaces composed of real number vectors;

[0014] Use the virtual leader method to solve the optimal time-varying lag formation control problem of the multi-agent system; in view of this, the virtual leader model is constructed as follows:

[0015] ;

[0016] wherein, represents the state vector of the virtual leader and satisfies , is a bounded nonlinear continuous function, and respectively represent dimensional and A space composed of n-dimensional real vectors.

[0017] Preferably, in step S2, the constructed disturbance observer is:

[0018] ;

[0019] Wherein, and respectively represent the state vector and control input of the th agent, and respectively represent the ideal neural network weight matrix and basis function vector, represents an auxiliary vector, and respectively represent the estimates of and , , and respectively represent -dimensional, -dimensional and -dimensional spaces composed of real vectors, represents -dimensional matrix space composed of real matrices.

[0020] Preferably, in step S2, the constructed neural network identifier is:

[0021] ;

[0022] Wherein, and respectively represent the state vector and control input of the th agent, and respectively represent the ideal neural network weight matrix and basis function vector, and respectively represent the estimates of and , represents the identifier state, is a Hurwitz matrix, , and respectively represent -dimensional, -dimensional and -dimensional spaces composed of real vectors, and respectively represent -dimensional and -dimensional matrix spaces composed of real matrices.

[0023] Preferably, in step S2, the steps of defining an error model and designing an adaptation rate to ensure that the observer error and the identifier error are bounded include:

[0024] Let 、 and , and the identifier error model and the observer error model are obtained as follows:

[0025] ;

[0026] Design the adaptation rate of as follows:

[0027] ;

[0028] Wherein, represents the approximation error, 、 and are positive constants, is a positive definite matrix satisfying the inequalities and , and are positive constants, is a Hurwitz matrix, represents dimensional identity matrix, represents matrix space composed of

[0029] For the error model, construct the following Lyapunov functional:

[0030] ;

[0031] Then, take the derivative of , and according to the Lyapunov stability theory, under the action of the adaptation rate, it can be obtained that 、 and are bounded.

[0032] Preferably, in step S3, the designed optimal time-varying lag formation control strategy is:

[0033] ;

[0034] Wherein, the compensation controller based on the disturbance observer and the neural network identifier is designed as:

[0035] ;

[0036] By utilizing the sliding mode control method and the actor-critic neural network reinforcement learning technique, the optimal controller is designed as follows:

[0037] ;

[0038] Among them, and ;

[0039] , , and respectively represent the state vector, control input, external disturbance, and sliding mode surface vector of the th agent, represents the time-varying relative state information, and respectively represent the ideal neural network weight matrix and basis function vector; and respectively represent the estimates of and , represents the recognizer state, represents the recognition error, and respectively represent the critic and actor neural network weight vectors; represents the communication relationship between the virtual leader and the agent . If the agent can obtain the state information of the leader, then , otherwise, ; The edge set of the signed graph is , represents the neighbor set of the agent , represents the adjacency matrix, and the element satisfies:

[0040] ;

[0041] In addition, , , , , and are positive constants, is the sliding mode surface parameter that makes the characteristic polynomial strictly Hurwitz, and respectively represent the critic learning rate and the actor learning rate, is a gain constant, and respectively represent the identity matrix of dimension and is a Hurwitz matrix, is a positive definite matrix satisfying and ; represents a set of real numbers, , and respectively represent the space composed of real vectors of dimension and ; , , and respectively represent the matrix space composed of real matrices of dimension , and .

[0042] Preferably, in step S4, for the closed-loop system of the sliding mode surface vector, construct the following Lyapunov functional:

[0043] ;

[0044] where , , , ;

[0045] Then, take the derivative of , and then according to the Lyapunov stability theory, it is obtained that under the action of the compensation controller and the actor-critic neural network reinforcement learning algorithm, the multi-agent system can achieve the optimal time-varying lag formation.

[0046] Beneficial effects achieved by the present invention:

[0047] For the first time, the sliding mode control method is used to solve the optimal actual time-varying lag formation control problem of high-order nonlinear multi-agent systems. This method can easily manage state variables by defining the sliding mode surface vector, which is essentially different from the backstepping technique in the existing literature; second, a neural network identifier and a disturbance observer are introduced. By designing a suitable neural network weight adaptation rate, the unknown system dynamics and external disturbances can be well compensated; third, based on the sliding mode control method and the reinforcement learning technique, the optimal actual time-varying lag formation criterion for high-order nonlinear multi-agent systems is derived, and this criterion has the characteristics of simple structure and convenient implementation. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a flowchart of an optimal time-varying lag formation control method for a high-order nonlinear multi-agent system.

[0049] Figure 2 It is a communication topology diagram of the multi-agent system provided by the embodiment of the present invention.

[0050] Figure 3 It is a schematic diagram of the formation trajectory provided by the embodiment of the present invention.

[0051] Figure 4 It is a schematic diagram of the position formation tracking error curve provided by the embodiment of the present invention.

[0052] Figure 5 It is a schematic diagram of the velocity formation tracking error curve provided by the embodiment of the present invention.

[0053] Figure 6 It is a schematic diagram of the sliding mode surface vector curve provided by the embodiment of the present invention.

[0054] Figure 7 It is a schematic diagram of the control input curve provided by the embodiment of the present invention.

[0055] Figure 8 It is a schematic diagram of the compensation controller curve provided by the embodiment of the present invention. Specific implementation manners

[0056] The technical solution of the present invention will be described in detail below with reference to specific drawings.

[0057] Please refer to Figure 1 , the embodiment of the present invention provides an optimal time-varying lag formation control method for a high-order nonlinear multi-agent system, and the method includes:

[0058] S1. Construct a mathematical model and a virtual leader model of the high-order nonlinear multi-agent system;

[0059] S2. Construct an interference observer and a neural network identifier, define an error model, and design a suitable adaptation rate to ensure that the observer error and the identifier error are bounded;

[0060] S3. Define a sliding mode surface vector and a performance index function to obtain the HJB equation, and use the sliding mode control method and the actor-critic neural network reinforcement learning technology to design an optimal time-varying lag formation control strategy;

[0061] S4. For the closed-loop system of the sliding mode surface vector, construct a Lyapunov functional to ensure that the multi-agent system realizes an optimal time-varying lag formation.

[0062] In this embodiment, in step S1, a mathematical model of the high-order nonlinear multi-agent system is constructed. The The mathematical model of an agent is as follows:

[0063] Equation 1: ;

[0064] where, and respectively represent the state vector and control input of the -th agent, is an unknown nonlinear continuous function, is an unknown continuous external disturbance and satisfies and , and respectively represent -dimensional and -dimensional spaces composed of real vectors;

[0065] The virtual leader method is used to solve the optimal time-varying lag formation control problem of multi-agent systems; in view of this, the virtual leader model is constructed as follows:

[0066] Equation 2: ;

[0067] where, represents the state vector of the virtual leader and satisfies , is a bounded nonlinear continuous function, and respectively represent -dimensional and -dimensional spaces composed of real vectors.

[0068] In step S2 of this embodiment, first, can be approximated by the following neural network:

[0069] Equation 3: ;

[0070] where, and respectively represent the ideal neural network weight matrix and basis function vector, is the approximation error that satisfies , and the compact set is defined as:

[0071] Equation 4: ;

[0072] where, ; , and are positive constants, are constant matrices, which will be determined later, and respectively represent -dimensional and -dimensional spaces composed of real vectors, and respectively denote -dimensional and -dimensional matrix spaces composed of real matrices.

[0073] From Formulas 1 and 3, we can obtain:

[0074] Formula 5: ;

[0075] In step S2 of this embodiment, the constructed disturbance observer is:

[0076] Formula 6: ;

[0077] wherein, ; and respectively represent the state vector and control input of the th agent, and respectively represent the ideal neural network weight matrix and basis function vector, represents an auxiliary vector, and respectively denote the estimations of and , , and respectively represent -dimensional, -dimensional and -dimensional spaces composed of real vectors, denotes -dimensional matrix space composed of real matrices.

[0078] In step S2, the constructed neural network identifier is:

[0079] Formula 7: ;

[0080] wherein, and respectively represent the state vector and control input of the th agent, and respectively represent the ideal neural network weight matrix and basis function vector, and respectively denote the estimations of and , Denote the recognizer state, is a Hurwitz matrix, , and respectively represent -dimensional, -dimensional and -dimensional spaces composed of real vectors, and respectively represent -dimensional and -dimensional matrix spaces composed of real matrices.

[0081] In step S2 of this embodiment, the steps of defining an error model and designing an adaptation rate to ensure that the observer error and the recognizer error are bounded include:

[0082] Let , and . The recognizer error model and the observer error model are obtained from formulas 4 and 6 as follows:

[0083] Formula 8: ;

[0084] In formula 8, design the adaptation rate of as follows:

[0085] Formula 9: ;

[0086] where , and are positive constants, is a positive definite matrix satisfying the inequality:

[0087] Formula 10: ;

[0088] Formula 11: ;

[0089] where and are positive constants, is a Hurwitz matrix, represents -dimensional identity matrix, represents -dimensional matrix space composed of real matrices;

[0090] For the error model, construct the following Lyapunov functional:

[0091] ;

[0092] Then, for Taking the derivative, we have:

[0093] Formula 12: ;

[0094] where is a positive definite matrix the largest eigenvalue of , , , ;

[0095] Based on the definition of and Formula 12, it is obvious that:

[0096] ;

[0097] Therefore, it can be concluded that , and are bounded, which indicates that under the action of the adaptive rate, the designed disturbance observer and neural network identifier are very effective.

[0098] In step S3 of this embodiment, let ; from Formulas 1 and 2, we can obtain:

[0099] Formula 13: ;

[0100] where ;

[0101] Let , the multi-agent system (Formula 1) can achieve actual time-varying lag formation if there exist positive constants such that:

[0102] ;

[0103] where and represent the time-varying relative state information and the time delay between agent and the virtual leader respectively, represents the set of real numbers, represents the space composed of

[0104] Define Formula 14: ;

[0105] According to Formulas 13 and 14, it can be concluded that:

[0106] Formula 15: ;

[0107] where , ; represents the communication relationship between the virtual leader and the agent If the agent can obtain the status information of the leader, then , otherwise, ; The edge set of the signed graph is , represents the neighbor set of the agent , represents the adjacency matrix, and the element satisfies:

[0108] ;

[0109] represents the Laplacian matrix, where , obviously, is a positive definite matrix.

[0110] Next, define the sliding mode surface vector as follows:

[0111] Equation 16: ;

[0112] where is the sliding mode surface parameter that makes the characteristic polynomial strictly Hurwitz.

[0113] From Equations 15 and 16, we can obtain:

[0114] Equation 17: ;

[0115] where ;

[0116] For Equation 17, design the control input as:

[0117] Equation 18: ;

[0118] where, to compensate for the effects of external disturbances and unknown system dynamics on a single agent, design the compensation controller as follows:

[0119] Equation 19: ;

[0120] is an optimal controller that will be designed later.

[0121] Define the performance index function as follows:

[0122] Equation 20: ;

[0123] Among them, ;

[0124] Based on the compensation controller formula 19, find an optimal admissible controller (where ), so that the multi-agent system (formula 1) can achieve actual time-varying lag formation and the performance index function is minimized. That is, it represents the optimal performance index function:

[0125] Formula 21: ;

[0126] Among them, ; represents the optimal performance index function.

[0127] From formula 21, the HJB equation can be obtained as:

[0128] Formula 22: ;

[0129] According to formula 22, it can be obtained that:

[0130] Formula 23: ;

[0131] From formulas 22 and 23, the HJB equation can be rewritten as:

[0132] Formula 24: ;

[0133] Obviously, due to the complex non-linearity in formula 24, it is very difficult to obtain by directly solving formula 24. In view of this fact, the present invention adopts an adaptive reinforcement learning method to solve this problem.

[0134] To achieve optimal time-varying lag formation, in formula 23 can be decomposed into:

[0135] Formula 25: ;

[0136] Among them, ; is a gain constant.

[0137] According to formulas 23 and 25, it can be obtained that:

[0138] Formula 26: ;

[0139] Then, can be approximated by the following neural network:

[0140] Formula 27: ;

[0141] where, and represent the ideal neural network weight matrix and basis function vector respectively, represents the approximation error and is bounded, and the compact set is defined as:

[0142] ;

[0143] where, is a positive constant and will be determined later.

[0144] From Formulas 25 - 27, we can obtain:

[0145] Formula 28: ;

[0146] Formula 29: ;

[0147] Since is an unknown constant matrix, the optimal control Formula 29 is not available. Next, the optimal controller is derived using the actor - critic neural network reinforcement learning algorithm.

[0148] The critic neural network is designed as follows:

[0149] Formula 30: ;

[0150] where, represents the estimation of , represents the critic neural network weight vector and is updated by the following adaptive rate:

[0151] Formula 31: ;

[0152] where, represents the critic learning rate, is a positive constant.

[0153] The actor neural network is designed as follows:

[0154] Formula 32: ;

[0155] where, represents the actor neural network weight vector and is updated by the following adaptive rate:

[0156] Formula 33: ;

[0157] Among them, ; is the actor learning rate.

[0158] Therefore, based on the compensation controller (Equation 19) and the actor-critic neural network reinforcement learning algorithm (Equations 30-33), the designed optimal time-varying lag formation control strategy is:

[0159] ;

[0160] Among them, the compensation controller based on the disturbance observer and the neural network identifier is designed as:

[0161] ;

[0162] By using the sliding mode control method and the actor-critic neural network reinforcement learning technique, the design of the optimal controller is specifically as follows:

[0163] ;

[0164] Among them, , ;

[0165] , , and respectively represent the state vector, control input, external disturbance, and sliding mode surface vector of the th agent, represents the time-varying relative state information, and respectively represent the ideal neural network weight matrix and basis function vector; and respectively represent the estimations of and , represents the identifier state, represents the identification error, and respectively represent the critic and actor neural network weight vectors; represents the communication relationship between the virtual leader and the agent .

[0166] As a preferred embodiment of the present invention, if the design parameters , and can satisfy the following conditions:

[0167] Equation 34: ;

[0168] Formula 35: ;

[0169] wherein, , , , , then, under the action of the optimal time-varying lag formation control strategy, the multi-agent system (i.e., Formula 1) can achieve the actual time-varying lag formation, and the performance index function can be minimized;

[0170] In this embodiment, in step S4, for the closed-loop system of the sliding mode vector, the following Lyapunov functional is constructed:

[0171] Formula 36: ;

[0172] wherein, , , , ;

[0173] According to Formula 16, it can be derived that:

[0174] Formula 37: ;

[0175] wherein, , ;

[0176] Since is a Hurwitz matrix, for any positive constant , there always exists a positive definite matrix such that holds. Therefore, it can be obtained that:

[0177] ;

[0178] wherein, , , ;

[0179] ;

[0180] Each term in is bounded, that is, there exists a positive constant such that

[0181] Therefore, Formula 38 is obtained:

[0182] ;

[0183] Among them, , ;

[0184] Based on Equation 38, it can be obtained that:

[0185] Equation 39: ;

[0186] That is , , and are bounded.

[0187] According to Equations 36 and 39, it can be obtained that:

[0188] Equation 40: ;

[0189] Equation 41: ;

[0190] Based on and Equation 37, it can be obtained that:

[0191] Equation 42: ;

[0192] Equation 43: ;

[0193] Among them, , , , , ;

[0194] Then, from Equations 40 - 43, it can be obtained that:

[0195] Equation 44: ;

[0196] Also from Equations 42 and 44, it can be obtained that:

[0197] ;

[0198] Among them, ;

[0199] Therefore, the multi - agent system (Equation 1) can achieve actual time - varying lag formation, and can approach 0 by choosing sufficiently large parameters , , and . Next, prove that: and ;

[0200] According to Formulas 36 and 39, for any it can be obtained that:

[0201] Formula 45: ;

[0202] Then, from Formulas 42, 43, and 45, it can be concluded that:

[0203] ;

[0204] Also, because , , obviously,

[0205] ;

[0206] Therefore, it can be concluded that: if , then and if , then .

[0207] The effectiveness of the proposed control scheme is verified through a multi-agent system composed of four unmanned vessels. The communication topology between the four unmanned vessels and the virtual leader is as Figure 1 shown. The kinematic and dynamic descriptions of the th unmanned vessel are as follows:

[0208] Formula 46: ;

[0209] where , represents the position, represents the yaw angle, , , and respectively represent the forward speed, side drift speed, and yaw angular velocity; represents the control input; , , and respectively represent the rotation matrix, inertia matrix, Coriolis force matrix, and damping matrix, where:

[0210] ;

[0211] ;

[0212] and the relevant parameters are provided in Table 1; represents the environmental disturbance, where , , ;

[0213] Table 1 Unmanned Ship Parameters

[0214]

[0215] Select the reference trajectory , , , the control objective is to design a suitable controller such that holds, where represents the time-varying relative state information, is a positive constant.

[0216] Next, use the optimal time-varying lag formation control strategy based on the sliding mode control method and reinforcement learning technology proposed in the present invention to solve the above problems. Let , and , the system (Equation 46) can be rewritten as:

[0217] ;

[0218] where , , , , ;

[0219] Let , and , we can get:

[0220] ;

[0221] where , obviously, is a bounded continuous function.

[0222] Select , , ;

[0223] The parameters in the adaptation rate are selected as:

[0224] , , , , , , , , , and the initial values of the auxiliary vector and the identifier state are selected as , , ;

[0225] Select , , ;

[0226] The parameters in the optimal controller (and Equation 32) are selected as: , , , , , , , , , , , ;

[0227] The initial values of the critic neural network weight vector and the actor neural network weight vector are selected as , , ;

[0228] Select the initial values:

[0229] , , ;

[0230] , , ;

[0231] , , , , ;

[0232] The simulation results are as Figures 3 - 8 shown. Figure 3 It shows the trajectories of four unmanned ships and the virtual leader. It can be seen from Figure 3 that the multi-agent system (46) can achieve actual time-varying lag formation; Figure 4 and Figure 5 show that the errors and can converge to zero; Figure 6 and Figure 7 show the trajectories of the sliding mode surface vector and the control input ; It can be clearly seen from Figure 8 that the compensation controller can approximate well , which means that the disturbance observer (6) and the neural network identifier (7) are very effective.

[0233] It should be noted that in this article, the term "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of additional identical elements in the process, method, article, or device including that element.

[0234] The above are only the preferred embodiments of the present invention and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall similarly be included within the patent protection scope of the present invention.

Claims

1. An optimal time-varying hysteresis formation control method for high-order nonlinear multi-agent systems, characterized in that: The method comprises: S1. Construct mathematical models and virtual leader models of high-order nonlinear multi-agent systems; S2, construct the disturbance observer and neural network identifier, define the error model, and design a suitable adaptive rate to ensure that the observer error and the identifier error are bounded; S3, define the sliding surface vector and performance index function to obtain the HJB equation, and use the sliding mode control method and actor-critic neural network reinforcement learning technology to design the optimal time-varying lag formation control strategy; S4. For the closed-loop system of sliding mode vector, a Lyapunov functional is constructed to ensure that the multi-agent system achieves the optimal time-varying lag formation.

2. The optimal time-varying hysteresis formation control method for a high-order nonlinear multi-agent system according to claim 1, characterized in that: In step S1, a mathematical model of a high-order nonlinear multi-agent system is constructed. The mathematical model of an agent is: ; in, and Respectively represent The state vector and control input of each agent, is an unknown nonlinear continuous function, is an unknown continuous external disturbance and satisfies and , and Respectively represent Peacekeeping The space composed of dimensional real vectors; The virtual leader method is used to solve the optimal time-varying lag formation control problem of multi-agent systems; in view of this, the virtual leader model is constructed as follows: ; in, represents the state vector of the virtual leader and satisfies , is a bounded nonlinear continuous function, and Respectively represent Peacekeeping dimensional real vector space.

3. The optimal time-varying hysteresis formation control method for a high-order nonlinear multi-agent system according to claim 1, characterized in that: In step S2, the constructed disturbance observer is: ; in, and Respectively represent The state vector and control input of each agent, and represent the ideal neural network weight matrix and basis function vector respectively, represents an auxiliary vector, and Respectively and The estimate, , and Respectively represent dimension, Peacekeeping dimensional real vector space, express dimensional real number matrices.

4. The optimal time-varying hysteresis formation control method for a high-order nonlinear multi-agent system according to claim 3, characterized in that: In step S2, the constructed neural network identifier is: ; in, and Respectively represent The state vector and control input of each agent, and represent the ideal neural network weight matrix and basis function vector respectively, and Respectively and The estimate, Indicates the state of the recognizer. is a Hurwitz matrix, , and Respectively represent dimension, Peacekeeping dimensional real vector space, and Respectively Peacekeeping dimensional real number matrices.

5. The optimal time-varying hysteresis formation control method for a high-order nonlinear multi-agent system according to claim 4, characterized in that: In step S2, the steps of defining the error model and designing the adaptation rate to ensure that the observer error and the identifier error are bounded include: make , and , the error model of the identifier and the error model of the observer are as follows: ; design The adaptive rate is as follows: ; in, represents the approximation error, , and is a normal number, is a positive definite matrix satisfying the inequality and , and is a normal number, is a Hurwitz matrix, express dimensional identity matrix, represent Matrix space composed of dimensional real number matrices; For the error model, the following Lyapunov functional is constructed: ; Then, yes Derivative, and then according to Lyapunov stability theory, under the action of adaptive rate, we can get , and is bounded.

6. The optimal time-varying hysteresis formation control method for a high-order nonlinear multi-agent system according to claim 1, characterized in that: In step S3, the optimal time-varying hysteresis formation control strategy is designed as follows: ; Among them, the compensation controller based on disturbance observer and neural network identifier Designed to: ; By using sliding mode control method and actor-critic neural network reinforcement learning technology, the optimal controller The design is as follows: ; in, , ; , , and Respectively represent The state vector, control input, external disturbance and sliding surface vector of each agent, represents the time-varying relative state information, and Represent the ideal neural network weight matrix and basis function vector respectively; and Respectively and The estimate, Indicates the state of the recognizer. Represents the recognition error, and Represent the critic and actor neural network weight vectors respectively; Represents virtual leaders and agents If the communication relationship between agents Can obtain the leader's status information, then ,otherwise, ; The edge set of the signed graph is , Representing an Agent The neighbor set of represents an adjacency matrix, and its elements satisfy: ; also, , , , , and is a normal number, is such that the characteristic polynomial Strictly Hurwitz sliding surface parameters, and Represent the critic learning rate and actor learning rate respectively, is a gain constant, and Respectively Peacekeeping dimensional identity matrix, is a Hurwitz matrix, is a positive definite matrix satisfying and ; represents a set of real numbers, , and Respectively represent dimension, Peacekeeping dimensional real vector space, , , and Respectively represent dimension, dimension, Peacekeeping dimensional real number matrices.

7. The optimal time-varying hysteresis formation control method for a high-order nonlinear multi-agent system according to claim 1, characterized in that: In step S4, for the closed-loop system of the sliding mode vector, the following Lyapunov functional is constructed: ; in, , , , ; Then, yes By taking the derivative and based on the Lyapunov stability theory, it is found that under the action of the compensation controller and the actor-critic neural network reinforcement learning algorithm, the multi-agent system can achieve the optimal time-varying lag formation.

Citation Information

Cited By

  • Formation control method and device for multi-agent system, equipment and medium

    CN122526289A