Multi-agent safety optimization formation control method based on preset performance function
Through obstacle functions and preset performance functions, the collision and connectivity constraints in multi-agent formation control are converted into error constraints. Combined with reinforcement learning and simplified neural network estimation faults, the stability and performance problems of the multi-agent system in complex environments are solved, and safe and optimized formation control is achieved.
Patent Information
- Application Number
- CN202510672071.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-23
AI Technical Summary
Existing multi-agent system formation control methods cannot effectively cope with limited communication range and actuator failures when dealing with collision avoidance and connectivity maintenance constraints, resulting in system instability or performance degradation.
By constructing an obstacle function, the safety distance indicator under the collision constraint and connectivity maintenance constraint is converted into a formation error constraint. Combining the preset performance function and reinforcement learning method, a safety optimization controller is designed. A simplified evaluation neural network structure is used to estimate actuator failures, achieving formation error constraints and performance optimization within a limited time.
The formation error is limited to an adjustable set within a limited time, which reduces the computational burden, ensures that the system can still maintain formation performance when a failure occurs, and improves the system's stability and resource utilization efficiency.
Smart Images

Figure CN120686894A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multi-agent system formation control, and relates to a multi-agent safety optimization formation control method based on a preset performance function. Background Art
[0002] Formation control of multi-agent systems has attracted increasing attention due to its ability to collaboratively handle complex tasks through communication and environmental interaction. For practical formation operations, it is also necessary to consider scenarios such as collisions between agents due to proximity and communication loss due to distance. Most research on formation control for multi-agent systems assumes a consistently connected communication graph and infinite communication ranges for all agents. However, this assumption is unsatisfactory in many practical applications. Each agent has a limited communication range, and if the distance between two agents exceeds the permitted communication range, the communication link between them may be lost. Furthermore, since each agent has a geometric dimension rather than a point mass, collisions may occur when the distance between two agents is very small. Therefore, collision avoidance and connectivity maintenance constraints must be considered in multi-agent formation control. Current key technologies for addressing this issue include artificial potential field methods and preset performance function constraints. Artificial potential field methods simulate physical potential fields to guide the movement of agents. By calculating the distance between agents and surrounding obstacles, corresponding repulsive or attractive forces are designed. Compared to the artificial potential field method, the preset performance function offers advantages because it can constrain the dynamic response of system errors by properly selecting the function's initial value. This method transforms collision avoidance and connectivity maintenance constraints into boundary constraints on the system state and ensures that these constraints are met through a preset performance control method. By predefining performance boundaries relevant to the requirements of the actual mission environment—that is, defining time-varying functions related to control performance such as convergence rate and steady-state error—the original system is transformed into a new system, ensuring the consistent boundedness of the transformed system state to solve the system control problem.
[0003] In existing research, finite-time performance functions achieve a steady-state error within a finite time by introducing power terms of the error, shortening the system's convergence time. Increasing design parameters and introducing trigonometric functions can construct finite-time performance functions of different structures. Using sign functions to adaptively adjust the performance function can adjust the system's overshoot and ensure good transient performance. Most performance functions are only applicable to situations where the system reaches a steady state and is not affected by external interference or system failures. Considering that disturbances can cause the error to violate the performance boundary, resulting in the failure of the preset performance control, by introducing error information into the performance function, a self-regulating performance function that can automatically adjust the constraint boundary can be constructed, which can improve the system's control performance in response to interference or failures.
[0004] Another important issue in achieving safe control is how the system performs under unknown and dangerous faults. As control systems become larger and more complex, and as real systems operate in complex environments or for extended periods, the likelihood of equipment failure increases significantly. These failures negatively impact control system performance and may even lead to system instability. Distributed fault-tolerant control systems are needed to enable intelligent agents to cope with potential actuator failures, thereby improving reliability and maintaining performance. Multi-agent fault-tolerant control approaches include passive and active control. In passive control approaches, each agent's controller is robust against a class of hypothetical actuator failures and does not require online fault information. In contrast, active fault-tolerant control designs offer additional approaches to proactively detect, identify, and estimate actuator failures, including adaptive control and fault accommodation strategies. Some studies treat actuator failures as unknown constants, but these failures may vary over time. Therefore, designing controllers for unknown, time-varying failures is more versatile. Existing research uses virtual auxiliary control variables to establish a relationship between faults and system state signals, and then formulates backstepping fault-tolerant control strategies. A finite-time extended state observer is designed to compensate for the overall uncertainty of the model caused by actuator failures and modeling deviations, and then a non-singular fast terminal sliding mode robust fault-tolerant controller is designed.
[0005] Reinforcement learning, a powerful machine learning technique, focuses on systematically adjusting agent behavior to maximize cumulative reward signals through interactive learning between agents and their environments. This technique can obtain optimal strategies that minimize cost metrics, thereby conserving control resources while meeting the performance requirements of a given system. In recent decades, reinforcement learning has become an effective tool for solving optimal control problems, showing great potential in optimizing complex systems with high-dimensional state-action spaces, such as cyber-physical power systems and wireless network control systems. As resource consumption becomes increasingly severe in the engineering field, it is necessary to consider both the transient and steady-state performance of control systems and the consumption of control resources when solving multi-agent formation control problems. In existing research, there are few solutions that consider the possibility of unknown time-varying faults in actuators of multi-agent systems with collision avoidance constraints and connectivity maintenance constraints, and design optimized controllers for their formation operation problems based on reinforcement learning methods. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to provide a multi-agent safety optimization formation control method based on a preset performance function. When the multi-agent is actually in formation, it is also necessary to consider situations such as collisions between agents due to being too close, and communication disconnection due to being too far away. The existing methods for dealing with collision avoidance and connectivity constraint problems mostly use artificial potential field methods. The preset performance function method can constrain the system's dynamic responses such as convergence time and steady-state error by reasonably selecting the initial value of the function, which is more advantageous. The present invention converts the safety distance index under the collision constraint and connectivity maintenance constraint into a formation error constraint through the norm inequality, and limits it to an adjustable set within a finite time with the help of the preset performance function, thereby constraining the position state of each agent within a safe range.
[0007] A barrier function is constructed to handle the constraints. This function approaches positive and negative infinity only when the formation position error approaches the upper and lower bounds of the constraints. This barrier function converts the constrained formation error into an unconstrained variable. Subsequent controller design only needs to ensure that the converted variable is bounded to achieve the formation position error constraint. Compared to the Lyapunov barrier method that ensures the formation error remains within the constraints, the barrier function method allows for the conversion of a constrained system into an unconstrained system through system transformation, allowing for controller design. This method is more versatile.
[0008] As control systems become larger and larger, and when operating in complex environments or for extended periods, the likelihood of equipment failure increases significantly. These failures negatively impact control system performance and may even lead to system instability. Considering the potential for actuator failure in multi-agent formation control, a safe optimization controller is designed for the system to achieve safe formation control while obtaining an optimal strategy that minimizes a performance cost function. A design-assisted system is used to estimate actuator failures using a simplified optimization learning method that replaces the traditional actuator-evaluator dual network structure with an evaluator neural network, thereby reducing the computational burden. The cost function, composed of local control inputs, system errors, and fault estimates, enables rapid response and suppression to actuator failures, ensuring that the system maintains formation performance even when a failure occurs.
[0009] In order to achieve the above object, the present invention provides the following technical solutions:
[0010] A multi-agent safety optimization formation control method based on a preset performance function comprises the following steps:
[0011] S1: According to the safety range s of agent i i and the perception radius c i , determine the upper and lower thresholds of the safety distance constraint between any two agents i and j:
[0012] l i,j,l =s i +s j j
[0013] l i,j,r =min{c i +s j ,c j +s i}
[0014] S2: The safety distance constraint is converted into a formation position error constraint through the norm inequality, specifically the formation position error ξ i,1 Perform inequality scaling to obtain the bounds:
[0015]
[0016] S3: Build preset performance functions:
[0017]
[0018] Among them, b i >0, 0<λ i ≤1 is the design parameter, ρ i,0 is the initial error bound, is the steady-state error bound, is the preset convergence time;
[0019] S4: Construct barrier function:
[0020]
[0021] Convert the constrained formation position error into an unconstrained variable, where k i,l 、k i,r is the error boundary adjustment coefficient;
[0022] S5: Design-assisted system to estimate actuator fault δ i (t), the auxiliary system includes a fault estimator and the state observer θ i , update the fault estimation value through the adaptive law;
[0023] S6: Construct an evaluation neural network based on the reinforcement learning framework, design a value function that includes local control input, system error, and fault estimation, and calculate the optimal control strategy through a single evaluation network structure;
[0024] S7: Use the backstepping method to design a multi-level virtual controller, introduce the barrier function conversion variable and fault estimation value in the design of each level controller, and finally generate the actual control input u i .
[0025] Furthermore, the preset performance function in S3 satisfies the initial condition And the constraint boundary satisfies:
[0026]
[0027] Furthermore, the auxiliary system in S5 is specifically:
[0028]
[0029] Among them, K i >0 is the observer gain, Γ i >0 is the adaptive rate parameter, q i >0 is the damping coefficient.
[0030] Furthermore, the value function in S6 is defined as:
[0031]
[0032] Among them, z i,n (t) is the nth level intermediate error variable, u i (t) is the control input, is the estimated fault value.
[0033] Furthermore, the design of the virtual controller in S7 includes:
[0034] S71: For the τth subsystem, τ=1,2,…,n-1, define the intermediate error variable:
[0035]
[0036] S72: Constructing a Gaussian function The neural network approximates the optimal value function; where k = 1, 2, ..., q, q is the number of neurons;
[0037] S73: The optimal virtual control law is obtained by solving the Hamilton-Jacobi-Bellman equation:
[0038]
[0039] in, is the neural network weight estimation, g i,τ are known system parameters.
[0040] Furthermore, the neural network weights are updated using an adaptive law:
[0041]
[0042] Among them, η i,τ >0 is the learning rate, e iτis the residual of the HJB equation,
[0043] Furthermore, the barrier function in S4 has the following properties:
[0044] and
[0045] Furthermore, the safety distance constraint conversion in S2 satisfies:
[0046]
[0047] in, is the expected position deviation.
[0048] Furthermore, the multi-agent system adopts a leader-follower architecture, including a virtual leader dynamics model:
[0049]
[0050] Among them, x d is the leader state, f d (·) is a known nonlinear function.
[0051] Furthermore, the actuator fault model is:
[0052]
[0053] Among them, δ i (t) is the time-varying fault signal, satisfying ||δ i (t)||≤δ max And there are boundaries.
[0054] The beneficial effects of the present invention are:
[0055] (1) This paper uses an inequality scaling method to convert the safe distance index under the collision constraint and connectivity maintenance constraint into a formation error constraint. Consider the situation where the agents collide due to being too close, and the situation where the communication is disconnected due to being too far away. The safe distance constraint between the agents is determined based on the safe operation radius range and the perception ability radius range. The relevant properties of the norm inequality are used to scale the multi-agent formation error and the distance constraints between the agents to obtain the constraint range of the formation position error.
[0056] (2) The present invention uses a barrier function based on preset performance to handle formation position error constraints. The preset performance function is used to limit the formation error to an adjustable set within a finite time, and the system's dynamic responses such as convergence time and steady-state error are constrained by reasonably selecting the initial value of the function. Then, a barrier function with the formation error and the preset performance function as variables is constructed. The barrier function approaches an infinite value when and only when the formation position error approaches the constraint boundary constructed by the preset performance function. The constructed barrier function converts the constrained formation error into an unconstrained variable. Therefore, in the subsequent controller design, the converted variable is guaranteed to be bounded, so that the formation position error can be constrained within a safe range.
[0057] (3) The present invention designs an auxiliary system for estimating actuator failures and uses a simplified optimization learning method to design a safe optimization controller for the system. While achieving safe formation control, an optimal strategy is obtained that minimizes the performance cost function. This method only uses an evaluation neural network to replace the traditional actuator-evaluator dual network structure, thereby reducing the computational burden. The cost function is composed of local control input, system error, and fault estimate, which can suppress actuator failures and ensure that the system can maintain formation performance when a failure occurs.
[0058] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:
[0060] Figure 1 This is a schematic diagram of multi-agent safety distance constraints;
[0061] Figure 2 It is a scaling method for safety distance constraint inequality;
[0062] Figure 3 It is a barrier function constraint processing method based on preset performance;
[0063] Figure 4 Block diagram of the formation control architecture for multi-agent security optimization. DETAILED DESCRIPTION
[0064] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the features in the following embodiments and embodiments can be combined with each other without conflict.
[0065] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.
[0066] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0067] 1. Multi-agent system leader-follower formation model
[0068] Consider a multi-agent system consisting of a virtual leader and N followers. Use a directed graph G = (ν, ε, A) to describe the information interaction between all agents. All agent units correspond to the node set v = (1, 2, ..., N). The communication relationship between agents is abstracted as the edges in the graph, which constitute the set of edges The adjacency matrix is A=[a i,j ]∈R N ×N . (v i ,v j ) represents the edge between node i and node j. If agent i can send information to agent j, then (v i ,v j )∈ε,a i,j =1, otherwise Assume that the virtual leader is node v0, and define the leader adjacency matrix A0 = diag{a1,0 ,a 2,0 ,...,a N,0}, agent i can directly obtain the leader information, then a i,0 =1, otherwise a i,0 =0.
[0069] The multi-agent system model is described as follows:
[0070]
[0071] Where τ=1,2,...,n-1 represents the system order, i=1,2,...,N,x i,τ ∈R m Represents the state vector of agent i. i ∈R m is the system output, is the control vector under actuator failure, δ i (t) is the unknown time-varying actuator fault, the system dynamic function Known, g i,τ Known.
[0072] During formation movement, agents should avoid collisions while maintaining connectivity of the communication link. The distance between any two agents should be within the safety distance constraint at all times, and the desired formation distance should also meet the safety distance condition. The diagram of the multi-agent safety distance constraint is shown in the figure below. Figure 1 As shown. Assume that the safe area of each agent is a circle centered on itself, and define s i and s j are the safe distance ranges of agents i and j respectively. The distance between agents should satisfy:
[0073] l i,j (t)=||x i,1 -x j,1 ||2>s i +s j
[0074] Definition i,0 and l j,0 Define x as the deviation vector between agents i and j and the desired formation position. d is the desired virtual leader position, then definition is the expected distance between agents i and j, then it satisfies
[0075] Similar to collision avoidance, the agent should maintain effective information flow during movement, i.e., connectivity maintenance, c iand c j Represent the perception radius range of agents i and j respectively, then the expected formation distance between agents needs to satisfy The actual formation distance between agents should also satisfy l i,j (t)<min{c i +s j ,c j +s i Therefore, to meet the requirements of collision avoidance and connectivity maintenance, the distance between agents has the following constraints:
[0076] l i,j,l <l i,j (t)<l i,j,r
[0077] where l i,j,l =s i +s j , l i,j,r =min{c i +s j ,c j +s i} respectively represent the upper and lower thresholds of the distance constraint. The parameters satisfy the condition c i >s i , expected distance when designing time-varying formation tasks The above constraints must also be met.
[0078] 2. Design of preset performance function based on formation error inequality scaling
[0079] like Figure 3 As shown, the position error of the i-th agent formation is expressed as follows:
[0080]
[0081] The expected position deviation error of agent i and agent j is defined as
[0082] Use the norm inequality to scale the formation error and distance constraints:
[0083]
[0084] Maintaining constraints d based on collision avoidance and connectivity i,j (t)=||x i,1 -x j.1 ||2∈[l i,j,l ,l i,j,r ], then there is Therefore, we can get
[0085] Scaling the formation position error by inequality:
[0086]
[0087] Through the above inequality scaling transformation, the safety distance constraint representing the collision avoidance and connectivity maintenance problem is converted into a formation position error constraint. The framework of the safety distance constraint processing method using inequality scaling technology is as follows: Figure 2 shown.
[0088] Construct boundary piecewise function ρ i , used to constrain the system formation position error:
[0089]
[0090] where b i >0,0<λ i ≤1 is the design parameter, ρ i,0 and is the design parameter related to the system performance index, then the initial value of the performance function For preset time, is the upper bound of the steady-state error.
[0091] If the system formation position error satisfies the constraint:
[0092] -k i,l ρ i (t)<ξ i,1 <k i,r ρ i (t)
[0093] where k i,l and k i,r is the designed adjustable parameter.
[0094] Taking advantage of the monotonically decreasing property of the function, select the initial value of the boundary function as shown below:
[0095]
[0096] Then the formation position error constraint can be satisfied, as shown in the following formula:
[0097]
[0098] A preset performance function is used to constrain the formation position error so that it is stable within a preset time and the error is within a preset error boundary, while satisfying collision avoidance and connectivity maintenance constraints.
[0099] 3. Design of safety formation optimization controller
[0100] Construct the following obstacle function to handle the formation position error constraint:
[0101]
[0102] For the constructed universal barrier function ω i,1 , if and only if ξ i,1 →k i,r ρ i When ω i,1 →+∞, if and only if ξ i,1 →-k i,l ρ i When ω i,1 →-∞. The constructed obstacle function converts the constrained formation error into an unconstrained variable. In the subsequent controller design, it is only necessary to ensure that the converted variable is bounded to ensure that the formation position error is constrained within a safe range.
[0103] Next, for the n-order subsystem, the virtual controller is optimized based on the optimization performance index function at each step, and finally the actual optimization controller is designed. From step 1 to step n, the error dynamics ω is considered in turn. i,1 ,z i,τ ,z i,n . Design an approximately optimal virtual controller in the τth step In the final step, the actual optimization controller is designed. For problems with collision avoidance and connectivity maintenance constraints, an inequality scaling technique is used to convert the safety distance indicator into a formation error constraint. A preset performance function is used to limit the formation error to within the constraint range within a finite time. An obstacle function is constructed to convert the constrained formation error into an unconstrained variable. An auxiliary system is designed to estimate actuator failures. A simplified optimization learning method is used that only uses an evaluation neural network to reduce the computational burden. The controller structure diagram is shown in the figure below. Figure 4 shown.
[0104] first step:
[0105] Taking the derivative of the barrier function, we get:
[0106]
[0107] Based on the designed barrier function ω i,1 Perform formation position error dynamic equation conversion:
[0108]
[0109] in
[0110] The performance indicator function in the first subsystem is defined as:
[0111]
[0112] in is the cost function, α i,1 is a virtual control signal. represents the optimal virtual control law, then the optimal performance index function is expressed as:
[0113]
[0114] where Ω i,1 is the virtual control law α i,1 The admissible set of .
[0115] Take x i,2 As the optimal virtual control input Taking the derivative of both sides of the optimal performance index function, we get the HJB equation:
[0116]
[0117] Solving differential equations Get the optimal virtual controller:
[0118]
[0119] Using neural network approximation
[0120]
[0121] in is the weight value of the neural network, and q is the number of neurons. ∈ J,i1 is the approximate error. is the basis function, which consists of the following Gaussian functions:
[0122]
[0123] Where x is the neural network input, r k is the center of the Gaussian function, w k is the width of the Gaussian function.
[0124] So About Omega i,1 Partial derivative of It can be expressed as:
[0125]
[0126] Use approximations Approximately optimal neural network weight values You can get:
[0127]
[0128] Therefore, the approximately optimal virtual controller is designed as:
[0129]
[0130] use approximate Then, the HJB equation is:
[0131]
[0132] Definition i,1 For subsequent virtual controller design:
[0133]
[0134] make definition The neural network weight adaptation law is designed as follows:
[0135]
[0136] where η i,1 >0 is the design parameter.
[0137] Step τ (τ=2,...,n-1):
[0138] Define the intermediate error variable:
[0139]
[0140] The error dynamic equation is obtained by differentiation:
[0141]
[0142] The performance index function in the τth subsystem is defined as:
[0143]
[0144] in is the cost function, α i,τ is a virtual control signal. represents the optimal virtual control law, then the optimal performance index function is expressed as:
[0145]
[0146] where Ω i,τ is the virtual control law α i,τ The admissible set of .
[0147] Take x i,τ+1 As the optimal virtual control input Taking the derivative of both sides of the optimal performance index function, we get the HJB equation:
[0148]
[0149] Solving differential equations Get the optimal virtual controller:
[0150]
[0151] Using neural network approximation
[0152]
[0153] in is the weight value of the neural network, and q is the number of neurons. ∈ J,iτ is the approximate error. is the basis function, which consists of the following Gaussian functions:
[0154]
[0155] Where x is the neural network input, r k is the center of the Gaussian function, w k is the width of the Gaussian function.
[0156] So About z i,τ Partial derivative of It can be expressed as:
[0157]
[0158] Use approximations Approximately optimal neural network weight values You can get:
[0159]
[0160] Therefore, the approximately optimal virtual controller is designed as:
[0161]
[0162] use approximate Then, the HJB equation is:
[0163]
[0164] make The neural network weight adaptation law is designed as follows:
[0165]
[0166] where η i,τ >0 is the design parameter.
[0167] Step n:
[0168] Design assistance system to obtain the estimated value of actuator fault
[0169]
[0170] where θ i is used to approximate the state x i,n The intermediate variable, K i >0,Γ i >0,q i >0 is the designed parameter.
[0171] Define the intermediate error variable:
[0172]
[0173] The error dynamic equation is obtained by differentiation:
[0174]
[0175] The performance index function in the nth subsystem is defined as:
[0176]
[0177] in is the cost function, u i is the actual control signal. represents the optimal control law, then the optimal performance index function is expressed as:
[0178]
[0179] where Ω i,n is the virtual control law u i The admissible set of .
[0180] Taking the derivative of both sides of the optimal performance index function, we get the HJB equation:
[0181]
[0182] Solving differential equations The actual optimal controller is:
[0183]
[0184] Using neural network approximation
[0185]
[0186] in is the weight value of the neural network, and q is the number of neurons. ∈ J,in is the approximate error. is the basis function, which consists of the following Gaussian functions:
[0187]
[0188] Where x is the neural network input, r k is the center of the Gaussian function, w k is the width of the Gaussian function.
[0189] So About z i,n Partial derivative of It can be expressed as:
[0190]
[0191] Use approximations Approximately optimal neural network weight values You can get:
[0192]
[0193] Therefore, the approximate actual optimal controller design is:
[0194]
[0195] use approximate Then, the HJB equation is:
[0196]
[0197] make The neural network weight adaptation law is designed as follows:
[0198]
[0199] where η i,n >0 is the design parameter.
[0200] This paper addresses the problem of multi-agent formation control with collision avoidance and connectivity maintenance constraints. It employs an inequality scaling technique to convert a safety distance metric into a formation error constraint. A preset performance function is used to constrain the formation error within a finite time within an adjustable set. By constructing a barrier function, the constrained formation error is converted into an unconstrained variable. Considering the potential for actuator failure in multi-agent formation control, an auxiliary system is designed to estimate actuator failures. A simplified optimization learning method is employed that employs only an evaluation neural network to reduce the computational burden. The algorithm flow is shown in the pseudocode below:
[0201]
[0202] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.
Claims
1. A multi-agent safety optimization formation control method based on a preset performance function, characterized by: Including the following step: S1: According to the safety range s of agent i i and the perception radius c i , determine the upper and lower thresholds of the safety distance constraint between any two agents i and j: l i,j,l =s i +s j j l i,j,r =min{c i +s j ,c j +s i } S2: The safety distance constraint is converted into a formation position error constraint through the norm inequality, specifically the formation position error ξ i,1 Perform inequality scaling to obtain the bounds: S3: Build preset performance functions: Among them, b i >0, 0<λ i ≤1 is the design parameter, ρ i,0 is the initial error bound, is the steady-state error bound, is the preset convergence time; S4: Construct barrier function: Convert the constrained formation position error into an unconstrained variable, where k i,l 、k i,r is the error boundary adjustment coefficient; S5: Design-assisted system to estimate actuator fault δ i (t), the auxiliary system includes a fault estimator and the state observer θ i , update the fault estimation value through the adaptive law; S6: Construct an evaluation neural network based on the reinforcement learning framework, design a value function that includes local control input, system error, and fault estimation, and calculate the optimal control strategy through a single evaluation network structure; S7: Use the backstepping method to design a multi-level virtual controller, introduce the barrier function conversion variable and fault estimation value in the design of each level controller, and finally generate the actual control input u i .
2. The multi-agent safety optimization formation control method based on a preset performance function according to claim 1 is characterized by: The preset performance function in S3 satisfies the initial conditions And the constraint boundary satisfies:
3. The multi-agent safety optimization formation control method based on a preset performance function according to claim 1 is characterized by: The auxiliary systems in S5 are specifically: Among them, K i >0 is the observer gain, Γ i >0 is the adaptive rate parameter, q i >0 is the damping coefficient.
4. The multi-agent safety optimization formation control method based on a preset performance function according to claim 1 is characterized by: The value function in S6 is defined as: Among them, z i,n (t) is the nth level intermediate error variable, u i (t) is the control input, is the estimated fault value.
5. The multi-agent safety optimization formation control method based on a preset performance function according to claim 1 is characterized by: The design of the virtual controller in S7 includes: S71: For the τth subsystem, τ=1,2,…,n-1, define the intermediate error variable: S72: Constructing a Gaussian function The neural network approximates the optimal value function; where k = 1, 2, ..., q, q is the number of neurons; S73: The optimal virtual control law is obtained by solving the Hamilton-Jacobi-Bellman equation: in, is the neural network weight estimation, g i,τ are known system parameters.
6. The multi-agent safety optimization formation control method based on a preset performance function according to claim 5 is characterized by: The neural network weight update adopts the adaptive law: Among them, η i,τ >0 is the learning rate, e iτ is the residual of the HJB equation, 7. The multi-agent safety optimization formation control method based on a preset performance function according to claim 1 is characterized by: The barrier function in S4 has the following properties: and 8. The multi-agent safety optimization formation control method based on a preset performance function according to claim 1 is characterized by: The safety distance constraint conversion in S2 satisfies: in, is the expected position deviation.
9. The multi-agent safety optimization formation control method based on a preset performance function according to claim 1, characterized in that: The multi-agent system adopts a leader-follower architecture and includes a virtual leader dynamics model: Among them, x d is the leader state, f d (·) is a known nonlinear function.
10. The multi-agent safety optimization formation control method based on a preset performance function according to claim 1, characterized in that: The actuator fault model is: Among them, δ i (t) is the time-varying fault signal, satisfying ||δ i (t)||≤δ max And it has boundaries.