Fault-tolerant control method for preset performance of nonlinear multi-agent system
By using RBF neural networks and Actor-Critic reinforcement learning in a multi-agent system, the problem of actuator failure in high-order nonlinear systems is solved, and adaptive compensation and optimal control of multiplicative and additive time-varying faults are achieved, ensuring the system's preset performance and stability.
Patent Information
- Application Number
- CN202511951587.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-03-20
AI Technical Summary
Existing fault-tolerant control methods struggle to effectively handle multiplicative and additive time-varying faults when dealing with actuator failures in multi-agent systems. Furthermore, traditional methods rely on accurate models and cannot simultaneously guarantee the transient and steady-state performance of high-order nonlinear systems.
An online identifier is constructed using an RBF neural network, a sliding surface is designed and a preset performance control is introduced, and combined with Actor-Critic reinforcement learning, adaptive compensation and optimal control for actuator faults are achieved.
Without relying on an accurate model, optimal consistency tracking of high-order nonlinear multi-agent systems was achieved, ensuring the system's preset performance constraints in both transient and steady states, and enhancing the system's reliability and control performance.
Smart Images

Figure CN121704535A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-agent cooperative control technology, and in particular to a method for preset performance fault-tolerant control of nonlinear multi-agent systems. Background Technology
[0002] Multi-agent systems (MASs) have wide applications in fields such as UAV swarms and underwater vehicle clusters. In practical applications, agents inevitably encounter actuator failures (such as jamming, efficiency loss, or deviation faults), which severely impact system stability and performance. Furthermore, real-world systems are often highly nonlinear, and model parameters are difficult to obtain precisely. Existing fault-tolerant control (FTC) methods mostly assume additive faults are constant, neglecting the complexity of time-varying faults. Simultaneously, to meet high-precision control requirements, it is necessary not only to ensure the convergence of the system's steady-state error but also to constrain the system's transient performance (such as overshoot and convergence speed). Preset performance control (PPC) is an effective method, but research combining it with fault-tolerant and optimal control is insufficient. Moreover, traditional methods for solving the Hamilton-Jacobi-Bellman (HJB) equations to achieve optimal control are computationally complex and rely on accurate models. Therefore, designing a control method that can guarantee preset transient / steady-state performance while achieving adaptive optimal fault tolerance for high-order nonlinear multi-agent systems with actuator failures and model uncertainties has significant theoretical and practical value. Summary of the Invention
[0003] To address the aforementioned problems, the present invention aims to provide a pre-set performance fault-tolerant control method for nonlinear multi-agent systems, which addresses the issues of optimal consistency tracking and pre-set performance constraints in high-order nonlinear multi-agent systems under conditions of unknown model and actuator failure.
[0004] To achieve the above objectives, the present invention adopts the following technical solution: A method for pre-defined performance-tolerant control of a nonlinear multi-agent system includes the following steps: S1: Considering the multiplicative and additive time-varying faults of the actuators, establish a high-order nonlinear multi-agent system model; S2: For unknown nonlinear functions and fault terms in high-order nonlinear multi-agent system models, an online identifier is constructed using an RBF neural network to reconstruct the system state and estimate the fault value, while eliminating the influence of model uncertainty. S3: Based on the state information provided by the online identifier, a sliding surface is designed to reduce the system order. Pre-defined performance techniques are introduced to constrain and transform the consistency error of the sliding surface, ensuring that the tracking error is always within the predefined performance envelope. S4: Based on the transformed error system, define the performance index function, derive the HJB equation and the theoretical optimal control law, and transform the fault-tolerant control into an optimal regulation problem; S5: Actor-Critic reinforcement learning is used to solve the optimal regulation problem, obtain the optimal control, and output the final control quantity to act on the system.
[0005] Furthermore, considering the multiplicative and additive time-varying faults of the actuators, a high-order nonlinear multi-agent system model is established, as follows: Consider by N A multi-agent system consisting of one follower and one leader, whose communication topology is an undirected graph. G Description, the first k The dynamic model of a follower agent is as follows: in, For the first k The state vector of each agent. The system state vector; function For unknown continuous nonlinear dynamic functions; This is the actual control input after a fault occurs; Considering that the actuator simultaneously exhibits multiplicative faults and additive time-varying deviation faults, the fault model is established as follows: in, Represents actual control input; diagonal matrix For the actuator efficiency factor matrix, It is a bounded time-varying additive fault.
[0006] The Leader's Reference Trajectory It is generated by the following dynamic equation: in, For the nonlinear dynamic function of the leader.
[0007] Furthermore, the adaptive state identifier is designed as follows: ; in, These are estimates of the neural network weights. This is a fault estimate. It is an estimate of the basis function vector of the neural network. It is the observer gain; the corresponding adaptive update law is: in, To identify errors, and It is a positive definite gain matrix. and All are normal numbers.
[0008] Furthermore, based on the state information provided by the online identifier, a sliding surface is designed to reduce the system order, as follows: Defined sliding mode variables and preset performance conversion error The details are as follows: Define tracking error ,in This is the state estimate. To establish a sliding mode variable for the leader's reference trajectory: ; in, The designed polynomial satisfies the Hurwitz characteristic polynomial. . a normal number.
[0009] Define neighbor consistency error Introduce a preset performance function: ; in, and They represent respectively The limiting value and convergence rate, .
[0010] Define normalization error The conversion error is obtained through the conversion function. : in, and These are performance boundary parameters.
[0011] Furthermore, the performance index function is: ; in, It is the cost function.
[0012] The derived ideal optimal control law is in the form of: ; in, The normalized derivative of the transformation function, and the topology-dependent parameters. , .
[0013] Furthermore, Actor-Critic reinforcement learning includes critic networks and actor networks: The critic network is used to approximate the optimal performance index function. Its output is: ; Actor networks are used to approximate optimal control strategies. Its output is: ; in, and These are the weights of the Critic and Actor networks, respectively. These are control parameters.
[0014] Design a gradient descent-based weight update law to update network weights online using Bellman residuals: ; in, , and These represent the learning rates of critics and actors, respectively.
[0015] The present invention has the following beneficial effects: 1. This invention simultaneously handles the multiplicative time-varying fault and the additive time-varying deviation fault of the actuator in a single control structure, thereby enhancing the reliability of the system. Furthermore, by introducing preset performance control, it strictly ensures that the consistency tracking error of the multi-agent system meets the predefined performance boundary in both the transient and steady-state phases, thus avoiding excessive overshoot. 2. This invention utilizes the Actor-Critic reinforcement learning algorithm to achieve near-optimal control without requiring a precise system dynamics model, balancing control performance and energy consumption; and the designed adaptive flagger can effectively identify unknown nonlinear dynamics and faults, achieving proactive compensation for uncertainties. Attached Figure Description
[0016] Figure 1 This is a system control principle block diagram provided in an embodiment of the present invention; Figure 2 These are the state components of each follower agent in the embodiments of the present invention. Reference signals for leaders Tracking effect diagram; Figure 3 These are the state components of each follower agent in the embodiments of the present invention. Reference signals for leaders Tracking effect diagram; Figure 4 These are the state components of each follower agent in the embodiments of the present invention. Reference signals for leaders Tracking effect diagram; Figure 5 This is the consistency error in the embodiments of the present invention. Evolution curves relative to the preset performance boundary; Figure 6 The weights of the Actor neural network in this embodiment of the invention. The online adaptive convergence curve; Figure 7 The weights of the Critic neural network in this embodiment of the invention. The online adaptive convergence curve; Figure 8 The adaptive identifier neural network weights in this embodiment of the invention The online adaptive convergence curve; Figure 9 This is the estimation effect of the adaptive identifier on the additive time-varying deviation fault of the actuator in the embodiment of the present invention. The online adaptive convergence curve; Figure 10 These are the control input signals for each intelligent agent in the embodiments of the present invention. Change curve graph; Figure 11 This is a flowchart of the control method provided in an embodiment of the present invention. Detailed Implementation
[0017] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments: The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0018] refer to Figure 1 This embodiment includes N Taking a high-order nonlinear multi-agent system with one follower and one leader as an example, the effectiveness of the pre-set performance fault-tolerant control method based on adaptive reinforcement learning proposed in this invention is verified. The specific steps are as follows: S1: Construct a high-order nonlinear multi-agent system model with actuator faults, specifically: Consider by N A multi-agent system consisting of one follower and one leader, whose communication topology is an undirected graph. G Description. (Page number missing) k The dynamic model of a follower agent is as follows: in, The system state vector; function For unknown continuous nonlinear dynamic functions; This is the actual control input after a fault occurs.
[0019] This embodiment considers that the actuator simultaneously has multiplicative faults and additive time-varying deviation faults, and the fault model is established as follows: in, Represents the actual control input. Matrix For actuator efficiency factor matrix (multiplicative fault). It is a bounded time-varying additive fault.
[0020] The Leader's Reference Trajectory It is generated by the following dynamic equation: in, For the nonlinear dynamic function of the leader.
[0021] Assumption 1: Leader State sum function Bounded.
[0022] Definition 1: If the states in a high-order typical multi-agent system (1) satisfy... This indicates that the multi-agent system can achieve consistency.
[0023] Control objective: For a typical nonlinear multi-agent system, design a controller that: 1) All closed-loop signals are semi-globally consistent and eventually bounded (SGUUB). 2) Followers achieve preset accuracy tracking of the leader's corresponding state.
[0024] S2: Design an adaptive state flagr based on a neural network, as follows: To address the unknown nonlinear terms and unknown time-varying faults in the system, an adaptive state identifier based on a radial basis function neural network is designed for online approximation and estimation.
[0025] Construct a status identifier of the following form: in, For state The estimated value; These are estimates of the neural network weights; The Gaussian function vector; This is a fault estimate. It is the observer gain.
[0026] To ensure the convergence of identification errors, the following adaptive law is designed to adjust the network weights and fault estimates online: in, To identify errors, and It is a positive definite gain matrix. and This is a damping parameter used to improve the robustness of the algorithm. Through this identifier, the system achieves proactive perception and compensation for model uncertainties and faults.
[0027] S3: Constructing the sliding surface and preset performance error conversion. To reduce the control difficulty of high-order systems and introduce preset performance constraints, sliding mode control and preset performance control techniques are combined, as follows: Defining tracking error and sliding surface: Defining tracking error based on observed state Constructing generalized sliding mode variables : in, The parameters are used to make the polynomial Hurwitz stable.
[0028] Define a preset performance conversion error: To constrain consistency error, define a local consistency error based on neighbor information. Introducing a preset performance function with exponential decay. and define performance boundaries and Introducing normalization error And it uses a logarithmic error transformation function to map it into an unconstrained transformation error. : Define the normalized derivative of the transformation function as: By controlling Boundedness guarantees the original error. It always remains within the preset performance envelope.
[0029] S4: Based on optimal control theory, the ideal control law is derived. To balance the system's tracking performance and control energy consumption, an infinite time-domain performance index function of the following form is defined: Based on the Bellman optimality principle, the Hamiltonian function is defined. According to the extreme value condition The ideal optimal fault-tolerant control law is derived. The format is: in, For parameters related to communication topology, This is the gradient term for the performance metric. Because... Analytical solutions are difficult to obtain, and unknown fault parameters exist in the system. Therefore, reinforcement learning methods are needed to approximate it.
[0030] S5: Design an Actor-Critic reinforcement learning controller, as follows: Critic Network Design: Using Critic Networks to Approximate the Gradient Term of the Optimal Performance Metric Output of the Critic network Its weight update law is designed as follows: in, For Critic network weights, For learning rate, For regularization terms; Actor Network Design: Approximating the Optimal Control Strategy using Actor Networks The output of the Actor network Its weight update law is designed as follows: in, For Actor network weights, The learning rate is used. The Actor network utilizes the evaluation information provided by the Critic to minimize the Bellman residual using gradient descent, thereby continuously optimizing the control strategy.
[0031] Final control law: The network output It is applied as an actual control signal to the multi-agent system.
[0032] To verify the effectiveness of the method of the present invention, a high-order nonlinear multi-agent simulation system containing 6 follower agents and 1 leader was built.
[0033] Simulation condition settings: 1) No. k A follower intelligent agent ( k= The third-order typical dynamic model of (1,…,6) is described as follows: Wherein, the coefficient vector of the nonlinear function is set as , The Leader's Reference Trajectory Generated dynamically from the following: .
[0034] 2) The communication connections between multiple agents in the system are described by adjacency matrix A, and the communication between agents and the leader is described by matrix B. The specific values are set as follows: , This topology indicates that only the second agent can directly obtain information about the leader.
[0035] 3) Actuator failure Including multiplicative efficiency loss and additive time-varying bias, the fault model is defined as follows: Among them, the actuator health factor matrix Simulate different levels of efficiency loss.
[0036] 4) A radial basis function neural network is used to approximate the unknown dynamics, with a certain number of network nodes. basis function width center point It is uniformly distributed in the interval [-8, 8].
[0037] The initial system state is set as follows: The initial values of the neural network weights are set to... , ; 5) To ensure system stability and convergence, the following key control parameters are selected: adaptive law parameters. 4.1 , Reinforcement learning parameters: Preset performance parameters: , , , .
[0038] Simulation Result Analysis: 1) Tracking effect: such as Figure 2 , Figure 3 and Figure 4 As shown, despite severe actuator failures and model uncertainties, the states of each follower agent... All can quickly and accurately track the leader's reference trajectory. This achieves system consistency.
[0039] 2) Preset performance retention: such as Figure 5 As shown, the system's consistency error Always strictly maintain the preset performance envelope (by...) Within the defined upper and lower bounds, and with fast convergence speed, the effectiveness of the preset performance control was verified.
[0040] 3) Parameter convergence: such as Figure 6 , Figure 7 As shown, the weights of the Actor and Critic neural networks The ability to converge and remain stable in a short time indicates that the reinforcement learning algorithm has successfully learned an approximately optimal control strategy.
[0041] 4) Fault estimation: such as Figure 8 , Figure 9 As shown, adaptive parameters and The rapid convergence indicates that the unknown dynamics and faults of the system are effectively estimated and compensated, and the control law tends to be optimal.
[0042] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0043] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0044] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0045] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0046] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A pre-defined performance fault-tolerant control method for a nonlinear multi-agent system, characterized in that, Includes the following steps: S1: Considering the multiplicative and additive time-varying faults of the actuators, establish a high-order nonlinear multi-agent system model; S2: For unknown nonlinear functions and fault terms in high-order nonlinear multi-agent system models, an online identifier is constructed using an RBF neural network to reconstruct the system state and estimate the fault value. S3: Based on the state information provided by the online identifier, a sliding surface is designed to reduce the system order. Pre-defined performance techniques are introduced to constrain and transform the consistency error of the sliding surface, ensuring that the tracking error is always within the predefined performance envelope. S4: Based on the transformed error system, define the performance index function, derive the HJB equation and the theoretical optimal control law, and transform the fault-tolerant control into an optimal regulation problem; S5: Actor-Critic reinforcement learning is used to solve the optimal regulation problem, obtain the optimal control, and output the final control quantity to act on the system.
2. The method for preset performance fault-tolerant control of a nonlinear multi-agent system according to claim 1, characterized in that, Considering the multiplicative and additive time-varying faults of the actuators, a high-order nonlinear multi-agent system model is established, as follows: Consider by N A multi-agent system consisting of one follower and one leader, whose communication topology is an undirected graph. G Description, the first k The dynamic model of a follower agent is as follows: in, For the first k The state vector of each agent. The system state vector; function For unknown continuous nonlinear dynamic functions; This is the actual control input after a fault occurs; Considering that the actuator simultaneously exhibits multiplicative faults and additive time-varying deviation faults, the fault model is established as follows: in, Represents actual control input; diagonal matrix For the actuator efficiency factor matrix, It is a bounded, time-varying additive fault; The Leader's Reference Trajectory It is generated by the following dynamic equation: in, For the nonlinear dynamic function of the leader.
3. The method for preset performance fault-tolerant control of a nonlinear multi-agent system according to claim 2, characterized in that, The adaptive state identifier is designed as follows: in, These are estimates of the neural network weights. It is an estimate of the basis function vector of the neural network. This is a fault estimate. It is the observer gain; the corresponding adaptive update law is: in, To identify errors, and It is a positive definite gain matrix. and All are normal numbers.
4. The pre-set performance fault-tolerant control method for a nonlinear multi-agent system according to claim 3, characterized in that, Based on the state information provided by the online identifier, a sliding surface is designed to reduce the system order, as detailed below: Defined sliding mode variables and preset performance conversion error The details are as follows: Define tracking error ,in This is the state estimate. To establish a sliding mode variable for the leader's reference trajectory: in, The designed polynomial satisfies the Hurwitz characteristic polynomial. positive numbers; Define neighbor consistency error Introduce a preset performance function: in, and They represent respectively The limiting value and convergence rate, ; Define normalization error The conversion error is obtained through the conversion function. : in, and These are performance boundary parameters.
5. A pre-set performance fault-tolerant control method for a nonlinear multi-agent system according to claim 4, characterized in that, The performance index function is: in, It is the cost function; The derived ideal optimal control law is in the form of: in, The normalized derivative of the transformation function, and the topology-dependent parameters. , .
6. The method for preset performance fault-tolerant control of a nonlinear multi-agent system according to claim 5, characterized in that, The Actor-Critic reinforcement learning includes a critic network and an actor network: Critics networks are used to approximate the optimal performance metric function. Its output is: ; Actor networks are used to approximate optimal control strategies. Its output is: in, and These are the weights of the Critic and Actor networks, respectively. For control parameters; Design a gradient descent-based weight update law to update network weights online using Bellman residuals: in, , and These represent the learning rates of critics and actors, respectively.