An observer-based multi-agent system reinforcement learning control method and control system

By constructing an adaptive observer and an execution-evaluation reinforcement learning framework, the problems of unpredictable state and time-varying faults in adversarial multi-agent systems are solved, thereby improving the stability and cooperative performance of the system and making it suitable for diverse control scenarios.

CN122219094APending Publication Date: 2026-06-16LIAONING UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LIAONING UNIVERSITY OF TECHNOLOGY
Filing Date
2026-03-25
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively address the issues of unpredictable states, time-varying faults, and full-state constraints in adversarial multi-agent systems, leading to system instability and decreased collaborative performance.

Method used

By establishing a dynamic model of a strictly feedback nonlinear multi-agent system, transforming it into an unconstrained pure feedback model, constructing an adaptive observer and combining it with an execution-evaluation reinforcement learning framework, designing an optimal controller, approximating the Hamilton-Jacobi-Isaacs equations, optimizing virtual control strategies and disturbance strategies, and verifying stability.

Benefits of technology

It solves the problems of full state constraints, unpredictable states, and time-varying faults simultaneously within a zero-sum game framework, improving the system's stability and collaborative performance. It is applicable to diverse control scenarios and provides reliable state information and efficient real-time control capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122219094A_ABST
    Figure CN122219094A_ABST
Patent Text Reader

Abstract

The application relates to an observer-based multi-agent system reinforcement learning control method and a control system, which are applied to the technical field of multi-agent system control, and the method comprises the following steps: a first dynamic model is established and is converted into a second dynamic model without constraints and pure feedback; an adaptive observer and an updating law are constructed; the stability of the observer is verified by using a Lyapunov function; the system is reconstructed based on the observer, and an optimal controller is designed; a consensus error dynamic is constructed; an execution-evaluation reinforcement learning framework is established to approximate a Hamilton-Jacobi-Isaacs equation; an optimal virtual control strategy and a worst-case disturbance strategy are generated; a network weight updating law is designed to iteratively optimize the control strategy; the stability is verified by using a second Lyapunov function; and finally, optimal cooperative control of the multi-agent system is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of multi-agent system control technology, and in particular to an observer-based reinforcement learning control method and control system for multi-agent systems. Background Technology

[0002] Among related technologies, multi-agent systems have become one of the core research directions in control theory due to their wide application in formation control, UAV swarms, and satellite formation. Through information interaction and collaborative decision-making among multiple agents, the system can achieve collective goals such as consistent control, synchronization, and formation maintenance. In practical applications, multi-agent systems often face multiple challenges, including strong nonlinear coupling, unpredictable states, time-varying faults, full-state constraints, and adversarial interactions among multiple agents, making it difficult for traditional control methods to meet the requirements of high-performance control.

[0003] Reinforcement learning, as a model-free intelligent control technique, does not rely on prior system models and optimizes control strategies through interaction with the environment, providing an effective approach for the control of nonlinear multi-agent systems. Existing research largely focuses on cooperative multi-agent systems, with relatively little attention paid to adversarial scenarios. In zero-sum games, multiple agents have conflicting objectives, where one agent's gain is equivalent to another's loss. Agents must optimize their own performance while resisting interference from adversaries. The core challenge lies in solving the Hamilton-Jacobi-Isax equations, whose complexity increases exponentially with the system order, making direct solutions difficult.

[0004] Furthermore, full-state constraints and time-varying bias faults further increase the difficulty of control design. If not effectively addressed, constraint violations may lead to system instability, while faults can degrade cooperative performance or even cause task failure. Existing fault-tolerant control methods largely rely on the assumption of state measurability, but for adversarial multi-agent systems with unmeasurable states and full-state constraints, there is a lack of effective integrated control schemes. Therefore, there is an urgent need to propose a control method that can simultaneously address adversarial interactions, unmeasurable states, time-varying faults, and full-state constraints. Summary of the Invention

[0005] To address or partially address the problems existing in related technologies, this application provides an observer-based reinforcement learning control method and control system for multi-agent systems. This method can solve problems such as adversarial interaction, unpredictable state, time-varying faults, and full-state constraints in adversarial multi-agent systems with unpredictable state and full-state constraints.

[0006] The first aspect of this application provides a reinforcement learning control method for a multi-agent system based on an observer, comprising: Establish the first dynamic model of a rigorous feedback nonlinear multi-agent system containing full-state constraints, unmeasurable states, and time-varying bias faults; The first dynamic model is transformed into a second dynamic model with unconstrained pure feedback. Based on the second dynamic model, an adaptive observer is constructed; Based on the adaptive observer, the adaptive update law is derived, and the adaptive observer is optimized. The stability of the adaptive observer is verified by constructing a first Lyapunov function. The second dynamic model is reconstructed based on the adaptive observer, the optimal controller is designed, and the consensus error dynamics among multiple agents are constructed. An execution-evaluation reinforcement learning framework is established to approximate the optimal value function of the Hamilton-Jacobi-Isaacs equation. Based on the consensus error, the optimal virtual control strategy and worst-case perturbation strategy of the multi-agent system are dynamically generated. Based on the optimal value function, the optimal multi-agent system virtual control strategy, and the worst-case perturbation strategy, a network weight update law is designed to iteratively optimize the optimal multi-agent system virtual control strategy. A second Lyapunov function is constructed to verify the stability of the virtual control strategy of the optimal multi-agent system. The multi-agent system is controlled based on the optimal multi-agent system virtual control strategy.

[0007] As one implementation of the first aspect, the transformation of the first dynamic model containing full-state constraints into a second dynamic model with unconstrained pure feedback includes: The first dynamic model is subjected to a one-to-one nonlinear mapping process to obtain an unconstrained pure feedback system. Based on the aforementioned unconstrained pure feedback system, a radial basis function neural network is used to approximate the first unknown nonlinear function. By disassembling the first dynamic model and incorporating the multi-agent interaction characteristics, faults, and disturbances, a second dynamic model is obtained.

[0008] As one implementation of the first aspect, constructing the adaptive observer includes: Obtain the unmeasurable states and bias faults of the second dynamic model; Based on the unmeasurable state and the bias fault, a radial basis function neural network is used to approximate the second unknown nonlinear function; Based on the radial basis function neural network approximation form of the second unknown nonlinear function, and combined with the reconstructed system dynamics, an adaptive observer is constructed for estimating unmeasurable states and time-varying bias faults.

[0009] As one implementation of the first aspect, the optimization of the adaptive observer based on deriving the adaptive update law includes: Define the gain and approximation error boundary of the adaptive observer; The effects of the bias fault are offset by feedforward compensation. The estimated value is used to replace the unmeasurable state; The adaptive update law is derived by combining the learning rate, activation function, and positive definite matrix.

[0010] As one implementation of the first aspect, a first Lyapunov function is constructed to verify the stability of the adaptive observer, including: Determine the estimation error of the adaptive observer, the approximate weight error of the unknown nonlinear function, and the estimation error of the bias fault, and construct the first Lyapunov function containing the inverse of the positive definite matrix; The effectiveness of the adaptive observer is verified by taking the derivative of the first Lyapunov function and scaling it using inequalities.

[0011] As one implementation of the first aspect, the second dynamic model is reconstructed based on the adaptive observer, an optimal controller is designed, and a consensus error dynamic among multiple agents is constructed, including: The second dynamic model is reconstructed; Define coordinate transformations and construct a first-order filter; The second dynamic model, the coordinate transformation, and the first-order filter are input into a zero-sum game framework, and the optimal controller is designed hierarchically using p-step backstepping techniques. Determine the agent's local tracking error and neighboring tracking error; Determine adjacency weights based on multi-agent communication topology; The consensus error is dynamically calculated by weighting the local tracking error and the adjacent tracking error with the adjacency weights respectively.

[0012] Construct the dynamic consensus error among multiple agents; As one implementation of the first aspect, generating the optimal multi-agent system virtual control strategy and the worst-case perturbation strategy includes: Construct execution networks and evaluation networks; The execution-evaluation reinforcement learning framework is established based on the execution network and the evaluation network; Collect operational data from multi-agent systems; The execution network and the evaluation network are iteratively updated based on the operational data. The optimal value function of the Hamilton-Jacobi-Isaac equation is approximated based on the evaluation network described above. Based on the optimal value function of the Hamilton-Jacobi-Isax equation, the optimal virtual control strategy and the worst-case perturbation strategy of the multi-agent system are generated through the execution network. As one implementation of the first aspect, the optimal multi-agent system virtual control strategy is iteratively optimized, including: Based on the optimal value function, the optimal multi-agent system virtual control strategy, and the worst-case perturbation strategy, a network weight update law is designed. The optimal virtual control strategy for the multi-agent system is iteratively optimized using the network weight update law.

[0013] As one implementation of the first aspect, the second Lyapunov function is constructed to verify the stability of the virtual control strategy of the optimal multi-agent system, including: The consensus error among multiple agents, the relative error caused by coordinate transformation, and the approximation error between the execution network and the evaluation network are obtained. The second Lyapunov function is constructed based on the consensus error, the relative error, and the approximation error; The stability of the virtual control strategy for the optimal multi-agent system is verified based on the second Lyapunov function.

[0014] A second aspect of this application provides an observer-based multi-agent system reinforcement learning control system, comprising: The model building module is used to establish the first dynamic model of a strictly feedback nonlinear multi-agent system containing full-state constraints, unmeasurable states, and time-varying bias faults. The conversion module is used to convert the first dynamic model with full-state constraints into a second dynamic model with unconstrained pure feedback. An observer building module is used to build an adaptive observer based on the second dynamic model; An observer update module is used to derive the adaptive update law of the adaptive observer; The first Lyapunov function construction module is used to construct the first Lyapunov function to verify the stability of the adaptive observer. The controller design module is used to design the optimal controller and construct the consensus error dynamics among multiple agents. The network approximation and update module is used to generate the optimal virtual control strategy and worst-case perturbation strategy for the multi-agent system; The second Lyapunov function construction module is used to construct a second Lyapunov function to verify the stability of the virtual control strategy of the optimal multi-agent system.

[0015] The technical solution provided in this application can include the following beneficial effects: by organically integrating one-to-one nonlinear mapping processing, adaptive observer and execution-evaluation reinforcement learning framework, an integrated control framework is constructed. For the first time, under the zero-sum game framework, the full state constraints, unmeasurable state and time-varying bias faults of strictly feedback nonlinear multi-agent systems are solved simultaneously, filling the gap in existing technology and being more adaptable to diverse control scenarios.

[0016] Furthermore, the adaptive observer, combined with the adaptive update law, can accurately estimate unmeasurable states and time-varying faults, providing reliable state information for subsequent control design.

[0017] Furthermore, an execution-evaluation reinforcement learning framework is established to efficiently approximate the optimal solution of the Hamilton-Jacobi-Isaacs equations without the need for offline solution of complex equations, making it suitable for real-time control scenarios. Two Lyapunov stability analyses ensure that all error signals are bounded, resulting in high system output tracking accuracy and strong fault tolerance.

[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0019] The above and other objects, features and advantages of this application will become more apparent from the more detailed description of exemplary embodiments thereof in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components in the exemplary embodiments thereof.

[0020] Figure 1 This is a flowchart illustrating a reinforcement learning control method for a multi-agent system according to an embodiment of this application; Figure 2 This is a diagram illustrating the architecture of a multi-agent system reinforcement learning control system in an embodiment of this application.

[0021] Symbol explanation: 1-Model building module; 2-Transformation module; 3-Observer building module; 4-Observer update module; 5-First Lyapunov function building module; 6-Controller design module; 7-Network approximation and update module; 8-Second Lyapunov function building module. Detailed Implementation

[0022] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to make this application more thorough and complete, and to fully convey the scope of this application to those skilled in the art.

[0023] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0024] It should be understood that although the terms "first," "second," "third," etc., may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0025] Multi-agent systems, with their wide application in formation control, UAV swarms, and satellite formation, have become one of the core research directions in control theory. However, existing technologies lack effective integrated control schemes for adversarial multi-agent systems with unpredictable states and full state constraints. Therefore, there is an urgent need to propose a control method that can simultaneously address adversarial interactions, unpredictable states, time-varying faults, and full state constraints.

[0026] To address the aforementioned issues, this application provides an observer-based reinforcement learning control method for multi-agent systems. This method can simultaneously solve problems such as full-state constraints, unmeasurable states, and time-varying bias faults in strictly feedback nonlinear multi-agent systems within a zero-sum game framework, filling a gap in existing technologies and being more adaptable to diverse control scenarios.

[0027] This application provides an observer-based reinforcement learning control method for a multi-agent system, the flowchart of which is shown below. Figure 2 As shown, it includes: S1: Establish the first dynamic model of a strictly feedback nonlinear multi-agent system containing full-state constraints, unmeasurable states, and time-varying bias faults.

[0028] Specifically, the first dynamic model is Equation (1).

[0029] in, Represents the dimension of the intelligent agent. Represents the agent's ID. In order to be in t Time of the first k The one-dimensional state change rate of an agent In order to be in t Time of the first k A single agent l Rate of change of state in dimensional state In order to be in t Time of the first k A single agent m Rate of change of state in a given dimension It is the first k Input of an agent, For the first k An intelligent agent at time t One-dimensional state variable, For the first k An intelligent agent at time t Two-dimensional state variables, For the first k An intelligent agent at any given moment l +1 dimension state variable, For the first k The first-order state vector of each agent. , indicating the first k One intelligent agent to l A state vector of order (state set form). , indicating the first k A smart agent l arrive m A state vector of order (state set form). For the system input affected by time-varying actuator bias fault, For time t, the first k External perturbations in the one-dimensional state of an agent for t Time of the first k A single agent l External disturbances in a dimensional state. For time t, the first k A single agent m External disturbances in a dimensional state. For the first k A first-order unknown continuous nonlinear function for an intelligent agent. For the first k A single agent l Unknown continuous nonlinear function of order, For the first k A single agent m Unknown continuous nonlinear function of order, f Marking the control input section unaffected by time-varying actuator bias faults. qThe total number of intelligent agents. m It is the total dimension of the state variables.

[0030] Furthermore, the unknown continuous nonlinear function is an inherent nonlinear characteristic of multi-agent systems. k A single agent m The rate of change of dimensional state constitutes the first dynamic model of the output.

[0031] The input to the first dynamic model is equation (2).

[0032] in, For the first k The actual control input of an agent with a time-varying bias fault, distinct from the ideal control input, is the real input signal acting on the system after the actuator's time-varying bias fault. This parameter is input into the first dynamic model to clarify the form of the time-varying bias fault's influence on the control input, providing a core research object and design basis for subsequent design of fault-tolerant control strategies and mitigating the fault's effects.

[0033] Indicates the actual control input. It is a bounded fault vector that satisfies >0, This represents the upper limit of the fault range, indicating that the fault magnitude is limited and will not increase indefinitely.

[0034] Before constructing the first dynamic model, it is necessary to define the system's dynamic variables, nonlinear functions, fault models, and external disturbances, clarify the full-state constraint range, unmeasurable state types, bounded characteristics of fault vectors, and disturbance attributes, determine the multi-agent topology, communication rules, and core properties, construct a rigorous feedback nonlinear framework, and clarify the state measurability range and nonlinear function characteristics.

[0035] The above parameters were determined entirely based on existing experience. Through continuous trial and error in simulation experiments, suitable parameters were finally obtained.

[0036] System dynamic variables characterize the state changes of each agent; nonlinear functions reflect the inherent nonlinear characteristics of the system; the fault model quantifies actuator bias faults; external disturbances are environmental interferences; the full-state constraint range is the state safety boundary; unmeasurable states are second-order and above states; the bounded property of the fault vector indicates that the fault amplitude is finite; the disturbance attribute is unknown, time-varying, and bounded; the multi-agent topology and communication rules define the information interaction relationships between agents; and the core properties include... Lipschitz The continuous and strictly feedback nonlinear framework is the standard expression of the system. The state measurability range is limited to the measurability of only the first-order state, and the nonlinear function is continuously differentiable.

[0037] Based on the above parameters, we first define the inherent characteristics and engineering constraints of the multi-agent system, then quantify the faults and disturbances and incorporate them into a strict feedback framework to construct the first dynamic model.

[0038] S2: Transform the first dynamic model into a second dynamic model with unconstrained pure feedback.

[0039] Specifically, the first dynamic model is transformed into the second dynamic model based on one-to-one nonlinear processing.

[0040] S3: Construct an adaptive observer based on the second dynamic model.

[0041] S4: Based on the adaptive observer, derive the adaptive update law and optimize the adaptive observer.

[0042] S5: Construct the first Lyapunov function to verify the stability of the adaptive observer.

[0043] S6: Reconstruct the second dynamic model based on the adaptive observer, design the optimal controller, and construct the consensus error dynamics among multiple agents.

[0044] S7: Establish an execution-evaluation reinforcement learning framework to approximate the optimal value function of the Hamilton-Jacobi-Isaacs equation, and dynamically generate the optimal virtual control strategy and worst-case perturbation strategy for the multi-agent system based on consensus error.

[0045] Among them, the consensus error dynamics among multiple agents is the "target anchor point" of the whole method. Based on this, the optimal virtual control strategy and worst-case perturbation strategy for the multi-agent system are determined through the execution-evaluation reinforcement learning framework.

[0046] S8: Based on the optimal value function, the optimal multi-agent system virtual control strategy, and the worst-case perturbation strategy, design a network weight update law to iteratively optimize the optimal multi-agent system virtual control strategy.

[0047] S9: Construct a second Lyapunov function to verify the stability of the virtual control strategy for the optimal multi-agent system.

[0048] Specifically, during the stability verification process, the second Lyapunov function optimizes the optimal virtual control strategy of the multi-agent system in real time, and the stability verification ends after the optimization is completed.

[0049] S10: Control the multi-agent system based on the optimal multi-agent system virtual control strategy.

[0050] Furthermore, the iteratively optimized virtual control strategy of the multi-agent system is fused with the output of the adaptive observer. Dynamic surface control technology is used to alleviate computational complexity and generate the actual control inputs for each agent, ensuring that the multi-agent system achieves accurate consensus tracking under adversarial disturbances, state constraints, and fault conditions.

[0051] In this embodiment, an integrated control framework is constructed by organically integrating one-to-one nonlinear mapping processing, an adaptive observer, and an execution-evaluation reinforcement learning framework. For the first time, this framework simultaneously solves problems such as full-state constraints, unmeasurable states, and time-varying bias faults in strictly feedback nonlinear multi-agent systems within a zero-sum game framework, filling a gap in existing technologies and offering greater adaptability to diverse control scenarios. Furthermore, the adaptive observer, combined with an adaptive update law, can accurately estimate unmeasurable states and time-varying faults, providing reliable state information for subsequent control design. In addition, the execution-evaluation reinforcement learning framework efficiently approximates the optimal solution of the Hamilton-Jacobi-Isax equations without requiring offline solution of complex equations, making it suitable for real-time control scenarios. Two Lyapunov stability analyses ensure that all error signals are bounded, resulting in high system output tracking accuracy and strong fault tolerance.

[0052] In embodiments of this application, transforming the first dynamic model containing full-state constraints into a second dynamic model with unconstrained pure feedback includes: S21: Perform a one-to-one nonlinear mapping on the first dynamic model to obtain an unconstrained pure feedback system.

[0053] An unconstrained pure feedback system has no hard boundary restrictions on any of its states, and its values ​​range across the entire real number domain. An unconstrained pure feedback system is an equivalent system obtained by transforming the first dynamical model (a strict feedback system with full state constraints) through a one-to-one nonlinear mapping. Its core function is to eliminate the interference of hard constraints on control design.

[0054] Furthermore, the one-to-one nonlinear mapping process is completed based on equation (3).

[0055] Equation (3) in, For the first k The l-th initial state of an agent. For the k-th agent, the first... l An initial state, For the first k The first agent of the intelligent agent m An initial state, For based on The first-dimensional unbounded state obtained by the transformation, For based on The converted firstl Unbounded state of order, For based on The converted first m Unbounded state of order, For the first k The l-th positive natural exponential function of an intelligent agent. For the first k The first intelligent agent l Positive natural exponential function, For the first k The first intelligent agent m Positive natural exponential function, For the first k The l-th negative natural exponential function of an agent. For the first k The first intelligent agent l Negative natural exponent function For the first k The first intelligent agent m Negative natural exponent function For the first k The lower bound threshold of the l-th state of an agent. For the first k The first intelligent agent l The lower bound threshold of each state, For the first k The first intelligent agent m The lower bound threshold of each state. For the first k The upper bound threshold of the l-th state of an agent. For the first k The first intelligent agent l The upper bound threshold of each state, For the first k The first intelligent agent m The upper bound threshold of each state, For the first k The l-th state of an agent is related to time. t The derivative, For the first k The first intelligent agent l A state related to time t The derivative, For the first k The first intelligent agent m A state related to time t The derivative of .

[0056] Furthermore, the original state is used to describe the original dynamics of the system (such as the position and velocity of the UAV), but is subject to physical constraints on the upper and lower bound thresholds of the corresponding agent state.

[0057] For the firstl The nonlinear gain coefficient for softening state constraints has a numerator that is a smooth, bounded nonlinear function, and a denominator that is the sum of the upper and lower boundaries of the state. This transforms the "hard constraint" into a nonlinear term in the dynamics, avoiding the nonsmoothness problem caused by the constraint. The nonlinear gain coefficient for softening the first state constraint and the... m The same applies to the nonlinear gain coefficients for state-constrained softening.

[0058] Transforming the original state into an unbounded state can eliminate the constraints of the original state, reconstruct the first dynamic model into a pure feedback unconstrained system, and reduce the complexity of control design.

[0059] Specifically, full-state constraints refer to the strict upper and lower bounds on all states of all agents. The upper and lower bounds are the corresponding upper and lower bound thresholds. Full-state constraints cover all states of all agents and all orders, and are hard safety boundaries that must be satisfied by multi-agent systems.

[0060] Based on equation (3), design the mapping relationship for each constraint state of the first dynamic model and derive the inverse mapping. Substitute the inverse mapping of the first dynamic model into equation (1) to obtain the unconstrained pure feedback system.

[0061] S22: Based on the unconstrained pure feedback system, the first unknown nonlinear function is approximated by combining a radial basis function neural network.

[0062] Among them, radial basis function neural networks (RBFNNs) have universal approximation characteristics and can approximate unknown nonlinear functions and uncertainties of the system with high accuracy. The first unknown nonlinear function is inherent in the nonlinear multi-agent system itself, such as the nonlinear dynamic characteristics caused by air resistance and its own weight distribution when a drone is flying.

[0063] S23: Deconstruct the first dynamic model, incorporate multi-agent interaction characteristics and fault and disturbance terms to obtain the second dynamic model.

[0064] Specifically, the first dynamic model is decomposed into a model targeting... k The first dynamic model with =1 is processed, and the full-state constraints are handled through one-to-one nonlinear mapping. The constrained strict feedback form is transformed into an unconstrained pure feedback form. The time-varying bias fault of the actuator and the external bounded disturbance are integrated, and the consensus tracking error and the virtual error surface are defined. Finally, the normalization construction of the second dynamic model is completed.

[0065] The second dynamic model is Equation (4).

[0066] in, In order to be int Time of the first k The one-dimensional rate of change of an agent's unconstrained state In order to be in t Time of the first k An agent in an unconstrained state l Rate of change of state in a given dimension In order to be in t Time of the first k An agent in an unconstrained state m Rate of change of state in a given dimension For the first k An intelligent agent at time t Unconstrained two-dimensional state variables, For the first k An agent's unconstrained state at any given moment l+ 1-dimensional state variable, For the first k An unconstrained first-order state vector of an agent. , indicating the first k One intelligent agent to l An unconstrained state vector of order (in state set form). , indicating the first k A smart agent l arrive m An unconstrained state vector of order (in state set form). For the first k A first-order unknown continuous nonlinear function reconstructed by an intelligent agent For the first k Reconstruction of individual agents l Unknown continuous nonlinear function of order, For the first k Reconstruction of individual agents m An unknown continuous nonlinear function.

[0067] Reconstructed l Unknown continuous nonlinear function of order Replace the original inherent l Unknown continuous nonlinear function of order ,and The input order is adapted to the reconstructed dynamic correlation logic. and Similarly.

[0068] In this embodiment, a strictly feedback system with full-state constraints is transformed into an unconstrained pure feedback system through a one-to-one nonlinear mapping. The inherent unknown nonlinear function of the system is approximated using a radial basis function neural network. By integrating single-agent dynamics decomposition, multi-agent interaction characteristics, actuator time-varying bias faults, and external disturbance terms, a normalized second dynamic model is constructed. This not only eliminates the non-smooth interference and control complexity of hard constraints on control design, but also effectively handles system uncertainties by leveraging the general approximation capability of neural networks. This lays a decoupled and easily expandable model foundation for subsequent adaptive observer design, optimal controller construction, and reinforcement learning strategy iteration, thereby achieving effective support for the safe and collaborative control of multi-agent systems under complex constraints.

[0069] In embodiments of this application, constructing an adaptive observer includes: S31: Based on the second dynamic model, obtain the unmeasurable state and bias fault of the observed target.

[0070] Specifically, by using an adaptive observer combined with an RBF neural network to approximate nonlinearities and dynamically adjusting the estimated values ​​using an adaptive update law, the estimated values ​​of unmeasurable states and bias faults can be obtained. During observation, the unmeasurable state is targeted at the state of the unconstrained pure feedback system, and the fault is targeted at the state of the system's highest-order control input layer (i.e., the m-dimensional state). The overall system is in a multi-agent interactive state with disturbances and faults, and the observation is completed based on the first-order measurable state.

[0071] S32: Based on the unmeasurable state and the bias fault, a radial basis function neural network is used to approximate the second unknown nonlinear function.

[0072] Specifically, we first define the second unknown nonlinear function to be approximated as... (depending on unconstrained state set vector) (Inheriting the inherent nonlinear characteristics of the original system), based on the universal approximation theorem, radial basis neural networks can approximate the continuous nonlinear function with arbitrary precision on compact sets; subsequently, a three-layer radial basis network of "input layer-hidden layer-output layer" is constructed, with the input being an unconstrained state set vector. The hidden layer uses a Gaussian kernel function to... Map the input to a high-dimensional feature vector (i.e., basis function vector), the network output is represented as the ideal weight vector. With basis function vectors The product plus a small approximation error Then, the weights are estimated in real time using an adaptive update law. By gradually reducing the approximation error, the approximation accuracy can be ultimately achieved. The precise quantification is approaching.

[0073] The second unknown nonlinear function is a nonlinear term conceived by humans. Since some system states cannot be directly observed, in order to enable the constructed adaptive observer to work accurately, the second unknown nonlinear function is specifically constructed for these system states.

[0074] Specifically, the process of approximating the second unknown nonlinear function using a radial basis function neural network is given by equation (5).

[0075] Equation (5) in, The nonlinear dynamics inherent in multi-agent systems For the second unknown nonlinear function, From 1 to l+ The first-order unconstrained state vector.

[0076] and These are the ideal weight vectors, and their weight update rates are different. and For the basis function vector, It is a specialized tool for one-dimensional nonlinearity, responsible only for handling unknown one-dimensional nonlinearities. yes l Specialized tools for 2D nonlinearity, only responsible for handling 2D to 3D. l Unknown nonlinearity of dimension, This is the first approximation error. The second approximation error is represented by positive numbers. and The boundary is satisfied. , .

[0077] S33: Construct an adaptive observer based on the second unknown nonlinear function and the given form of the observer.

[0078] Based on the radial basis function neural network approximation form of the second unknown nonlinear function, and combined with the reconstructed system dynamics, an adaptive observer is constructed for estimating unmeasurable states and time-varying bias faults.

[0079] Specifically, the adaptive observer is Equation (6): Equation (6) in, These are the weight estimates for the first dimension of the neural network. For the first l dimensional neural network weight estimates, For the first pdimensional neural network weight estimates, yes p Specialized tools for nonlinearity, only responsible for handling l Weizhi p Unknown nonlinear dimension For the first k The observation gain of the first dimension of the agent. For the k-th intelligent agent, the... l Dimensional observation gain, For the first k The first intelligent agent p Dimensional observation gain, It is the first k The first intelligent agent p dimensional unconstrained state variables, It is a control input with a fault.

[0080] In this embodiment, based on the second dynamic model, a radial basis function neural network is used to approximate the second unknown nonlinear function related to the unmeasurable state and the bias fault. By designing basis function vectors of a dedicated dimension to handle nonlinear characteristics of different orders, and combining an adaptive update law to adjust the ideal weight vector in real time to reduce the approximation error, the neural network weight estimate is embedded into the given form of the observer to achieve the collaborative estimation of the unmeasurable state and the time-varying bias fault of the actuator. Finally, the stability of the observation error system is rigorously verified by the first Lyapunov function, ensuring that in a multi-agent interaction environment with disturbances and faults, the full state information of the system can be accurately reconstructed based only on the first-order measurable state, providing a reliable observation basis for the subsequent optimal controller design.

[0081] In the embodiments of this application, deriving the adaptive update law of the adaptive observer includes: S41: Set the gain and approximation error boundary of the adaptive observer.

[0082] The gain of the adaptive observer is used to adjust the accuracy of state estimation, the speed of policy learning, and the ability to resist interference. The approximate error boundary of the adaptive observer is the upper limit of the neural network's approximation accuracy, which helps to prove the stability of the system and avoid excessive errors.

[0083] S42: Feedforward compensation is used to offset the effects of bias faults.

[0084] Among them, the time-varying bias fault of the actuator manifests as an additive disturbance term superimposed on the ideal control input. The pre-compensation estimates the fault value in real time and generates a reverse compensation quantity, which directly achieves accurate cancellation of the fault term at the control input layer, eliminating the impact of the fault on the system dynamics from the root.

[0085] Eliminating the effects of bias faults is to avoid distorted control commands and miscoordination of multi-agent cooperation, ensuring that the system can still perform tasks stably under fault conditions and conform to the design assumptions of the control strategy.

[0086] S43: Use an estimated value to replace an unmeasurable state.

[0087] The purpose of replacing unmeasurable states is to provide complete state information for control strategies and optimization algorithms, solve the bottleneck of "blind control", and provide necessary basis for stability analysis and optimal strategy iteration. Since the actual value of the unmeasurable state is unmeasurable, an estimated value is used as a substitute. The unmeasurable state after substitution has an error, but it is within an acceptable range.

[0088] S44: Derive the adaptive update law by combining the learning rate, activation function, and positive definite matrix.

[0089] In this process, the adaptive observer needs to obtain updated estimated weight values ​​through an adaptive update rate to observe the state of the unknown system, thereby optimizing the observation results.

[0090] Specifically, the adaptive update law is given by equation (7).

[0091] in, For 1-dimensional adjustment of positive parameters, for l Adjust the positive parameter. for p Adjust the positive parameter. A specialized tool for handling one-dimensional unknown nonlinearities. Specialized tools for handling unknown nonlinearities of 2 dimensions and above. The estimated value for the fault input. for t The first-dimensional neural network weight estimate at time step [time]. for t The first moment l Weight estimates for 3D neural networks For the estimated value of the control input with fault, To be updated 1 Weight estimates for 3D neural networks For the first k The agent is used to process the second unknown nonlinear function. l 3D activation function For the first k The first agent of the intelligent agent p dimensional unconstrained state variables, To be updated l Weight estimates for 3D neural networks It is a positive definite matrix that satisfies .

[0092] In this embodiment, estimation accuracy, learning rate, and anti-interference capability are balanced by setting the adaptive observer gain and approximation error boundary. Feedforward compensation is used to accurately offset the time-varying bias fault of the actuator, and the estimated value is used to replace the unmeasurable state to solve the bottleneck of "blind control". Furthermore, by combining the learning rate, activation function, and positive definite matrix to derive the adaptive update law, real-time collaborative estimation of unmeasurable state and bias fault is finally achieved. This eliminates the distortion and interference of faults on system dynamics from the root and provides complete and reliable state information support for multi-agent collaborative control.

[0093] In the embodiments of this application, constructing a first Lyapunov function to verify the stability of the adaptive observer includes: S51: Determine the estimation error of the adaptive observer, the approximate weight error of the unknown nonlinear function, and the estimation error of the bias fault, and construct the first Lyapunov function containing the inverse of the positive definite matrix.

[0094] Among them, the estimation error of the adaptive observer is used to quantify the state estimation accuracy, the approximate weight error of the unknown nonlinear function is used to measure the approximation effect of the neural network, and the estimation error of the bias fault is used to evaluate the accuracy of fault compensation. These are obtained through error dynamics derivation and stability analysis, and are all core components of the first Lyapunov function. Their boundedness is guaranteed by the negative definiteness of the derivative of the function. Their core role is to unify the constraints of the three types of errors and provide theoretical support for the stability of the adaptive observer.

[0095] The estimation error, approximate weight error, and estimation error all follow the principles of positive definiteness, radial unboundedness, and negative definite derivative, and are obtained by performing difference calculations between the true value and the estimated value / optimal value.

[0096] S52: Differentiate the first Lyapunov function and verify the effectiveness of the adaptive observer by scaling with inequalities.

[0097] The inequality scaling verification is based on Young's inequality, and the core is to scale the product terms into the sum of square terms, which is a key tool for handling cross-product terms in Lyapunov analysis. The first Lyapunov function is a positive definite "energy function" that characterizes system error. By taking its derivative and analyzing the negative definiteness of the derivative, we can determine whether the system energy decays, and thus verify the effectiveness of the observer / fault compensation / control strategy. This is the core application of the Lyapunov stability theorem.

[0098] Specifically, the first Lyapunov function is given by equation (8).

[0099] Equation (8) in, It is the first Lyapunov function. for The approximate error, for The approximate error, It is the inverse term of a positive definite matrix. For actuator bias fault estimation error, For state estimation error, q The total number of intelligent agents. p The state dimension of a single agent.

[0100] In this embodiment, by constructing a first Lyapunov function containing the inverse of a positive definite matrix, the estimation error of the adaptive observer, the approximate weight error of the neural network, and the bias fault estimation error are uniformly incorporated into the energy function framework, forming a collaborative constraint mechanism for the three types of errors. This ensures that the adaptive observer can still work effectively in complex environments with unmeasurable states, unknown nonlinearities, and time-varying faults, providing a quantitative means for evaluating the state estimation accuracy, neural network approximation effect, and fault compensation accuracy of multi-agent systems.

[0101] In embodiments of this application, the second dynamic model is reconstructed based on an adaptive observer, an optimal controller is designed, and consensus error dynamics among multiple agents are constructed, including: S61: Reconstruct the second dynamic model.

[0102] In this process, after verifying the effectiveness of the adaptive observer based on the first Lyapunov function, the estimation error converges. The estimation error is then input into the second dynamic model to replace the unmeasurable state, thus completing the reconstruction.

[0103] S62: Define coordinate transformation and construct a first-order filter.

[0104] Specifically, the coordinate transformation is defined based on equation (9).

[0105] Equation (9) in, For virtual control signals, For virtual error surface, As auxiliary state variables, Compensation error generated by coordinate transformation The leader is used to lead a multi-agent system, where each agent tracks the leader.

[0106] The first dimension tracking error, For the first l 3D tracking error, For the firstk One-dimensional state estimate of an agent , For the first k A single agent l Dimensional state estimate, and l= 2, ...... , p .

[0107] The constructed first-order filter is given by equation (10).

[0108] Equation (10) in, These are design parameters used to achieve the control objective, namely system stability.

[0109] for exist t The value when =0, for exist t The value when =0, For the first l- 1-dimensional virtual control signals exist t The l-th dimension virtual control signal at time t.

[0110] Furthermore, and The relationship between them is given by equation (11).

[0111] in, This is to compensate for the error caused by coordinate transformation. = ,and It is a continuous function.

[0112] S63: Input the second dynamic model, coordinate transformation and first-order filter into the zero-sum game framework, and design the optimal controller in a hierarchical manner using p-step backstepping technique.

[0113] The optimal controller, which is the controller of the multi-agent system, takes the estimated value of the adaptive observer as input and designs virtual control signals from low to high order through a backstepping recursive method. Finally, it generates an actual control law that integrates tracking control, fault compensation, and robust suppression at the highest order, which directly acts on the highest-order dynamics of the system to achieve multi-agent consensus tracking. In each step of the p-step backstepping technique, the corresponding optimal value function and Hamiltonian function must be constructed, specifically including: Step 1: Obtain the first optimal value function of the multi-agent system according to equation (12), and obtain the first Hamiltonian function of the multi-agent system according to equation (13).

[0114] in, The first optimal value function and the first cost function = , Representative and the k The first optimal virtual control input related to the agent. To influence the first worst-case perturbation affecting the k-th agent, The first interference suppression level coefficient, This represents the first-dimensional consensus error. The strategy of the minimizer, i.e., the controller. The maximizer, i.e., the perturbation strategy, ds Differential form of the integral variable Equation (13) in, For the first-dimensional Hamilton-Jacobi-Isax operator, The time derivative of the performance index. For the first k The observer gain of the first dimension of the agent. For the first k An agent and its neighboring agents h The communication weight (a non-negative real number) between them represents the strength of their information exchange. The square of the first-dimensional consensus error. Representative and the k The square of the first optimal virtual control input associated with each agent. To affect the first k The square of the first worst-case perturbation for each agent , For adjacency matrix elements, intelligent agent k The weight of communication with leaders For the leader's output.

[0115] No. l Step 1, according to equation (14), the second optimal value function of the multi-agent system is obtained, and according to equation (15), the second Hamiltonian function of the multi-agent system is obtained.

[0116] Specifically, the first l Step conforms ( l =2, K , p -1).

[0117] in, The second optimal value function, the second cost function , Representative and the k The second optimal virtual control input related to the agent. To affect the first k The second worst-case perturbation for an agent. For the first k The first intelligent agent l The perturbation suppression coefficient of dimension, For the first k The first intelligent agent l dimensional value function For the first k An agent and its neighboring agents l The communication weight (a non-negative real number) between them represents the strength of their information exchange. For the first l A dimensional value function is obtained by applying the maxima-mina principle to this value function. To obtain the optimal value function, No. k The first agent of the intelligent agent l External disturbance input, t Represents time.

[0118] in, For the first l The Vihamington-Jacobi-Isaks operator, For the first k The first intelligent agent l 3D virtual error surface, For the first k The first intelligent agent l The observer gain of the dimension, To affect the first k The square of the worst-case perturbation in the l-th dimension of an agent. For the first l The square of the tracking error, Representative and the k The first agent related to the intelligent agent l The square of the optimal virtual control input. For the first l Dimensional auxiliary state variables.

[0119] No. p Step: According to equation (16), the third optimal value function of the multi-agent system is obtained, and according to equation (17), the third Hamiltonian function of the multi-agent system is obtained.

[0120] in, For the third optimal value function, the... Cost function , Representative and the k The third optimal virtual control input related to the agent. To influence the third worst-case perturbation affecting the k-th agent, This represents the third level of interference suppression coefficient. For the first k The first intelligent agent p dimensional value function For the first p dimensional value function, For the first k Control input for an intelligent agent For the first k The first intelligent agent p A virtual error surface of dimension.

[0121] in, For the first p The Vihamington-Jacobi-Isaks operator, For the first k The optimal control input for each agent. For the first p The square of the tracking error, For the first k The square of the optimal control input for each agent. For the first k The first agent of the intelligent agent p The square of the worst-case perturbation. For the estimated value of the control input with fault, For the first k The first intelligent agent p The observer gain of the dimension, For the first p Dimensional auxiliary state variables.

[0122] S64: Determine the agent's local tracking error and neighbor tracking error.

[0123] S65: Determine adjacency weights and leader communication weights based on multi-agent communication topology.

[0124] S66: The local tracking error and the adjacent tracking error are weighted and summed with the adjacency weights respectively, and the consensus error is dynamic.

[0125] Based on the coordinate transformation, the error dynamics are described using equation (18).

[0126] in, l= 2,..., p- 1, k= 1,..., q , No. k The first-dimensional disturbance input of the agent, No. k The first agent of the intelligent agent l Dimensional interference input, No. k The first agent of the intelligent agent p Dimensional interference input, No. k The first agent of the intelligent agent p The dimension is an unconstrained state variable.

[0127] Furthermore, the consensus error dynamics among multiple agents are constructed based on the error dynamics through equation (19).

[0128] in, For intelligent agents k The set of neighboring agents, , For adjacency matrix elements, intelligent agent k The weight of communication with leaders and respectively intelligent agents k and neighboring agents h Compared to the leader's observed output error, This represents the dynamic consensus error among intelligent agents.

[0129] Specifically, the local tracking error of each agent relative to the reference trajectory is first defined by coordinate transformation. Then, the adjacency weight and leader communication weight are determined according to the multi-agent communication topology, and the fusion coefficient is calculated. Finally, the tracking error difference between the current agent and the neighboring agents and the tracking error of the neighboring agents are weighted and summed according to the fusion coefficient and the adjacency weight, respectively, to construct the consensus error that reflects the global collaborative state of the multi-agent agents, thus completing the dynamic construction of the consensus error.

[0130] In this embodiment, the system reconstruction is completed by substituting the adaptive observer estimate into the second dynamic model, thus solving the "blind control" problem caused by unmeasurable states. Furthermore, a coordinate transformation is defined and a first-order filter is constructed to eliminate the computational complexity caused by virtual control differentiation. Based on this, the reconstructed model, coordinate transformation, and filter input zero-sum game framework are used. A p-step backstep recursive technique is employed to design the optimal controller hierarchically. Adjacency weights and leader communication weights are determined based on the multi-agent communication topology. The local tracking error and adjacent tracking error are weighted and summed to construct the consensus error dynamic. This observation-control closed-loop architecture ensures the stable operation of the system under unmeasurable states, time-varying faults, and external disturbances, providing a complete theoretical framework and feasible control strategies for the safe and collaborative control of multi-agent systems.

[0131] In embodiments of this application, generating the optimal virtual control strategy and the worst-case perturbation strategy for a multi-agent system includes: S71: Construct the execution network and evaluation network.

[0132] The execution network is used to evaluate the optimal virtual control strategy for the multi-agent system, while the evaluation network is used to evaluate the optimal value function.

[0133] S72: Establish an execution-evaluation reinforcement learning framework based on the execution network and the evaluation network.

[0134] S73: Collects operational data from multi-agent systems.

[0135] The runtime data was obtained from a MATLAB simulation system, which is an unmanned aerial vehicle (UAV) system.

[0136] S74: Iteratively update the execution network and the evaluation network based on the running data; S75: Approximation of the optimal value function of the Hamilton-Jacobi-Isax equation based on an evaluation network.

[0137] Since the optimal value function of a multi-agent nonlinear system has no analytical solution, it cannot be directly used for control design. Therefore, it needs to be approximated by a neural network. The optimal value function itself is the evaluation standard of minimax optimal performance, the core variable for solving the Hamilton-Jacobi-Isax equation, and the link for multi-frame fusion. After approximation, it becomes a computable adaptive design parameter, which is the core key to realizing the fusion control of "robust optimal + adaptive + backstep recursion".

[0138] S76: Based on the optimal value function of the Hamilton-Jacobi-Isaacs equation, the optimal virtual control strategy and worst-case perturbation strategy of the multi-agent system are generated by the execution network. Among them, the role of the optimal multi-agent virtual control strategy is to drive the system to satisfy full-state constraints and offset the impact of bias faults under the zero-sum game framework in order to achieve fault-tolerant consensus tracking. It is generated by dynamically optimizing the optimal solution of the Hamilton-Jacobi-Isaacs equation and the consensus error through the execution network, according to the preset learning rate.

[0139] The worst-case perturbation strategy is used to quantify the maximum impact of adversarial interference to support the disturbance-resistant design. It is generated by the execution network based on the gradient of the evaluation network's value function, the perturbation attenuation coefficient, and the RBFNN approximation results.

[0140] Furthermore, the optimal virtual control strategy and worst-case perturbation strategy for the multi-agent system are generated based on p-step backstepping technology.

[0141] Specifically, step 1 is completed based on equation (20).

[0142] in, For the first-dimensional optimal virtual control estimation, For the worst-case perturbation estimation in the first dimension, To implement the network's first-dimensional weight update rate, To evaluate the update rate of the first dimension weights of the network, For the first dimension of positive design parameters, To perform the first-dimensional weight estimation of the network, To evaluate the estimation of the first dimension weights of the network, and The two terms, learning rate and learning rate, have the same meaning but different numerical values. and Both are activation functions; they have the same meaning but different values. To evaluate the estimation of the first dimension weights of the network, To perform the first-dimensional weight estimation of the network, The activation function for executing and evaluating the network.

[0143] No. l The step is completed based on equation (21).

[0144] in, For the first l 3D optimal virtual control estimation, For the first l worst-case disturbance estimate, To execute the network l Dimension weight update rate To evaluate the network l Dimension weight update rate For the first l Optimized design parameters To execute the network l Dimensional weight estimation, To evaluate the network l Dimensional weight estimation.

[0145] No. p The steps are completed based on equation (22).

[0146] in, For the optimal virtual control strategy of multi-agent systems, For the first p worst-case disturbance estimate, To execute the network p Dimension weight update rate To evaluate the network p Dimension weight update rate For the first p Optimized design parameters To execute the network p Dimensional weight estimation, To evaluate the network pDimensional weight estimation.

[0147] Backstepping techniques are used throughout multi-agent systems. p In the full recursive process, the optimal multi-agent system virtual control strategy and the worst-case disturbance strategy are generated and embedded in each stage of the virtual control design. The optimal multi-agent system virtual control strategy and the worst-case disturbance strategy are designed to make the control strategy "robust optimal". and The two weight update rates are used to adaptively compensate for nonlinear errors. The core difference between the two weight update rates is that the backstep order, input variables, and convergence targets are different. Essentially, they are order-wise compensation for nonlinearity. Each backstep revolves around "order recursion, error compensation, optimal robustness, and stable convergence", ultimately achieving robust optimal control with multi-agent fault-tolerant consensus.

[0148] In this embodiment, the optimal value function of the Hamilton-Jacobi-Isaacs equation is approximated by an execution-evaluation network, breaking through the theoretical bottleneck of nonlinear optimal control having no analytical solution. At the same time, the nonlinear error is compensated step by step by p-step backstep recursion and a zero-sum game framework is embedded, so that the virtual control strategy has both optimality and robustness. Finally, it realizes full state constraint satisfaction, bias fault adaptive compensation and high-precision fault-tolerant consensus tracking of multiple agents.

[0149] In the embodiments of this application, the optimal virtual control strategy for a multi-agent system is iteratively optimized, including: S81: Design a network weight update law based on the optimal value function, the optimal multi-agent system virtual control strategy, and the worst-case perturbation strategy.

[0150] After obtaining the network weight update law, it is necessary to clarify its design parameters and learning rate, and iterate and optimize the network weights periodically until convergence.

[0151] Since the initial network weight update law may not meet the stability requirements of a multi-agent system, the optimal weights are found through repeated iterations. The design parameters are the parameters that participate in the update rate design, and the learning rate is the core positive parameter that can be designed to adjust the parameter update speed and convergence.

[0152] S82: Iteratively optimize the virtual control strategy of the optimal multi-agent system through the network weight update law.

[0153] Specifically, the optimal solution of the Hamilton-Jacobi-Isaacs equation is estimated by using an RBF neural network. Combined with dynamic surface control technology, the optimal virtual control strategy for the multi-agent system is obtained through iterative cooperation between the execution network and the comment network.

[0154] Specifically, the network parameters and learning rate are first initialized. An initial virtual control strategy is designed based on the execution network. The stability of the initial virtual control strategy is evaluated by judging the approximation error of the network according to the Hamilton-Jacobi-Isax equation. If the stability does not meet the standard, the network weights are adjusted based on the preset update rules. The three steps of "generating strategy - stability evaluation - adjusting weights" are repeated until the stability of the virtual control strategy meets the standard. The virtual control strategy at this time is taken as the optimal virtual control strategy for the multi-agent system.

[0155] In this embodiment, an adaptive online optimization of the optimal virtual control strategy is achieved through a dual-network collaborative iterative mechanism of execution and evaluation. Its core beneficial effect lies in the following: by utilizing the continuous approximation capability of the RBF neural network to the optimal solution of the Hamilton-Jacobi-Isax equation, and through a closed-loop iterative process of "generating strategy-stability evaluation-weight adjustment", the network weight update law can adaptively adjust the learning rate and design parameters. While ensuring the convergence of the algorithm, it overcomes the limitation of insufficient stability of the initial strategy, and finally obtains the optimal virtual control strategy that meets the strict stability requirements of multi-agent systems, realizing data-driven adaptive optimization control under nonlinear uncertainty environment.

[0156] In the embodiments of this application, the stability verification of the virtual control strategy of the optimal multi-agent system is performed by constructing a second Lyapunov function, including: S91: Obtain the consensus error among multiple agents, the relative error caused by coordinate transformation, and the approximation error between the execution network and the evaluation network.

[0157] Among them, consensus error is used to quantify consistency, relative error is used to quantify coordinate transformation consistency, and approximation error is used to quantify the approximation effect of all unknowns that need to be approximated by the neural network.

[0158] S92: Construct a second Lyapunov function based on consensus error, relative error, and approximation error.

[0159] The second Lyapunov function is constructed based on equation (23).

[0160] in, For the first k The second Lyapunov function of the l-th dimension of the agent For the first k The first intelligent agent l The second Lyapunov function of dimensionality, For the first k The first intelligent agent pThe second Lyapunov function of dimensionality, To combine positive definite matrices, For the observation output error of the intelligent agent, For the first k The first intelligent agent b The square of the first-order tracking error, This is the square of the compensation error caused by the coordinate transformation. For the first k The first intelligent agent b Estimation of network weights at the order of execution. For the first k The first intelligent agent b Evaluation network weight estimation of order, For the first k Weight estimation of the first-order evaluation network for each agent. For the first k Weight estimation of the first-order execution network for each agent. For the first k The first intelligent agent b Estimation of network weights at the order of execution. For the first k Weight estimation of the first-order evaluation network for each agent. For the first k Weight estimation of the first-order execution network for each agent. For the first k The first intelligent agent b Evaluation network weight estimation of order ( b =1,2,...,l).

[0161] Furthermore, specifically, we first incorporate consensus error, dynamic surface error, and approximation error into the function as squared / transposed product terms, respectively, in the corresponding formulas. , and weighting error The core term, combined with the Lipschitz matrix L and matrix B, constitutes... The consensus error is weighted positive definite. This process not only conforms to the topological characteristics and Lipschitz continuity of the multi-agent system, but also ensures the positive definiteness of the consensus error energy term, thereby integrating all error terms. By standardizing the energy term by assigning a coefficient of 1 / 2 to each error term, and performing a two-level summation based on the number of agents and the order of the subsystem, a second Lyapunov function is finally constructed that includes three types of errors, energizes the total system energy, and satisfies the positive definite core requirement. S93: Stability verification of virtual control strategy for optimal multi-agent system based on second Lyapunov function.

[0162] Specifically, the optimal virtual control strategy of the multi-agent system is substituted into the system dynamics to calculate its time derivative to verify the negative qualitative nature. Finally, the Lyapunov stability criterion is used to determine whether the system is asymptotically stable globally.

[0163] In this embodiment, by combining consensus error, dynamic surface relative error, and neural network approximation error with graph theory Laplacian matrix and Lipschitz properties to construct a weighted positive definite term for topology awareness, the stability analysis is made to fit the multi-agent network topology and satisfy the nonlinear Lipschitz continuity assumption. Through energy term standardization and double-layer summation, a complete stability criterion is established that simultaneously covers consensus convergence, coordinate transformation accuracy, and approximation error. This rigorously proves the asymptotic stability of the closed-loop system under the fusion control of backstepping recursion and reinforcement learning, providing a theoretical guarantee for the reliability of fault-tolerant consensus tracking.

[0164] A second aspect of this application provides an observer-based multi-agent system reinforcement learning control system, the system architecture of which is as follows: Figure 2 As shown, it includes: Model building module 1 is used to establish the first dynamic model of a strictly feedback nonlinear multi-agent system containing full-state constraints, unmeasurable states, and time-varying bias faults.

[0165] The conversion module 2 is used to convert the first dynamic model with full state constraints into a second dynamic model with unconstrained pure feedback.

[0166] Observer building module 3 is used to build an adaptive observer based on the second dynamic model.

[0167] The observer update module 4 is used to derive the adaptive update law of the adaptive observer.

[0168] The first Lyapunov function construction module 5 is used to construct the first Lyapunov function to verify the stability of the adaptive observer.

[0169] Controller design module 6 is used to design the optimal controller and construct the consensus error dynamics among multiple agents.

[0170] The network approximation and update module 7 is used to generate the optimal virtual control strategy and worst-case perturbation strategy for the multi-agent system.

[0171] The second Lyapunov function construction module 8 is used to construct the second Lyapunov function to verify the stability of the virtual control strategy of the optimal multi-agent system.

[0172] In this embodiment, functions such as state constraint processing, adaptive observation, reinforcement learning optimization, and dual Lyapunov stability verification are organically integrated into multiple functional modules, forming a closed-loop control system of "modeling-transformation-estimation-verification-optimization-re-verification". This system not only eliminates the limitations of full-state constraints on control through coordinate transformation, but also solves the problems of unmeasurable state and fault estimation with the help of adaptive observers. At the same time, it uses an execution-evaluation network to approximate the optimal value function online to achieve data-driven optimization. Finally, the dual convergence of observation error and policy stability is strictly guaranteed through a dual-layer Lyapunov function. This provides a complete engineering solution for complex nonlinear multi-agent systems that combines constraint satisfaction, fault tolerance compensation, optimal performance, and robust stability.

[0173] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A reinforcement learning control method for a multi-agent system based on observers, characterized in that, include: Establish the first dynamic model of a rigorous feedback nonlinear multi-agent system containing full-state constraints, unmeasurable states, and time-varying bias faults; The first dynamic model is transformed into a second dynamic model with unconstrained pure feedback. Based on the second dynamic model, an adaptive observer is constructed; Based on the adaptive observer, the adaptive update law is derived, and the adaptive observer is optimized. The stability of the adaptive observer is verified by constructing a first Lyapunov function. The second dynamic model is reconstructed based on the adaptive observer, the optimal controller is designed, and the consensus error dynamics among multiple agents are constructed. An execution-evaluation reinforcement learning framework is established to approximate the optimal value function of the Hamilton-Jacobi-Isaacs equation. Based on the consensus error, the optimal virtual control strategy and worst-case perturbation strategy of the multi-agent system are dynamically generated. Based on the optimal value function, the optimal multi-agent system virtual control strategy, and the worst-case perturbation strategy, a network weight update law is designed to iteratively optimize the optimal multi-agent system virtual control strategy. A second Lyapunov function is constructed to verify the stability of the virtual control strategy of the optimal multi-agent system. The multi-agent system is controlled based on the optimal multi-agent system virtual control strategy.

2. The multi-agent system reinforcement learning control method according to claim 1, characterized in that, The process of transforming the first dynamic model with full-state constraints into a second dynamic model with unconstrained pure feedback includes: The first dynamic model is subjected to a one-to-one nonlinear mapping process to obtain an unconstrained pure feedback system. Based on the aforementioned unconstrained pure feedback system, a radial basis function neural network is used to approximate the first unknown nonlinear function. By disassembling the first dynamic model and incorporating the multi-agent interaction characteristics, faults, and disturbances, a second dynamic model is obtained.

3. The multi-agent system reinforcement learning control method according to claim 2, characterized in that, The construction of the adaptive observer includes: Obtain the unmeasurable states and bias faults of the second dynamic model; Based on the unmeasurable state and the bias fault, a radial basis function neural network is used to approximate the second unknown nonlinear function; Based on the radial basis function neural network approximation form of the second unknown nonlinear function, and combined with the reconstructed system dynamics, an adaptive observer is constructed for estimating unmeasurable states and time-varying bias faults.

4. The multi-agent system reinforcement learning control method according to claim 3, characterized in that, The process of deriving the adaptive update law based on the adaptive observer and optimizing the adaptive observer includes: Define the gain and approximation error boundary of the adaptive observer; The effects of the bias fault are offset by feedforward compensation. The estimated value is used to replace the unmeasurable state; The adaptive update law is derived by combining the learning rate, activation function, and positive definite matrix.

5. The multi-agent system reinforcement learning control method according to claim 1, characterized in that, The stability of the adaptive observer is verified by constructing a first Lyapunov function, including: Determine the estimation error of the adaptive observer, the approximate weight error of the unknown nonlinear function, and the estimation error of the bias fault, and construct the first Lyapunov function containing the inverse of the positive definite matrix; The effectiveness of the adaptive observer is verified by taking the derivative of the first Lyapunov function and scaling it using inequalities.

6. The multi-agent system reinforcement learning control method according to claim 1, characterized in that, Based on the adaptive observer, the second dynamic model is reconstructed, an optimal controller is designed, and the consensus error dynamics among multiple agents are constructed, including: The second dynamic model is reconstructed; Define coordinate transformations and construct a first-order filter; The second dynamic model, the coordinate transformation, and the first-order filter are input into a zero-sum game framework, and the optimal controller is designed hierarchically using p-step backstepping techniques. Determine the agent's local tracking error and neighboring tracking error; Determine adjacency weights based on multi-agent communication topology; The consensus error is dynamically calculated by weighting the local tracking error and the adjacent tracking error with the adjacency weights respectively.

7. The multi-agent system reinforcement learning control method according to claim 6, characterized in that, Generating the optimal multi-agent system virtual control strategy and the worst-case perturbation strategy includes: Construct execution networks and evaluation networks; The execution-evaluation reinforcement learning framework is established based on the execution network and the evaluation network; Collect operational data from multi-agent systems; The execution network and the evaluation network are iteratively updated based on the operational data. The optimal value function of the Hamilton-Jacobi-Isaac equation is approximated based on the evaluation network described above. Based on the optimal value function of the Hamilton-Jacobi-Isax equation, the optimal virtual control strategy and the worst-case perturbation strategy of the multi-agent system are generated through the execution network.

8. The multi-agent system reinforcement learning control method according to claim 7, characterized in that, Iterative optimization of the virtual control strategy for the optimal multi-agent system includes: Based on the optimal value function, the optimal multi-agent system virtual control strategy, and the worst-case perturbation strategy, a network weight update law is designed. The optimal virtual control strategy for the multi-agent system is iteratively optimized using the network weight update law.

9. The multi-agent system reinforcement learning control method according to claim 1, characterized in that, Constructing the second Lyapunov function to verify the stability of the virtual control strategy of the optimal multi-agent system includes: The consensus error among multiple agents, the relative error caused by coordinate transformation, and the approximation error between the execution network and the evaluation network are obtained. The second Lyapunov function is constructed based on the consensus error, the relative error, and the approximation error; The stability of the virtual control strategy for the optimal multi-agent system is verified based on the second Lyapunov function.

10. An observer-based multi-agent system reinforcement learning control system, applied to the multi-agent system reinforcement learning control method according to any one of claims 1-9, characterized in that, include: The model building module is used to establish the first dynamic model of a strictly feedback nonlinear multi-agent system containing full-state constraints, unmeasurable states, and time-varying bias faults. The conversion module is used to convert the first dynamic model with full-state constraints into a second dynamic model with unconstrained pure feedback. An observer building module is used to build an adaptive observer based on the second dynamic model; An observer update module is used to derive the adaptive update law of the adaptive observer; The first Lyapunov function construction module is used to construct the first Lyapunov function to verify the stability of the adaptive observer. The controller design module is used to design the optimal controller and construct the consensus error dynamics among multiple agents. The network approximation and update module is used to generate the optimal virtual control strategy and worst-case perturbation strategy for the multi-agent system; The second Lyapunov function construction module is used to construct a second Lyapunov function to verify the stability of the virtual control strategy of the optimal multi-agent system.