Heterogeneous multi-agent distributed Nash equilibrium optimization adaptive fault-tolerant control method

By employing a hierarchical game control framework and an adaptive fault-tolerant control method based on radial basis function neural networks, the problems of unknown nonlinearity and actuator failure in heterogeneous multi-agent systems are solved, achieving stable tracking and improved robustness of the system.

CN122085707APending Publication Date: 2026-05-26ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANHUI UNIV
Filing Date
2026-04-22
Publication Date
2026-05-26

Smart Images

  • Figure CN122085707A_ABST
    Figure CN122085707A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multi-agent system control, solves the technical problem of lack of an effective active compensation mechanism for unknown nonlinear dynamics and actuator fault problems in an underlying physical system in the traditional technology, and particularly relates to an adaptive fault-tolerant control method for heterogeneous multi-agent distributed Nash equilibrium optimization. By constructing a hierarchical game control framework, hierarchical decoupling is carried out on high-level game strategy optimization of a decision-making layer and bottom-layer trajectory tracking of a control layer; and establishing an adaptive fault-tolerant control law, and configuring adaptive update law parameters based on a Lyapunov stability theory, so that a control layer can track a Nash equilibrium reference strategy. According to the method, distributed game control of the heterogeneous multi-agent system can be realized under the condition that nonlinear uncertainty and actuator faults exist, so that tracking errors are converged to a small tight set, accurate tracking of an optimal strategy is realized, and the fault-tolerant capability and robustness of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-agent system control technology, and in particular to an adaptive fault-tolerant control method for heterogeneous multi-agent distributed Nash equilibrium optimization. Background Technology

[0002] In recent years, multi-agent systems have been widely applied in various engineering fields, such as electricity market bidding systems, networked unmanned vehicle systems at sea, and intelligent transportation systems. In these systems, multiple agents optimize their respective performance indicators through local communication and information exchange. The resulting non-cooperative behavior makes the distributed Nash equilibrium optimization problem an important research direction in control theory and optimization.

[0003] Most existing research is based on the simplified assumption of homogeneous agent dynamics, which assumes that all agents have the same dynamic structure and state dimension. However, in real-world engineering systems, multi-agent systems are often composed of subsystems with different physical properties, different control input structures, and different state dimensions, thus exhibiting obvious heterogeneous characteristics.

[0004] Furthermore, in the actual physical control layer, the system is typically affected by unknown nonlinear dynamics, unmodeled dynamics, and external disturbances. These uncertainties can lead to a decline in the system's tracking performance. Simultaneously, during long-term operation, actuators may experience failures, reduced efficiency, or biasing. When an actuator fails, the agent's actual control capability will deviate from the optimal strategy calculated by the decision layer, thus affecting local optimization results and potentially propagating error information through the communication network, further impacting the overall system's game equilibrium state.

[0005] Most existing methods focus on optimizing algorithm design, lacking effective proactive compensation mechanisms for unknown nonlinear dynamics in the underlying physical system and actuator failures. Therefore, it is necessary to propose a hierarchical adaptive control method that can simultaneously handle system heterogeneity, compensate for unknown nonlinearities, and address the impact of actuator failures. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides an adaptive fault-tolerant control method for heterogeneous multi-agent distributed Nash equilibrium optimization, which solves the technical problems of traditional technologies lacking effective active compensation mechanisms for unknown nonlinear dynamics in the underlying physical system and actuator failures.

[0007] To address the aforementioned technical problems, this invention provides the following technical solution: an adaptive fault-tolerant control method for heterogeneous multi-agent distributed Nash equilibrium optimization, comprising the following steps: S1. Construct a dynamic model to describe the heterogeneous multi-agent system containing unknown nonlinear dynamics and actuator failure physical characteristics, and define the local cost function of the non-cooperative game at the decision-making level. S2. Based on the dynamic model and local cost function, a hierarchical game control framework including a decision layer and a control layer is constructed to decouple the high-level game strategy optimization of the decision layer from the low-level trajectory tracking of the control layer, including: At the decision-making level, a consensus-based gradient descent protocol is used to achieve distributed Nash equilibrium optimization with the local cost function as the objective, so that each agent searches for the Nash equilibrium reference strategy only through local communication information interaction. In the control layer, a radial basis function neural network is used to approximate the unknown nonlinear dynamics of the heterogeneous multi-agent system, and a linear parameterization method is used to describe the actuator fault. S3. Establish an adaptive fault-tolerant control law to compensate for unknown nonlinear dynamics and actuator failures, and combine it with... The adaptive parameter update law of the correction term is used to estimate the weights and fault parameters of the radial basis function neural network online; S4. Based on Lyapunov stability theory, configure control parameters for the decision-making and control layers, enabling the control layer to track the Nash equilibrium reference strategy output by the decision-making layer, and prove that all signals in the closed-loop system are consistent and eventually bounded.

[0008] Furthermore, the expression for the dynamic model is: ; in, For the first Each agent's state vector The time derivative; For the first The state vectors of each agent; For the first Control input for each intelligent agent; and The system matrix is ​​known. Representing unknown nonlinear dynamics; This indicates an actuator malfunction or external disturbance. The expression for the local cost function is: ; in, Indicates the first Local cost function of each agent; For the first Decision variables for each agent; For local reference signals; Represents intelligent agents The set of neighbors; For the first The decision variables of an agent.

[0009] Furthermore, achieving the distributed Nash equilibrium optimization includes: decision variables Mapped to the underlying reference state used to implement the connection between the decision-making layer and the control layer. The expression is: ; in, It is a constant mapping matrix; Based on the underlying reference state Define the tracking error of the control layer for: ; in, For the first The state vectors of each agent; A consensus-based gradient descent protocol, with a local cost function as the objective, achieves distributed Nash equilibrium optimization at the decision level. Its update law is expressed as: ; in, To optimize step size; This is a locally estimated vector; For the first Time derivatives of decision variables of an agent; Let be the gradient operator, indicating that for the , Decision variables of an agent Find the partial derivative; To estimate the vector locally The local cost function for the variable; Introducing auxiliary variables that enable each agent to obtain an estimate of the Nash equilibrium reference policy of other agents. Its dynamic update law is: ; in, Represents intelligent agents For intelligent agents Estimates of decision variables; These are the adjacency matrix elements of the communication topology; Consensus gain; The Kronecker function; For intelligent agents For intelligent agents The rate of evolution of the estimated values ​​of decision variables; For intelligent agents For intelligent agents Estimates of decision variables; The total number of intelligent agents; For the first The decision variables of an agent.

[0010] Furthermore, the unknown nonlinear dynamics of the heterogeneous multi-agent system are approximated, and actuator faults are described using a linear parameterization method, including: Using radial basis function neural networks to study unknown nonlinear dynamics in heterogeneous multi-agent systems To approximate, its representation is as follows: ; in, This is the weight matrix of the radial basis function neural network; The radial basis function vector; To approximate the error; Using linear parameterization method for actuator faults To perform modeling, that is: ; in, The fault parameter matrix is ​​unknown; The regression vector is known.

[0011] Furthermore, the adaptive fault-tolerant control law controls the input. The expression is:

[0012] in, This is the feedback gain matrix; This is the feedforward gain matrix; and The weights of the radial basis function neural network are respectively and fault parameters The estimated value; For tracking error in the control layer; For decision variables; The radial basis function vector; The regression vector is known.

[0013] Furthermore, the adaptive parameter update law is used to adjust the tracking error. Dynamically update the weight estimates of the radial basis function neural network and fault parameter estimates Its expression is:

[0014] in, and These are the time derivatives of the radial basis function neural network weight estimates and the time derivatives of the actuator fault parameter estimates, respectively. and The learning rate matrix is ​​positive definite. and For robust damping parameters; It is a positive definite matrix; The system matrix is ​​known.

[0015] Furthermore, the control parameter configurations of the decision-making layer and the control layer satisfy the following conditions: In configuring control parameters at the decision-making level, optimize the step size. With consensus gain It satisfies the algebraic constraints, namely: ; in, It is the smallest non-zero eigenvalue of the Laplace matrix corresponding to the communication topology; The Lipschitz constant is the pseudo-gradient of the local cost function; In the control parameter configuration of the control layer, the positive definite matrix It satisfies the algebraic Lyapunov equation, that is: ; in, Let be a given positive definite symmetric matrix; and The system matrix is ​​known. To make the feedback gain matrix of the nominal closed-loop matrix Herwitz; By configuring the control parameters, the derivative of the overall Lyapunov function of the closed-loop system is made... satisfy: ; in, For global Lyapunov functions; and , which are the exponential convergence rate constant and the bounded residual term of the system, respectively.

[0016] Furthermore, the configuration control parameters also include: In the configuration of learning rate and robustness parameters, the positive definite learning rate matrix and and robust damping parameters and All are constants greater than zero.

[0017] Furthermore, the global Lyapunov function Including Lyapunov functions selected at the decision-making level And the Lyapunov function selected in the control layer The expression is: ; ; ; in, For intelligent agents Quantity; and The learning rate matrix is ​​positive definite. It is a positive definite matrix; For tracking error in the control layer; Let be the policy error vector of all agents in the decision-making layer relative to the Nash equilibrium point; This is the consensus estimation error vector among agents regarding each other's policy states; Represents the trace operation of a matrix; This is the estimation error matrix for the weights of the radial basis function neural network; This is the estimation error matrix for actuator fault parameters.

[0018] By employing the above technical solution, the present invention provides an adaptive fault-tolerant control method for heterogeneous multi-agent distributed Nash equilibrium optimization, which has at least the following beneficial effects: 1. This invention enables distributed game control of heterogeneous multi-agent systems under conditions of nonlinear uncertainty and actuator failure, allowing tracking errors to converge to small compact sets, achieving accurate tracking of the optimal strategy, and improving the system's fault tolerance and robustness.

[0019] 2. This invention constructs a hierarchical game control architecture, decouples high-level strategy optimization from low-level physical trajectory tracking, and uses radial basis function neural networks and adaptive fault estimation mechanisms to compensate for unknown nonlinearities and actuator faults, thereby ensuring that the heterogeneous multi-agent system can still achieve stable tracking of the Nash equilibrium point in complex and uncertain environments.

[0020] 3. This invention reduces the coupling complexity of the controller design by using a hierarchical design of high-level distributed Nash equilibrium optimization and low-level heterogeneous physical system trajectory tracking.

[0021] 4. This invention can effectively compensate for nonlinear uncertainties and actuator failures in the system by introducing a radial basis function neural network to approximate unknown nonlinearities and combining it with an adaptive fault estimation mechanism.

[0022] 5. This invention designs control laws and parameter update laws within the framework of Lyapunov stability theory, ensuring that all signals within the closed-loop system remain consistent and eventually bounded, thereby guaranteeing that the heterogeneous multi-agent system can stably track the Nash equilibrium point. Attached Figure Description

[0023] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of the adaptive fault-tolerant control method in this invention; Figure 2 This is a topology diagram of the multi-agent communication network in this invention; Figure 3 This is a simulation curve showing the convergence of the strategies of each intelligent agent in this invention; Figure 4 This is a simulation curve of the tracking error of the physical system in this invention; Figure 5 This is a simulation curve of adaptive fault parameter estimation in this invention. Detailed Implementation

[0024] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. This will allow for a full understanding of how the present application uses technical means to solve technical problems and achieve technical effects, and to facilitate its implementation.

[0025] This embodiment proposes an adaptive fault-tolerant control method for heterogeneous multi-agent distributed Nash equilibrium optimization. By constructing a hierarchical game-theoretic control framework, the high-level game strategy optimization of the decision layer and the low-level trajectory tracking of the control layer are designed in a hierarchical manner. Radial basis function neural network approximation and adaptive fault estimation methods are used to compensate for unknown nonlinear dynamics and actuator faults in the heterogeneous multi-agent system, thereby achieving stable tracking of the Nash equilibrium reference strategy. Figure 1 As shown, the method includes the following steps: S1. Construct a dynamic model to describe the heterogeneous multi-agent system containing unknown nonlinear dynamics and actuator failure physical characteristics, and define the local cost function of the non-cooperative game at the decision level.

[0026] Suppose a heterogeneous multi-agent system consists of The system consists of several agents, each with a different dynamic structure, thus classifying it as a heterogeneous multi-agent system. The dynamic model of an agent is represented as follows: ; in, For the first Each agent's state vector The time derivative; For the first The state vectors of each agent; For the first Control input for each intelligent agent; and The system matrix is ​​known. Representing unknown nonlinear dynamics; This indicates an actuator malfunction or external disturbance.

[0027] To achieve distributed Nash equilibrium optimization, a local cost function is constructed for each agent at the decision layer, which takes the following form: ; in, Indicates the first Local cost function of each agent; For the first Decision variables for each agent; For local reference signals; Represents intelligent agents The set of neighbors; For the first The first term is the decision variable for each agent. The second term is a coordination penalty term constructed based on a consensus mechanism, used to constrain the policy differences between neighboring agents.

[0028] S2. Based on the dynamic model and local cost function, a hierarchical game control framework including a decision layer and a control layer is constructed to decouple the high-level game strategy optimization of the decision layer from the low-level trajectory tracking of the control layer, including: At the decision-making level, a consensus-based gradient descent protocol is used to achieve distributed Nash equilibrium optimization with the local cost function as the objective, enabling each agent to search for the Nash equilibrium reference policy through local communication information interaction.

[0029] To establish a connection between the decision-making and control levels, decision variables will be... Mapped to underlying reference state The expression is: ; in, It is a constant mapping matrix.

[0030] Further based on the underlying reference state Define the tracking error of the control layer for: ; in, For the first The state vector of each agent.

[0031] At the decision-making level, to achieve distributed Nash equilibrium optimization, a consensus-based gradient descent protocol is designed with a local cost function as the objective. Its update law is expressed as: ; in, To optimize step size; This is a locally estimated vector; For the first Time derivatives of decision variables of an agent; Let be the gradient operator, indicating that for the , Decision variables of an agent Find the partial derivative; To estimate the vector locally The local cost function for the variable; Meanwhile, to enable each agent to obtain an estimate of the Nash equilibrium reference policy of other agents, auxiliary variables are introduced. Its dynamic update law is:

[0032] in, Represents intelligent agents For intelligent agents Estimates of decision variables; These are the adjacency matrix elements of the communication topology; Consensus gain; The Kronecker function; For intelligent agents For intelligent agents The rate of evolution of the estimated values ​​of decision variables; For intelligent agents For intelligent agents Estimates of decision variables; The total number of intelligent agents; For the first The decision variables of an agent.

[0033] In the control layer, a radial basis function neural network is used to approximate the unknown nonlinear dynamics of the heterogeneous multi-agent system, and a linear parameterization method is used to describe actuator faults.

[0034] To handle unknown nonlinear dynamics in heterogeneous multi-agent systems This invention utilizes a radial basis function neural network to approximate it, and its representation is as follows: ; in, This is the weight matrix of the radial basis function neural network; The radial basis function vector; This is the approximation error.

[0035] Meanwhile, to describe the impact of actuator failure, a linear parameterization method is used to analyze the actuator failure. To perform modeling, that is: ; in, The fault parameter matrix is ​​unknown; The regression vector is known.

[0036] In this embodiment, the dynamic model provides the decision-making layer with state prediction capabilities, supporting the accuracy of policy optimization. The dynamic model and the local cost function form a closed-loop control through state variables and optimization objectives. The hierarchical game control framework achieves the separation of policy optimization and trajectory tracking through hierarchical decoupling between the decision-making layer and the control layer. This not only improves the robustness and computational efficiency of heterogeneous multi-agent systems, but also provides a structured solution for handling complex problems such as actuator failure and unknown nonlinear dynamics.

[0037] S3. Establish an adaptive fault-tolerant control law to compensate for unknown nonlinear dynamics and actuator failures, and combine it with... The adaptive parameter update law of the correction term is used to estimate the weights and fault parameters of the radial basis function neural network online.

[0038] This embodiment designs an adaptive fault-tolerant control law to compensate for the unknown nonlinear dynamics and actuator faults of a heterogeneous multi-agent system. Its control input... The expression is: ; in, To make the feedback gain matrix of the nominal closed-loop matrix Herwitz; To make the feedforward gain matrix of the nominal closed-loop matrix Herwitz; and The weights of the radial basis function neural network are respectively and fault parameters The estimated value; For tracking error in the control layer; For decision variables; The radial basis function vector; The regression vector is known.

[0039] To enable online updates of unknown control parameters, a design with... The adaptive parameter update law for the correction term is used to adjust the tracking error. Dynamically update the weight estimates of the radial basis function neural network and fault parameter estimates Its expression is: ; in, and These are the time derivatives of the radial basis function neural network weight estimates and the time derivatives of the actuator fault parameter estimates, respectively. and The learning rate matrix is ​​positive definite. and For robust damping parameters; It is a positive definite matrix; The system matrix is ​​known.

[0040] S4. Based on Lyapunov stability theory, configure control parameters for the decision-making and control layers, enabling the control layer to track the Nash equilibrium reference strategy output by the decision-making layer, and prove that all signals in the closed-loop system are consistent and eventually bounded.

[0041] In this embodiment, the control parameter configurations of the decision-making layer and the control layer satisfy the following conditions: In configuring control parameters at the decision-making level, optimize the step size. With consensus gain It satisfies the algebraic constraints, namely: ; in, It is the smallest non-zero eigenvalue of the Laplace matrix corresponding to the communication topology; Let be the Lipschitz constant of the pseudo-gradient of the local cost function.

[0042] In the control parameter configuration of the control layer, the positive definite matrix It satisfies the algebraic Lyapunov equation, that is: ; in, Let be a given positive definite symmetric matrix; and The system matrix is ​​known. To make the feedback gain matrix of the nominal closed-loop matrix Herwitz.

[0043] In the configuration of learning rate and robustness parameters, the positive definite learning rate matrix and and robust damping parameters and All are constants greater than zero.

[0044] With the above control parameter configuration, the overall Lyapunov function derivative of the closed-loop system is made... satisfy: ; in, For global Lyapunov functions; and , where are the exponential convergence rate constant and the bounded residual term of the system, respectively. The tracking error of the heterogeneous multi-agent system under unknown nonlinear dynamics and actuator failures is constrained to a consistent final bounded set.

[0045] To verify the stability of heterogeneous multi-agent systems, candidate functions are constructed based on Lyapunov stability theory. A Lyapunov function is selected at the decision level. for: ; in, Let be the policy error vector of all agents in the decision-making layer relative to the Nash equilibrium point; This is the consensus estimation error vector among agents regarding each other's policy states; By utilizing the strong monotonicity of the pseudo-gradient of the cost function and the Lipschitz continuity, it can be proved that under suitable control parameters, the policy error and the estimation error are consistent and eventually bounded.

[0046] Select a Lyapunov function in the control layer. for: ; in, Represents the trace operation of a matrix; This is the estimation error matrix for the weights of the radial basis function neural network; This is the estimation error matrix for actuator fault parameters.

[0047] Further construct the global Lyapunov function for: ; in, The number of intelligent agents.

[0048] By analyzing the global Lyapunov function The time derivative can be used to prove the derivative of the global Lyapunov function. satisfy: ; in, and , where are the exponential convergence rate constant and the bounded residual term of the system, respectively. This demonstrates that all signals within the closed-loop system are uniformly eventually bounded.

[0049] Therefore, the adaptive fault-tolerant control method proposed in this invention can achieve stable tracking of the Nash equilibrium strategy by a heterogeneous multi-agent system under the presence of unknown nonlinearity and actuator failure.

[0050] Next, a simulation experiment was conducted in Python to demonstrate the adaptive fault-tolerant control method proposed in this invention, as follows: System parameters and communication topology initialization: Set a configuration that includes... A heterogeneous multi-agent system with [number] agents. Its communication topology is described by a ring graph, such as... Figure 2 As shown. The associated Laplace matrix. The definition is as follows: ; Dynamic model and local cost function settings: Part 1 The heterogeneous dynamics equations for each agent are set as follows: ; The second-order system model is the aforementioned generalized matrix model, namely: The specific physical implementation form of.

[0051] Let the state vector be: ; The corresponding system matrix is: ; ; in, and The first The first and second dimension components of the state vector of each agent; and Constant parameters characterizing the physical properties of each intelligent agent; For control input; To represent unknown nonlinear dynamics, specifically defined as follows: The control parameters for each agent are configured as follows: Agent 1: ; Agent 2: ; Agent 3: ; Agent 4: ; Agent 5: ; The local cost function of a non-cooperative game is defined as: ; in, Indicates the first Local cost function of each agent; and The first Decision variables of each agent and excluding agents The set of decision variables for all other agents besides; For the first Local reference signals for each agent; For the first An intelligent agent (i.e., an intelligent agent) The decision variables are those of the neighbors. Expanded, they take the following form: ; ; Based on the above settings, the theoretical Nash equilibrium point of the system can be calculated as follows: .

[0052] Fault Injection and Controller Parameter Configuration: To verify the fault tolerance of the control layer, a physical fault was injected into the actuator of agent 1 at t=15s during the simulation. This fault was modeled as a constant deviation (e.g., offset caused by actuator jamming), and its expression is as follows: ; At the same time, configure the key parameters of the control system: the optimization step size of the decision layer. The consensus gain parameter is set to 10. The learning rate of the adaptive fault estimation mechanism in the control layer is set to... .

[0053] The adaptive fault-tolerant control method of the present invention was simulated according to the control parameters designed above. Figure 3 The strategies (actions) of each agent are displayed. The convergence variation is shown, where the dashed line represents the theoretically calculated optimal Nash equilibrium point. Figure 4 and Figure 5 The dynamic response of the physical system's tracking error and the adaptive fault parameter estimates are displayed respectively. The changes, among which Figure 5 The black dashed line represents the actual actuator failure value. The system actuator failure occurred at 15 seconds. At the moment of failure, the tracking error of agent 1 showed significant fluctuations and jumps, and the local tracking performance of the system experienced a short-term decline, such as... Figure 4 As shown. This is because the actuator suddenly experienced a constant deviation fault with a value of 2.0, causing the actual physical control quantity to deviate significantly from the optimal reference strategy issued by the decision-making layer. After the fault occurred, this invention began to actively compensate for the drive fault, and the adaptive fault estimator in the control layer reacted quickly. Figure 5 Fault parameter estimates of agent 1 (Solid line) rises rapidly and precisely matches the actual fault value.

[0054] Through real-time compensation using an adaptive fault-tolerant control law, the damaged system performance is restored, the tracking error is rapidly suppressed, and it reconverges to a near-zero compact set. Simultaneously, local faults do not cause malignant propagation in the ring communication topology, and the strategies of each heterogeneous agent eventually converge steadily to the global Nash equilibrium state. Simulation results verify the real-time effectiveness and strong robustness of the method designed in this invention, providing a new perspective for subsequent research and practical engineering applications.

[0055] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0056] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. Since the above embodiments are fundamentally similar to the method embodiments, their descriptions are relatively simple; relevant parts can be found in the descriptions of the method embodiments.

Claims

1. An adaptive fault-tolerant control method for heterogeneous multi-agent distributed Nash equilibrium optimization, characterized in that, The method includes the following steps: S1. Construct a dynamic model to describe the heterogeneous multi-agent system containing unknown nonlinear dynamics and actuator failure physical characteristics, and define the local cost function of the non-cooperative game at the decision-making level. S2. Based on the dynamic model and local cost function, a hierarchical game control framework including a decision layer and a control layer is constructed to decouple the high-level game strategy optimization of the decision layer from the low-level trajectory tracking of the control layer, including: At the decision-making level, a consensus-based gradient descent protocol is used to achieve distributed Nash equilibrium optimization with the local cost function as the objective, so that each agent searches for the Nash equilibrium reference strategy only through local communication information interaction. In the control layer, a radial basis function neural network is used to approximate the unknown nonlinear dynamics of the heterogeneous multi-agent system, and a linear parameterization method is used to describe the actuator fault. S3. Establish an adaptive fault-tolerant control law to compensate for unknown nonlinear dynamics and actuator failures, and combine it with... The adaptive parameter update law of the correction term is used to estimate the weights and fault parameters of the radial basis function neural network online; S4. Based on Lyapunov stability theory, configure control parameters for the decision-making and control layers, enabling the control layer to track the Nash equilibrium reference strategy output by the decision-making layer, and prove that all signals in the closed-loop system are consistent and eventually bounded.

2. The adaptive fault-tolerant control method according to claim 1, characterized in that, The expression for the dynamic model is: ; in, For the first Each agent's state vector The time derivative; For the first The state vectors of each agent; For the first Control input for each intelligent agent; and The system matrix is ​​known. Representing unknown nonlinear dynamics; This indicates an actuator malfunction or external disturbance. The expression for the local cost function is: ; in, Indicates the first Local cost function of each agent; For the first Decision variables for each agent; For local reference signals; Represents intelligent agents The set of neighbors; For the first The decision variables of an agent.

3. The adaptive fault-tolerant control method according to claim 2, characterized in that, Implementing the distributed Nash equilibrium optimization includes: decision variables Mapped to the underlying reference state used to implement the connection between the decision-making layer and the control layer. The expression is: ; in, It is a constant mapping matrix; Based on the underlying reference state Define the tracking error of the control layer for: ; in, For the first The state vectors of each agent; A consensus-based gradient descent protocol, with a local cost function as the objective, achieves distributed Nash equilibrium optimization at the decision level. Its update law is expressed as: ; in, To optimize step size; This is a locally estimated vector; For the first Time derivatives of decision variables of an agent; Let be the gradient operator, indicating that for the , Decision variables of an agent Find the partial derivative; To estimate the vector locally The local cost function for the variable; Introducing auxiliary variables that enable each agent to obtain an estimate of the Nash equilibrium reference policy of other agents. Its dynamic update law is: ; in, Represents intelligent agents For intelligent agents Estimates of decision variables; These are the adjacency matrix elements of the communication topology; Consensus gain; The Kronecker function; For intelligent agents For intelligent agents The rate of evolution of the estimated values ​​of decision variables; For intelligent agents For intelligent agents Estimates of decision variables; The total number of intelligent agents; For the first The decision variables of an agent.

4. The adaptive fault-tolerant control method according to claim 1, characterized in that, The unknown nonlinear dynamics of the heterogeneous multi-agent system are approximated, and actuator faults are described using a linear parameterization method, including: Using radial basis function neural networks to study unknown nonlinear dynamics in heterogeneous multi-agent systems To approximate, its representation is as follows: ; in, This is the weight matrix of the radial basis function neural network; The radial basis function vector; To approximate the error; Using linear parameterization method for actuator faults To perform modeling, that is: ; in, The fault parameter matrix is ​​unknown; The regression vector is known.

5. The adaptive fault-tolerant control method according to claim 1, characterized in that, The adaptive fault-tolerant control law controls the input. The expression is: ; in, This is the feedback gain matrix; This is the feedforward gain matrix; and The weights of the radial basis function neural network are respectively and fault parameters The estimated value; For tracking error in the control layer; For decision variables; The radial basis function vector; The regression vector is known.

6. The adaptive fault-tolerant control method according to claim 5, characterized in that, The adaptive parameter update law is used to adjust the tracking error. Dynamically update the weight estimates of the radial basis function neural network and fault parameter estimates Its expression is: ; in, and These are the time derivatives of the radial basis function neural network weight estimates and the time derivatives of the actuator fault parameter estimates, respectively. and The learning rate matrix is ​​positive definite. and For robust damping parameters; It is a positive definite matrix; The system matrix is ​​known.

7. The adaptive fault-tolerant control method according to claim 1, characterized in that, The control parameter configurations of the decision-making layer and the control layer satisfy the following conditions: In configuring control parameters at the decision-making level, optimize the step size. With consensus gain It satisfies the algebraic constraints, namely: ; in, It is the smallest non-zero eigenvalue of the Laplace matrix corresponding to the communication topology; The Lipschitz constant is the pseudo-gradient of the local cost function; In the control parameter configuration of the control layer, the positive definite matrix It satisfies the algebraic Lyapunov equation, that is: ; in, Let be a given positive definite symmetric matrix; and The system matrix is ​​known. To make the feedback gain matrix of the nominal closed-loop matrix Herwitz; By configuring the control parameters, the derivative of the overall Lyapunov function of the closed-loop system is made... satisfy: ; in, For global Lyapunov functions; and , which are the exponential convergence rate constant and the bounded residual term of the system, respectively.

8. The adaptive fault-tolerant control method according to claim 7, characterized in that, The configuration control parameters also include: In the configuration of learning rate and robustness parameters, the positive definite learning rate matrix and and robust damping parameters and All are constants greater than zero.

9. The adaptive fault-tolerant control method according to claim 8, characterized in that, The global Lyapunov function Including Lyapunov functions selected at the decision-making level And the Lyapunov function selected in the control layer The expression is: ; ; ; in, For intelligent agents Quantity; and The learning rate matrix is ​​positive definite. It is a positive definite matrix; For tracking error in the control layer; Let be the policy error vector of all agents in the decision-making layer relative to the Nash equilibrium point; This is the consensus estimation error vector among agents regarding each other's policy states; Represents the trace operation of a matrix; This is the estimation error matrix for the weights of the radial basis function neural network; This is the estimation error matrix for actuator fault parameters.

Citation Information

Patent Citations

  • Multi-agent system formation strategy based on hierarchical differential game

    CN116360265A

  • Heterogeneous agent distributed multi-alliance game control method, device and equipment and medium

    CN121284029A