Multi-agent system double-layer q learning control method and system based on wasserstein distance

By employing a two-layer Q-learning control method for multi-agent systems based on Wasserstein distance, inner and outer layer controllers were designed to address complex scenarios involving actuator failures and random disturbances in multi-agent systems. This approach achieves stable and consistent control of the system under unknown models, reducing the reliance on precise models.

CN120871645BActive Publication Date: 2025-12-05NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511403591.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2025-12-05
Estimated Expiration
2045-09-29

AI Technical Summary

Technical Problem

Traditional Q-learning methods struggle to handle complex scenarios involving actuator failures and random disturbances in multi-agent systems, especially when the system model is unknown and fault characteristics and disturbance statistics are missing. Existing control methods rely on accurate models and lack efficient handling capabilities for distributed uncertain disturbances, leading to conservative controller design and limited performance.

Method used

A two-layer Q-learning control method based on Wasserstein distance for multi-agent systems is adopted. By constructing an inner fault-tolerant control system and an outer robust control system, a distributed consensus controller is designed. The Q-learning algorithm is used for online policy iterative training to optimize the gain and bias terms of the inner and outer controllers, thereby achieving dynamic decoupling between faults and disturbances.

Benefits of technology

It achieves system stability and asymptotic consistency under actuator failure and random disturbances, reduces the dependence on accurate models and disturbance statistics, and ensures H-infinity performance of the system when failures and disturbances coexist.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120871645B_ABST
    Figure CN120871645B_ABST
Patent Text Reader

Abstract

The application discloses a multi-agent system double-layer Q learning control method and system based on a Wasserstein distance, which comprises the following steps: establishing a multi-agent system state space model under faults and unknown distribution disturbances; constructing an inner fault-tolerant control system and an outer robust control system; designing internal fault-tolerant control gains based on a Q learning algorithm; designing outer robust control gains and bias terms based on a Wasserstein distance and a zero-sum game framework; and designing a distributed consistency protocol in combination with neighborhood state information. The method can effectively reduce the influence of compensation faults and disturbances under the condition that a system model is unknown, an actuator has additive time-varying faults, external disturbance probability distribution is uncertain, and only limited samples are available, and can realize asymptotic state synchronization of all agents while meeting H-infinity performance constraints, thereby ensuring system stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent optimization control technology, and in particular to a two-layer Q-learning control method and system for multi-agent systems based on Wasserstein distance. Background Technology

[0002] Q-learning, a data-driven reinforcement learning method, constructs a Q-function to evaluate the long-term value of state-action pairs, thereby deriving the optimal control policy without relying on an explicit dynamic model of the system. This method has been widely applied in the cooperative control of multi-agent systems to address issues such as consistency and containment control. However, traditional Q-learning methods are mostly designed for single types of uncertainty, failing to fully consider complex scenarios where actuator failures and random disturbances coexist, and lacking efficient mechanisms for handling distributed uncertainty disturbances.

[0003] In practical applications, multi-agent systems often face the dual challenges of actuator failure and external random disturbances. Actuator failure can lead to performance degradation or even loss of control, and is both unavoidable and potentially catastrophic. On the other hand, random disturbances such as environmental noise and load variations often have unknown statistical characteristics. Traditional H-infinity control methods are typically designed for deterministic disturbances with bounded energy, making it difficult to handle random disturbances with unknown statistical characteristics, and they usually require a precise mathematical model of the system. Adding to the complexity, there is a dynamic coupling between failures and disturbances. Individually designed fault-tolerant or robust controllers are insufficient to achieve optimal overall system performance, necessitating a collaborative control architecture capable of handling both types of uncertainty simultaneously.

[0004] Despite some progress in fault-tolerant and robust control, existing technologies still suffer from the following significant drawbacks: First, traditional methods heavily rely on accurate system models, fault characteristics, and disturbance statistics, which are difficult to obtain in practice. Second, traditional H-infinity control methods primarily target energy-bounded disturbances and have limited ability to handle stochastic disturbances with unknown statistical characteristics. Third, most Q-learning methods can only handle single types of uncertainty and lack optimization capabilities in scenarios where faults and disturbances coexist. Furthermore, the coupling effect between faults and disturbances is not sufficiently decoupled, leading to conservative controller design and limited performance. Summary of the Invention

[0005] The purpose of this invention is to provide a two-layer Q-learning control method and system for multi-agent systems based on Wasserstein distance, which can achieve consistent control of multi-agent systems under the premise of satisfying the H-infinity performance index, even when the system model is unknown and the actuator fault characteristics and external disturbance distribution information are missing.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A two-layer Q-learning control method for multi-agent systems based on Wasserstein distance includes:

[0008] Step 1: Establish a mathematical model of a discrete-time multi-agent system with actuator failures and random disturbances with unknown probability distributions;

[0009] Step 2: Based on the mathematical model of the discrete-time multi-agent system, construct an inner-layer fault-tolerant control system and an outer-layer robust control system;

[0010] Step 3: Design the inner fault-tolerant control system controller and local value function, and solve for the optimal gain of the inner fault-tolerant control system controller based on the Q-learning algorithm to minimize the local value function;

[0011] Step 4: Design the local performance index function of the outer robust control system based on Wasserstein distance and zero-sum game framework. Solve for the optimal gain and bias terms through Q-learning algorithm to minimize the local performance index function. Combine neighborhood state information, gain and bias terms to design the controller of the outer robust control system.

[0012] Step 5: Combine the inner fault-tolerant control system controller and the outer robust control system controller to design a distributed consensus controller for the control of the multi-agent system.

[0013] Furthermore, the mathematical model of the discrete-time multi-agent system described in step 1 is as follows:

[0014] ;

[0015] in, , and They represent the first The state of the agent, the actuator fault input signal, and the unknown probability distribution External random disturbances, where A, B, and E are unknown time-invariant system matrices. , and They are respectively represented as dimension, peacekeeping 3D column vector, Let this be the discrete time step at the current moment. For the next discrete time step, the actuator fault input signal is:

[0016] ;

[0017] in, For an unknown time-varying additive actuator fault, This serves as the input for the distributed consensus control to be designed. For energy-limited space, The number of agents.

[0018] Furthermore, the inner-layer fault-tolerant control system described in step 2 is as follows:

[0019] ;

[0020] in, , It is the state of the outer robust control system. The controller for the inner-layer fault-tolerant control system to be designed;

[0021] The outer robust control system is:

[0022] ;

[0023] in, The outer robust control system controller to be designed.

[0024] Furthermore, the inner-layer fault-tolerant control system controller is:

[0025] ;

[0026] in, Adjacency matrix The Line 1 The values ​​in the column represent the agent. With intelligent agents Communication weights between them, adjacency matrix It includes the interconnections between all intelligent agents. , Let be the controller gain of the inner-layer fault-tolerant control system to be solved.

[0027] Furthermore, the local value function introduces the Bellman optimality equation, which is:

[0028] ;

[0029] in, and All are positive definite matrices. For a given level of interference suppression, This is an additive fault signal test signal. To test the control input signal, To test the scaling factor of the control input signal.

[0030] Furthermore, the Q-function value of the Q-learning algorithm for solving the optimal gain of the controller in the inner fault-tolerant control system is:

[0031] ;

[0032] Expressing the Q-function value in parameterized form, the optimal inner fault-tolerant controller gain is obtained using the Q-learning algorithm:

[0033] ;

[0034] in, Scaling factor for The optimal solution, parameters , .

[0035] Furthermore, based on the Wasserstein distance and zero-sum game framework, the local performance index function of the outer robust control system is designed as follows:

[0036] ;

[0037] in, , , , For penalty parameters, It is a discount factor. To test the input signal, To test random interference input signals, The control gain of the outer robust control system to be solved is... To test the gain matrix of random interference input signals, Expressed as the mathematical expectation of a random variable, Indicates the distance to Wasserstein. , It focuses on interfering samples Dirac measure, The number of interference samples, and This is a bias term.

[0038] Furthermore, the Q-function of the Q-learning algorithm for solving the optimal gain and bias terms is:

[0039] ;

[0040] Among them, matrix , and constant for:

[0041] , , ;

[0042] in, , and It is a constant; and The submatrices are: , , , , , , , , , , , , , It is a linear coefficient vector, specifically , and The calculation expression is:

[0043] ,

[0044] ,

[0045] ;

[0046] The optimal control gain and bias term are obtained by using the Q-learning online policy iterative algorithm:

[0047] , ,

[0048] , ;

[0049] in, for The optimal solution. for The optimal solution;

[0050] The outer robust control system controller is further designed as follows:

[0051] .

[0052] A two-layer Q-learning control system for multi-agent systems based on Wasserstein distance, comprising:

[0053] The mathematical model building unit for multi-agent systems establishes a mathematical model for discrete-time multi-agent systems with actuator failures and random perturbations with unknown probability distributions.

[0054] The control system construction unit is based on the mathematical model of discrete-time multi-agent system to construct an inner fault-tolerant control system and an outer robust control system.

[0055] The inner fault-tolerant control system controller design unit designs the inner fault-tolerant control system controller and the local value function. Based on the Q-learning algorithm, it solves the optimal gain of the inner fault-tolerant control system controller to minimize the local value function.

[0056] The outer robust control system controller design unit designs the local performance index function of the outer robust control system based on the Wasserstein distance and zero-sum game framework. The optimal gain and bias terms are solved by the Q-learning algorithm. Combined with neighborhood state information, gain and bias terms, the outer robust control system controller is redesigned.

[0057] The consensus controller design unit combines the inner fault-tolerant control system controller and the outer robust control system controller to design a distributed consensus controller for the control of a multi-agent system.

[0058] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present application designs a data-driven two-layer Q-learning architecture, realizes dynamic decoupling of actuator faults and random disturbances by constructing an auxiliary system, and derives the explicit expressions and online update laws of the inner and outer controllers; (2) In order to reduce the dependence on precise mathematical models, fault prior information and disturbance statistical characteristics, the present application introduces the Q-learning framework in reinforcement learning, designs a distributed robust controller based on Wasserstein distance and zero-sum game, and obtains the optimal H-infinity control policy through online policy iterative training; (3) The present application rigorously proves the consistency of the state of the multi-agent system through Lyapunov stability theory, and proves that the proposed safety control method can ensure that the system remains stable and reaches asymptotic consensus under the condition of actuator faults and random disturbances with unknown probability distribution; (4) In the case of unknown system model, additive time-varying faults of actuators, uncertain probability distribution of external disturbances and only limited samples, the present method can effectively reduce the impact of compensation faults and disturbances, realize the asymptotic state synchronization of all agents and H-infinity performance constraints, and finally make the system stable. Attached Figure Description

[0059] Figure 1 This is a flowchart of the method provided in this application.

[0060] Figure 2 This is a block diagram of the distributed system provided in this application.

[0061] Figure 3 This is a topology diagram provided in the embodiments of this application.

[0062] Figure 4 This is a schematic diagram of random interference distribution provided in an embodiment of this application.

[0063] Figure 5 This is a schematic diagram of the controller gain matrix provided in an embodiment of this application.

[0064] Figure 6 This is an internal tracking error trajectory diagram provided in the embodiments of this application.

[0065] Figure 7 This is a system status trajectory diagram provided in the embodiments of this application. Detailed Implementation

[0066] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0067] Combination Figure 1 This application provides a two-layer Q-learning control method for multi-agent systems based on Wassstein distance, and the specific implementation method is as follows:

[0068] Step 1: Establish a mathematical model of a discrete-time multi-agent system with actuator failures and random disturbances with unknown probability distributions;

[0069] Combination Figure 2 Specifically, a mathematical model of a multi-agent system subjected to actuator failures and random disturbances is established:

[0070] ;

[0071] in, , and They represent the first The state of the agent, the actuator fault input signal, and the unknown probability distribution External random disturbances, where A, B, and E are unknown time-invariant system matrices. , and They are respectively represented as dimension, peacekeeping 3D column vector, Let this be the discrete time step at the current moment. For the next discrete time step, the actuator fault input signal is:

[0072] ;

[0073] in, For an unknown time-varying additive actuator fault, This serves as the input for the distributed consensus control to be designed. For energy-limited space, The number of intelligent agents.

[0074] Step 2: Construct an inner fault-tolerant control system and an outer robust control system; specifically including:

[0075] Design an inner-layer fault-tolerant control system:

[0076] ;

[0077] in, , It is the state of the outer robust control system. The controller for the inner-layer fault-tolerant control system to be designed;

[0078] Design an outer-layer robust control system:

[0079] ;

[0080] in, It is the state of the outer robust control system;

[0081] The total control input signal is:

[0082] ;

[0083] in It is an inner-layer fault-tolerant controller used to compensate for actuator failures. It serves as an outer robust controller to suppress random disturbances.

[0084] Step 3: Design the controller of the inner fault-tolerant control system and solve for the optimal control gain of the inner controller based on the Q-learning algorithm;

[0085] The inner fault-tolerant controller is designed as follows:

[0086] ;

[0087] in, Adjacency matrix The Line 1 The values ​​in the column represent the agent. With intelligent agents Communication weights between them, adjacency matrix It includes the interconnections between all intelligent agents. , Let be the controller gain of the inner-layer fault-tolerant control system to be solved.

[0088] Design local value functions:

[0089] ;

[0090] in and All are positive definite matrices. Given the level of interference suppression;

[0091] Using the Bellman optimality principle, the local value function is re-expressed in parameterized form:

[0092] ;

[0093] in, and All are positive definite matrices. For a given level of interference suppression, This is an additive fault signal test signal. To test the control input signal, To test the scaling factor of the control input signal, This represents the total value of the local value function.

[0094] Furthermore, the Q-function value of the Q-learning algorithm for solving the optimal gain of the controller in the inner fault-tolerant control system is:

[0095] ;

[0096] Expressing the Q-function value in parameterized form, the optimal inner fault-tolerant controller gain is obtained using the Q-learning algorithm:

[0097] ;

[0098] in, Scaling factor for The optimal solution, parameters , .

[0099] By solving and The optimal fault-tolerant control gain is obtained:

[0100] ;

[0101] The specific iterative steps of the Q-learning algorithm are as follows:

[0102] Step 3.1: Initialize the number of iterations Choose the initial Q-function matrix Arbitrary initialization and ;

[0103] Step 3.2, in the... In this iteration, system data is collected. ;

[0104] Step 3.3: Update the Q-function matrix using the least squares method. :

[0105] ;

[0106] Among them, data ,

[0107] , and They are matrices The line, number Column elements and vectors The One component;

[0108] Step 3.4, based on the updated Calculate the new control gain:

[0109] ;

[0110] ;

[0111] Step 3.5, if ( If the value is a very small constant, then stop the iteration to obtain the optimal fault-tolerant control gain, denoted as . Otherwise, Return to step 3.2.

[0112] Step 4: Design the outer robust control gain and bias terms based on Wasserstein distance and zero-sum game framework, and design the optimal outer robust controller by combining neighborhood state information, gain, and bias terms; specifically including:

[0113] To address random disturbances, a partial Bruker optimization method based on Wassstein distance is adopted, and a performance index function is designed:

[0114] ;

[0115] in, , , , For penalty parameters, It is a discount factor. To test the input signal, To test random interference input signals, The Wasserstein distance is used to measure the empirical distribution of perturbed samples. With the true distribution The deviation between them is:

[0116] ;

[0117] in, ,and Indicates marginal distribution and The joint distribution set. Furthermore, by collecting only a finite number... There are 1 interfering sample, denoted as _____. The empirical distribution expression is as follows:

[0118] ;

[0119] in, It is focused on Dirac measure;

[0120] Using Kantorovich duality theory, the above problem is transformed into a deterministic zero-sum game problem, and the Q-function is defined:

[0121] ;

[0122] Among them, matrix , and constant It manifests as:

[0123] , , ;

[0124] in, , and It is a constant; and The submatrices are: , , , , , , , , , , , , , It is a linear coefficient vector, specifically , and The calculation expression is:

[0125] ,

[0126] ,

[0127] ;

[0128] The optimal control gain and bias term are obtained by using the Q-learning online policy iterative algorithm:

[0129] , ,

[0130] , ;

[0131] in, for The optimal solution. for The optimal solution;

[0132] The specific steps of the online policy iteration algorithm include:

[0133] Step 4.1: Initialize the number of iterations Select initial parameters and ;

[0134] Step 4.2, in the... In this iteration, system data is collected. ;

[0135] Step 4.3: Update the parameter vector using the least squares method. :

[0136] ;

[0137] in, , ;

[0138] Step 4.4, based on the updated The new control gain and bias terms are calculated as follows:

[0139] ,

[0140] ,

[0141] ,

[0142] ;

[0143] in, ;

[0144] Step 4.5, if If the result is positive, then stop the iteration to obtain the optimal control gain and bias term; otherwise, let... Return to step 4.2.

[0145] The design objective of the external robust controller is to ensure that the system meets the H-infinity performance index under worst-case disturbance conditions, meaning that for all energy-bounded disturbances, the ratio of the system's output energy to the disturbance energy does not exceed a given attenuation level. This performance metric is guaranteed by the following inequality:

[0146] ;

[0147] in, Let H be an infinite performance index.

[0148] By combining neighborhood state information, control gain, and bias term, the external robust controller is redesigned:

[0149] ;

[0150] Step 5: Combining the optimal controller of the inner fault-tolerant control system and the optimal outer robust controller, design a distributed consensus control protocol.

[0151] The distributed consensus control protocol for a single intelligent agent is as follows:

[0152] ;

[0153] This control protocol aims to enable distributed H-infinite consistency control in multi-agent systems under the presence of actuator failures and random disturbances, ensuring that the states of all agents eventually reach asymptotic consistency.

[0154] This invention also provides a two-layer Q-learning control system for multi-agent systems based on Wasserstein distance, comprising:

[0155] The mathematical model building unit for multi-agent systems establishes a mathematical model for discrete-time multi-agent systems with actuator failures and random perturbations with unknown probability distributions.

[0156] The control system construction unit is based on the mathematical model of discrete-time multi-agent system to construct an inner fault-tolerant control system and an outer robust control system.

[0157] The inner fault-tolerant control system controller design unit designs the inner fault-tolerant control system controller and the local value function. Based on the Q-learning algorithm, it solves the optimal gain of the inner fault-tolerant control system controller to minimize the local value function.

[0158] The outer robust control system controller design unit designs the local performance index function of the outer robust control system based on the Wasserstein distance and zero-sum game framework. The optimal gain and bias terms are solved by the Q-learning algorithm. Combined with neighborhood state information, gain and bias terms, the outer robust control system controller is redesigned.

[0159] The consensus controller design unit combines the inner fault-tolerant control system controller and the outer robust control system controller to design a distributed consensus controller for the control of a multi-agent system.

[0160] The external robust controller is trained iteratively through online policy and dynamically updates the control gain using real-time data. and bias terms This method achieves optimal suppression of random disturbances. It does not rely on prior knowledge of the disturbance distribution and requires only a finite number of samples to achieve distributed bar control.

[0161] Example

[0162] The effectiveness of the present invention will be further illustrated below with specific examples.

[0163] A multi-agent system consisting of six High-Maneuverability Aircraft Televisions (HiMATVs) was constructed for simulation testing of theoretical results. The six aircraft were numbered 1 to 6. The system communication topology is as follows: Figure 3 As shown, it is a directed graph structure that satisfies the conditions of Assumption 1.

[0164] Consider a system with the following parameters:

[0165] ;

[0166] Sampling time set to For the first There are HiMATVs, and their state vectors are defined as follows: The initial system state is set as follows: , , , , , .

[0167] according to Figure 3 The communication topology and adjacency matrix shown for:

[0168] ;

[0169] In addition, the actuator additive fault is set as , .

[0170] External random disturbances The probability distribution is unknown. To simulate the uncertainty of the distribution, an interference term is defined. They conform to a Gaussian mixture distribution and have equal probabilities, respectively. and ,like Figure 4 As shown. Meanwhile, independent noise. From normal distribution Extraction from the middle. The controller design can only acquire a limited number of disturbance samples; in this example, the number of samples in each group is set to... That is, for each intelligent agent Collect 40 random samples The mean values ​​of the perturbation samples of each agent are calculated as follows: , , , , , .

[0171] The parameters for the Q-learning algorithm are set as follows: discount factor Punishment factor The weight matrix is ​​set as follows: Fault-tolerant controller scaling factor ;parameter .

[0172] After setting the initial system state, learning parameters, and communication topology, as shown in the figure... Figure 5 As shown, after approximately 8 policy iterations, the fault-tolerant control gain... Converging to a stable value After approximately 100 policy iterations, robust control gain was achieved. Converging to a stable value After algorithm iteration and updates, the trajectories of the internal state tracking error and the original system state are as follows: Figure 6 and Figure 7 As shown, the tracking errors of all agents eventually converge to near 0, indicating that the inner fault-tolerant controller effectively compensates for actuator failures and achieves consistency in internal state tracking. Despite actuator failures and random disturbances of unknown distribution, the states of all agents eventually reach consistency and stabilize near 0, demonstrating that the proposed two-layer control architecture effectively achieves asymptotic consistency of the system state. Figures 4 to 7 The coordinate axes are all dimensionless units.

[0173] This invention addresses the H-infinity consistency control problem of discrete-time multi-agent systems under actuator failure and random disturbances with unknown probability distributions. To decouple and compensate for actuator failures and random disturbances, this application designs a data-driven control architecture based on two-layer Q-learning, derives fault-tolerant and robust control gain matrices, and ensures asymptotic consistency of the system state when both failures and disturbances coexist. Furthermore, to reduce dependence on the statistical characteristics of disturbances, this application introduces the Wasserstein distance and zero-sum game framework, obtaining the optimal minimax control strategy through online policy iterative training. Finally, this application derives the uniform convergence of the system state and tracking error, proving that the proposed control method can ensure system stability under actuator failures and random disturbances. The two-layer Q-learning control method for multi-agent systems under failure and distribution uncertainty proposed in this invention can be applied to practical scenarios such as UAV formation, smart grid collaborative scheduling, and multi-robot collaborative operations.

[0174] It should be understood that the above description of the preferred embodiments is quite detailed, but it should not be considered as a limitation on the scope of protection of this invention. Those skilled in the art, under the guidance of this invention, can make substitutions or modifications without departing from the scope of protection of the claims of this invention, and all such substitutions or modifications fall within the scope of protection of this invention. The scope of protection of this invention should be determined by the appended claims.

Claims

1. A two-layer Q-learning control method for multi-agent systems based on Wasserstein distance, characterized in that: The application relates to a method for designing a distributed fault-tolerant and robust controller for a multi-agent system with actuator failures and random disturbances with unknown probability distribution. The method comprises the following steps: Step 1, establishing a discrete-time multi-agent system mathematical model with actuator failures and random disturbances with unknown probability distribution; Step 2, constructing an inner fault-tolerant control system and an outer robust control system based on the discrete-time multi-agent system mathematical model; Step 3, designing an inner fault-tolerant control system controller and a local value function, and solving the optimal gain of the inner fault-tolerant control system controller based on a Q-learning algorithm to minimize the local value function; Step 4, designing a local performance index function of the outer robust control system based on a Wasserstein distance and a zero-sum game framework, solving the optimal gain and bias term based on a Q-learning algorithm to minimize the local performance index function, and combining the neighborhood state information, the gain and the bias term to design an outer robust control system controller; Step 5, combining the inner fault-tolerant control system controller and the outer robust control system controller to design a distributed consensus controller and control the multi-agent system. ; wherein, , and respectively represent the state of the th agent, the actuator fault input signal and the external random disturbance with unknown probability distribution , A, B and E are unknown time-invariant system matrices, , and respectively represent -dimensional, -dimensional and -dimensional column vectors, is the current time step, is the next time step, and the actuator fault input signal is: ; wherein, is an unknown time-varying additive actuator fault, is a distributed consensus control input to be designed, is an energy-limited space, is the number of agents; The discrete-time multi-agent system mathematical model in step 1 is as follows: ; wherein the variables , are the outer robust control system states, is the inner fault-tolerant control system controller to be designed; The inner fault-tolerant control system in step 2 is as follows: ; wherein, is the outer robust control system controller to be designed; The outer robust control system is as follows: ; in, , , , For penalty parameters, It is a discount factor. To test the input signal, To test random interference input signals, The control gain of the outer robust control system to be solved is... To test the gain matrix of random interference input signals, Expressed as the mathematical expectation of a random variable, Indicates the distance to Wasserstein. , It focuses on interfering samples Dirac measure, The number of interference samples, and For bias terms; The local performance index function of the outer robust control system designed based on the Wasserstein distance and the zero-sum game framework is as follows: ; where the matrix , and the constant are given by: , , ; wherein , , are constants; , each sub-matrix of , , , , , , , , , , , , , is a linear coefficient vector, the parameters , and are calculated as follows: , , ; The Q function of the Q-learning algorithm for solving the optimal gain and the bias term in step 4 is as follows: , , , ; wherein is the optimal solution of is the optimal solution of 2.The multi-agent system double-layer Q-learning control method based on the Wasserstein distance according to claim 1, wherein: The optimal control gain and the bias term solved by the Q-learning online policy iteration algorithm are as follows: ; wherein, is the value of the adjacency matrix at the row and the column, representing the communication weight between agent and agent , , is the inner fault-tolerant control system controller gain to be solved.

3. The multi-agent system double deep Q-learning control method based on the Wasserstein distance according to claim 2, characterized in that: The inner fault-tolerant control system controller is as follows: ; wherein and are positive definite matrices, is a given disturbance suppression level, is an additive fault signal test signal, is a test control input signal, is a scaling factor for the test control input signal, is a local cost function value.

4. The multi-agent system double deep Q-learning control method based on the Wasserstein distance according to claim 3, characterized in that: The local value function introduces the Bellman optimality equation and is as follows: ; The Q function value of the Q-learning algorithm in step 3 is as follows: ; wherein is a scaling factor, is optimal solution of , .

5. The multi-agent system double deep Q-learning control method based on the Wasserstein distance according to claim 1, characterized in that: The Q function value is expressed in a parameterized form, and the optimal inner fault-tolerant controller gain is solved by the Q-learning algorithm and is as follows: 。 6. A Wasserstein distance based multi-agent system double deep Q-learning control system implementing the method of any one of claims 1-5. The outer robust control system controller is designed as follows: The application relates to a method for designing a distributed fault-tolerant and robust controller for a multi-agent system with actuator failures and random disturbances with unknown probability distribution. The method comprises the following steps: A multi-agent system mathematical model construction unit is configured to establish a discrete-time multi-agent system mathematical model with actuator failures and random disturbances with unknown probability distribution; A control system construction unit is configured to construct an inner fault-tolerant control system and an outer robust control system based on the discrete-time multi-agent system mathematical model; An inner fault-tolerant control system controller design unit is configured to design an inner fault-tolerant control system controller and a local value function, and solve the optimal gain of the inner fault-tolerant control system controller based on a Q-learning algorithm to minimize the local value function; An outer robust control system controller design unit is configured to design a local performance index function of the outer robust control system based on a Wasserstein distance and a zero-sum game framework, solve the optimal gain and bias term based on a Q-learning algorithm, and combine the neighborhood state information, the gain and the bias term to design an outer robust control system controller; A consensus controller design unit is configured to combine the inner fault-tolerant control system controller and the outer robust control system controller to design a distributed consensus controller and control the multi-agent system.

Citation Information

Patent Citations

  • Industrial process fault-tolerant control method based on data-driven Q-learning

    CN114035523A

  • Multi-agent space-time dynamic system intrusion tolerance and fault tolerance cooperative security consistency control method

    CN119668180A