An optimal control method for human-in-the-loop multi-agent systems under FDI attacks

By constructing a non-fully autonomous leader system and a dynamic event triggering mechanism, combined with a neural network optimization control strategy, the performance optimization and communication resource consumption problems of the multi-agent system under FDI attacks are solved, thereby improving the system security and reliability.

CN119758816BActive Publication Date: 2025-09-19INST OF ELECTRONICS & INFORMATION ENG OF UESTC IN GUANGDONG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411845812.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-09-19
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

Existing technologies cannot effectively design strategies for attackers and defenders under FDI attacks, resulting in difficulties in optimizing the performance of multi-agent systems when attacked by false data injection and high consumption of communication resources.

Method used

Build a non-fully autonomous leader system, combine neural networks and dynamic event triggering mechanisms, design performance functions and Nash equilibrium points, and optimize control strategies and reduce communication resource consumption through zero-sum game and adaptive dynamic programming theory.

Benefits of technology

Dynamic event-triggered game optimization control of multi-agent systems under FDI attacks is realized, which improves system security and reliability while saving communication resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119758816B_ABST
    Figure CN119758816B_ABST
Patent Text Reader

Abstract

The present invention provides an optimization control method for a human-in-the-loop multi-agent system under FDI attack, belonging to the technical field of multi-agent system control. The control method constructs a non-fully autonomous leader system and takes into account the FDI attack from the follower controller to the actuator end. In view of the antagonistic relationship between the attacker and the defender, a performance function is constructed and the Nash equilibrium point, that is, the optimal strategy of the attacker and the defender, is obtained under the unified design framework of zero-sum game and adaptive dynamic programming theory; at the same time, in order to reduce the consumption of network resources, a dynamic event triggering mechanism is designed; combined with a single evaluation network structure based on a neural network, the Nash optimal solution is obtained. The control method of the present invention realizes dynamic event-triggered game optimization control under the influence of FDI attack, constructs a method for simultaneously designing the optimal strategies of the attacker and the defender under a unified framework, and saves communication resources while achieving game performance optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multi-agent system control, and in particular relates to an optimization control method for a human-in-the-loop multi-agent system under FDI attack. Background Art

[0002] With the rapid development of technologies such as artificial intelligence, the Internet of Things, and big data, multi-agent systems have been widely applied in fields such as automation, intelligent transportation, robotic collaboration, smart grids, and financial market analysis. Exploring these systems holds significant strategic significance and positive economic prospects. Because the security and reliability of fully autonomous systems cannot be fully guaranteed, the control of human-in-the-loop multi-agent systems is crucial. In a network communication topology composed of agents, malicious network attacks, such as false data injection attacks (FDI attacks), can significantly impact the accurate transmission of information between agents. Furthermore, existing research on human-in-the-loop multi-agent control (CN202410193921.8, CN202410031096.1) lacks a solution for simultaneously designing attacker and defender strategies under FDI attacks, making it impossible to achieve a compromise in optimizing system performance when subjected to FDI attacks. Furthermore, effectively reducing network communication resource overhead is an inevitable requirement for the increasing complexity of multi-agent systems.

[0003] Therefore, under the trend of high informatization, collaboration and intelligence, it is of practical significance to design a dynamic event-triggered game optimization method for human-in-the-loop multi-agent systems under FDI attacks. Summary of the Invention

[0004] In response to the problems existing in the background technology, the purpose of the present invention is to provide an optimization control method for a human-in-the-loop multi-agent system under FDI attack. This control method constructs a non-fully autonomous leader system and takes into account the FDI attack on the follower controller to the actuator end. In view of the antagonistic relationship between the attacker and the defender, a performance function is constructed and the Nash equilibrium point, that is, the optimal strategy of the attacker and the defender, is obtained under the unified design framework of zero-sum game and adaptive dynamic programming theory; at the same time, in order to reduce the consumption of network resources, a dynamic event triggering mechanism is designed; combined with a single evaluation network structure based on a neural network, the Nash optimal solution is obtained. The control method of the present invention realizes dynamic event-triggered game optimization control under the influence of FDI attack, constructs a method for simultaneously designing the optimal strategies of the attacker and the defender under a unified framework, and saves communication resources while achieving game performance optimization.

[0005] To achieve the above object, the technical solution of the present invention is as follows:

[0006] An optimization control method for a human-in-the-loop multi-agent system under FDI attack includes the following steps:

[0007] Step 1: Construct a dynamic model of a non-fully autonomous multi-agent system, including the dynamic equations of the externally controlled leader and followers, where the external human expert command signal is an unknown quantity;

[0008] Step 2: Based on the dynamic model of the multi-agent system constructed in step 1, consider the FDI attack on the follower controller to the actuator end, and construct a dynamic model of the agent system attacked by FDI;

[0009] Step 3: Based on the dynamic model of the agent system attacked by FDI obtained in step 2, design a performance function and find the Nash equilibrium point for each agent, that is, the Nash optimal strategy of the attacker and defender;

[0010] Step 4: Based on the Nash equilibrium point in step 3, design the HJI equation and Nash optimal strategy based on dynamic event triggering;

[0011] Step 5: Based on the HJI equation based on dynamic event triggering obtained in step 4, a dynamic event triggering mechanism is designed to save network resource overhead;

[0012] Step 6: Design a single-evaluation network model based on a neural network to approximate the performance function in step 3, and obtain an approximate solution to the Nash equilibrium point triggered by dynamic events;

[0013] Step 7: Based on the Bellman residual generated in the approximation process in step 6, use the gradient descent method to generate the weight change law of the single evaluation network in step 6 until convergence, and the control input can be obtained based on the weight result.

[0014] Furthermore, in step 1, a non-fully autonomous multi-agent system is established, including a leader and N followers. The multi-agent system is specifically:

[0015] Leaders:

[0016] Followers:

[0017] in, represents the derivative, the subscript i is the number of a single agent system, i = 1, 2, ..., N, represents the system state of agent i, represents the control input of agent i, represents the internal dynamics of agent i, represents the output dynamics of agent i; Represents the system state of the leader, represents an unknown and non-zero command signal sent by an external human expert to the leader, Represents the leader's internal dynamics, Represents the output dynamics of the leader.

[0018] Furthermore, in step 2, the specific process of establishing the dynamic model of the i-th agent system attacked by FDI is as follows:

[0019] Considering the FDI attack on the channel from the controller of agent i to the actuator, the control input of agent i that suffers from FDI attack is,

[0020]

[0021] Among them, u i for u i The abbreviation of (t), χ i Input for FDI attack;

[0022] The dynamic model of the i-th agent attacked by FDI can be expressed as:

[0023]

[0024] in, It is the state of the system under attack.

[0025] Furthermore, the attacker should meet the following conditions:

[0026] (1) All executive agencies may be attacked by attackers at any time.

[0027] (2) State variables can be monitored in real time and provided as information to attackers,

[0028] (3) False data and control signals can be injected into the actuator synchronously,

[0029] (4) The attacker's energy is limited.

[0030] Furthermore, the specific process of step 3 is:

[0031] The dynamic equation defining the bipartite consistency error of the i-th agent is:

[0032]

[0033] in, a ij is the connection weight between agent i and agent j, if a ij >0 means that agent i and agent j are in a cooperative relationship. If a ij <0 means that agent i and agent j are in a competitive relationship; b i is the connection weight between agent i and the leader, if b i>0 means that there is a cooperative relationship between agent i and the leader, otherwise it is a competitive relationship; represents the degree matrix and

[0034]

[0035] According to game theory, the expected performance of control input is improved while the expected performance of external FDI attack is reduced. There is an antagonistic relationship between the two. When the performance of the two reaches the optimal equilibrium, a Nash equilibrium point can be formed. Therefore, the design performance function is as follows:

[0036]

[0037] Among them, γ i represents the desired attack attenuation level, γ i >0, Q ii 、R ij 、R ii 、P ii and P ij are all symmetric positive definite matrices with matching dimensions, represents the neighboring agent of the i-th agent;

[0038] According to Pontryagin's maximum-minimum principle, the Nash equilibrium point The following inequality is satisfied:

[0039]

[0040] The Hamiltonian function is constructed as follows:

[0041]

[0042] in, represents the partial derivative of the performance function with respect to the bipartite consensus error;

[0043] Then the corresponding HJI (Hamilton-Jacobi-Isaacs) equation can be expressed as:

[0044] According to the adaptive dynamic programming theory, the specific form of the Nash equilibrium point is:

[0045]

[0046] Furthermore, the specific process of step 4 is:

[0047] definition is the kth triggering moment, and satisfies The current state With sampling status The error between can be expressed as:

[0048]

[0049] The bipartite consistency error based on dynamic event triggering has the following form:

[0050]

[0051] The dynamic event triggering error is expressed as:

[0052]

[0053] Considering that FDI attacks are external outputs and uncontrollable, the specific forms of the optimal control strategy based on dynamic event triggering and the optimal FDI attack strategy based on time triggering are:

[0054]

[0055] Combining the above results with the HJI equation, the specific form of the HJI equation based on dynamic event triggering is obtained as follows:

[0056]

[0057] Furthermore, the specific process of step 5 is: according to the HJI equation based on dynamic event triggering in step 4, the dynamic event triggering mechanism is designed as follows:

[0058]

[0059] Among them, T DED is the dynamic trigger threshold, and

[0060]

[0061] λ min (Q ii ) represents the matrix Q ii The smallest eigenvalue, ν i,1 is the design parameter, 0<ν i,1 <1,

[0062] Design dynamic parameters The change law of is as follows:

[0063]

[0064] in,

[0065]

[0066] Furthermore, the specific process of step 6 is:

[0067] A single evaluation network model based on a neural network is designed to approximate the performance function in step 3. The specific expression of the single evaluation network model based on a neural network is:

[0068]

[0069] in, is the ideal weight vector, h ci is the number of neurons in the hidden layer, S c (∈ i ) is a differentiable continuous activation function, o(∈ i ) is the approximation error;

[0070] The performance function gradient is expressed as:

[0071]

[0072] in, is the gradient of the activation function, is the gradient of the approximation error.

[0073] Considering that the ideal weight vector is an unknown quantity and cannot be directly obtained, the estimated form of the approximate value of the gradient value of the performance function is constructed as follows:

[0074]

[0075] in, W ci The estimated form of .

[0076] Furthermore, the estimated value of the optimal control strategy based on dynamic event triggering and the approximate value of the optimal FDI attack strategy based on time triggering in step 4 is obtained as follows:

[0077]

[0078] Combining the above results with the HJI equation based on dynamic event triggering in step 4, the estimated form of its approximate value is obtained as:

[0079]

[0080] in,

[0081]

[0082] Furthermore, the specific process of step 7 is:

[0083] The Bellman residual generated in the approximation process in step 6 is The Bellman residual is defined as:

[0084]

[0085] The design optimization objective function is:

[0086]

[0087] Based on this function, using the gradient descent method, the specific form of the change law of the weight of the single evaluation network in step 6 can be obtained as follows:

[0088]

[0089] Where 0<μ i <1 is the design parameter;

[0090] Substitute the obtained weight value into the expression form of control input,

[0091]

[0092] The control input of the system can be obtained.

[0093] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0094] This method utilizes trajectory adjustment command signals sent by external human experts to the leader to respond to emergencies in real time, improving system safety and reliability. It also considers FDI attacks between the intelligent agent's controller and actuator, establishing a system model under attack. Based on this model, and considering the antagonistic relationship between attackers and defenders, within the unified design framework of zero-sum game theory and adaptive dynamic programming theory, a performance function is constructed to obtain the Nash equilibrium point, i.e., the optimal strategy for attackers and defenders, effectively achieving a balance in the system's game optimization performance. More importantly, considering the need to conserve communication resources in practical applications, a dynamic event triggering mechanism is designed to ensure that communication resources are conserved while optimizing game performance, effectively improving control quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] Figure 1 Flowchart of the control method of the present invention.

[0096] Figure 2 FIG. 1 is a schematic diagram of the communication topology adopted by the present invention.

[0097] Figure 3 A graph showing the state consistency of the intelligent agent of the present invention.

[0098] Figure 4 Graph showing the bipartite consistency error of the agent of the present invention.

[0099] Figure 5 This is the FDI attack curve of the intelligent agent of the present invention.

[0100] Figure 6 is the optimal control input curve diagram of the intelligent agent of the present invention,

[0101] Among them, (a) is the optimal control input curve diagram of intelligent agent 1, (b) is the optimal control input curve diagram of intelligent agent 2, (c) is the optimal control input curve diagram of intelligent agent 3, (d) is the optimal control input curve diagram of intelligent agent 4, and (e) is the optimal control input curve diagram of intelligent agent 5.

[0102] Figure 7 This is a graph showing the weight changes of the intelligent agent single evaluation neural network of the present invention.

[0103] Figure 8 A curve diagram of the triggering interval of the intelligent agent dynamic event triggering of the present invention. DETAILED DESCRIPTION

[0104] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with the implementation methods and drawings.

[0105] Figure 1 is a flow chart of the control method of the present invention, as shown in Figure 1 As shown, the present invention discloses an optimization control method for a human-in-the-loop multi-agent system under FDI attack, which specifically includes the following steps:

[0106] Step 1: Establish a dynamic model of a non-fully autonomous multi-agent system. The specific process is as follows:

[0107] Leaders:

[0108] Followers:

[0109] Among them, the subscript i is the number of a single follower agent system, represents the system state of agent i, represents the control input of agent i, represents the internal dynamics of agent i, Represents the output dynamics of agent i. Represents the system state of the leader, represents an unknown and non-zero command signal sent by an external human expert to the leader, Represents the leader's internal dynamics, Represents the leader’s output dynamics;

[0110] Step 2: Considering the FDI attack on the follower controller to the actuator end, a dynamic model of the intelligent agent system attacked by FDI is constructed. The specific process is as follows:

[0111] For model building, it is assumed that the attacker has the following capabilities:

[0112] (1) All executive agencies may be attacked by attackers at any time.

[0113] (2) State variables can be monitored in real time and provided as information to attackers,

[0114] (4) The attacker's false data and control signals can be injected into the actuator synchronously.

[0115] (4) The attacker’s energy is limited.

[0116] The control input of agent i under FDI attack becomes:

[0117]

[0118] Among them, χ i Input for FDI attack;

[0119] Then the dynamic model of the i-th agent attacked by FDI can be expressed as

[0120]

[0121] in, is the state of the system being attacked;

[0122] Step 3: Based on the attacked system model obtained in Step 2, design a performance function and find the Nash equilibrium point for each agent, that is, the Nash optimal strategy of the attacker and defender. The specific process is as follows:

[0123] Define the bipartite consistency error of the i-th agent as:

[0124]

[0125] Among them, a ij is the connection weight between agent i and agent j, if a ij >0 means that agent i and agent j are in a cooperative relationship. If a ij <0 means that agent i and agent j are in a competitive relationship; b i is the connection weight between agent i and the leader, if b i >0 means that there is a cooperative relationship between agent i and the leader, otherwise it is a competitive relationship; represents the degree matrix and

[0126] Then the dynamic equation of the bipartite consistency error of the i-th agent is:

[0127]

[0128] in,

[0129] According to game theory, the expected performance of control input increases while the expected performance of external FDI attack decreases. There is an antagonistic relationship between the two. The point where the two reach a balance of performance is the Nash equilibrium point, and the performance at the Nash equilibrium point is the Nash optimal.

[0130] Control input u i Not only must consistent final bounded stability be guaranteed, but also L2 gain must be satisfied, that is, the attack energy is limited, specifically expressed as:

[0131]

[0132] Among them, γ i >0 indicates the desired attack attenuation level;

[0133] Therefore, the design performance function is as follows:

[0134]

[0135] Among them, Q ii 、R ij 、R ii 、P ii and P ij are all symmetric positive definite matrices with matching dimensions, represents the neighboring agent of the i-th agent;

[0136] According to Pontryagin's maximum-minimum principle, the Nash equilibrium point The following inequality is satisfied:

[0137]

[0138] The Hamiltonian function is constructed as follows:

[0139]

[0140] in, represents the partial derivative of the performance function with respect to the bipartite consensus error;

[0141] Then the corresponding HJI equation can be expressed as:

[0142] According to the adaptive dynamic programming theory, the specific form of the Nash equilibrium point is:

[0143]

[0144] According to the expression of Nash equilibrium point, the HJI equation can be specifically expressed as:

[0145]

[0146] in,

[0147] Step 4: Based on the design of the Nash equilibrium point in step 3, design the HJI equation and optimal Nash strategy based on dynamic event triggering. The specific process is as follows:

[0148] definition is the kth triggering moment and satisfies The current state With sampling status The error between can be expressed as:

[0149]

[0150] Furthermore, the bipartite consistency error triggered by dynamic events has the following form:

[0151]

[0152] The dynamic event triggering error is expressed as:

[0153]

[0154] Considering that FDI attacks are external outputs and uncontrollable, the specific forms of the optimal control strategy based on dynamic event triggering and the optimal FDI attack strategy based on time triggering are:

[0155]

[0156] Combining the above results with the HJI equation in step 3, the specific form of the HJI equation based on dynamic event triggering is obtained as follows:

[0157]

[0158] Step 5: The specific process of designing a dynamic event triggering mechanism is as follows:

[0159] According to the HJI equation based on dynamic event triggering in step 4, the dynamic event triggering mechanism is designed as follows:

[0160]

[0161] Among them, T DED is the dynamic trigger threshold, and

[0162]

[0163] Design dynamic parameters The change law of is as follows:

[0164]

[0165] in,

[0166]

[0167] Step 6: Design a single-evaluation network model based on a neural network to approximate the performance function in step 3, and obtain an approximate solution to the Nash equilibrium point triggered by dynamic events. The specific process is as follows:

[0168] The form of constructing a single evaluation network based on a neural network is:

[0169]

[0170] in, is the ideal weight vector, h ci is the number of neurons in the hidden layer, S c (∈ i ) is a differentiable continuous activation function, o(∈ i ) is the approximation error;

[0171] The performance function gradient is expressed as:

[0172]

[0173] in, is the gradient of the activation function, is the gradient value of the approximation error;

[0174] Considering that the ideal weight vector is an unknown quantity and cannot be directly obtained, the estimated form of the approximate value of the gradient value of the performance function is constructed as follows:

[0175]

[0176] in, W ci The estimated form of .

[0177] Then the estimated value of the optimal control strategy based on dynamic event triggering and the approximate value of the optimal FDI attack strategy based on time triggering in step 4 is:

[0178]

[0179] Combining the above results with the HJI equation based on dynamic event triggering in step 4, the estimated form of its approximate value is obtained as:

[0180]

[0181] in,

[0182]

[0183] Step 7: Based on the Bellman residual generated in the approximation process in step 6, use the gradient descent method to train the weight change law of the single evaluation network generated in step 6 until convergence. The specific process is:

[0184] The Bellman residual generated during the approximation process in step 6 is Define it as:

[0185]

[0186] The design optimization objective function is:

[0187] Based on this function, using the gradient descent method, the specific form of the change law of the weight of the single evaluation network in step 6 can be obtained as follows:

[0188]

[0189] Where 0<μ i <1 is the design parameter;

[0190] Substitute the obtained weight value into the expression form of control input,

[0191]

[0192] The control input of the system can be obtained.

[0193] The system of the present invention is also stable under event triggering conditions. The specific reasons are:

[0194] Define the approximation error of the neural network weights as Then the dynamic equation of the error system is:

[0195]

[0196] in, Used to ensure continuous motivation conditions for the neural network learning process;

[0197] Considering that the optimal control input is sampled based on the dynamic event trigger mechanism, the continuous closed-loop system becomes a jump system; define the augmented variable The specific form of the dynamic equation of augmented variables is:

[0198]

[0199] in,

[0200] It can be seen from the above formula that even when the system is at the jump time point, the augmented variable of the system tends to a fixed value, that is, when the control input is updated, the jump point will not make the system unstable.

[0201] Example 1

[0202] A multi-agent system consisting of five followers and one leader is used. Assume that the input command of the external human expert is in the form of:

[0203]

[0204] The schematic diagram of the communication topology is as follows Figure 2 As shown in Figure 2, agents are divided into cooperative and competitive relationships. Red and black represent two groups of agents, respectively. Agents of the same color are in a cooperative relationship, while agents of different colors are in a competitive relationship.

[0205] Figure 3 This is a graph of the agent state consistency. As can be seen from the figure, Agents 1 and 2 can achieve consistency with the leader's state. The remaining agents' states are consistent in magnitude but opposite in sign, achieving bipartite consistency. Furthermore, external human experts can send commands to the leader, allowing it to adjust its trajectory in a timely manner. This allows the multi-agent system to avoid unexpected obstacles, improving system safety and reliability.

[0206] Figure 4 This is the bipartite consistency error curve of the agent. It can be seen that the bipartite consistency error can converge to a very small area near zero, thus ensuring the realization of bipartite consistency.

[0207] Figure 5 and Figure 6 (a)-(e) are the FDI attack curves and optimal control input curves of the intelligent agent, respectively. It can be seen that the designed optimal attack strategy is bounded and convergent, and conforms to the characteristics of time triggering; the optimal control input is in sampling form, which conforms to the characteristics of dynamic event triggering, effectively reduces the number of triggers, and saves communication resources.

[0208] Figure 7 The graph shows the weight changes of the single-evaluation neural network of the agent, which shows the boundedness and convergence of the gain. This shows that the convergence of the game optimization algorithm based on neural network approximation in this invention can be guaranteed, which is also a necessary condition for the Nash equilibrium point to converge to the optimal one.

[0209] Figure 8This graph shows the trigger intervals for the agent's dynamic events. As can be seen, the dynamic event triggering mechanism designed in this invention can effectively reduce the number of triggers, save communication resources, and effectively avoid the Zeno phenomenon. This control scheme remains effective and superior even under FDI attacks.

[0210] The above description is only a specific embodiment of the present invention. Any feature disclosed in this specification, unless otherwise stated, can be replaced by other equivalent or alternative features with similar purposes; all disclosed features, or all steps in the methods or processes, except for mutually exclusive features and / or steps, can be combined in any way.

Claims

1. An optimization control method for a human-in-the-loop multi-agent system under FDI attack, characterized in that: The following steps are involved: Step 1: Construct a dynamic model of a non-fully autonomous multi-agent system, including the dynamic equations of the leader and followers under external control, where the external human expert command signal is unknown; The non-fully autonomous multi-agent system includes a leader and N followers. Specifically, the multi-agent system is: Leaders: Followers: in, represents the derivative, the subscript i is the number of a single agent system, i = 1, 2, ..., N, represents the system state of agent i, represents the control input of agent i, represents the internal dynamics of agent i, represents the output dynamics of agent i; Represents the system state of the leader, represents an unknown and non-zero command signal sent by an external human expert to the leader, Represents the leader's internal dynamics, Represents the output dynamics of the leader; Step 2: Based on the dynamic model of the multi-agent system constructed in step 1, consider the FDI attack on the follower controller to the actuator end, and construct the dynamic model of the agent system attacked by FDI. The specific process is as follows: Considering the FDI attack on the channel from the controller of agent i to the actuator, the control input of agent i that suffers from FDI attack is, Among them, u i for u i The abbreviation of (t), χ i Input for FDI attack; The dynamic model of the i-th agent attacked by FDI can be expressed as: in, is the state of the system being attacked; Step 3: Based on the dynamic model of the agent system attacked by FDI obtained in Step 2, design a performance function and find the Nash equilibrium point for each agent, that is, the Nash optimal strategy of the attacker and defender. The specific process is as follows: The dynamic equation defining the bipartite consistency error of the i-th agent is: in, a ij is the connection weight between agent i and agent j, if a ij >0 means that agent i and agent j are in a cooperative relationship. If a ij <0 means that agent i and agent j are in a competitive relationship; b i is the connection weight between agent i and the leader, if b i >0 means that there is a cooperative relationship between agent i and the leader, otherwise it is a competitive relationship; represents the degree matrix and According to game theory, the expected performance of control input is improved while the expected performance of external FDI attack is reduced. There is an antagonistic relationship between the two. When the performance of the two reaches the optimal equilibrium, a Nash equilibrium point can be formed. Therefore, the design performance function is as follows: Among them, γ i represents the desired attack attenuation level, γ i >0, Q ii 、R ij 、R ii 、P ii and P ij are all symmetric positive definite matrices with matching dimensions, represents the neighboring agent of the i-th agent; According to Pontryagin's maximum-minimum principle, the Nash equilibrium point The following inequality is satisfied: The Hamiltonian function is constructed as follows: in, represents the partial derivative of the performance function with respect to the bipartite consensus error; Then the corresponding HJI (Hamilton-Jacobi-Isaacs) equation can be expressed as: According to the adaptive dynamic programming theory, the specific form of the Nash equilibrium point is: Step 4: Based on the Nash equilibrium point in step 3, design the HJI equation and Nash optimal strategy based on dynamic event triggering; Step 5: Based on the HJI equation based on dynamic event triggering obtained in step 4, a dynamic event triggering mechanism is designed to save network resource overhead; Step 6: Design a single-evaluation network model based on a neural network to approximate the performance function in step 3, and obtain an approximate solution to the Nash equilibrium point triggered by dynamic events; Step 7: Based on the Bellman residual generated in the approximation process in step 6, use the gradient descent method to generate the weight change law of the single evaluation network in step 6 until convergence, and the control input can be obtained based on the weight result.

2. The optimization control method for a human-in-the-loop multi-agent system under FDI attack according to claim 1, characterized in that: The attacker must meet the following conditions: (1) All executive agencies may be attacked by attackers at any time. (2) State variables can be monitored in real time and provided as information to attackers, (3) False data and control signals can be injected into the actuator synchronously, (4) The attacker's energy is limited.

3. The optimization control method for a human-in-the-loop multi-agent system under FDI attack according to claim 1, characterized in that: The specific process of step 4 is: definition is the kth triggering moment, and satisfies The current state With sampling status The error between can be expressed as: The bipartite consistency error based on dynamic event triggering has the following form: The dynamic event triggering error is expressed as: Considering that FDI attacks are external outputs and uncontrollable, the specific forms of the optimal control strategy based on dynamic event triggering and the optimal FDI attack strategy based on time triggering are: Combining the above results with the HJI equation, the specific form of the HJI equation based on dynamic event triggering is obtained as follows:

4. The optimization control method for a human-in-the-loop multi-agent system under FDI attack according to claim 3, characterized in that: The specific process of step 5 is: according to the HJI equation based on dynamic event triggering in step 4, the dynamic event triggering mechanism is designed as follows: Among them, T DED is the dynamic trigger threshold, and λ min (Q ii ) represents the matrix Q ii The smallest eigenvalue, ν i,1 is the design parameter, 0<ν i,1 <1, Design dynamic parameter θ i The change law of is as follows: in, 5. The optimization control method for a human-in-the-loop multi-agent system under FDI attack according to claim 4, characterized in that: The specific process of step 6 is: A single evaluation network model based on a neural network is designed to approximate the performance function in step 3. The specific expression of the single evaluation network model based on a neural network is: in, is the ideal weight vector, h ci is the number of neurons in the hidden layer, S c (∈ i ) is a differentiable continuous activation function, o(∈ i ) is the approximation error; The performance function gradient is expressed as: in, is the gradient of the activation function, is the gradient value of the approximation error; Considering that the ideal weight vector is an unknown quantity and cannot be directly obtained, the estimated form of the approximate value of the gradient value of the performance function is constructed as follows: in, W ci The estimated form of .

6. The optimization control method for a human-in-the-loop multi-agent system under FDI attack according to claim 5, characterized in that: The estimated form of the approximate value of the optimal control strategy based on dynamic event triggering and the optimal FDI attack strategy based on time triggering in step 4 is: Combining the above results with the HJI equation based on dynamic event triggering in step 4, the estimated form of its approximate value is: in, 7. The optimization control method for a human-in-the-loop multi-agent system under FDI attack according to claim 6, characterized in that: The specific process of step 7 is: The Bellman residual generated in the approximation process in step 6 is The Bellman residual is defined as: The design optimization objective function is: Based on this function, using the gradient descent method, the specific form of the change law of the weight of the single evaluation network in step 6 can be obtained as follows: Where 0<μ i <1 is the design parameter; Substitute the obtained weight value into the expression form of control input, The control input of the system can be obtained.

Citation Information

Patent Citations

  • Design method of event-triggered memory-equipped DOF controller under random FDI attack

    CN110865616A

  • Multi-missile consensus self-learning differential game cooperative guidance method based on event triggering

    CN118550323A