A model-free control method for human-in-the-loop multi-agent systems under denial-of-service attacks
By constructing a leader system and a specified time performance observer, combined with Q learning and neural networks, the control problem of multi-agent systems under DoS attacks is solved, optimal control and error convergence are achieved, and the security and reliability of the system are improved.
Patent Information
- Application Number
- CN202411845811.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-12-16
AI Technical Summary
Existing human-in-the-loop multi-agent systems find it difficult to strike a balance between control performance and control cost under denial-of-service attacks. In addition, the system model is complex and changeable, making it difficult to obtain accurately, which affects the control effect.
A non-fully autonomous leader system is constructed, and a leader output observer with fully distributed specified time performance is designed. Model-free optimal control is implemented through the Q-learning algorithm, combined with a neural network model to approximate the Q function to respond to DoS attacks in real time.
Model-free optimal control of the multi-agent system is achieved under DoS attacks, which improves the security and reliability of the system and ensures the control quality and the convergence performance of the consistency error within the specified time.
Smart Images

Figure CN119717614B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-agent system control, and in particular relates to a model-free control method for a human-in-the-loop multi-agent system under a denial of service attack. Background Art
[0002] With the rapid development of technologies such as artificial intelligence, the Internet of Things, and big data, multi-agent systems have been widely applied in fields such as automation, intelligent transportation, robotic collaboration, smart grids, and financial market analysis. Exploring these systems holds significant strategic significance and positive economic prospects. Because the safety and reliability of fully autonomous systems cannot be fully guaranteed, the control of human-in-the-loop multi-agent systems is crucial.
[0003] In a network communication topology composed of multiple agents, malicious network attacks such as denial of service attacks (DoS attacks) can have a huge impact on the information transmission between agents. At the same time, in the existing human-in-the-loop multi-agent control results [1][2], few consider the trade-off between control performance and control cost under DoS attacks. The cost of use is high, which is not conducive to engineering implementation. In addition, under the trend of high informatization, the system model is complex and changeable, and it is difficult to obtain accurately. Therefore, it is of practical significance to design a model-free control method for the communication channel of a human-in-the-loop multi-agent system under DoS attacks.
[0004] [1] Liu Lu, Liu Xinwei, Zhou Jiaxing, et al. A human-in-the-loop unmanned ship platoon collaborative control system and its method [P]. Liaoning Province: CN202410193921.8, 2024-06-04.
[0005] [2]Xu Yong, Sun Jian, Mei Di, et al. Human-in-the-loop reinforcement learning collaborative tracking control method for unmanned swarm systems[P]. Beijing: CN202410031096.1, 2024-05-24. Summary of the Invention
[0006] In response to the problems existing in the background technology, the purpose of the present invention is to provide a model-free control method for a human-in-the-loop multi-agent system under a denial of service attack. This control method constructs a non-fully autonomous leader system, and based on the time-varying switching topology caused by the DoS attack, designs a leader output observer with fully distributed specified time performance, and constructs an augmented system containing the follower system dynamics and the observer dynamics based on this observer; on the basis of this system, an optimization performance function is designed and a model-free Q-learning algorithm is derived based on this function. The control method of the present invention realizes model-free optimal control of a human-in-the-loop multi-agent system under the influence of a DoS attack, while taking into account the specified time convergence performance of the consistency error, thereby improving the control quality.
[0007] To achieve the above object, the technical solution of the present invention is as follows:
[0008] A model-free control method for a human-in-the-loop multi-agent system under a denial-of-service attack comprises the following steps:
[0009] Step 1: Construct a dynamic model of a non-fully autonomous multi-agent system, including the dynamic equations of the followers and the dynamic equations of the externally controlled leader, where the external human expert command signal is an unknown quantity;
[0010] Step 2: Based on the system dynamic equations constructed in step 1, consider the communication between intelligent agents suffering from DoS attacks and construct a switching topology model;
[0011] Step 3: Based on the switching topology model obtained in step 2, design a leader output observer with fully distributed specified time performance for each follower;
[0012] Step 4: Based on the follower system equation in step 1 and the leader output observer in step 3, construct the local output error and obtain the local output error dynamic system, namely the augmented state dynamic model;
[0013] Step 5: Based on the augmented state dynamic model obtained in step 4, design the optimal value function and Hamiltonian function, and obtain the Q function on this basis;
[0014] Step 6: Construct a single-evaluation network model based on a neural network to approximate the Q function in step 5. The state and input values of the augmented dynamic model obtained in step 4 are collected in real time to form a data set. Based on the collected data set, the weights of the single-evaluation network model are trained until convergence.
[0015] Step 7: Based on the weights of the single-evaluation neural network obtained in step 6, the optimal controller is obtained according to the gradient descent method, and the multi-agent system is controlled based on the optimal controller.
[0016] Furthermore, the non-fully autonomous multi-agent system in step 1 includes 1 leader and N followers. The dynamic model of the multi-agent system is specifically:
[0017] Leaders:
[0018] Followers:
[0019] in, Indicates derivation, the subscript i is the i-th follower agent, represents the system state of agent i, represents the control input of agent i, represents the output of agent i, represents the internal dynamics of agent i, represents the output dynamics of agent i, represents the output matrix of agent i; Represents the system state of the leader, represents an unknown and non-zero command signal sent by an external human expert to the leader, Represents the output of the leader and can only be obtained by some followers, Represents the leader's internal dynamics, Represents the output dynamics of the leader, Represents the output matrix of the leader.
[0020] Furthermore, the specific process of constructing the switching topology model in step 2 is as follows:
[0021] The communication topology containing N followers is composed of switching time-varying topologies express, Among them, the subscript is a piecewise continuous function caused by DoS attack, m represents the topology switching caused by the mth attack, represents the vertex set of N followers, The edge set representing N followers; is the weight matrix, represent is the neighbor set of the ith agent, definition To switch topology The Laplace matrix of is the degree matrix, definition Indicates that the i-th agent can obtain the leader's information.
[0022] Furthermore, the specific process of step 3 is:
[0023] The observation consistency error of the i-th agent is defined as:
[0024] j is the jth agent;
[0025] Design a fully distributed specified time performance leader output observer as follows:
[0026]
[0027] Among them, κ i is the output of the observer, ν i is the input signal of the observer Specify a time function for the time-varying , and T is the expected time, τ0 is the time constant; ρ is the observer gain, and ρ>0; is the time interval; in each time interval
[0028] picture It is a fixed topology with a directed spanning tree, the root node receives the output signal of the external human expert; if t∈[t l ,t l+1 ),but and is the switching moment of the topology in the time interval (0,T], and has In addition, the attack duration must satisfy t l+1 -t l >τ2>0, where τ2 is the constant dwell time.
[0029] Furthermore, the time-varying specified time function is preferably w is a constant.
[0030] Furthermore, the specific process of step 4 is:
[0031] According to the follower system equation in step 1 and the observer system equation in step 3, construct the local output error z i (t), is:
[0032]
[0033] Among them, the augmented output vector is The augmented state is X i =[x i κ i ] T .
[0034] Then the dynamic model of the local output error dynamic system, that is, the augmented system, is described as:
[0035]
[0036] in represents the zero vector of p rows and 1 column, represents a zero matrix with p rows and m columns, represents a zero matrix with n rows and p columns, Represents the identity matrix with p rows and p columns.
[0037] Furthermore, the specific process of step 5 is:
[0038] According to the adaptive dynamic programming theory, for the dynamic model of the augmented system in step 4, the optimal value function is designed as:
[0039]
[0040] Among them, the superscript * indicates the optimal, k i is the discount factor, k i >0, Q i and R i is a symmetric positive definite matrix with matching dimensions;
[0041] Based on the optimal value function, the corresponding Hamiltonian function is:
[0042]
[0043] in represents the partial derivative of the optimal value function with respect to the augmented state;
[0044] Combining the optimal value function and the Hamiltonian function, according to reinforcement learning theory, the expression of the Q function can be established as:
[0045]
[0046] Based on the established Q function, the analytical form of the optimal controller that does not depend on the model is obtained as:
[0047]
[0048] Furthermore, the specific process of step 6 is:
[0049] Design a single-evaluation network model based on a neural network to approximate the Q function in step 5:
[0050]
[0051] in, is the ideal weight vector, k ci is the number of neurons in the hidden layer, ψ(X i ,u i ) is a differentiable continuous activation function, σ μ (X i ,u i ) is the approximation error, μ represents the μth iteration;
[0052] Considering that the ideal weight vector is unknown and cannot be directly obtained, the estimated form of the approximate value of the Q function is constructed as follows:
[0053]
[0054] in, for The estimated form of .
[0055] The real-time state of the augmented model obtained in step 4 and the input value are collected in real time to form a data set. The data set size of agent i is According to the least squares method, the specific form of the weight of the single evaluation network can be obtained as follows:
[0056]
[0057] in,
[0058]
[0059] Furthermore, the specific process of step 7 is:
[0060] According to the gradient descent method, the estimated form of the approximate value of the optimal controller that does not depend on the model is:
[0061]
[0062] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0063] The proposed method utilizes trajectory adjustment command signals sent by external human experts to the leader to respond to emergencies in real time, improving system safety and reliability. Furthermore, it models continuous DoS attacks as a time-varying switching topology model and designs a fully distributed time-specific leader output observer based on this time-varying topology model. This effectively mitigates the adverse effects of DoS attack jumps on observation performance and achieves time-specific convergence of observation errors. More importantly, considering that precise system models cannot be established and obtained in actual industrial processes, an optimization algorithm based on model-free reinforcement learning was developed. This ensures optimal performance under completely unknown system dynamics, effectively improving control quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 Flowchart of the control method of the present invention.
[0065] Figure 2 Schematic diagram of topology changes caused by a DoS attack sequence.
[0066] Figure 3 FIG. 4 is a graph showing the observer consistency error of the intelligent agent of the present invention.
[0067] Figure 4 Graph showing the observation error of the intelligent agent of the present invention.
[0068] Figure 5 This is the observer state and leader output curve diagram of the present invention.
[0069] Figure 6Graph showing the neural network weight norm of the agent of the present invention.
[0070] Figure 7 Graph showing the output of the agent and the leader of the present invention.
[0071] Figure 8 Graph showing the optimal controller for the intelligent agent of the present invention.
[0072] Figure 9 It is a cost function curve diagram of the control method of the present invention. DETAILED DESCRIPTION
[0073] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with the implementation methods and drawings.
[0074] Figure 1 is a flow chart of the control method of the present invention, as shown in Figure 1 As shown, the present invention discloses a model-free control method for a human-in-the-loop multi-agent system under a denial-of-service attack, which specifically includes the following steps:
[0075] Step 1: Establish a dynamic model of a non-fully autonomous multi-agent system, which includes a leader and N followers. The specific process is as follows:
[0076] Leaders:
[0077] Followers:
[0078] Among them, the corner mark is the number of the single follower agent system, represents the system state of agent i, represents the control input of agent i, represents the output of agent i, represents the internal dynamics of agent i, represents the output dynamics of agent i, Represents the output matrix of agent i. Represents the system state of the leader, represents an unknown and non-zero command signal sent by an external human expert to the leader, Represents the output of the leader and can only be obtained by some followers, Represents the leader's internal dynamics, Represents the output dynamics of the leader, Represents the output matrix of the leader.
[0079] Step 2: Based on the system dynamic equation constructed in step 1, consider the communication between intelligent agents suffering from DoS attack and build a switching topology model. The specific process is as follows:
[0080] In the process of directed communication between agents, the communication channel may be subject to continuous DoS attacks. When such DoS attacks occur, the communication between agents may be interrupted. Therefore, the communication topology containing N followers is changed from a switching time-varying topology to a Indicates that is a piecewise continuous function caused by DoS attack, represents a vertex set, represents the edge set of N followers; Defined as a weight matrix, in the consistency control of a multi-agent system, if there is a connection between agents, the weight is 1, and if there is no connection, the weight is 0, where represent The neighbor set of the i-th agent is defined as definition To switch topology The Laplace matrix of is a degree matrix and definition and Indicates that the i-th agent can obtain the leader's information.
[0081] Step 3: Based on the switching topology model obtained in step 2, design a leader output observer with fully distributed specified time performance for each follower. The specific process is as follows:
[0082] Based on the time-varying switching topology established in step 2, it should be noted that in each time interval Where t0=0, It is a fixed topology with a directed spanning tree, the root node receives the output signal of the external human expert; if t∈[t l ,t l+1 ),but and is the switching moment of the topology in the time interval (0, T], and has In addition, the attack duration must satisfy t l+1 -t l >τ2>0, where τ2 is the constant dwell time.
[0083] The specific form of establishing a time-varying specified time function is:
[0084]
[0085] The observation consistency error of the i-th agent is defined as:
[0086]
[0087] Design a fully distributed specified time performance leader output observer as follows:
[0088]
[0089] in is the output of the observer, and the input signal of the observer Expressed as:
[0090]
[0091] in, is the time interval;
[0092] Step 4: Based on the follower system equation in step 1 and the observer system equation in step 3, construct the local output error, which is in the form of:
[0093] Local output error
[0094] The augmented output vector is The augmented state is X i =[x i κ i ] T .
[0095] Then the dynamics of the local output error system, that is, the augmented system, is described as:
[0096]
[0097] in
[0098] Step 5: Based on the adaptive dynamic programming theory, design the value function for the augmented system in step 4. Its specific form is:
[0099]
[0100] Wherein, the discount factor k i >0, Q i and R i is a symmetric positive definite matrix with matching dimensions. Therefore, the specific form of the corresponding optimal value function is:
[0101]
[0102] Based on the optimal value function, the corresponding Hamiltonian function is:
[0103]
[0104] in Represents the partial derivative of the optimal value function with respect to the augmented state.
[0105] Combining the optimal value function and the Hamiltonian function, according to the reinforcement learning theory, the specific expression of the Q function can be established as:
[0106]
[0107] in,
[0108] Based on the established Q function, the analytical form of the optimal controller that does not depend on the model is obtained as:
[0109]
[0110] Specifically, the Q learning algorithm can be expressed as follows:
[0111] Given an initial allowed control protocol Start iteration from the 0th time, that is, μ = 0, μ represents the μth iteration; solve according to the following expression
[0112]
[0113] Furthermore, the optimal controller is updated as follows:
[0114]
[0115] Therefore, the strategy for checking whether the optimal controller iteration process converges is: if the difference between the Q functions of two iterations is less than a small constant threshold set in advance The iteration stops; otherwise, it continues to the next round of iteration. Its expression is:
[0116]
[0117] Step 6: Construct a single evaluation network model based on a neural network, collect the real-time status and input values of the augmented model obtained in step 4 in real time and form a data set. The data set size of agent i is Based on the single evaluation network model, the Q function in step 5 is approximated as follows:
[0118]
[0119] in, is the ideal weight vector, k ci is the number of neurons in the hidden layer, ψ(X i ,u i ) is a differentiable continuous activation function, σμ (X i ,u i ) is the approximation error.
[0120] Considering that the ideal weight vector is unknown and cannot be directly obtained, the estimated form of the approximate value of the Q function is constructed as follows:
[0121]
[0122] in, for The estimated form of .
[0123] The Bellman residual generated during the neural network approximation process can be expressed as follows:
[0124]
[0125] In order to express the Bellman residual conveniently, the following definition is given:
[0126]
[0127] Therefore, the Bellman residual can be simplified as follows:
[0128]
[0129] Taking minimizing the Bellman residual as the optimization goal, according to the least squares method, the specific form of the weight of the single evaluation network in step 6 can be obtained as:
[0130]
[0131] in,
[0132]
[0133] Specifically, the neural network-based model-free Q-learning algorithm is expressed as follows:
[0134] Collect the real-time status and input value of the augmented model of each agent in step 4 in real time and form a data set in is the dataset size of agent i; calculate The value of; Given the initial evaluation neural network weight vector Start iteration from the 0th time, that is, μ=0;
[0135] Solve the neural network weights according to the following expression:
[0136]
[0137] Therefore, the strategy for testing whether the weight iteration process of the optimal controller based on neural network approximation converges is: if the difference between the weights of two iterations is less than a small constant threshold set in advance The iteration stops and the corresponding Q function value and the optimal controller value are calculated according to this weight; otherwise, the next round of iteration is continued; its expression is:
[0138]
[0139] Step 7: Based on the weights of the single evaluation network neural network trained in step 6, the approximate value of the optimal controller that does not depend on the model is obtained according to the gradient descent method:
[0140]
[0141] Example 1
[0142] A multi-agent system consisting of five followers and one leader is used, and the specific form of the input command from the external human expert is assumed to be:
[0143]
[0144] The conditions for switching to DoS attack mode are as follows:
[0145]
[0146] The schematic diagram of the topology changes caused by this DoS attack time series is as follows: Figure 2 As shown in the figure, 0 represents the leader, when When represents no attack.
[0147] Figure 3 and Figure 4 The observer consistency error curve of the agent and the observation error curve of the agent are plotted respectively. Figure 4 middle, This is the difference between the follower's own output and the observed leader's output. As can be seen from the figure, within the specified time of 1 second, both the consistency error and the observation error converge to a very small range near zero, ensuring that consistent observation can be achieved even under DoS attacks, achieving the specified time convergence performance.
[0148] Figure 5 The observer state and leader output curves are plotted. It can be seen that the designed observer can accurately estimate the leader's output while eliminating the adverse effects of topology switching caused by DoS attacks on the observation results of each agent, and can effectively observe the leader's output in the form of bounded gain.
[0149] Figure 6 The norm of the agent's neural network weights is plotted, demonstrating the boundedness and convergence of the gains. This demonstrates that the convergence of the Q-function algorithm based on neural network approximation in this invention can be guaranteed, which is also a necessary condition for the model-free optimal controller to converge to the optimal state.
[0150] Figure 7 The output of the agents and the output of the leader are plotted. It can be seen that the present invention can achieve output synchronization. In addition, the external human expert can send instructions to the leader to adjust the leader's trajectory in a timely manner, so that the multi-agent system can avoid sudden obstacles and improve the safety and reliability of the system. In addition, it can be concluded that if there is no external human expert guidance (y0 h (without HiTL)), the system security will be reduced to a certain extent.
[0151] Figure 8 The optimal controller for the agent is plotted. It can be seen that the optimal control strategy is bounded and converges. Notably, at t = 0.5s and t = 3s, the DoS pattern changes simultaneously with the commands of the external human expert, but the designed control scheme still works well, demonstrating its effectiveness and superiority under sustained DoS attacks.
[0152] Figure 9 The convergence curve of the cost function is plotted. It can be seen that under the proposed control method, the cost function can quickly converge to a small neighborhood near 0, indicating that the proposed optimization control scheme can achieve a good balance between performance and cost, improving control quality.
[0153] The above description is only a specific embodiment of the present invention. Any feature disclosed in this specification, unless otherwise stated, can be replaced by other equivalent or alternative features with similar purposes; all disclosed features, or all steps in the methods or processes, except for mutually exclusive features and / or steps, can be combined in any way.
Claims
1. A model-free control method for a human-in-the-loop multi-agent system under denial-of-service attacks, characterized in that: The following steps are involved: Step 1: Construct a dynamic model of a non-fully autonomous multi-agent system, including the dynamic equations of the followers and the dynamic equations of the leader under external control; Among them, the non-fully autonomous multi-agent system includes 1 leader and N followers, so the dynamic model of the multi-agent system is specifically: Leaders: Followers: in, Indicates derivation, the subscript i is the i-th follower agent, Represents the system state of agent i, written as x i ; Represents the control input of agent i, written as u i ; Represents the output of agent i, written as y i ; represents the internal dynamics of agent i, represents the output dynamics of agent i, represents the output matrix of agent i; Represents the system state of the leader, represents an unknown and non-zero command signal sent by an external human expert to the leader, Represents the output of the leader and can only be obtained by some followers, Represents the leader's internal dynamics, Represents the output dynamics of the leader, represents the output matrix of the leader; the external human expert command signal is an unknown quantity; Step 2: Based on the system dynamic equations constructed in step 1, consider the communication between intelligent agents suffering from DoS attacks and construct a switching topology model; Step 3: Based on the switching topology model obtained in step 2, design a leader output observer with fully distributed specified time performance for each follower; Step 4: Based on the follower's dynamic equation in step 1 and the leader's output observer in step 3, construct the local output error and obtain the local output error dynamic system, namely the augmented state dynamic model; Step 5: Based on the augmented state dynamic model obtained in step 4, design the optimal value function and Hamiltonian function, and obtain the Q function on this basis; The specific process is: According to the adaptive dynamic programming theory, for the dynamic model of the augmented system in step 4, the optimal value function is designed as: Among them, the superscript * indicates the optimal, k i is the discount factor, k i >0, Q i and is a symmetric positive definite matrix with matching dimensions; X i It is an augmented state; Based on the optimal value function, the corresponding Hamiltonian function is: in represents the partial derivative of the optimal value function with respect to the augmented state; represents the zero vector of p rows and 1 column, represents a zero matrix with p rows and m columns, represents a zero matrix with n rows and p columns, represents the identity matrix with p rows and p columns; Combining the optimal value function and the Hamiltonian function, according to reinforcement learning theory, the expression of the Q function can be established as: Based on the established Q function, the analytical form of the optimal controller that does not depend on the model is obtained as: Step 6: Construct a single-evaluation network model based on a neural network to approximate the Q function in step 5. The state and input values of the augmented dynamic model obtained in step 4 are collected in real time to form a data set. Based on the collected data set, the weights of the single-evaluation network model are trained until convergence. The specific process is: Design a single-evaluation network model based on a neural network to approximate the Q function in step 5: in, is the ideal weight vector, k ci is the number of neurons in the hidden layer, ψ(X i ,u i ) is a differentiable continuous activation function, σ μ (X i ,u i ) is the approximation error, μ represents the μth iteration; Considering that the ideal weight vector is unknown and cannot be directly obtained, the estimated form of the approximate value of the Q function is constructed as follows: in, for The estimated form of Then, the real-time state of the augmented model obtained in step 4 and the input value are collected in real time to form a data set. The data set size of agent i is According to the least squares method, the specific form of the weight of the single evaluation network model can be obtained as follows: in, Step 7: Based on the weights of the single-evaluation neural network obtained in step 6, the optimal controller is obtained according to the gradient descent method, and the multi-agent system is controlled based on the optimal controller.
2. The model-free control method for a human-in-the-loop multi-agent system under a denial-of-service attack as claimed in claim 1, wherein: The specific process of building the switching topology model in step 2 is as follows: The communication topology containing N followers is composed of switching time-varying topologies express, Among them, the subscript is a piecewise continuous function caused by DoS attack, m represents the topology switching caused by the mth attack, represents the vertex set of N followers, The edge set representing N followers; is the weight matrix, represent is the neighbor set of the ith agent, definition To switch topology The Laplace matrix of is the degree matrix, definition and Indicates that the i-th agent can obtain the leader's information.
3. The model-free control method for a human-in-the-loop multi-agent system under a denial-of-service attack as claimed in claim 2, wherein: The specific process of step 3 is: The observation consistency error of the i-th agent is defined as: j is the jth agent; Design a fully distributed specified time performance leader output observer as follows: Among them, κ i is the output of the observer, ν i is the input signal of the observer Specify a time function for the time-varying , and T is the expected time, τ0 is the time constant; ρ is the observer gain, and ρ>0; is the time interval; in each time interval [t l ,t l+1 ),l=0,1,2,…,l,l+1,…,t0=0; picture It is a fixed topology with a directed spanning tree, the root node receives the output signal of the external human expert; if t∈[t l ,t l+1 ),but And t1,…,t l is the switching moment of the topology in the time interval (0,T], and has tt l >τ1>0 assumption.
4. The model-free control method for a human-in-the-loop multi-agent system under a denial-of-service attack as claimed in claim 3, wherein: The duration of the DOS attack must meet t l+1 -t l >τ2>0, where τ2 is the constant dwell time.
5. The model-free control method for a human-in-the-loop multi-agent system under DoS attack according to claim 3, wherein: The time-varying specified time function is w is a constant.
6. The model-free control method for a human-in-the-loop multi-agent system under a denial-of-service attack as claimed in claim 3, wherein: The specific process of step 4 is: According to the follower system equation in step 1 and the observer system equation in step 3, construct the local output error z i (t), is: Among them, the augmented output vector is The augmented state is X i =[x i κ i ] T . Then the dynamic model of the local output error dynamic system, that is, the augmented system, is described as:
7. The model-free control method for a human-in-the-loop multi-agent system under a denial-of-service attack as claimed in claim 6, wherein: The specific process of step 7 is: According to the gradient descent method, the estimated form of the approximate value of the optimal controller that does not depend on the model is:
Citation Information
Patent Citations
Design method of multi-agent system event trigger controller when DoS attack exists
CN109491249A
Multi-unmanned aerial vehicle event triggering formation control method with switching topology under DoS attack
CN118331327A