Optimal consistency control method for random multi-agent system with unknown control gain
By combining reinforcement learning and dynamic event triggering mechanisms with neural networks and observer design, the problem of unknown control gain in stochastic multi-agent systems is solved, achieving efficient and consistent control of the system, improving robustness and control performance, and reducing system energy consumption.
Patent Information
- Application Number
- CN202511176310.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-11-25
AI Technical Summary
In the presence of random disturbances and unknown control gain, traditional control strategies struggle to achieve effective and consistent control of multi-agent systems. Furthermore, existing methods exhibit poor robustness and slow convergence in noisy environments.
By adopting a dynamic event triggering mechanism based on reinforcement learning, combined with neural network and observer design, optimal consistent control of stochastic multi-agent systems is achieved through adaptive identifiers and dynamic threshold parameters. The control strategy is optimized by using critic and actor neural networks to reduce system energy consumption and communication burden.
Achieving state consistency in multi-agent systems under noisy and random perturbation environments improves system robustness and control performance, reduces energy consumption and communication resource consumption, and simplifies algorithm complexity.
Smart Images

Figure CN121008477A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of stochastic multi-agent system control, and more particularly to an optimal uniform control method for a second-order stochastic multi-agent system with unknown control gain. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] With the rapid development of modern automation technology and cyber-physical systems, multi-agent systems have been widely applied in fields such as UAV formation, distributed sensor networks, and intelligent manufacturing. Achieving state consistency among agents is one of the core issues in multi-agent cooperative control. In recent years, significant research has been conducted on consistency control strategies for deterministic multi-agent systems, especially with relatively mature controller design methods under the premise of known models and well-defined communication topologies. However, in practical applications, due to factors such as environmental noise, communication delays, and sensor errors, the dynamic behavior of agents is often affected by random disturbances. The behavior of such systems is more complex, making traditional deterministic control strategies difficult to apply directly.
[0004] Furthermore, in many real-world scenarios, the control input channels of an agent may contain uncertainties, such as actuator aging and changes in drive efficiency, leading to unknown control gains. Uncertainty in control gains can severely impact controller performance and even cause system instability. While some studies have employed adaptive estimation methods to approximate unknown control gains, these methods often suffer from slow convergence and poor robustness in the presence of random disturbances, necessitating the design of more adaptive and stable control strategies.
[0005] To improve system resource utilization efficiency and reduce unnecessary communication and control updates, event-triggered mechanisms have been introduced into the consensus control of multi-agent systems. This mechanism, by setting appropriate triggering conditions, ensures that agents only perform control updates or information exchanges when specific conditions are met, thereby effectively reducing system energy consumption and communication burden. Traditional static event-triggered mechanisms rely on fixed threshold functions, which can easily lead to excessively frequent or sparse triggering. To address this, dynamic event-triggered mechanisms proposed in recent years introduce auxiliary dynamic variables, making triggering conditions more flexible and adaptable. These mechanisms can automatically adjust the triggering frequency based on changes in system state, thereby further optimizing resource utilization while ensuring system performance.
[0006] In summary, this invention proposes a reinforcement learning-based consensus control strategy by combining a dynamic event triggering mechanism to achieve optimal consensus control of a stochastic multi-agent system with completely unknown control gain. Summary of the Invention
[0007] The purpose of this invention is to provide an optimal consensus control method based on reinforcement learning in order to achieve optimal consensus in a second-order stochastic multi-agent system with unknown control gain.
[0008] To achieve the above objectives, the technical solution of the present invention is as follows:
[0009] This invention is an optimal consensus control method for stochastic multi-agent systems with unknown control gain, comprising the following steps:
[0010] Step 1: Consider a class of... A second-order stochastic multi-agent system consisting of one follower and one leader is constructed to establish the consistency error between the states of the followers and the leader, and the dynamic system of the consistency error is derived.
[0011] No. The dynamic equations of a follower are modeled as the following second-order differential equations:
[0012] ,
[0013] in and Followers ( Position and velocity states, , , and They are all continuous and bounded nonlinear functions. The control gain function is unknown but bounded. Furthermore, It is an independent Wiener process.
[0014] The dynamic equation of the leader is:
[0015] ,
[0016] in and These are the leader's position and speed status, , , and They are all continuous and bounded nonlinear functions.
[0017] The consistency error between the follower state and the leader state is:
[0018] ,
[0019] in Followers The set of neighboring intelligent agents, if ,but ,otherwise If followers If the navigator's information can be received, then ,otherwise .remember .
[0020] make Then its dynamic equation can be derived as follows:
[0021]
[0022] in , , , .remember .
[0023] Step 2: Define the value function as a performance index for evaluating the control effect on a stochastic multi-agent system, and obtain the HJB equation based on the uniform error system and the value function.
[0024] The value function is designed as follows:
[0025] .
[0026] The HJB equation is expressed as:
[0027] ,
[0028] in It is the optimal control strategy. .
[0029] Step 3: Construct a dynamic event-triggered sampling error and threshold function. When the sampling error exceeds the trigger threshold, the event is triggered, and the controller is updated.
[0030] Event-triggered sampling error design is , among which when hour, , For followers The A trigger moment.
[0031] The trigger threshold is designed as follows:
[0032] ,
[0033] in It is a positive parameter, a positive constant. make satisfy Auxiliary dynamic variables The dynamic equations are designed as follows:
[0034] ,
[0035] in and It is a positive constant and satisfies ,also, initial value . It is a dynamic threshold parameter, and its dynamic equation is designed as follows:
[0036] ,
[0037] in It is a normal number.
[0038] Step 4: Design an adaptive identifier to approximate the random terms and unknown nonlinearities in the system.
[0039] The adaptive identifier is designed as follows:
[0040] ,
[0041] in The state of the identifier, This indicates the estimation error of the identifier. These are parameters that need to be designed. and These are the weight estimates and activation function of the labeler neural network, respectively. The update rules are designed as follows:
[0042] ,
[0043] in These are parameters that need to be designed.
[0044] Step 5: Design an observer to reconstruct the unknown control gain.
[0045] The observer is designed as follows:
[0046] ,
[0047] in It is the state of the observer. This represents the observer's estimation error. These are parameters that need to be designed. and These are the weight estimates and activation function of the labeler neural network, respectively. The update rules are designed as follows:
[0048] ,
[0049] in These are parameters that need to be designed.
[0050] Step Six: Construct a two-layer neural network model containing a critic neural network and an actor neural network, and design the weight update rule of the neural network based on the gradient descent algorithm. The critic neural network is used to approximate the value function of each follower to evaluate the control effect, and the output of the actor neural network is the control policy of the corresponding follower.
[0051] The value function fitted by the critic neural network is:
[0052] ,
[0053] in These are the weight estimates of the critic neural network. It is an activation function.
[0054] The control strategy fitted by the actor neural network is as follows:
[0055] ,
[0056] in These are the weight estimates for the actor neural network.
[0057] and The update rules are designed as follows:
[0058] ,
[0059] ,
[0060] in and These are the learning rates for the critic neural network and the actor neural network, respectively.
[0061] The beneficial effects of this invention are:
[0062] 1. This invention proposes a neural network-based observer that can achieve effective control of stochastic multi-agent systems without relying on prior information of control gain;
[0063] 2. In response to the noise and random disturbances that are common in real-world environments, this invention designs a consistent controller based on a reinforcement learning algorithm within the identifier-evaluation-execution framework, which can ensure that the system can still achieve state consistency under the influence of random disturbances.
[0064] 3. This invention integrates a dynamic event triggering mechanism. By introducing auxiliary dynamic variables and dynamic threshold parameters, it avoids the problem of excessively high or low triggering frequency that may occur in traditional static event triggering mechanisms. Thus, while ensuring control performance, it reduces the number of control updates, thereby reducing system energy consumption and communication burden.
[0065] 4. This invention designs the neural network weight update rule by using a positive function related to the Bellman residual for gradient descent. This not only ensures the stability of the system, but also greatly reduces the complexity of the algorithm. Attached Figure Description
[0066] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0067] Figure 1 This is a flowchart of an optimal consensus control method for a stochastic multi-agent system with unknown control gain, according to the present invention. Detailed Implementation
[0068] The technical solutions of the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. However, it should be understood that these practical details are not intended to limit the invention. That is, in some embodiments of the invention, these practical details are not essential.
[0069] This invention provides an optimal consensus control method for a stochastic multi-agent system with unknown control gain. The control steps are illustrated in the flowchart below. Figure 1 As shown.
[0070] I. System Modeling and Problem Description
[0071] contain The topology of a random multi-agent system with followers is represented by a directed graph. To indicate, among which It is a set of nodes. It is an edge set. It is an adjacency matrix. If the nodes are paired... Then there exists a node To the node Information flow, ,otherwise .when When, it is called a node It is a node The neighbors of the node The neighbor set is recorded as .make ,in ,use To represent the diagram The Laplacian matrix. Define the diagonal matrix. This indicates communication between the leader node and the follower nodes, where information from the leader can be directly transmitted to the follower nodes. ,but ;otherwise Assuming at least one follower node is directly connected to the leader node, then .remember .
[0072] Consider a class of A second-order stochastic multi-agent system consisting of 1 follower and 1 leader, the... The dynamic equations of a follower are modeled as the following second-order differential equations:
[0073] ,
[0074] in and Followers ( Position and velocity states, , , and They are all continuous and bounded nonlinear functions. The control gain function is unknown but bounded. It is an independent Wiener process;
[0075] The dynamic equation of the leader is:
[0076] ,
[0077] in and These are the leader's position and speed status, , , and They are all continuous and bounded nonlinear functions.
[0078] The consistency error between the follower state and the leader state is:
[0079] ,
[0080] make Then its dynamic equation can be derived as follows:
[0081] ,
[0082] in , , , .remember .
[0083] Control objective: For the second-order stochastic multi-agent system with unknown control gain, design an event-triggered optimal consensus controller based on reinforcement learning algorithm to achieve leader-follower consensus of the system and ensure that all error signals of the closed-loop system are bounded.
[0084] II. Optimal Consistent Control Algorithm Design
[0085] For the optimal uniform control problem of stochastic MASs, the goal is to achieve optimal uniform control for each agent. Find the optimal controller While ensuring consistency among all agents in the system, minimize the following performance metric function.
[0086] .
[0087] Therefore, the optimal performance index function can be expressed as:
[0088] .
[0089] Therefore, the corresponding HJB equation can be expressed as follows:
[0090] ,
[0091] in .
[0092] According to optimal control theory, the optimal controller It can be solved To obtain, that is:
[0093] .
[0094] It should be noted that here... Currently, events are triggered based on time, which often leads to unnecessary resource consumption. To avoid this problem, this invention will introduce a dynamic event triggering mechanism in the next step.
[0095] III. Design of Dynamic Event Triggering Mechanism
[0096] Using monotonically increasing sequences Indicates the first The first follower At each trigger moment, the sampled state can be represented as: , To determine the trigger time, the following sampling error is defined:
[0097] ,
[0098] When the sampling error exceeds the trigger threshold defined later, the control strategy is updated, and the sampling error is reset to 0. Therefore, the event-triggered control strategy can be expressed as:
[0099] .
[0100] The dynamic event triggering conditions are as follows:
[0101] ,
[0102] in It is a positive parameter, a positive constant. make satisfy Auxiliary dynamic variables The dynamic equations are designed as follows:
[0103] ,
[0104] in and It is a positive constant and satisfies ,also, initial value . It is a dynamic threshold parameter, and its dynamic equation is designed as follows:
[0105] ,
[0106] in It is a normal number.
[0107] IV. Adaptive Identifier Design
[0108] Due to the presence of random disturbances and unknown nonlinear dynamics in the system, the following identifier is designed:
[0109] ,
[0110] in The state of the identifier, This indicates the estimation error of the identifier. These are parameters that need to be designed. and These are the weight estimates and activation function of the labeler neural network, respectively. The update rules are designed as follows:
[0111] ,
[0112] in These are parameters that need to be designed.
[0113] It should be noted that the controller gain in the identifier expression designed here... Since it is still an unknown nonlinear function, the next step of this invention will be to use a neural network to design an observer to reconstruct the unknown control gain.
[0114] V. Observer Design
[0115] The observer is designed as follows:
[0116]
[0117] in It is the state of the observer. This represents the observer's estimation error. These are parameters that need to be designed. and These are the weight estimates and activation function of the labeler neural network, respectively. The update rules are designed as follows:
[0118] ,
[0119] in These are parameters that need to be designed.
[0120] VI. Controller Design
[0121] Based on the above design, and considering both the identifier and the observer, the subsequent discussion can use an observer system instead of the original uniform error system. Therefore, when designing the controller, we use... replace , replace etc.
[0122] The corresponding HJB equation is
[0123] ,
[0124] A two-layer neural network model consisting of a critic neural network and an actor neural network is constructed. The critic neural network is used to approximate the value function of each agent, and the actor neural network is used to approximate the control policy function of each agent.
[0125] The value function fitted by the critic neural network is expressed as follows: ,in These are the ideal weight values for a critic neural network. It is an activation function. This represents the estimation error of the neural network, and it contains positive constants. Make it satisfy The approximate value function is:
[0126] ,
[0127] in These are the weight estimates for the critic neural network.
[0128] The control strategy fitted by the actor neural network is expressed as follows: ,in These are the ideal weight values for the actor neural network. Therefore, the approximate control strategy is:
[0129] ,
[0130] in These are the weight estimates for the actor neural network.
[0131] We designed weight update rules for the critic neural network and the actor neural network based on the gradient descent algorithm, and obtained the optimal control strategy by adjusting the weight parameters.
[0132] and The update rules are designed as follows:
[0133] ,
[0134] ,
[0135] in and These are the learning rates for the critic neural network and the actor neural network, respectively.
[0136] Furthermore, after steps one through six above, the optimal control strategy is obtained. Using this strategy to perform leader-follower consistent control on the stochastic multi-agent system can achieve the control objective described in step one.
[0137] This invention proposes a dynamic event-triggered control algorithm based on reinforcement learning, aiming to solve the cooperative control problem in complex environments such as unknown nonlinearities, random disturbances, and uncertain control gains in stochastic multi-agent systems. This method integrates neural networks, observer design, and the actor-critic framework from reinforcement learning to construct a complete adaptive optimal consensus control strategy. First, by introducing a neural network-based identifier structure, this invention effectively approximates the random disturbances and unknown nonlinear characteristics in the system, thereby improving the system's robustness to external disturbances and model uncertainties. Based on this, a neural observer without prior control gain information is further designed, enabling online reconstruction of the system's internal state and control gain, significantly enhancing the controller's adaptability and practicality. Regarding the control strategy, this invention combines the actor-critic algorithm from reinforcement learning to design an optimal consensus controller, which can ensure consistent convergence of the multi-agent system state in real-world environments where noise and random disturbances are prevalent, improving the overall control performance of the system. Furthermore, this invention innovatively introduces a dynamic event triggering mechanism, effectively overcoming the problem of excessively high or low triggering frequencies that may occur in traditional static event triggering mechanisms by constructing auxiliary dynamic variables and adaptive threshold parameters. This mechanism not only ensures good control performance but also significantly reduces the number of control signal updates, thereby reducing system energy consumption and communication resource consumption. Finally, to improve the learning efficiency and stability of the algorithm, this invention employs a positive function based on Bellman residuals for gradient descent optimization and designs a neural network weight update rule, significantly reducing the complexity of the algorithm implementation while ensuring the stability of the entire closed-loop system. In summary, this invention is significant in both theoretical innovation and engineering application, suitable for efficient cooperative control of multi-agent systems in complex environments, and has broad application prospects and promotional value.
[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An optimal consensus control method for a stochastic multi-agent system with unknown control gain, characterized in that, Includes the following steps: Step 1: Establish a first-order nonlinear stochastic multi-agent system model with a directed graph topology. The multi-agent system consists of... It consists of several intelligent agents, constructs the consistency error between the follower state and the leader state, and derives the consistency error dynamic system; Step 2: Define the value function as a performance index for evaluating the control effect on a stochastic multi-agent system, and obtain the HJB equation based on the uniform error system and the value function; Step 3: Construct a dynamic event-triggered sampling error and threshold function. When the sampling error exceeds the threshold function, the event is triggered, and the control strategy is updated. Step 4: Design an adaptive identifier to approximate the random terms and unknown nonlinearities in the system; Step 5: Design an observer to reconstruct the unknown control gain; Step 6: Construct a two-layer neural network model containing a critic neural network and an actor neural network, and design the corresponding neural network weight update rules to obtain the final controller.
2. The optimal consensus control method for a stochastic multi-agent system with unknown control gain according to claim 1, characterized in that, In step one, in the nonlinear stochastic multi-agent system, the first... The dynamic equations of a follower are modeled as the following second-order differential equations: , in and Followers ( Position and velocity states, , , and They are all continuous and bounded nonlinear functions. The control gain function is unknown but bounded. It is an independent Wiener process; The dynamic equation of the leader is: , in and These are the leader's position and speed status, , , and They are all continuous and bounded nonlinear functions; The consistency error between the follower state and the leader state is: , in Followers The set of neighboring intelligent agents, if ,but ,otherwise If followers If the navigator's information can be received, then ,otherwise ,remember ; make Then its dynamic equation can be derived as follows: , in , , , ,remember .
3. The optimal consensus control method for a stochastic multi-agent system with unknown control gain according to claim 1, characterized in that, In step two, the value function is designed as follows: ; The HJB equation is expressed as: , in It is the optimal control strategy. .
4. The optimal consensus control method for a stochastic multi-agent system with unknown control gain according to claim 1, characterized in that, In step three, the event-triggered sampling error is designed as follows: , among which when hour, , For followers The One trigger moment; The trigger threshold is designed as follows: , in It is a positive parameter, a positive constant. make satisfy Auxiliary dynamic variables The dynamic equations are designed as follows: , in and It is a positive constant and satisfies ,also, initial value . It is a dynamic threshold parameter, and its dynamic equation is designed as follows: , in It is a normal number.
5. The optimal consensus control method for a stochastic multi-agent system with unknown control gain according to claim 1, characterized in that, In step four, the adaptive identifier is designed as follows: , in The state of the identifier, This indicates the estimation error of the identifier. These are parameters that need to be designed. and These are the weight estimates and activation function of the labeler neural network, respectively. The update rules are designed as follows: , in These are parameters that need to be designed.
6. The optimal consensus control method for a stochastic multi-agent system with unknown control gain according to claim 1, characterized in that, In step five, the observer is designed as follows: , in It is the state of the observer. This represents the observer's estimation error. These are parameters that need to be designed. and These are the weight estimates and activation function of the labeler neural network, respectively. The update rules are designed as follows: , in These are parameters that need to be designed.
7. The optimal consensus control method for a stochastic multi-agent system with unknown control gain according to claim 1, characterized in that, In step six, the value function fitted by the critic neural network is: , in These are the weight estimates of the critic neural network. It is an activation function; The control strategy fitted by the actor neural network is as follows: , in These are the weight estimates of the actor neural network; and The update rules are designed as follows: , , in and These are the learning rates for the critic neural network and the actor neural network, respectively.
Citation Information
Cited By
Method and system for controlling consistency of reinforcement learning multi-agent triggered by event on time mark
CN121763784A
Event-triggered reinforcement learning multi-agent consensus control method and system on time scale
CN121763784B