A Multi-Agent Event-Triggered Control Method and System Based on Deep Reinforcement Learning
By constructing a dynamic model and consistency controller of the agent system, combined with the deep reinforcement learning training event triggering strategy, the problem of waste of communication resources and uncertain convergence time in multi-agent systems is solved, and efficient predefined time consistency control in complex systems is achieved.
Patent Information
- Application Number
- CN202510593057.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The event trigger control method of existing multi-agent systems is designed to be conservative in complex systems, resulting in waste of communication and computing resources, and the convergence time of existing consistency control algorithms is uncertain or conservative when the initial state is unknown.
Build a dynamic model and consistency controller of the agent system, introduce convergence time penalty term and time-varying function, train the agent's event triggering strategy using deep reinforcement learning, and obtain the optimal event triggering strategy through deep neural networks to achieve communication savings and fast consistency.
Reduces the design conservatism of event triggering conditions, realizes more adaptable predefined time consistency control in complex systems, and improves the robustness and communication efficiency of the controller.
Smart Images

Figure CN120122458B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi - agent cooperative control, and particularly relates to a multi - agent event - triggered control method and system based on deep reinforcement learning. Background Technique
[0002] The cooperative control technology of multi - agent systems has attracted extensive attention in the past few decades. As a fundamental problem in cooperative control, consensus control aims to ensure that multiple agents reach a state of agreement through local communication, which has become the focus of research scholars. This is due to its wide application in different fields, including multi - robot formation control, power management, and distributed optimization. The key objective of the consensus control problem is to design appropriate strategies to drive all agents to reach a common target state. Although early research mainly focused on asymptotic consensus, these methods require an infinite amount of time to reach consensus, making them impractical for real - world applications.
[0003] Finite - time consensus has developed rapidly due to its faster convergence speed compared to asymptotic consensus. Although finite - time consensus algorithms achieve a faster convergence speed, the estimated convergence time depends on the initial state of the multi - agent system. When the initial state is unknown or unavailable, the results of finite - time consensus may be limited. To solve this problem, fixed - time consensus methods have been proposed, where the estimated settling time is independent of the initial state of the agents. Although the convergence time of fixed - time consensus does not depend on the initial state of the agents, the upper bound of the estimated settling time is still conservative and still affected by the controller parameters. To overcome this challenge, predefined - time consensus control has emerged because its settling time can be set in advance without conservative estimation.
[0004] The above - mentioned consensus control algorithms require agents to continuously update the received information, which may lead to waste of communication and computing resources, especially for agents with limited capabilities. To minimize unnecessary communication, event - triggered control for multi - agent systems has been proposed. This method only runs by updating the controller when the measured state error exceeds a certain threshold. However, for complex systems, the design of the triggering condition is still a difficult problem. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi - agent event - triggered control method and system based on deep reinforcement learning.
[0006] In the first aspect, the present invention provides a multi - agent event - triggered control method based on deep reinforcement learning, which includes the following steps:
[0007] Construct the dynamic model and consensus controller of the agent system; set the observations, actions, and rewards of each agent in the agent system, and introduce a convergence time penalty term into the rewards; use whether the current agent communicates with adjacent agents as the action of the current agent; construct an experience storage area based on the observations, actions, and rewards at different times, and use the experience storage area to train the constructed deep neural network; obtain the optimal event-triggering strategy according to the trained deep neural network and the ε-greedy strategy; use the consensus controller to complete the real-time control of each agent according to the optimal event-triggering strategy.
[0008] Preferably, the rewards of each agent are obtained as follows:
[0009]
[0010] where and are the position states of the current agent i and the adjacent agent j at time t respectively; N i is the set of adjacent agents of the current agent i; ; is the action of the current agent i at time t; is the convergence time penalty term.
[0011] Preferably, if the agent system remains consistent before the specified convergence time, the time when the agent system achieves consistency is used as the convergence time penalty term; if the agent system does not remain consistent before the specified convergence time, a preset time is used as the convergence time penalty term.
[0012] Preferably, the observations of each agent include the position state of the current agent at the current time and the position states of the current agent and adjacent agents at the previous event-triggering time.
[0013] Preferably, the method for obtaining the optimal event-triggering strategy is as follows: obtain a random probability value; if the random probability value is less than the preset value ε, randomly set the action selected by the agent; if the random probability value is greater than or equal to the preset value ε, use the action obtained by the deep neural network according to the observations of the agent as the action selected by the agent.
[0014] Preferably, a consensus controller is constructed by introducing a specified convergence time and a time-varying function; the time-varying function has the following expression:
[0015]
[0016] where is the specified convergence time; β is a preset parameter.
[0017] Preferably, adjacent agents that communicate with the current agent are obtained by constructing the network topology of the multi-agent system.
[0018] In a second aspect, the present invention provides a multi-agent event-triggered control system based on deep reinforcement learning, including a plurality of agents and a controller for controlling the agents; the multi-agent event-triggered control system is used to execute the above-mentioned multi-agent event-triggered control method; the multi-agent event-triggered control system further includes an observation module and an event-triggered control module; the observation module is used to obtain the observation data of the agents; the event-triggered control module is used to select the optimal event-triggered strategy according to the observation data.
[0019] In a third aspect, the present invention provides a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor; the memory stores the computer program; the processor executes the above-mentioned multi-agent event-triggered control method.
[0020] In a fourth aspect, the present invention provides a readable storage medium storing a computer program; when the computer program is executed by a processor, it is used to implement the above-mentioned multi-agent event-triggered control method.
[0021] The beneficial effects of the present invention are as follows:
[0022] 1. The present invention trains the event-triggering conditions for agents to communicate with their neighbors through a reinforcement learning method. Compared with the traditional event-triggered control mechanism, the present invention reduces the design conservatism of the event-triggering conditions and can achieve predefined time consistency at a lower communication frequency, providing a more adaptable solution for event-triggered control in complex systems.
[0023] 2. The present invention introduces a convergence time penalty term into the constructed reward, so that the obtained optimal time-triggered strategy can control the multi-agent system to reach consensus faster; at the same time, the present invention introduces a time-varying function and a specified convergence time into the consensus controller, which can dynamically adjust the control parameters according to the specified convergence time, adapt to the complex and changeable operating environment, and improve the robustness of the controller. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a flowchart of the present invention.
[0025] Figure 2 is a graph showing the change of the positions of five agents in the present invention over time.
[0026] Figure 3 is a graph of the time intervals of the event triggering of the first agent.
[0027] Figure 4Time interval diagram triggered by the second agent event.
[0028] Figure 5 Time interval diagram triggered by the third agent event.
[0029] Figure 6 Time interval diagram triggered by the fourth agent event.
[0030] Figure 7 Time interval diagram triggered by the fifth agent event. Detailed implementation mode
[0031] The present invention will be further described below in conjunction with the accompanying drawings.
[0032] As Figure 1 shown, a multi-agent event-triggered control method based on deep reinforcement learning includes the following steps:
[0033] Step 1: Establish the dynamic model of the agent system
[0034] The dynamic model of the first-order multi-agent is expressed as:
[0035]
[0036] Among them, is the derivative of the position state of the agent at time ; is the control input of the th agent.
[0037] Step 2: Construct the consensus controller of the multi-agent system
[0038] The expression of the adjacency matrix A of the network topology structure of the multi-agent system is:
[0039]
[0040] Among them, the element in the adjacency matrix represents the weight of the edge between the agent and the agent ; if the agent can receive information from the agent , then the agent is the adjacent agent of the agent , ; if the agent cannot receive information from the agent , then ; ; .
[0041] Based on the adjacency matrix Construct a consensus controller for the multi-agent system , and its expression is as follows:
[0042]
[0043] where is an element in the adjacency matrix; and respectively represent the position states of the current agent and the adjacent agent at time ; are all preset parameters, , ; is the specified convergence time; N i is the set of adjacent agents of the current agent ; is a time-varying function, and its expression is:
[0044]
[0045] Step 3. Determine the distributed event-triggered control framework based on deep reinforcement learning
[0046] Reinforcement learning is a method for solving sequential decision-making processes by interacting with the environment through trial and error. At time , the agent observes the state of the current environment and selects the corresponding action according to the policy ; subsequently, the agent receives a reward and transfers from the state to the next state with the state transition probability. The goal of reinforcement learning is to find an optimal policy to maximize the expected discounted reward ; where, G t is the total return starting from time t; represents the discount factor; is the time step passed since time ; is the total time step. In reinforcement learning, the value function is a criterion for measuring the quality of each state or state-action. The state value function corresponding to the policy is denoted as , and the action value function is denoted as .
[0047] The Q-learning algorithm is a main type of reinforcement learning method based on the Markov decision process (MDP), and its action value function can be learned through the following formula:
[0048]
[0049] where, o t is the observation of the agent at time , and α is the learning rate.
[0050] The Deep Q-Network (DQN) algorithm is used to learn the event-triggered policy. It is a variant of the Q-learning algorithm. In the DQN algorithm, the agent will use a deep neural network to approximate the optimal action value function . Construct the tuple representation of the DQN algorithm ; where, o is the observation of the agent; a is the action of the agent; r is the reward of the agent. Considering the partially observable scenario of the multi-agent system, each agent in it is independent and learns its event-triggered policy by receiving its own local observation information and communicating with neighboring agents. The definitions of each element in the tuple representation for a single agent are described respectively:
[0051] ① Observation o: The observation of agent at time is:
[0052]
[0053] where, is the position state of agent at time ; and represent the position states of the current agent and the target agent at time and time respectively; and are the previous event trigger times of the current agent and the target agent respectively.
[0054] ② Action a: The action of agent at time is defined as ; where, represents the discrete action space of agent . When When an event is triggered, the agent will communicate with its neighbors; conversely, if the event is not triggered, the agent will not communicate with its neighbors.
[0055] ③ Reward r: To balance the consensus control performance and communication energy consumption of the multi-agent system, the reward of the agent is designed as follows: where
[0056]
[0057] ; ; is the L2 norm symbol; is the L1 norm symbol; is the convergence time penalty term, expressed as:
[0058]
[0059] where is the time for the system to achieve consensus; t e is the preset time; T0 is the indicator function, T0 = 1 indicates that consensus is maintained before the specified convergence time T s ; T0 = 0 indicates that consensus is not maintained before the specified convergence time T s ;
[0060] In this embodiment, the preset time t e is 500 s.
[0061] Step 4. Obtain the optimal event-triggering strategy
[0062] 4-1. Initialize the reinforcement learning environment.
[0063] Use a deep neural network with parameters to model the event-triggering strategy of each agent and set the action-value function of each agent and the experience replay area for storing historical data.
[0064] 4-2. At each time step, the agent selects an action based on the ε-greedy strategy according to the local observation information , and the specific process is as follows:
[0065] At each time step, obtain the random probability value P and compare the random probability value P with the preset value ; if the random probability value P is less than the preset value ε, randomly set the agent Selected action ; if the random probability value P is greater than or equal to the preset value ε, then the deep neural network is used to select the action according to the observation obtained as the action selected by the agent .
[0066] 4-3. After the agent executes the action , the control input quantity is provided by the consensus controller of the multi-agent system in step two .
[0067] 4-4. Calculate the reward of the agent at time t, and the state of the environment will change to .
[0068] 4-5. Put the historical data of each agent into its respective experience storage area .
[0069] 4-6. After the number of historical data in the experience storage area reaches the set batch size, each agent will select a batch of data M from the experience storage area to train the deep neural network until the maximum number of training rounds is reached; during the training process, the target network value function estimate y of the agent n is:
[0070]
[0071] where r m is the reward corresponding to the m-th data in the batch of data M; γ is the discount factor; Q m is the action value function corresponding to the m-th data in the batch of data M; is the observation corresponding to the (m + 1)-th data in the batch of data M; a m is the action corresponding to the m-th data in the batch of data M; θ m is the network parameter corresponding to the m-th data in the batch of data M;
[0072] The loss function L is expressed as:
[0073]
[0074] where is the observation of the agent in the m-th data; is the action of the agent in the m-th data.
[0075] 4-7. Use the trained deep neural network and ε-greedy strategy to obtain the optimal event-triggering strategy.
[0076] Step Five: According to the obtained optimal event-triggering strategy, use the controller to complete the real-time control of each agent, so that the system states are consistent and the purpose of communication saving is achieved at the same time.
[0077] The goal to be achieved by the multi-agent system is that the position states of all agents can reach consistency within a predetermined time with a low communication frequency. Among them, for a known system, in any initial state and a given predefined convergence time , for any agent and agent in the system, if the following conditions are all satisfied, it means that this system converges within the predefined time:
[0078]
[0079] Embodiment 1
[0080] Taking the multi-agent system to complete the consensus task within the predefined time as an example, a multi-agent event-triggering control method based on deep reinforcement learning is as follows:
[0081] First, use the existing event-triggered predefined-time consensus controller to compare and prove the effectiveness of the framework proposed by the present invention. The controller of this method is the same as the controller of the present invention, and the event-triggering condition of the th agent is defined as follows:
[0082]
[0083] Among them, represents the next triggering moment of the th agent; represents the difference between the state at the previous triggering moment and the state at the current moment of the th agent; , ; is an external dynamic function, and its expression is:
[0084]
[0085] Among them, ; ; .
[0086] It should be noted that the communication topology of the system, the initial state of the agents, and the parameters of the consensus controller are all the same as those of the above method. The only difference is that the present invention uses the DQN algorithm to autonomously learn the event triggering conditions. The predefined convergence time is set to , and the sampling time is , and the total simulation duration is .
[0087] To verify the reliability of the present invention, the initial positions of five agents are randomly generated within the range of . The variation of the positions of the five agents with time is as shown in Figure 2 . The ordinate represents the position of the agent , and the abscissa represents time . Figure 2 It reflects the convergence process of the positions of the five robots. It can be seen from the figure that the multi-agent system quickly reaches consensus within the predefined time range, which reflects the efficiency of the present invention in the predefined time consensus control method.
[0088] The time intervals of event triggering for the five agents are respectively as shown in Figures 3 to 7 . In Figures 3 to 7 , the ordinate represents the size of the time interval , and the abscissa represents time . The comparison of the stabilization time between the method provided in this embodiment and the existing predefined time consensus control method based on event triggering is shown in Table 1 below:
[0089] Table 1 Average event triggering rate
[0090] Method Average event triggering rate The present invention 3.53 Event-triggered predefined time consistency control method 3.78
[0091] It can be seen from Table 1 that the method provided in this embodiment can greatly improve the speed at which the multi-agent system meets the consensus requirements, verifying that the event-triggered control framework based on the deep reinforcement learning algorithm can achieve predefined time consensus control.
[0092] The above embodiments are only the preferred implementation schemes of the present invention, rather than limitations thereto. It should be pointed out that for those skilled in the art: they can still modify and improve the solutions proposed in the foregoing embodiments, and these should also be regarded as the protection scope of the present invention, and they will not affect the implementation effect of the present invention and the practicability of the patent.
Claims
1. A multi-agent event-triggered control method based on deep reinforcement learning, characterized in that: It includes the following steps: Construct the dynamic model and consensus controller of the agent system; set the observations, actions, and rewards of each agent in the agent system, and introduce a convergence time penalty term into the rewards; use whether the current agent communicates with adjacent agents as the action of the current agent; construct an experience storage area based on the observations, actions, and rewards at different times, and use the experience storage area to train the constructed deep neural network; use the trained deep neural network and ε-greedy strategy to obtain the optimal event-triggering strategy; use the consensus controller according to the optimal event-triggering strategy to complete the real-time control of each agent; The rewards of each agent are obtained as follows: ; Among them, and are the position states of the current agent i and the adjacent agent j respectively; N i is the set of adjacent agents of the current agent i; ; is the action of the current agent i; is the convergence time penalty term.
2. The multi-agent event-triggered control method based on deep reinforcement learning according to claim 1, characterized in that: If the agent system remains consistent before the specified convergence time, use the time when the agent system achieves consistency as the convergence time penalty term; if the agent system does not remain consistent before the specified convergence time, use the preset time as the convergence time penalty term.
3. A multi-agent event-triggered control method based on deep reinforcement learning according to claim 1, characterized in that: The observations of each agent include the position state of the current agent at the current moment and the position states of the current agent and adjacent agents at the previous event-triggering moment.
4. A multi-agent event-triggered control method based on deep reinforcement learning according to claim 1, characterized in that: The method for obtaining the optimal event-triggering strategy is as follows: obtain a random probability value; if the random probability value is less than the preset value ε, randomly set the action selected by the agent; if the random probability value is greater than or equal to the preset value ε, use the action obtained by the deep neural network based on the observations of the agent as the action selected by the agent.
5. A multi-agent event-triggered control method based on deep reinforcement learning according to claim 1, characterized in that: Construct a consistency controller through a time-varying function; the time-varying function has the following expression: ; wherein, is the specified convergence time; β is a preset parameter.
6. A multi-agent event-triggered control method based on deep reinforcement learning according to claim 1, characterized in that: Obtain the adjacent agents that communicate with the current agent by constructing the network topology structure of the multi-agent system.
7. A multi-agent event-triggered control system based on deep reinforcement learning, including multiple agents and a controller for controlling the agents; characterized in that: This multi-agent event-triggering control system is used to execute the multi-agent event-triggering control method described in claim 1; this multi-agent event-triggering control system further includes an observation module and an event-triggering control module; The observation module is used to obtain the observation data of the agent; The event-triggering control module is used to select the optimal event-triggering strategy according to the observation data.
8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: The memory stores a computer program; the processor executes a multi-agent event-triggering control method based on deep reinforcement learning as described in any one of claims 1-6.
9. A readable storage medium stores a computer program; characterized in that: When the computer program is executed by the processor, it is used to implement a multi-agent event-triggering control method based on deep reinforcement learning as described in any one of claims 1-6.
Citation Information
Patent Citations
Heterogeneous multi-agent-oriented multi-task strategy game method
CN117575220A